# Section: Blog --- title: "Fivetran + dbt Labs Announces New Capabilities to Make Enterprise Data Agent-Ready at dbt Summit 2026" description: "dbt v2, dbt State are GA and Fivetran + dbt Labs debuts Fivetran Context Layer, dbt Charts and a new open lakehouse vision" url: "https://www.getdbt.com/blog/fivetran-dbt-labs-announces-new-capabilities-to-make-enterprise-data-agent-ready-at-dbt-summit" date: "2026-09-16" authors: ["Elaine Green"] categories: ["Press"] --- # Fivetran + dbt Labs Announces New Capabilities to Make Enterprise Data Agent-Ready at dbt Summit 2026 _Fivetran + dbt Labs makes dbt v2 and dbt State generally available, alongside the debut of Fivetran Context Layer, dbt Charts and a new open lakehouse vision for greater flexibility across storage and compute_ **LAS VEGAS – September 16, 2026** – Fivetran + dbt Labs today announced the general availability of dbt v2 and dbt State, delivering new levels of speed and cost optimization, and introduced Fivetran Context Layer, new dbt Wizard experiences, dbt Charts and an open lakehouse vision for greater flexibility across storage and compute. As enterprises deploy AI agents into production, they need trusted data, context the business controls, and the flexibility to work across the platforms, models and tools already in use. To meet these demands, Fivetran + dbt Labs is advancing its vision for Open Data Infrastructure: a vendor-neutral, interoperable architecture that lets organizations independently choose and evolve their storage, compute, data movement, transformation and visualization technologies at every layer. This enables AI systems to work across platforms using data and context the enterprise owns, not a single vendor. "Every model our customers have built, every test they've written, every metric they've defined already captures the context AI agents need to do meaningful work," said Anjan Kundavaram, Chief Product Officer, Fivetran + dbt Labs. "What we're delivering now is the open infrastructure to put that context to work across systems, while giving organizations the freedom to choose how their data is stored, moved, transformed and used as AI evolves." **A faster, more efficient engine with deeper SQL understanding** [dbt v2](https://docs.getdbt.com/blog/dbt-v2-is-ga) is a full Rust rewrite of the dbt engine built for the scale that teams run at today and for how agents write SQL. It parses a 10,000-model project up to 10x faster than v1 and gives teams and their agents accurate real-time feedback, surfacing errors, column checks, and lineage before anything runs. With this release, the two engine era of Core and Fusion ends. Now, dbt is one engine with two versions: dbt Core v1, the python implementation, is dbt v1. Fusion, the Rust implementation, has become dbt v2. Both versions remain Apache 2.0-licensed and security-supported. [dbt State](https://www.getdbt.com/blog/dbt-state-is-ga) determines what has changed by checking warehouse metadata and model SQL, then builds, skips, clones or defers each run accordingly. This simplifies orchestration and allows engineers to iterate faster without complex development rituals, while reducing unnecessary warehouse compute. "dbt State has been a paradigm shift for how we work,” said Gordon Curzon, Head of Analytics Engineering, Virgin Media O2. “With freshness codified, simpler orchestration, and freed-up developer capacity, we focus more time on initiatives that add value to our business on top of the 25% savings on both job run time and BigQuery compute costs." **More flexibility across storage and compute** Open Data Infrastructure centers on a customer-owned data layer built on open formats, avoiding vendor lock-in for storage and compute so organizations can store data once and access it through different engines for different use cases. Fivetran's Managed Data Lake Service, already generally available, organizes, structures and maintains data as managed Apache Iceberg™ tables in customers' own cloud storage. Lake Compute, now in Private Beta, is a single-node SQL engine built on DuckDB and runs dbt models directly against Apache Iceberg™ tables, built and priced specifically for transformation, not general-purpose compute. Together, the two give teams the flexibility to run each workload on whichever engine fits best, optimizing cost without re-platforming. **An open standard for agent context** Agents are only as trustworthy as the context they can access. Fivetran Context Layer (Private Beta) unifies the data and metadata needed to give LLMs and AI agents relevant context, building on dbt’s structured context and adding unstructured knowledge, like docs and Slack threads. This service uses Agents Schema, an open source standard, that centralizes context in a structured, extensible format directly in the data warehouse. The context is accessible to teams via preferred MCP or AI tools, including generally available integrations through AI marketplaces including Anthropic and a plugin in ChatGPT. **One agent, grounded in your dbt project – wherever you work** Coding agents can now write SQL as well as most engineers, but writing code isn't the same as understanding a governed dbt project, including its lineage, its tests, its contracts, and what breaks when something changes. dbt Wizard in the dbt platform (Public Preview) is built to close that gap. It's natively connected to your project, knows which tool to call, pulls the right context automatically, and proactively validates changes before they ship. Wizard is also expanding beyond the dbt platform with Wizard CLI (Public Beta), bringing the project-grounded agent directly into the terminal, and Wizard Desktop (Private Beta), a dedicated local workspace for longer, more complex work. Wizard Explore Mode (Public Preview) brings conversational analytics to business users, enabling them to ask questions in plain language and get answers grounded in the same dbt project the data team maintains. When an answer falls short, those questions can also surface what the data team should improve next. **A shared language for BI, built for humans and agents** [dbt Charts](https://docs.dbtcharts.com/) (Public Beta) brings governed BI alongside the models it depends on. Instead of governance living in a separate, closed tool, they are defined as YAML and version-controlled alongside the dbt models they reference, creating a shared, declarative format that both humans and AI agents can read, write and review. **Customers building with Fivetran + dbt Labs** “Since rolling out dbt State, we’ve reduced warehouse costs by 59% on scheduled jobs in dbt platform,” said Chris Shepherd, Principal Data Engineer, RxBenefits. “That’s $8,173.23 in the first 60 days alone on top of a Snowflake adaptive warehouse. We’ve reused 716k models instead of rebuilding, which reduced query run time a total of 14 days, 11 hours, and 15 minutes over the same period.” “We've been impressed by the flexibility Lake Compute gives us. Now we can choose where each dbt workload runs, and use whichever engine actually fits the job,” said Tyson Doberneck, Senior Data Engineer, Obie. "dbt Wizard is changing how we work. Instead of hand-coding everything, we draft logic with an agent that already has full context on our jobs and our codebase. No more finding and uploading a manifest file just to explain myself, that step used to slow down every request. Now we're pointing it at sales and marketing data too, so analysts get answers themselves instead of waiting on my team," said Farin Fukunaga, Data Engineering Lead, Paylocity. To learn more about the product innovations unveiled at dbt Summit 2026, read the [recap](https://www.getdbt.com/blog/dbt-summit-2026-product-announcements) or register to watch the full keynote: [https://www.getdbt.com/dbt-summit/registration/online](https://www.getdbt.com/dbt-summit/registration/online) **About Fivetran + dbt Labs** Fivetran + dbt Labs deliver the data infrastructure layer that makes agents trustworthy – from the moment data moves, through every transformation, to the context an agent reasons from. The Fivetran platform moves, manages, and transforms data from every system a business runs on into a secure, reliable foundation engineered to evolve, with the flexibility to work across clouds, engines, and tools. With Fivetran, analytics, operations, and AI run on data you trust and control. Thousands of organizations worldwide, including OpenAI, LVMH, Pfizer, and Verizon, rely on Fivetran to turn data into a competitive advantage. Learn more at Fivetran.com, or follow Fivetran on LinkedIn. Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Globally, more than 100,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at getdbt.com, and follow dbt Labs on LinkedIn, X, Instagram, and YouTube. --- --- title: "Everything we announced at dbt Summit and why it matters" description: "Every product announced at dbt Summit, from dbt v2 and dbt State to dbt Wizard and dbt Charts, and why each one matters." url: "https://www.getdbt.com/blog/dbt-summit-2026-product-announcements" date: "2026-09-16" authors: ["Corinne Hallander"] categories: ["Product"] --- # Everything we announced at dbt Summit and why it matters A decade ago, dbt gave people a name for something they were already trying to do: write SQL like software engineers, version it, test it, document it, trust it. That idea became a new practice: analytics engineering. The people who built their careers on it became some of the most capable data professionals in the industry. Now, those same people are asking a harder question: what impact will AI have on the practice of analytics engineering? AI needs an engine that's fast enough to keep up, context that's‌ trustworthy, and agents that understand your business instead of guessing at it. The question isn't whether analytics engineers still matter; it’s how do they level up for this new era? This week at [dbt Summit](https://www.getdbt.com/dbt-summit/registration/online), we welcomed 2,000 data professionals to Las Vegas (and thousands more online) to discuss the future of analytics in the AI age. Fivetran and dbt Labs are coming together around one thesis: the data foundation that makes analytics trustworthy is the same foundation that makes AI trustworthy. To help our users level up on both fronts, we announced a series of new features across the dbt and Fivetran product portfolios. ## Level up the engine ### One dbt, one engine, built to move as fast as you do dbt has always evolved with growing data workloads. Last year, we introduced the dbt Fusion engine, a full rewrite of dbt in Rust with native SQL comprehension and dramatically faster performance than the original Python-based standard. But maintaining two engines created real friction, both for us and for the thousands of teams trying to figure out which one to build on. **dbt v2** ends that. Now GA, [dbt v2](https://docs.getdbt.com/blog/dbt-v2-is-ga) is one modern, Rust-based engine powering all of dbt, whether you work locally or in the dbt platform. It's the same workflow, on a faster foundation—thanks to all the innovation in Fusion over the past 18 months. ```json { "_key": "c077b514b103", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/_6ZMtNWheiQ" } ``` We publish [two distributions of v2](https://docs.getdbt.com/blog/comparing-dbt-and-dbt-oss): - The superset `dbt `includes all parts of the framework, super-charged with SQL comprehension features. It’s free to use, with additional optional paid features (for example, dbt State). - The subset distribution `dbt-oss` includes _only_ the components that have an Apache 2 license dbt Core isn't going anywhere either. It’s still Apache 2.0, still open source, simply renamed as the previous version, dbt v1. Adapter availability: - BigQuery, Databricks, DuckDB, Redshift, and Snowflake are GA - ClickHouse and Spark are in beta and available to install and test locally - Athena, Fabric, and Postgres are coming soon > “With the new dbt v2 engine, the performance improvements showed up across the entire development experience. Teams spent less time waiting for processes to complete, moved changes through the pipeline faster, and could focus more of their time on building and delivering data products.” — Vishesh Jain, Delivery Lead, Data Analytics Platform at RMIT University [Get started](https://docs.getdbt.com/docs/local/install-dbt?install-method=pip&version=2) on v2 today. ## Build what's changed, skip what hasn't with dbt State, now GA **Now officially GA**,** [dbt State](https://www.getdbt.com/blog/dbt-state-is-ga)** checks your warehouse metadata and model SQL for what's changed, then builds, skips, clones, or defers each run accordingly. That intelligence results in an average 15-30%+ reduction in warehouse compute and removes the manual syntax and workarounds teams built to avoid running more than they needed to. > “Since rolling out dbt State, we’ve reduced warehouse costs by 59% on scheduled jobs in the dbt platform, which runs on top of a Snowflake adaptive warehouse. We’ve reused over 700k models instead of rebuilding, which reduced query run time by two weeks over a 60 day period.” — Chris Shepherd, Principal Data Engineer, RxBenefits With dbt State, freshness moves from the job to the model. Every model carries its own freshness requirement in code, a `lag_tolerance` that says how stale it's allowed to be, simplifying orchestration. And because codified freshness rules and automatic reuse constrain what any run can cost, the guardrails now sit in the infrastructure instead of with whoever, or whichever agent, issues the command. The result is felt daily in development. dbt State removes complex dev setup rituals and risks of accidental builds for faster, safer dev work. > “dbt State has been a paradigm shift for how we work. With freshness codified, simpler orchestration, and ultimately, freed-up developer capacity, we focus more time on initiatives that add value to our business. And that’s on top of the 25% savings on both job run time and BigQuery compute costs.” — Gordon Curzon, Head of Analytics Engineering, Virgin Media O2 ```json { "_key": "eb8188a221e6", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/mczQsM8pl78" } ``` dbt State runs on dbt v1.7 through v2, wherever you run dbt: locally, in the dbt platform, with your own orchestrator, and across Snowflake, BigQuery, Databricks, and Redshift. [Get started on dbt State today](https://docs.getdbt.com/docs/deploy/dbt-state-setup?version=2) so you can optimize costs, save time, and level up wherever you run dbt. ## Own your data. Align costs to value. Introducing Lake Compute, now in Beta Most companies run every dbt model, from massive joins to simple staging tables, on the same warehouse compute. With the rise of Apache Iceberg, you can store your data in open-table formats, and have the flexibility to choose the right compute engine for each workload. We now offer that choice with **Lake Compute (Private Beta)**,** **a single-node SQL engine built on DuckDB for dbt that runs transformations directly against Apache Iceberg tables. Tag one model or a hundred to run on Lake Compute. The rest keep running on your warehouse, and `ref`s keep working across both. Engine choice becomes a per-model decision instead of a replatforming program. With Lake Compute, we’re working towards the promise of an open data lakehouse: one where you can choose the right compute for the job. Learn more about Lake Compute [here](https://docs.getdbt.com/docs/lake-compute). Want to see it in action? Join our [upcoming webinar](https://www.getdbt.com/resources/webinars/cost-optimized-transformations-with-dbt-and-apache-iceberg-on-open-multi-engine-compute) on cost-optimized transformations with dbt, Apache Iceberg, and multi-engine compute. We'll walk through a live, dual-engine dbt project running partial Snowflake, partial Lake Compute, on the same Iceberg tables, so you can see exactly what "engine choice as a per-model decision" looks like in practice. [Save your seat here](https://www.getdbt.com/resources/webinars/cost-optimized-transformations-with-dbt-and-apache-iceberg-on-open-multi-engine-compute). ## Level up for AI: context, engineered AI agents don't necessarily need _more_ of your data. They need an understanding of what the data means, where it came from, whether it's fresh, and who owns it. That foundation is already in place: [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl?version=2) governs metrics, [Agents Schema](https://www.fivetran.com/blog/how-agents-schema-brings-trusted-business-context-to-ai) centralizes context, and [dbt MCP Server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2) exposes your models, metrics, lineage, and test results to any AI agent. Delivering structured dbt context to your favorite AI tools just got easier with new **out-of-the-box integrations with Anthropic and a plugin in ChatGPT (GA)**. No more multiple MCP servers to manage. Just one click, and your team can securely access structured context from your dbt project instantly. Structured context gets you closer to reliable AI answers. But most of what a business runs on doesn't live in structured tables at all. It lives in call recordings, support tickets, Slack threads, and it changes constantly. That's what [**Fivetran Context Layer**](https://fivetran.com/docs/context-layer) is built for: turning every source your business runs on—both structured sources like dbt and unstructured ones that BI tools never touched—into context for the AI tools you're already using, from Claude to Slack. Because Fivetran and dbt have visibility into your entire data estate, this context gets built and maintained as data moves and transforms, not stitched together after the fact. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0864156bd3668028d1019a580b859b1e953f9568-2048x1342.jpg) [Sign up](https://go.fivetran.com/signup/context-layer) for the early access program. Want to go deeper on this? Join our session, [_From Analytics Engineer to Context Engineer: A dbt Playbook_](https://www.getdbt.com/resources/webinars/from-analytics-engineer-to-context-engineer-a-dbt-playbook), where we make the case that context engineering is analytics engineering with a new last mile. We'll walk through the exact patterns dbt Labs uses internally to turn unstructured sources into governed, versioned context, and introduce the `dbt_context_engineering` package that puts those patterns into practice. [Save your seat here](https://www.getdbt.com/resources/webinars/from-analytics-engineer-to-context-engineer-a-dbt-playbook). ## Level up with AI AI is raising the bar on what it takes to ship trustworthy data: more models, more context, less room to guess. Most of that time gets lost relearning what should already be known: agents rediscovering a project's lineage and tests before they can start, visualizations living outside the codebase entirely. [**dbt Wizard**](https://www.getdbt.com/product/dbt-wizard) is an agent built specifically for analytics engineering. It’s grounded natively in your project, so it already knows which tools to call and which context to pull without any setup. It validates proactively: checking upstream and downstream impact, compiling and building the change before anyone sees the diff. And because data work is visual, you can review all of that against the full DAG. Wizard in the dbt platform is now in Public Preview. [**dbt Wizard Explore Mode**](https://docs.getdbt.com/docs/platform/wizard-home#ask-questions-in-explore-mode) (Public Preview) brings conversational analytics right where the data work already happens. It lets business users and data teams ask questions in plain language and get answers grounded in the same dbt project your data team already maintains. When an answer falls short, you're already one step from the model that needs fixing. [**Wizard CLI**](https://docs.getdbt.com/docs/dbt-ai/wizard-cli) (Public Beta) puts the same project-grounded agent in the terminal you're already running dbt from. No new app, no new tab, and no platform account required. And [**Wizard Desktop**](https://docs.getdbt.com/docs/dbt-ai/wizard-desktop?version=2) (Private Beta) picks up where the terminal runs out of room: a dedicated local workspace for longer, more complex work. You can run several tasks side by side instead of juggling windows with a visual preview of the code, data, and lineage before anything ships. Learn more [here](https://www.getdbt.com/product/dbt-wizard). ```json { "_key": "0f25474437ae", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/cEwRVyYZd1I" } ``` ## Closing the last gap: dashboards as code Asking questions is only one way people need to work with data. Sometimes what they need is a dashboard—one they'll come back to every day. And that's where the story gets uncomfortable: everywhere else in the stack, teams have brought in real engineering discipline—version control, code review, CI. dbt was the reason SQL went from copy-pasted queries to models you can trust. But the dashboard at the end of that pipeline—the thing an executive looks at—was still locked in, trapped inside whatever proprietary tool you bought, stored in someone else's format, disconnected from the code that produces the numbers. [dbt Charts](https://docs.dbtcharts.com/) (Public Beta) closes that gap. It brings discipline to dashboards: built as YAML, version-controlled right next to the models they depend on, living in the same repo, the same pull request, the same CI as the SQL underneath it. Because it's declarative, it becomes a shared language, one humans and agents can both read, write, and review, just as reliably as they do your SQL. [Check out dbt Charts](https://dbtcharts.com/) today. ## We’re all leveling up These announcements represent more than a set of new products. Together, they represent the forward direction of Fivetran + dbt Labs. AI is changing the game, and we are building the data foundation to help you succeed in the agentic AI era. Here’s a recap of how to get started with any of these new products: - [Upgrade to dbt v2](https://docs.getdbt.com/docs/local/install-dbt?install-method=pip&version=2) for faster development, SQL comprehension, and more. - Turn on [dbt State ](https://www.getdbt.com/product/dbt-state)and start simplifying your development and saving on compute. - Use [Lake Compute](https://docs.getdbt.com/docs/lake-compute) to run the right engine for each model, without a replatforming project. - Start developing agentically with [dbt Wizard](https://www.getdbt.com/product/dbt-wizard)—it’s available in the [dbt platform](https://docs.getdbt.com/docs/dbt-ai/wizard-ide?version=2), in a brand new [Desktop app](https://docs.getdbt.com/docs/dbt-ai/wizard-desktop?version=2), and via the [CLI](https://docs.getdbt.com/docs/dbt-ai/wizard-cli). - [Explore mode](https://docs.getdbt.com/docs/platform/wizard-home#ask-questions-in-explore-mode) in Wizard makes it easy for business users to chat with your data. [Add a read-only user](https://docs.getdbt.com/docs/platform/wizard-read-only-users?version=2#set-your-team-up-for-good-answers) at no cost, they'll learn about the data, and you'll learn a lot from what they ask. - Get on the waitlist for [Fivetran Context Layer](https://go.fivetran.com/signup/context-layer). - Try [dbt Charts](https://dbtcharts.com/), the first language-first BI layer. --- --- title: "We built dbt State to stop rebuilding what hadn't changed" description: "dbt State is now generally available everywhere you run dbt. See how it cuts compute costs and speeds up development." url: "https://www.getdbt.com/blog/dbt-state-is-ga" date: "2026-09-16" authors: ["David Macias"] categories: ["Product"] --- # We built dbt State to stop rebuilding what hadn't changed ```json { "_key": "191612bc0735", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/mczQsM8pl78" } ``` Today [dbt State](https://www.getdbt.com/product/dbt-state) is generally available, everywhere you run dbt. This includes your own orchestrator: Airflow, Dagster, GitHub Actions, or a laptop. And as of today, it’s available on Snowflake’s dbt projects. It works on Snowflake, BigQuery, Databricks, and Redshift. Most dbt jobs run on a schedule. dbt State makes them run on change instead. On every run, it reads your model SQL and your warehouse metadata, works out whether a model's result would be different, and then builds, skips, clones, or defers each node accordingly. There's no selection syntax to write, no manifest scripts, and no custom orchestration to maintain We built it on a straightforward idea: an enormous amount of redundant work happens in data pipelines, and reusing what hasn't changed would save teams hours of run time and warehouse compute. That has held true. > “Since rolling out dbt State, we’ve reduced costs by 59% on scheduled jobs in the dbt platform running on a Snowflake adaptive warehouse. We’ve reused over 700k models instead of rebuilding, which reduced query run time by two weeks over a 60 day period.” — Chris Shepherd, Principal Data Engineer, RxBenefits On average, more than 22 million models are built on dbt every single day. Most of the time the upstream data hasn't refreshed, so thousands of models rebuild to produce a result identical to the run before. Early adopters have averaged 15-30% compute savings by stopping that. Simply turning it on nets some benefits, but the bigger jump in compute efficiency happens when teams apply tuned configurations, or move to an adaptive warehouse on Snowflake. More on that below. ## The savings are real, and they're not what teams are talking about Compute savings are the easiest thing to measure, which is why they're usually the reason you turn dbt State on. They show up on the next bill, and they're what gets the project approved. But when we went back to the teams who have been running dbt State in production _and development_ for weeks and asked what changed, almost nobody led with the number. > "dbt State created a paradigm shift in how we work. With capacity freed up, freshness codified, and simpler, yet smarter orchestration, we can focus on initiatives that add value to our business. And that's on top of the 25% savings on both job run time and BigQuery compute costs." — Gordon Curzon, Head of Analytics Engineering and Data Modeling, Virgin Media O2 ## Orchestration stops being a scheduling problem Before dbt State, the way to hit a freshness SLA is to build a job for it. Then another job, on a different schedule, for the part of the DAG with different requirements. Then tags to trigger the right subsets, selectors to narrow them, and source freshness gating layered on top to stop jobs firing when nothing has landed. It works, but it also becomes brittle over time. dbt State changes moves the freshness from the job schedule to the model. Instead of a job deciding what runs, every model carries its own freshness requirement in code, a `lag_tolerance` that says how stale a model is allowed to be. Then dbt State decides, per model, per run, whether that requirement is met. Instead of asking which job a model belongs to, you ask how fresh it needs to be. That's a question an analytics engineer can answer, in a pull request, with review. At Custom Ink, a four-person analytics engineering team supports around 25 analysts, data scientists, and machine learning engineers across a project of more than 1,000 models. Consolidating that into a single deployment is the part they keep coming back to. > "We've been able to simplify our previous setup down to one run, which is amazing, because it's a very clean deployment. Everybody knows what's going to run. There's no confusion." — Brett Petersen, Director of Data and Metrics, Custom Ink ## Development is where you notice it every day The second thing we heard is that dbt State changed development just as much as it changed deployment. In dev, the old ritual is familiar to anyone who has worked in a large project. You want to edit one model in the middle of a long lineage. So you clone production into your dev schema, or you write a macro to do it, or you remember to turn on deferral, or you skip all of that, rebuild the upstream chain, and wait. As best as you can, you track how stale your copy has become. dbt State does the cloning and reuse automatically. You build what you're working on, and the rest is reused. > "Now you just dbt build --select what you're working on. dbt compares against prod state, reuses every fresh upstream at zero compute, and rebuilds only your change. The whole chain resolves in seconds. No clone command, no waiting, no dev-environment ritual, the state tracking you used to do in your head is now the tool's job. It saves so much development time." — Mykkel Ryda**l**, Lead Data Engineer, Joe & The Juice At Joe & The Juice, iteration on a core sales fact table of 400M+ rows used to take 15-25 minutes a cycle, and `--full-refresh` was blocked in code because it was too expensive to allow. Setup before new dev work went from 15-25 minutes to zero. ## The guardrails move into the infrastructure For years, the safety net in dev has been knowledge. Nothing technically stops a dbt build from kicking off an entire pipeline. What stops it is a developer knowing the right selection syntax, which means the guardrail is only as good as the least experienced person running the command. It becomes untenable and harder to scale as development work becomes evermore agentic. dbt State is a structural guardrail rather than a manually enforced one. Codified freshness rules and automatic reuse constrain what any run can cost, regardless of who or which AI agent issued it. > "With dbt State doing the cloning and reuse automatically, the guardrails are just in the infrastructure now, whether it's an analyst or an agent doing the work." — Brett Petersen, Director of Data and Metrics, Custom Ink Custom Ink leans heavily on AI-assisted development, as many teams do these days, and one clean job flow is part of what makes that workable. It’s consistent DAG behavior at scale. ## What dbt State GA means dbt State is out of preview and generally available. Setup is turning it on. Run a job with dbt State enabled twice and see `No-Op` (no operation) show up over and over in your logs. Declare freshness SLAs or other configurations with precise controls: `lag_tolerance`, `require_fresh_data_from`, `compare_unrendered_code`, and more. Use `dbt state explain` to see why any node was built, skipped, cloned, or deferred. Make sure the rest of the data team turns on dbt State in local development so they can iterate faster and lower the cognitive overhead it takes to minimize risk of costly or potentially breaking builds. We worked with Snowflake to bring dbt State’s intelligent reuse to Snowflake dbt Projects, so as of today, those teams can skip unchanged models without moving where they work. dbt State on platform includes Cost Insights and now has a richer `explain` experience for more insight into how to optimize and prove ROI. It also manages concurrent builds so two jobs don't clash on the same model at the same time. Pricing is consumption-based and tied to reuse rather than builds. You pay for daily active target tables (DATT), so each distinct model or test that dbt State skips, clones, or reuses on a given day. Every reuse after the first within the same day is free. Price is independent of table size and compute, so running jobs more often doesn't increase what you pay. In fact, running more often tends to save more. And every user gets a 30-day free trial. Custom pricing is available for upfront commits. See the [docs](https://docs.getdbt.com/docs/deploy/dbt-state-setup?version=2) to get started today. --- --- title: "Celebrating the 2026 dbt partner of the year winners" description: "Meet the 2026 dbt Labs Partner of the Year winners: phData, Snowflake, Cívica, Datum Studio, and 66degrees." url: "https://www.getdbt.com/blog/2026-partner-of-the-year-winners" date: "2026-09-15" authors: ["Yahsmene Butler"] categories: ["Partnerships"] --- # Celebrating the 2026 dbt partner of the year winners Our partners help data teams level up their dbt implementations: faster rollouts, cleaner data foundations, and AI initiatives that ship. At Partner Day during dbt Summit 2026, we recognized five partners who did that work at the highest level this year. Here’s who won, and why. ## Partner of the year: phData phData earned partner of the year for the fourth consecutive year for what its customers walk away with: dbt implementations and data foundations to build trusted agents on, and a certified bench ready to take on the hardest problems. As a Visionary-tier partner, phData is a team that shared customers trust with their most critical data programs. > “Winning dbt Labs Partner of the Year for the fourth year in a row says more about the depth of this partnership than any single project could. dbt has become the backbone of how we help clients turn raw data into something the business can actually trust and reason with. We’re proud of the work we’ve done together, and even more excited about where dbt is headed as it becomes central to how our clients build agentic and AI use cases on top of their data.” — Dustin Dorsey, Senior Director, Data Engineering, phData ## Technology partner of the year: Snowflake Snowflake earned technology partner of the year for the joint value that customers get out of the partnership. In every region, teams running dbt on Snowflake are standing up trusted, production-grade pipelines faster, with dbt Labs and Snowflake engineers working the same problems together rather than handing customers off to each other. It’s one of the most widely adopted foundations in the ecosystem. > “Snowflake and dbt Labs share a vision of empowering data teams to move faster with confidence, and this partnership continues to raise the bar for what our joint customers can achieve. Together, we’ve made it simpler than ever for organizations to build trusted, production-grade data pipelines on Snowflake. We’re proud to be named dbt Labs Partner of the Year and excited to deepen our collaboration to help even more customers turn their data into real business impact.” — Rodrigo Rocha, VP, Global ISV and Technology Partnerships, Snowflake ## EMEA partner of the year: Cívica Cívica earned EMEA partner of the year for the quality of what customers get, not just the volume. More customers in the region do their dbt work with Cívica than with any other partner, and every engagement is backed by a certified delivery team. It’s a Visionary-tier partner customers across EMEA rely on to get it right the first time. > “Being recognized as dbt Labs Partner of the Year in EMEA is a tremendous honor and a true testament to the hard work and dedication of our entire team. What excites us most about collaborating with dbt is the strength of its community: a vibrant ecosystem where open collaboration, innovation, and shared learning help drive the entire industry forward. We’re more motivated than ever to continue accelerating the future of data alongside this incredible network.” — Josep Roig, CEO, Cívica ## APJ partner of the year: Datum Studio Datum Studio earned APJ partner of the year for the depth of what it delivered for customers in Japan this year: more joint dbt work than any other partner in the region, as many customer wins as any partner globally, and one of the strongest certified delivery teams in APJ. In a market where technical depth is everything, Datum Studio was the team customers turned to this year. > “dbt has become the de facto standard for analytics engineering, and our role is to make that standard take root in the Japanese market, not just as a tool, but as a way of working. What excites me most about the partnership is where it’s heading: as AI agents start writing and reviewing transformations, the foundation dbt provides, lineage, testing, documentation, governance, and shared context, is what makes that work trustworthy at scale.” — Sohei Takechi, President and CEO, Datum Studio ## Emerging partner of the year: 66degrees 66degrees earned emerging partner of the year for how fast it got to real customer work. In under a year, 66degrees went from a standing start to delivering dbt projects for customers, bringing Google Cloud depth to shared accounts, and showing up alongside dbt Labs at field and virtual events. This is what a partnership looks like when a team decides to invest. > “66degrees is honored by this recognition as dbt Labs’ Emerging Partner of the Year, which celebrates our long-term Fivetran partnership and our years spent building dbt directly into our core accelerators, including Hub66 and Paradigm Data, our agentic data product delivery agents. Fivetran and dbt together define a transformative direction for data architecture, establishing the seamless pipeline required for modern analytics and AI. As both platforms innovate, we look forward to evolving our accelerators and driving the next era of intelligent data delivery right alongside them, giving our customers a best-in-class data platform.” — Daniel Zagales, SVP, Data and Analytics, 66degrees ## What this means for the dbt community These five partners represent five different ways of showing up for customers: depth of delivery, regional expertise, and the speed to build something new. Congratulations to phData, Snowflake, Cívica, Datum Studio, and 66degrees on the well-earned recognition. [Become a dbt Labs partner today.](https://www.getdbt.com/partners) --- --- title: "Why your AI pilot stalled at the context gap" description: "Many AI projects are stuck. Don’t blame the models. The problem is a lack of trusted context. Here’s how you solve it." url: "https://www.getdbt.com/blog/why-your-ai-pilot-stalled-at-the-context-gap" date: "2026-08-27" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Why your AI pilot stalled at the context gap Everyone has high hopes for the value they can derive from their agentic AI projects. The dream is to get them into production where they can assist users and make autonomous decisions that drive the business forward. But many projects get stuck in the pilot phase. The numbers are stark: - [Only 16% of companies have deployed agentic AI](https://www.infosys.com/newsroom/features/2026/enterprises-scaled-agentic-ai.html) at an enterprise scale. - [Nearly 80% of companies in one survey](https://www.cloudera.com/about/news-and-blogs/press-releases/2026-04-14-nearly-80-percent-of-enterprises-say-ai-is-held-back-by-data-access-challenges-cloudera-report-finds.html) said their AI projects are constrained by data access issues - [70% of those asked by Deloitte](https://www.deloitte.com/us/en/about/press-room/deloitte-survey-examines-ai-readiness-agentic-ai-success.html) said they don’t feel they can adequately trust and govern agents This results in an all-too-familiar situation where AI agents either guess about data or make it up completely. An agent that finds multiple conflicting definitions of “revenue” across data sources might arbitrarily pick one. The problem occurs when AI agents are deprived of the governed context they need to make informed decisions. Let’s look at this context problem, why ungoverned agents fail, and how the dbt platform enables you to build scalable AI agents that everyone in your company can trust. ## The four ways an ungoverned agent fails Without trusted, governed data, [an agentic AI solution that works under test conditions often fails](https://www.getdbt.com/blog/why-agentics-projects-fail-and-how-to-fix-them) when faced with real user questions. There are four common reasons why: **It writes unreliable SQL**. This isn’t often an outright syntactic failure; usually, the SQL parses and runs. The problem is that it’s selecting the wrong values. **It invents or misreads metric definitions**. Without sufficient context, an agent might use stale, missing, or fabricated metrics. This problem is exacerbated by a lack of reviews. **It has no guardrails and no audit trail**. No auxiliary processes check the agent’s work, and there’s no record you can check to verify how it reached its conclusions. **It creates rising compute costs due to inefficient work**. Left to their own devices, AI agents may use more compute than necessary, causing your data processing costs to spike. They’re getting the work done, but processing is eating your profits. None of these are “sometimes agents make a mistake” issues. These are predictable and, fortunately, fixable problems. You just need to take the correct approach to data. For example: - SQL generation can be validated through rigorous testing. You can also train your models and agents to learn how your data works. - Revenue can be centrally defined and shared across teams using a [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction), instead of spread across dozens of data stores. ## Machine-readable governance In the past, we relied on humans to manually review data and ensure its accuracy. We’re producing too much data for that to be a scalable approach in the AI age. To make agentic AI truly scalable, you need governance. But not the type where everything is written down in a large document no one reads. You need **computational** and **machine-readable** governance. In a computational model, governance is enforced using several capabilities: - [Contracts](https://roundup.getdbt.com/p/contracts-have-consequences). A contract is a machine-readable description of how the data is shaped, how it functions, and how it differs between releases, along with the endpoints used to access it. Data that doesn’t meet a contract fails to ship, keeping a class of errors out of production. The contract also enables agents to discover and use the data, particularly as its shape evolves over time. - [Tests](https://www.getdbt.com/blog/data-testing). Data needs to be tested the same way we test software. Tests that ensure correct data can be run when shipping new data transformation changes, and run periodically in production to ensure ongoing data health. - [Semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction). A semantic layer provides one central location for all metric definitions. This eliminates the agent from guessing what “revenue” means. Each metric provides additional metadata that’s invaluable to AI agents: where it came from, how it was derived, and its business purpose. Without this machine-readable approach to governance, you can’t guarantee that your AI agents will return accurate answers at scale. ## The shift to agent consumption of data Historically, the primary consumers of data have been humans. It’s quickly becoming AI agents, which we humans now rely on to help distill the vast amounts of information we keep generating. [Fivetran and dbt Labs realized that, together, we could do more to advance a new era of trusted, Open Data Infrastructure for AI at scale](https://www.getdbt.com/blog/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents). Your AI agents don’t need to languish in the prototype phase. With the dbt platform, you have the tools you need to create high-quality, governed, and trusted data that enables your agents to make accurate decisions. To learn how, [watch the full demo of the dbt platform in action](https://www.getdbt.com/resources/webinars/dbt-platform-live-demo). --- --- title: "Scaling AI is easy. Trusting it is hard." description: "As AI scales, cracks in data trust and governance start to show. Here's what breaks first, and what it takes to fix it." url: "https://www.getdbt.com/blog/scaling-ai-is-easy-trusting-it-is-hard" date: "2026-08-25" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Scaling AI is easy. Trusting it is hard. Organizations are moving beyond experimentation and embedding AI into everyday business operations. As adoption accelerates, many are discovering that scaling AI successfully requires far more than deploying increasingly powerful models. Over the next three years, [92% of companies plan to increase their AI investments, yet only 1% consider themselves mature in AI deployment](https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/superagency-in-the-workplace-empowering-people-to-unlock-ais-full-potential-at-work). As organizations scale AI and agents, trusted data infrastructure becomes the foundation for trusted AI—from copilots to autonomous agents. Without trusted data, governance, and business context, even the most advanced AI systems struggle to deliver reliable, explainable outcomes. We explored these foundational principles in [The Data Leader's Primer for Agentic AI.](https://www.getdbt.com/resources/the-data-leader-s-primer-for-agentic-ai) Here, we examine what starts to break when those foundations aren't in place and why AI maturity has become the next challenge organizations need to solve. What starts to break as AI scales? As AI becomes embedded across more workflows and business functions, existing weaknesses become more visible and more pronounced, turning what were once manageable issues into enterprise-wide challenges. Organizations often experience familiar operational challenges such as: - Data quality issues become amplified as AI consumes more data and influences more decisions. - Data governance becomes harder to maintain across teams, systems, and AI workflows. - Ownership becomes unclear as responsibility for data, metrics, AI outputs, and agent workflows spans multiple stakeholders, particularly as agents begin operating autonomously across team boundaries. - Infrastructure, compute, and operational costs become harder to manage as AI adoption grows. - Trust becomes harder to maintain as AI-generated outputs and agent-driven actions reach more employees and customers. - Teams struggle to explain how AI-generated answers were produced. These issues rarely occur in isolation. Together, they point to the same underlying challenge: organizations are scaling AI faster than their trusted data infrastructure can support it. ## Why these challenges matter Whether organizations are using off-the-shelf AI tools, sophisticated agent harnesses, or advanced agentic workflows, AI depends on trusted data, governance, and business context to produce reliable outcomes. As AI adoption grows, the operational challenges that have long affected analytics become even more visible and more consequential. The further organizations move toward autonomous, agentic systems, the higher the cost of getting these foundations wrong. The findings from the [2026 dbt Labs State of Analytics Engineering Report](https://www.getdbt.com/resources/state-of-analytics-engineering-2026) reinforce this reality: - **53%** report poor data quality as a top challenge. - **41%** cite ambiguous data ownership as an ongoing challenge. - **71%** are concerned about hallucinated or incorrect data reaching stakeholders. These findings reinforce that AI readiness depends on the trusted data infrastructure supporting it. As organizations scale AI and agents, the quality of that foundation increasingly determines whether AI can deliver reliable business outcomes at scale. ## Building the foundation for AI at scale Deploying more models is only part of what it takes to scale AI successfully. Organizations also need the trusted data infrastructure that allows AI to perform reliably over time. That foundation extends beyond data alone. It includes governance, ownership, business context, interoperability, and the operational efficiency required to help AI produce accurate, explainable, and consistent outcomes at scale. Together, these capabilities enable organizations to move beyond isolated AI initiatives and scale AI confidently across the business. Understanding your organization's current capabilities is the first step toward identifying operational gaps and prioritizing the investments that will have the greatest impact. Strengthening that foundation better positions organizations to support the next generation of AI systems and agents as technologies and use cases continue to evolve. ## What's next? The upcoming **Enterprise AI Data Maturity Model **provides a practical framework for assessing your organization's AI maturity, identifying capability gaps, and understanding where to focus next. The accompanying guide explores the capabilities organizations need to progress from trusted data to trusted AI and agents. Join dbt Labs Senior Director of Product Strategy Russell Christopher and Infinite Lambda Chief Product Officer Petyo Pahunchev on September 2 or 3 for a first look at the Enterprise AI Data Maturity Model—a five-stage framework for finding out where your organization stands and what it takes to move up. [Register for the webinar](https://www.getdbt.com/resources/webinars/from-data-chaos-to-agent-accessible-how-to-move-up-the-ai-data-maturity-curve). --- --- title: "Databricks processes your data. dbt defines what it means" description: "Your compute platform and your transformation logic are two separate decisions. Most executives approve them as one." url: "https://www.getdbt.com/blog/databricks-processes-your-data-dbt-defines-what-it-means" date: "2026-08-17" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # Databricks processes your data. dbt defines what it means When Databricks announced it was acquiring Neon in May 2025, it published a number from Neon’s own telemetry: over 80% of the databases on the platform were being created by AI agents rather than by people. That number is about Neon. The question underneath it applies to every platform, Databricks included. When an agent writes your revenue model, someone still has to own the definition. When your CFO asks how last quarter’s number was produced, the answer has to live somewhere you can reach. This is for data leaders whose teams already run on Databricks and are being asked whether dbt is still worth the overhead. Short answer: Databricks is an excellent platform, and we’re partners. The dbt-databricks adapter is maintained by the Databricks team. If you’re running ML pipelines, building a lakehouse, or doing serious AI feature engineering, it belongs in your stack. Thousands of data teams run [dbt and Databricks together in production](https://www.getdbt.com/data-platforms/databricks). What I want to challenge is a different assumption: that choosing Databricks as your compute platform and choosing Databricks to own your transformation logic are the same decision. Most companies make both at once without noticing. They’re not the same decision. ## Databricks is good, and that’s not the point Lakehouse compute, Delta tables, collaborative notebooks, ML pipelines. These are real capabilities that deliver real value. Databricks has also been building products that sit squarely in the transformation layer: Lakeflow Declarative Pipelines (formerly Delta Live Tables) for pipeline orchestration, Databricks SQL for query and analytics work, Unity Catalog for governance and lineage. Each is reasonable on its own. Together they represent a set of decisions about where your transformation logic lives. Choosing Databricks for compute is an infrastructure decision. Choosing Databricks to own your transformation logic is a strategic bet on where your data team’s institutional knowledge will live. Most executives sign off on the first without realizing they’ve also made the second. ## What owning the transformation layer actually means The risks of full consolidation aren’t theoretical. They become concrete the first time you try to leave, audit, or explain. Start with portability. If your transformation logic lives in Lakeflow pipelines and Databricks notebooks, moving it means rewriting it. That exit cost won’t appear in today’s contract negotiation. It’ll appear as a multi-month engineering project you didn’t budget for, triggered by a pricing change or a strategic pivot that seemed distant when you signed. Then auditability. When your CFO asks why Q3 revenue was revised, tracing the answer requires Databricks platform access rather than a git diff. Version control of transformation logic depends on the platform, which means it depends on a vendor’s access controls and pricing tier. And timing. Enterprises don’t feel vendor lock-in when things are working well. They feel it when pricing changes, when they want to run workloads somewhere else, or when the platform pivots. By then the transformation layer is load-bearing infrastructure. Databricks would counter that Unity Catalog provides lineage and governance. That’s true. But lineage in a proprietary catalog is a record of what happened. Transformation logic in open, version-controlled SQL gives you the ability to change what happens, and to do it independently, anywhere. ## What dbt gives you that Databricks can’t replace dbt solves a different problem than Lakeflow Declarative Pipelines. It puts transformation logic in open, version-controlled SQL that any engineer can read, test, and run, regardless of what’s underneath. Five things dbt gives you that a Lakeflow pipeline can’t. **One project, any warehouse.** The same dbt project runs on Databricks, Snowflake, BigQuery, DuckDB, and a [growing list of compatible platforms](https://docs.getdbt.com/docs/trusted-adapters). The transformation logic is yours. When your warehouse strategy shifts, you port the logic instead of rewriting the pipelines. **Open standards, and the layer above them.** On June 1, 2026, dbt Labs open sourced the dbt Fusion engine runtime, releasing it as [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here) in alpha under the Apache 2.0 license. Databricks has made its own open-source moves here too, donating its declarative pipelines framework to Apache Spark in 2025. The difference is what sits on top. An open runtime is table stakes. What your team actually needs to own is the layer above it: the tests, the contracts, the metric definitions, and the lineage that explain what a number means and prove it hasn’t quietly changed. **A semantic layer that travels.** Metric definitions in the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) are platform-agnostic by design. Define revenue once and the definition holds whether the query lands on Databricks or somewhere else. **A community organized around a standard.** More than 100,000 data teams [build on dbt](https://www.getdbt.com/community). That community coheres around a shared, open standard for transformation rather than around a single vendor’s roadmap, and that’s worth weighing when you decide where your team’s institutional knowledge will live. **Cost control that isn’t tied to a platform.** [dbt State](https://www.getdbt.com/product/dbt-state) works as a caching layer for your pipelines: it builds what’s changed and skips what hasn’t. Teams see warehouse compute drop by 30% or more. It runs on dbt Core 1.7+, as a plugin or out of the box in the dbt platform, so the savings don’t depend on consolidating your transformation layer anywhere in particular. This is the shape of the argument in miniature. The efficiency comes from the engine understanding your project, not from the platform owning it. ## Five questions executives should ask before consolidating 1. If we needed to move 30% of our data workloads off Databricks in 90 days, what would that require? 2. Can we produce a complete audit trail for every transformation that produced last quarter’s revenue number, without Databricks access? 3. Are our metric definitions in code, or in Databricks notebooks? 4. Does our data team own the transformation logic, or does it live inside a vendor product? 5. If the pricing on our transformation tooling changes next year, what’s our fallback? Each of these describes a situation organizations have already faced at scale. The executives who ended up in difficult positions weren’t naive. They made the two decisions together, without realizing they were separate. ## The right boundary: what each layer owns “dbt vs. Databricks” is the wrong frame. We’re partners, and the combination is production-proven. The risk shows up when companies blur the boundary between the two. Here’s the boundary that holds. Layer What it owns Databricks Compute, storage via Delta, ML pipelines, AI feature engineering, collaborative notebooks dbt Transformation logic, the semantic layer, data contracts, lineage in code, testing Databricks processes your data. dbt defines what it means. Those are two different jobs, and conflating them is where the risk comes from. The executives who regret full consolidation are the ones who approved it while everything was working, before they needed to explain a number to the board, migrate a workload, or watch their data team’s institutional knowledge walk out the door embedded in Databricks notebooks. dbt is how your organization keeps the ability to reason about its own data, independently of any single vendor. **See how dbt and Databricks work together.** [Explore the integration](https://www.getdbt.com/data-platforms/databricks) --- --- title: "dbt Core v1.12 is GA" description: "dbt Core v1.12: what's new and how to upgrade." url: "https://www.getdbt.com/blog/dbt-core-v1-12-is-ga" date: "2026-08-17" authors: ["Grace Goheen", "Sara Gawlinski"] categories: ["Product"] --- # dbt Core v1.12 is GA dbt Core v1.12 is a big one. It does two things at once. It delivers meaningful improvements for teams using dbt Core today, including UDF enhancements, simpler Iceberg catalog and Semantic Layer specs, and plenty of quality-of-life upgrades like a new `on_error` config for handling upstream failures, a dedicated `vars.yml` file, and ad hoc SQL through `dbt run-operation --sql`. It also introduces a new opt-in Rust-based parser, the same one that powers [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0), giving teams a practical, low-risk way to start preparing for the next major version of dbt. Check out the [v1.12 upgrade guide](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.12?version=2.0) for the full list of changes. If you’d rather hear it straight from the team, [join us for a live virtual recap and Q&A](https://www.getdbt.com/resources/webinars/dbt-core-v1-12-live). We’ll walk through what shipped, explain why it matters, share more context on the path to v2.0, and answer your questions live. Let’s get into what’s new in dbt Core. ### Take the first step toward dbt Core v2.0 with the v2 parser Before dbt can compile or run your project, it needs to read your project files, understand your resources and configurations, resolve dependencies, and construct the DAG. As projects grow, the time required to do that work can become a meaningful part of the development loop and startup time. dbt Core v1.12 introduces the opt-in `--use-v2-parser` flag, which delegates that work to the new Rust parser built for v2 instead of the Python parser used by dbt Core v1.x. On larger projects, the Rust parser can be 5–10× faster. The flag is entirely opt-in. Nothing changes unless you enable it, and you can return to the existing parser simply by removing the flag. That makes v1.12 a low-risk place to test the new parser against your real project, identify compatibility issues, and address them gradually rather than all at once. In other words: this is not just a performance improvement. It is a stepping stone to dbt Core v2.0. The [next major version of dbt Core](https://docs.getdbt.com/blog/dbt-core-v2-is-here) is being rebuilt in Rust on the same foundations as the dbt Fusion engine. It raises the baseline for parser performance, language validation, artifacts, documentation, and adapter development. Trying the parser in v1.12 gives you an early look at one of the most foundational parts of that new architecture without requiring you to move your whole project to dbt Core v2.0 today. And yes, we want your feedback. If you encounter a difference or edge case, open an issue and let us know. Real-world testing from the community is how we close those gaps before dbt Core v2.0 reaches GA. ## New dbt framework features ### More control when an upstream model fails This one has been a long time coming. In 2020, community member [@ian-whitestone proposed](https://github.com/dbt-labs/dbt-core/issues/2142) allowing downstream models to continue running in cases where an upstream failure did not make their results unusable. Let’s say for example an infrequently changing country or currency dimension: if that dimension fails to refresh, a daily rollup may still be able to process new transactions against its last successful version. The new [`on_error` config](https://docs.getdbt.com/reference/resource-configs/on_error?version=2.0) gives teams more control over whether downstream models should be skipped or allowed to continue after a failure. `on_error` accepts two values: - `skip_children` (the default): all downstream models are skipped, exactly as dbt behaves today. - `continue:` downstream models keep running instead of being skipped. ```sql -- models/dim_customers.sql {{ config( materialized='table', on_error='continue' ) }} ``` ### A cleaner home for project variables If your project contains a lot of variables, `dbt_project.yml` starts doing double duty: project configuration and variable storage, in one increasingly long file that everyone on the team edits. Back in 2020, [@benjaminsingleton suggested](https://github.com/dbt-labs/dbt-core/issues/2955) giving variables their own file to keep that file readable and cut down on merge conflicts. dbt Core v1.12 makes that possible with support for a dedicated `vars.yml` file at the project root. ```yaml # vars.yml vars: schema_name: analytics materialization: table ``` In addition to keeping `dbt_project.yml` cleaner, variables defined there are available while the project file is parsed, making them useful in project-level configuration as well. That means you can reference those variables inside your project file itself, and this is something you couldn't do when the variables lived in the same file they needed to configure: ```yaml # dbt_project.yml models: my_dbt_project: +schema: "{{ var('schema_name') }}" +materialized: "{{ var('materialization') }}" ``` ### Run ad hoc SQL without creating a macro The new `--sql` flag for `dbt run-operation` lets you execute a one-off database statement through dbt’s Jinja compilation context, without first creating a named macro. It is a simpler way to handle one-time operations while still using dbt’s existing connection and compilation behavior. ```sql dbt run-operation --sql "grant select on {{ ref('fct_orders') }} to role reporting" ``` ### Extend reusable logic with new UDF capabilities In dbt Core v1.11, user-defined functions (UDFs) officially became part of the dbt standard. That work was shaped by years of community experimentation and feedback and the community continued to help move it forward in v1.12. **JavaScript UDFs.** You can now define JavaScript UDFs for Snowflake and BigQuery directly within your dbt project. Drop the function body in a `.js` file under `functions/`: ``` // functions/is_positive_int.js return /^[0-9]+$/.test(a_string) ? 1 : 0; ``` Then define its arguments and return type in the corresponding properties file. A special thank you goes to @pempey, whose adapter override macro for experimenting with UDFs in additional languages gave the team a solid starting point for this work. **Python UDFs on Databricks.** Python UDFs aren't new, but starting in v1.12, they can run on Databricks (Unity Catalog required), joining Snowflake and BigQuery. **Multiple signatures with `overloads`.** The new `overloads` property lets one function accept several argument signatures, so you don't need a separate UDF for every input type. Each overload points to its own body file: ```yaml # functions/is_positive_int.yml functions: - name: is_positive_int arguments: - name: a_string data_type: string returns: data_type: integer overloads: - defined_in: is_positive_int_numeric arguments: - name: a_num data_type: numeric ``` **Third-party packages for Python UDFs.** Python UDFs can now declare public PyPI packages through the packages config. Your warehouse installs them when it creates the function: ```yaml # functions/is_positive_int.yml functions: - name: is_positive_int config: runtime_version: "3.11" entry_point: main packages: - numpy - pandas==1.5.0 ``` Together, these enhancements make reusable transformation logic easier to define, govern, and deploy alongside the rest of your dbt project. ### New Semantic Layer spec and Apache Ossie support v1.12 introduces the [latest dbt Semantic Layer YAML specification](https://docs.getdbt.com/docs/build/latest-metrics-spec?version=2), designed to make semantic definitions feel more closely connected to the models and columns they describe. Rather than defining a semantic model as a separate top-level resource, you can nest semantic information directly within a model. Entities and dimensions are defined at the column level, while simple metrics replace measures and can live alongside the model that provides their underlying data. Legacy spec: ```yaml semantic_models: - name: orders model: ref('orders') defaults: agg_time_dimension: ordered_at entities: - name: order type: primary expr: order_id - name: customer type: foreign expr: customer_id dimensions: - name: ordered_at type: time type_params: time_granularity: day - name: status type: categorical expr: order_status measures: - name: order_total agg: sum expr: amount metrics: - name: order_total type: simple type_params: measure: order_total ``` Latest spec: ```yaml models: - name: orders semantic_model: enabled: true agg_time_dimension: ordered_at columns: - name: order_id entity: type: primary name: order - name: customer_id entity: type: foreign name: customer - name: ordered_at granularity: day dimension: type: time - name: order_status dimension: type: categorical metrics: - name: order_total type: simple agg: sum expr: amount ``` This makes it easier to understand the relationship between a physical model and its semantic meaning without jumping between disconnected definitions. dbt Core v1.12 also adds support for defining semantic models using [Apache Ossie](https://www.getdbt.com/blog/osi-is-now-apache-ossie) documents (formerly the Open Semantic Interchange). dbt can parse Ossie-format JSON files alongside native dbt semantic models and generate an `osi_document.json` artifact representing your project’s Semantic Layer. Together, these changes move us forward on a path where semantic context is easier to author in dbt and more portable across the broader ecosystem. ## Adapter-specific features and enhancements As always, the Core release is only part of the story. Adapter maintainers have continued improving the experience across individual data platforms. Highlights include: - **Snowflake:** Iceberg v3 support and more control over dynamic-table scheduling, support for separate warehouses during initial builds, and transient tables. - **BigQuery:** Parallel microbatch execution, standard SQL for partition metadata, and per-resource execution timeouts. - **Redshift:** Support for the `query_group` session parameter, enabling better workload routing and query logging. - **Databricks:** Unity Catalog row filters and additive merging of tags across project configuration levels. Head to the [v1.12 upgrade guide](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.12?version=2.0#adapter-specific-features-and-functionalities) for the complete adapter-by-adapter breakdown. ### Quick hits Here's what else landed in v1.12. - **Native private packages:** Install packages from private GitHub, GitLab, or Azure DevOps repositories using your existing SSH configuration, without specifying a full Git URL or separately configuring a token. - **Expanded Iceberg support:** Use a simplified `catalogs.yml` specification, enable cross-platform dbt Mesh, and create Iceberg v3 tables on Snowflake. - **Composable selectors:** Reference a named YAML selector within `--select` or `--exclude`, making it easier to combine reusable selectors with other selection methods. - **Clearer errors:** More internal Python exceptions are now translated into useful dbt compilation and parsing errors, with cleaner default output and fewer mysterious stack traces. - **Latest version pointer:** Set the new `latest_version_pointer_enabled_by_default` flag to `true` and dbt automatically creates a pointer view for every versioned model in your project, always resolving to the latest version, without any per-model configuration. - and [MANY](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.12?version=2#quick-hits) others ### What’s next: dbt Core v2.0 [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0) is the next major version of dbt. It replaces the Python-based v1.x runtime with the high-performance, Rust-based foundation developed for the dbt Fusion engine, while keeping the dbt framework open source under the Apache 2.0 license. A major-version transition gives us an opportunity to remove deprecated behavior, enforce a more rigorous language specification, and establish a stronger foundation for the next era of dbt. It also means teams deserve a clear, gradual path to get there. That path starts with v1.12. Try the v2 parser. Resolve outstanding deprecations. See how it behaves with your macros, packages, configurations, and project structure. Tell us what works, and, more importantly, what does not. To learn more, join the dbt Core product, engineering, and developer experience teams for a live virtual release recap and Q&A. We’ll cover the most important changes in v1.12, demonstrate the new parser, discuss how the release fits into the path toward v2.0, and answer your questions live. And to everyone who filed an issue, contributed code, tested a prerelease, joined a discussion: thank you. --- --- title: "Model for the token, not the table" description: "We were burning through Gong's API to feed AI. Modeling the transcripts in the warehouse with dbt cut token costs 20x." url: "https://www.getdbt.com/blog/model-for-the-token-not-the-table" date: "2026-08-17" authors: ["Britton Stamper"] categories: ["Insights"] --- # Model for the token, not the table Gong is one of the highest-value AI context sources we own. When internal usage of Claude started scaling, teams wanted to connect Gong to start pulling from our large volume of call history. The shortest path for those requests flowed through a wrapper around the Gong MCP to the Gong REST API. Unfortunately, that also meant that the team was getting raw, unmodeled context that we couldn’t govern and improve. A single customer-history question fans out through an MCP wrapping Gong’s APIs using `list_calls` to collect IDs, burning 8,300 tokens per call transcript or around 240,000 tokens in context for one account history of ~30 calls. Multiplying this across our entire sales department researching multiple accounts is exactly how you redline the Gong API’s per-second and per-day rate ceilings, which hit our cap very quickly. Gong’s API constraints weren’t the real problem. Our core issue was using a transactional API as a high-throughput context layer for AI. Gong and other sources were never designed for these high-volume, operational use-cases. Databases are. **** ## Compression, context engineering, and cost savings We identified the root cause as the same historical problems that data teams face. Instead of going through our traditional data modeling and warehousing layer, users were hitting source data directly, going to Gong every time with the same "_give me all my transcripts, now summarize them_" request. Analyzing MCP connector usage showed us that Claude was reaching directly to sources that we already had in our data warehouse. Fivetran was already funneling Gong into Snowflake continuously, where we could model data so people would never need the raw source. Serving Gong data from the warehouse through the dbt MCP server removes the API ceiling and lets us model the rich, but often verbose context into exactly the shapes an agent needs. We could even benefit from the broader data warehouse and join Gong data to everything else we know about the account, creating a rich context layer for our team as they used agents like Claude to help them work. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/9788483fe10edcc2adb0a37bf301444b679579c6-2480x1520.png) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/cb1bb2a01332ce4a865edd12be3fe0f400600a0a-2120x1432.png) ## Higher-quality context at lower cost By leveraging our warehouse and dbt, we took care of our API constraint issue _and_ opened up significant potential token savings. Ingesting the raw Gong data into our data warehouse lets us use data modeling to summarize each call, eliminating noise and extracting only relevant signal. Given how valuable Gong data is at informing our business and how frequently it’s used, it made sense to set up pipelines on top of it that we could use for all the call transcript use-cases. Actively compressing transcripts shrank data volume by 20x or more, slashing a 60-minute call from 10,000+ tokens to just a few hundred. The same transcripts, summarized, consume 74x fewer tokens and are now 99% cheaper to serve to AI agents while preserving or improving context quality. We can run hundreds or thousands more Claude sessions a day at the same cost, making ROI skyrocket. The compression itself runs as an incremental dbt job as data enters the warehouse, with inference priced through Snowflake credits, Vertex batch pricing, or Databricks AI Function pricing, all materially cheaper than interactive API calls and paid once per call rather than every turn. The result is order-of-magnitude savings: the cost of compressing 1,000 calls in batch is in the low tens of dollars. Because compression is a one-time cost per call and the savings repeat on every read, the batch job pays for itself within the first meaningful agent session. ## What token efficiency looks like at enterprise scale _Based on Claude Sonnet 5 pricing as of August 2026, subject to change: $3 per million input tokens, $15 per million output tokens._ ## The pattern This pattern is exactly the work dbt is built to do: model the transcripts in the warehouse, compress them with warehouse-native AI functions, and serve the compressed form to agents through the dbt MCP server. We just needed to dogfood things internally to show that dbt transforms data for AI as effectively as it does for BI: 1. **Fivetran delivers the Gong transcript** into the warehouse when a call concludes using an existing connector, no new infrastructure required. 2. **A dbt model invokes a warehouse-native AI function** (Snowflake Cortex COMPLETE, Vertex AI batch prediction, or Databricks AI Functions) to produce a structured summary, capturing deal-relevant points, named entities, sentiment, action items, and objection types. A 10,000-token transcript typically compresses to 500-to-1,000 tokens of structured summary without losing the signal an agent needs for deal context. 3. **The dbt MCP server serves the compressed form** to the asking AI client. This is the same connection AI clients already use for any other dbt-modeled data, and the summary looks like just another well-modeled table. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/e19fcd728d5fb653a3c4da3eb92a7f40cf4aa9f3-1650x578.jpg) Gong was our initial test case, but the pattern generalizes to any token-heavy, durable data source. The same architecture works for engineering context from sources like email archives, Slack channel logs, support tickets, contracts, or marketing collateral. ## Three ways to serve Gong Head-to-head, here’s how using the official Gong MCP, an MCP wrapping Gong’s APIs Gong MCP, and the dbt MCP server stack up against each other for serving unstructured Gong data as AI context: **** ## Cost comparison: Gong vs MCP wrapper vs dbt There's no single best tool for all queries, but there _is_ a best tool per question type. Gong’s official MCP is efficient when all you want is a quick, server-synthesized, single-account summary; an MCP wrapper/gateway platform is good for pulling raw transcript text. Everything else—anything filtered, aggregated, joined, or trended, i.e. the majority of the analytical questions people actually ask—should go through dbt’s modeling and warehousing layer. _Note: Cost calculated using Opus pricing at $5/million tokens, anchored to 8,314 raw and 112 brief tokens per call._ **** ## Get started If you want Gong context wired into your AI workflows, the path forward is to stop connecting directly to Gong's API, partner with the data team to get the Fivetran + dbt + warehouse-AI pipeline modeled. Any AI client (Claude, agents, whatever surface you’re using) connects to the[ dbt MCP server ](https://github.com/dbt-labs/dbt-mcp)as a single governed entry point and gets access to dbt-modeled data through the same surface used for everything else in the warehouse. Analytics engineers have always transformed structured data to build tables for humans and BI tools. Now we can apply the same general principles to unstructured data to build models and engineer context for AI agents. Modeling Gong (and similar sources) for AI is among the highest-impact data-team work available right now, with token efficiency as a first-class design constraint. It’s time to expand beyond thinking what metrics can we serve and start treating call transcripts and other qualitative data as first-class, governed warehouse assets. Now the first question to ask of any source you wire to AI is _what does the raw form cost per question, and what is the smallest modeled form that still answers it?_ --- --- title: "Why agentic projects fail and how to fix them" description: "Why do some AI deployments yield great successes while so many still crash and burn? The answer is in the data." url: "https://www.getdbt.com/blog/why-agentics-projects-fail-and-how-to-fix-them" date: "2026-08-14" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Why agentic projects fail and how to fix them It’s amazing what deltas exist between AI implementations. [Wayfair](https://openai.com/index/wayfair/) built agents to support its suppliers and now automates 41,000 support tickets a month. [C.H. Robinson](https://www.langchain.com/blog/customers-chrobinson) built agents that read shipping-request emails, connect information across messages and attachments, fetch additional context, and automatically create more than 5,500 shipping orders per day, saving 600 person-hours per day. But then there's [Klarna](https://www.customerexperiencedive.com/news/klarna-says-ai-agent-work-853-employees/805987/), which one analyst called the "poster child for bad AI deployments." A customer service agent meant to do the work of over 850 employees contributed to quality issues and declining customer satisfaction. Same technology, wildly different outcomes. And the deciding factor is rarely the model. As we’ll discuss below, what separates the wins from the cautionary tales is the underlying data, and whether users can trust it. **** ## Why agentic AI projects fail without trusted data Agentic AI is undergoing the most aggressive technology adoption curve in a generation. According to the [2026 Gartner CIO and Technology Executive Survey](https://www.gartner.com/en/articles/hype-cycle-for-agentic-ai), more than 60% of organizations plan to deploy AI agents in the next two years, and only 17% have done so today. [McKinsey estimates](https://www.mckinsey.com/capabilities/risk-and-resilience/our-insights/deploying-agentic-ai-with-safety-and-security-a-playbook-for-technology-leaders) that gen AI could add $2.6 trillion to $4.4 trillion in value annually across enterprise use cases. Yet [Gartner also projects](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027) that more than 40% of agentic AI projects will be canceled by the end of 2027. A [2026 Fivetran report](https://www.fivetran.com/resources/reports/the-2026-agentic-ai-readiness-index) found that only 15% of organizations are fully ready, even as the vast majority have already invested millions. That gap between ambition and readiness has a cause, and it's a specific one. The key limiting factor for successful agentic AI implementation is generally not model quality, but poor data quality and governance. Frontier AI labs have released powerful foundational models. But if the data that feeds them is inaccurate, incomplete, or inconsistent, you get poor results. [Agentic AI](https://mitsloan.mit.edu/ideas-made-to-matter/agentic-ai-explained) extends the reasoning ability of generative AI to decisions and actions performed through software, not only producing information but also acting in the world. That shift is exactly what raises the stakes. Without adequate data and context, agentic AI can misanalyze a situation, choose the wrong response, perform the wrong action at scale, and cause cascading workflow errors. With poor security, governance, and accountability, including at the level of data assets, it can become difficult or impossible to trace the origin of a wrong decision and remediate it. The public examples are instructive. A chatbot deployed by [Air Canada](https://www.theguardian.com/world/2024/feb/16/air-canada-chatbot-lawsuit) gave a customer inaccurate information about bereavement pricing, failing to refer to and correctly cite the company's internal policies. [Replit's software development copilot](https://fortune.com/2025/07/23/ai-coding-tool-replit-wiped-database-called-it-a-catastrophic-failure/) deleted a production database during a code freeze and created false data in the process, the kind of incident that hard production controls should prevent, whether the actor is human or AI. We find it useful to organize the risks posed by agentic AI into three categories: 1. **Operational correctness risk:** the agent acts on stale, missing, or misunderstood data. 2. **Control-plane risk:** the agent has the wrong permissions, weak auditability, or poor security boundaries. 3. **Human-system risk:** humans overtrust, under-review, or cannot effectively supervise the agent. The through-line in all three is context. For agents, stale data is not merely an analytics problem; it can become an operational action taken on the wrong version of reality. A dashboard built on last week's numbers is a bad report. An agent acting on last week's numbers is a bad decision, executed automatically, at machine speed. **** ## Agents demand more from your data than dashboards ever did Agentic AI imposes far greater demands on an organization's data infrastructure than human-centric analytics workflows, especially reports and other forms of decision support. A human analyst consumes data intermittently; an agent consumes it continuously. A human can absorb tacit, tribal knowledge through experience and can intuit, remember, or investigate where a data asset came from. An agent needs explicit access to context, explicit governance, and access to data lineage. The key challenge is the lack of a governed, consistent context across the enterprise data estate. AI agents need more than raw data. They need reliable data movement, shared business logic, semantic context, lineage, and access controls across every source and consumer. Meeting that requirement rests on two pillars: automation and centralization. Centralizing data and ensuring that it's inventoried and defined once is essential for scaling access to trusted data, controlling infrastructure costs, and managing compliance risks. Once data is centralized, your team has to systematically transform, that is, model, it into a usable context layer for AI. That's what makes lineage, semantic layers, and governance critical: - [**Lineage**](https://www.getdbt.com/lp/data-lineage) shows users, including agents, where data came from and whether it can be trusted. - [**Semantic layers**](https://www.getdbt.com/product/semantic-layer) apply shared business meaning to tables, granting humans and agents alike a shared, consistent understanding of how data maps to real-world business concepts. - [**Governance**](https://www.getdbt.com/lp/data-governance) controls what users and agents can access, decide, and change. There's one more requirement that's easy to underrate: interoperability. As AI tooling continues to evolve, the best model, compute engine, orchestration layer, or activation channel for one workflow may not be the best for another. Interoperability gives teams the freedom to connect systems without duplicating data, rebuilding pipelines, or locking agent workflows into a single vendor. If each AI use case requires copying data into a proprietary silo, organizations lose governance, portability, and control. ## Where agentic AI actually works best Deciding where to point an agent matters as much as the infrastructure underneath it. AI has a very jagged ability profile due to its design, excelling at some tasks while deficient at others. It's strong at pattern recognition and completion for text and code, including drafting, editing, summarizing, translation, and coding. It's strong at ideation and brainstorming, especially when breadth is required, and the cost of a bad suggestion is low. It’s skilled at reasoning through problems with clear, specified premises and constraints. On the other hand, it's weak at discerning truth from plausibility when facts are obscure or highly specific and at knowing when not to answer. It struggles with open-ended problems that require causal reasoning from limited evidence and long-horizon planning, and with adversarial interactions. That profile points to a clear set of characteristics for good agentic use cases: - High volume - Repeatable structure - Text/code-heavy inputs - Clear success criteria - Low-cost human review - Reversible or low-risk actions - Available authoritative data Poor use cases have the inverse: ambiguous accountability, high legal or safety stakes, sparse data, adversarial users, long-horizon planning, and irreversible actions. It helps to remember that these systems tend to augment work rather than replace it. In 2016, [Geoffrey Hinton](https://www.nytimes.com/2025/05/14/technology/ai-jobs-radiologists-mayo-clinic.html), who would later win the 2024 Nobel Prize in Physics for his work on artificial neural networks, predicted that radiologists would be extinct as a profession by 2021 due to AI image recognition. By 2025, radiologists' pay, employment, and workloads had never been higher. AI is far likelier to augment complex workflows than eliminate roles. We've put these principles to work ourselves. Fivetran's Chief Product Officer uses agentic AI to perform conversational analytics on Jira data. Directly querying Jira's MCP server was untenable at scale, so the team moved Jira data into BigQuery via Fivetran, used a Claude Skill to query it, and produced product-ops insights in hours rather than multiple analyst sprints. Separately, our support team embedded a custom AI app in Zendesk to answer questions, draft responses, summarize handovers, and find similar tickets. Built using Fivetran and dbt, it centralizes knowledge from Zendesk, Slab, Jira, GitHub, Google Drive, Gong, Salesforce, and docs. The common thread: each agent is narrow, high-volume, and grounded in authoritative data it can‌ reach. **** ## Build an agent you can trust Once you've picked a workflow, the build itself is more approachable than most teams assume. Building agentic AI models from scratch is a complex undertaking that can cost many millions of dollars and months of development time. A more practical and less risky option is to augment a foundation model with your organization's unique, proprietary data using a RAG architecture. You can create specialized agents that perform specific tasks by interacting with your operations through the [dbt MCP server](https://www.getdbt.com/blog/mcp) and similar controlled interfaces. The safest way to roll that out is in tiers of progressively growing autonomy: - Start with read-only agents that retrieve and summarize information. - Then, build drafting agents that prepare outputs for human review and final implementation. - Next, build bounded write-back agents that act within strict, narrow limits. Potentially risky actions should require approval, while sensitive, irreversible, regulated, or safety-critical actions should remain prohibited from autonomous execution. Getting each of those tiers right depends on several details, such as your reference architecture, the discipline of [context engineering](https://www.getdbt.com/blog/bringing-structured-context-to-ai-with-dbt), and a concrete readiness checklist. We’ve laid out this blueprint in [The data leader's primer for agentic AI](https://www.getdbt.com/resources/the-data-leaders-primer-for-agentic-ai). It walks through the reference architecture step by step, the strengths-and-weaknesses map for choosing use cases, and the checklist we use to take an agent from idea to production so your project avoids ending up part of the 40% cancellation statistic. [**Download the full guide**](https://www.getdbt.com/resources/the-data-leaders-primer-for-agentic-ai) for the roadmap. --- --- title: "How dbt State cuts warehouse compute and speeds up every run" description: "How Fanatics cut warehouse compute by only rebuilding what's changed" url: "https://www.getdbt.com/blog/dbt-state-use-case" date: "2026-08-14" authors: ["Daniel Poppy"] categories: ["Product"] --- # How dbt State cuts warehouse compute and speeds up every run Every day we see about 859,000 dbt jobs run. Most of them run hourly. And most run hourly even when the data hasn't changed, so thousands of models get rebuilt to produce the same results. For a long time, that made sense. Cloud compute kept getting cheaper per unit, and it was easy to just use more of it. Then everyone started using more data: more dashboards, and thus more pipelines powering machine learning models and recommendation engines. Once agents arrived, anyone could write a complex SQL query, or even create a dbt model. What started as a reasonable way to run your pipelines got expensive. Those cheap cloud costs started to skyrocket. [dbt State](https://www.getdbt.com/product/dbt-state) is our answer. It takes what was a stateless engine, dbt, and makes it stateful: every time you run `dbt run` or `dbt build`, it only builds what's‌ changed and skips everything that hasn't. The tagline we keep coming back to is simple: build what's changed, skip what hasn't. Here’s how it works and how one of our customers uses dbt State to push model reuse rates as high as 25%. **** ## Why we built dbt State For all of our data people, this was always about more than compute. Saving 30% on the data warehouse bill is great, and we wouldn't turn it down. But the thing you feel every day is the time, the focus, and the trust it takes to keep a project running. Think about the friction: - Every time you develop a new model, you have to ask which selectors it belongs in. Managing 20 different selectors across 10,000 models becomes a problem in itself. - Every time you run in development, you wait for 100 or 200 models to build, 10 or 20 minutes, just to figure out whether you built the right thing. That costs focus. - As costs climb and the business doesn't see more value, because the data isn't refreshing any more often, people start asking whether it's worth it. dbt State started at our last [dbt Summit](https://www.getdbt.com/dbt-summit/), where we shipped [state-aware orchestration (SAO)](https://docs.getdbt.com/docs/deploy/state-aware-about) into preview. The feedback was clear. Customers were saving 30%+ on their data warehouse bill. They came back with two questions. First, why is this only in production, when I'd love to use it in staging, [continuous integration (CI)](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud), or development to make my runs faster? Second, why is it available only in the dbt platform on the [dbt Fusion engine](https://www.getdbt.com/product/fusion) when I want to run it anywhere? We listened. We took the best parts of SAO and made them work everywhere. The result is dbt State. ## What dbt State does The basic idea stays the same as SAO: build what's changed, skip anything that hasn't. That unlocks three things. - **Optimize costs.** Teams see 30% less warehouse compute on average by only building what's changed upstream and adhering to the freshness [service-level agreements (SLAs)](https://docs.getdbt.com/reference/resource-configs/freshness) you set, so the business isn't paying for fresh data more often than it uses it. - **Run freely.** When you kick off a `dbt run` or `dbt build`, you can be confident you won't build 100 models when only 10 have changed. You'll build 10. - **Speed up development.** You can increase the speed of every run without resorting to custom workarounds like sampling or using different data across development and production. You turn dbt State on, and it just works. ## Run it anywhere you run dbt dbt State is available natively in the dbt platform and locally in [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0&name=Fusion). If you're on v2.0, you can run `dbt login` and set up dbt State right away. If you're on v1, from v1.7 through v1.12, it's available as a plugin: run `pip install dbt-state` and you can be up and running across development and production in less than five minutes. There's no platform upgrade required. Anywhere really does mean **anywhere**. Run dbt in an orchestrator like [Dagster](https://dagster.io/) or [Airflow](https://airflow.apache.org/). Add it to your own scheduled setup in a Python environment on EC2. If you can run dbt, you can run dbt State. That last point is the part that matters most. SAO and dbt State were built on the same premise, skipping builds when the upstream data or code hadn't changed. What we added with dbt State was reach: no longer requiring Fusion, no longer limited to production, and a way to reuse data you've already built instead of paying to rebuild it. When the same logic and data already exist somewhere else, say in a staging or production environment, dbt State clones them in instead of rebuilding, and you won't even notice the difference. ## How dbt State works It helps to separate two things: the **data plane**, where you run dbt, and the **control plane**, where dbt State runs. In the data plane, every `dbt run` executes against your warehouse, where it has access to your data and builds your models. The control plane is where dbt State decides what actually needs building. It keeps a list of your tables and a hash of the last data state and code state. There's no actual data there, just a hash of what the state was. So on every run, dbt State asks a short series of questions for each model. Has it actually changed? If it has, does a version already exist somewhere else that we can clone in? Only when the answer to both is no do we build it. [**For a demonstration of dbt State in action, check out this on-demand webinar.**](https://www.getdbt.com/resources/webinars/dbt-state-build-what-s-changed-skip-what-hasn-t) ## Paying only for what you reuse All of this starts with a 30-day trial for every user. Run `pip install`, try it in development, run it in production, and tell us what you find. When you do start paying, we only want to charge you for the parts that produce value. Our metric is the daily active target table (DATT): a model or test that dbt State reused. We count unique reuses per day, so we don't double-charge you. If your `fact_orders` and `dim_customers` models rebuild six times a day but the data didn't change, you're not paying for six reuses each. You're paying for two, one per unique table, no matter how many times it was reused that day. ## How Fanatics Betting and Gaming put dbt State to work One of the teams that turned dbt State on early was [**Fanatics Betting and Gaming**](https://www.fanaticsinc.com/fanatics-betting-gaming), and they saw double-digit compute savings on the very first project they tried it on. Alvin Chai, a senior analytics engineer at Fanatics, runs a team responsible for building and maintaining gold-layer data models across the betting and gaming business, everything from sportsbook trading analytics and product insights to regulatory and financial reporting. Their stack was [Snowflake](https://snowflake.com) for the warehouse, dbt Core, and Airflow to orchestrate. About 18 months ago, they moved to the dbt platform to remove a bottleneck: refreshes and backfills used to require a specific data engineer to kick them off. Today they run about 11 projects, one per analytics domain, 8 of them on Fusion, with a couple thousand models in total. Democratizing scheduling gave every analytics team the freedom to schedule their own refreshes. The downside, as Alvin says, was overscheduling with little visibility into it: "Maybe you have an analytics team that might set up their models to run every 30 minutes. But the source that model is built on might run maybe only every hour or so." dbt State surfaced their model reuse rates and effectively put guardrails on overscheduling, refreshing models only when they actually need it instead of rebuilding the same output over and over. That made the whole thing more hands-off, so the team no longer had to pester other analytics domains about their costs getting too high. The rollout is a useful lesson in how to get value out of dbt State. Fanatics started with lower-risk work. Their pilot was a governance project supporting responsible gaming and anti-money laundering use cases, chosen because it had few downstream dependencies through [dbt Mesh](https://www.getdbt.com/blog/data-mesh-architecture-explained) and would be easy to roll back. They then worked down the hierarchy by risk, migrating everything except their most critical regulatory and financial reporting. Turning dbt State on with no configuration gave them a model reuse rate of about 0.2%, essentially nothing. The effectiveness came after they configured source freshness and declared model-level SLAs, what we now call lag tolerance. Their approach was to set a baseline model-level SLA for every model based on its current refresh frequency, add a more aggressive project-level SLA on top, and write a small script to translate their existing schedule tags into those SLAs. A model tagged "daily," for example, became a 24-hour refresh window. That took the pilot from 0.2% to roughly 15% model reuse. Historically, overscheduled projects later hit as high as 25%. Others landed nearer 5%. Across all their jobs, the average sits around 8%. The compute savings are real, but Alvin's take is that the biggest long-term win is operational simplicity: "The question kind of turns from, 'What job should this model belong to?' to 'How fresh does this model actually _need _to be?'" Instead of juggling separate jobs by cadence and managing tags for each one, Fanatics can move toward DAG-based jobs: point at a final consumption model, refresh everything upstream of it, and let the model-level SLAs handle the different freshness requirements underneath. The payoff is where the team spends its time. Rather than burning mental energy figuring out how to optimize schedules, they can focus on building new models and delivering more insight and business value. ## Try it yourself A decade of dbt has been about bringing engineering discipline to analytics. dbt State extends that discipline to the compute itself: stop paying to rebuild what hasn't changed, wherever you run dbt. You don't have to take our word for it. [**Create a free dbt account**](https://www.getdbt.com/lp/dbt-free-account) or `run pip install dbt-state` on your existing project, and watch your development runs get faster and your warehouse bill get smaller. [**To see dbt State in action, watch the on-demand webinar.**](https://www.getdbt.com/resources/webinars/dbt-state-build-what-s-changed-skip-what-hasn-t) --- --- title: "dbt Summit 2026: the keynotes and product sessions" description: "The dbt Summit 2026 keynotes and the product sessions behind them." url: "https://www.getdbt.com/blog/dbt-summit-2026-keynotes-product-sessions" date: "2026-08-10" authors: ["Daniel Poppy"] categories: ["Learn"] --- # dbt Summit 2026: the keynotes and product sessions There’s something we keep hearing from data teams. A chief data officer at a health insurance administrator framed it well recently. He has two data engineers he wants to fast-track onto AI work. He has executive support. He has budget. And he still can’t start. “Everybody's making noise around AI. Can you do something around AI? And I'm thinking, ‘But your data is still not at the level where you can put an AI agent on top of that.’” He already knows what most organizations are about to find out the hard way. It’s a gap this year’s dbt Summit content happens to address. [dbt Summit lands September 15-18 at The Cosmopolitan in Las Vegas.](https://www.getdbt.com/dbt-summit) ## You have done this before Ten years ago, dbt changed what it meant to be a data professional. The analytics engineer emerged, with production-grade pipelines at speed and scale, governed data, and real influence over decisions across the business. It’s easy to forget how that felt at the time. The shift was disorienting. Plenty of people weren’t sure where they’d land. It turned out to be one of the best things to happen to careers in this field. We’re at that moment again. AI is rewriting how data gets used, and it’s raising the stakes on every data decision made in the next 18 months. The consumers of data are shifting from analysts in a BI tool to agents working continuously, at machine speed, without a human checking every answer. The teams who thrive will be the ones who level up, which is why you need to be at dbt Summit this year. ## Two keynotes, four pillars The [dbt Summit keynote](https://www.getdbt.com/dbt-summit/sessions/keynote-level-up) is built on four pillars, and they line up closely with what we hear directly from data teams and what they are up against. **Level up the engine** with faster parsing, better scalability, and real cost control, wherever you run dbt. **Level up for AI and what your agents consume**. Structured, governed context is the missing piece between your warehouse and an agent whose answers you’d stake a decision on. **Level up with AI and what you consume** with [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) and the agentic workflows around it, so you build and ship faster while keeping control of quality and spend. **Level up the stack** with an open, flexible infrastructure. Your stack stays yours. Thursday brings the [Community Keynote](https://www.getdbt.com/dbt-summit/sessions/keynote-community-keynote) with Grace Goheen and Jeremy Cohen. This is a celebration of the dbt community, how dbt Core and the engine unite under one framework, and how the work continues in the open. ## Go deeper in the product sessions The keynotes are the headline. The product breakouts and roundtables are where you find out how it all works, straight from the product managers who built it. [**Optimizing your runs for lower compute, fresher data, and faster iteration with dbt State**](https://www.getdbt.com/dbt-summit/sessions/optimizing-your-runs-for-lower-compute-fresher-data-and-faster-iteration-with-dbt-state). Most dbt projects rebuild the same models every run, whether anything changed or not. [dbt State](https://www.getdbt.com/product/dbt-state) checks warehouse metadata and model SQL, works out whether the result would actually change, and then builds, skips, clones, or auto-defers to production. Average compute savings run 30%. Reuben McCreanor walks through freshness SLAs, auto-cloning from prod, and the orchestration logic you get to retire, whether you run dbt Core, the dbt platform, or something in between. [**dbt Wizard: your AI teammate for data development**](https://www.getdbt.com/dbt-summit/sessions/dbt-wizard-your-ai-teammate-for-data-development). dbt Wizard is the coding agent built for data, not software. Ani Venkateshwaran and Brandon Thomson walk through its three modes—develop, analyze, and discover—the agent skills framework that makes it extensible, and a live demo of the fully autonomous analytics engineering workflows dbt’s own team has built on top of it. [**Rebuilding dbt Core in the open: faster runtime, adapters, and docs v2**](https://www.getdbt.com/dbt-summit/sessions/rebuilding-dbt-core-in-the-open-faster-runtime-adapters-and-docs-v2). Years of dbt Core v1.x left teams with a Python runtime that slowed on big projects, fragmented adapters, and a docs experience that couldn’t keep up. Hope Watson unpacks the rebuild: a Rust-based Apache 2 engine, a cleaner adapter model on ADBC and the Arrow ecosystem, and parquet-backed artifacts. You’ll see side-by-side parse speed demos and get a real migration path from v1.x through the v1.12 parser on-ramp. [**The semantic layer is dead. Long live the semantic layer!**](https://www.getdbt.com/dbt-summit/sessions/the-semantic-layer-is-dead-long-live-the-semantic-layer) Semantic layers used to be a solved problem you built years ago for dashboards. Then agents showed up. Zach Mandell makes the case that your semantic layer is now the most load-bearing piece of your AI strategy, shows where MetricFlow and the dbt Semantic Layer are heading, and covers how the dbt MCP server turns governed metrics into context your agents can use. This is the session that answers the hallucination problem directly. [**Product roundtable: open data infrastructure in practice**](https://www.getdbt.com/dbt-summit/sessions/product-roundtable-open-data-infrastructure-in-practice-unlock-your-dbt-projects-with-apache-iceberg-and-mesh). Jack Lowery and Anna Lee host a working conversation on unlocking dbt projects with Apache Iceberg and mesh patterns. Roundtables are small and discussion-first, so come with the problem you’re actually stuck on. [**Apache Ossie: Realizing Semantic Layer Portability**](https://www.getdbt.com/dbt-summit/sessions/apache-ossie-realizing-semantic-layer-portability). Define a metric in one tool, redefine it in the next, and your dashboards, notebooks, and agents all end up disagreeing on what “active customer” means. Apache Ossie, the vendor-neutral, open-source semantic interchange spec co-led by dbt Labs, Snowflake, Salesforce, BlackRock, and RelationalAI, changes that contract. You’ll see how MetricFlow, now open-sourced under Apache 2.0, compiles governed metrics into a portable format so BI, conversational analytics, and AI agents finally agree on the same business logic. [**The next wave of data infrastructure is in the Lake**](https://www.getdbt.com/dbt-summit/sessions/the-next-wave-of-data-infrastructure-is-in-the-lake). The modern data stack got you here, but it won’t get you to where agents need to operate. Russell Christopher and Casey Karst lay out Open Data Infrastructure: land data in open formats like Iceberg in storage you control, then pick the right compute for any job and swap tools without rebuilding pipelines. You’ll leave with a framework for building open, portable, AI-ready infrastructure with Fivetran + dbt. [**We said I do. Now, we’re joined at the DAG**](https://www.getdbt.com/dbt-summit/sessions/we-said-i-do-now-were-joined-at-the-dag). Where did this data come from, what happened to it, and who’s using it? Roxi Pourzand shows how Fivetran and dbt unify ingestion, transformation, and consumption into a single lineage graph—one that falls out of the systems doing the work instead of being stitched together after the fact—unlocking real-cost visibility tied to usage, PII that stays governed downstream, and agents that reason with the full picture. [**With great context comes great autonomy: leveling up your agent context**](https://www.getdbt.com/dbt-summit/sessions/with-great-context-comes-great-autonomy-leveling-up-your-agent-context). An agent querying your warehouse shouldn’t be guessing what “active” means or which revenue table to trust. Ben Moser and Kevin Kim demo how dbt structured context and the Fivetran context layer turn raw metadata, governed semantic models, and real usage into the agent schema that powers accurate results—and how to let agents reason beyond fixed definitions when the question calls for it. More sessions are landing between now and September. Browse [the full agenda](https://www.getdbt.com/dbt-summit/sessions) to build the rest of your week. ## Come find out what you get to build next The organizations that win the AI era won’t be the ones with the best models. They’ll be the ones whose data foundation was ready when it mattered. That foundation already has a name, and you built it. dbt Summit 2026 September 15-18, 2026 The Cosmopolitan, Las Vegas Your next level starts here. [Register for dbt Summit](https://www.getdbt.com/dbt-summit/registration). --- --- title: "From analytics engineer to context engineer" description: "First in a series on the shift from modeling data for dashboards to modeling context for agents. We start with our own Gong data." url: "https://www.getdbt.com/blog/from-analytics-engineer-to-context-engineer" date: "2026-08-06" authors: ["Britton Stamper"] categories: ["Insights"] --- # From analytics engineer to context engineer Most companies are taking shortcuts in their enterprise AI rollouts today, and falling into a classic trap their data teams are already deeply familiar with. Taking the “easy route”—wiring AI straight to MCPs provided by vendors, connecting them directly to source data providers—is collectively costing companies billions in token consumption, vendor API charges, and lack of data ownership and portability. At dbt Labs, we initially were no different. We used vendor MCPs until we discovered how significant these hidden costs actually were. That's when we realized that the most efficient, scalable way to connect AI to data was what we’ve evangelized all along: to centralize our data and engineer context directly within our data warehouse, giving us a massive opportunity to own our AI’s context layer and expand our data team’s remit. ## How it started As Claude rolled out in our organization, one of the top user groups was sales connecting to Gong, Salesforce and other rich context sources to analyze their deals. When our GTM teams requested customer insights to prevent churn and identify high-value accounts, of course we agreed and allowed them to analyze Gong data by connecting AI agents directly via the native MCP. With hundreds of salespeople using this data many times a day, costs added up quickly. Analyzing a single call consumed up to 50,000 tokens; that's $0.25 in token cost just for AI to read a single transcript. We had 500M+ tokens worth of Gong transcript data that we could tap into. With many salespeople running many Claude sessions throughout the day, our total AI spend went up significantly at scale. We knew we needed a new approach. That’s when we decided to ingest the raw Gong data into our data warehouse and used data modeling to summarize each call. By actively context engineering the transcript data to eliminate noise and extract signal, we shrank data volume by 20x and slashed token consumption on a 60-minute call from tens of thousands to just a few hundred token. The question was, if we can transform any of our data into more meaningful context, how could we engineer the smallest, most meaningful context layer that still answers most queries? ## How dbt reduced token costs When our data team had previously modeled Gong data for BI, the ROI wasn't there. Now, though, Claude and ChatGPT agents unlocked new use cases that let us do qualitative data analysis at scale, and the call transcript table became one of the highest value assets in the data warehouse. Sourcing context through the data warehouse is vastly more efficient than pulling it straight from MCPs. Serving that call transcripts from warehouse summaries dropped token costs by roughly 98% while maintaining or improving context quality. In this blog series, we will share core patterns for efficiently modeling trusted context for AI, an end-to-end technical walkthrough of our Gong pilot, the fundamentals of reading data once to serve every agent, and why you can transform traditional analytics engineering into context engineering without changing your stack. ## Qualitative data analysis at scale with engineered context Before AI, working with text data at meaningful scale required either simple and ineffective regex techniques like keyword matching or advanced statistics and data science. Maintaining complex transformations for qualitative data sources like call recordings or support tickets was just not feasible for data teams. As our Gong breakthrough shows, though, we have crossed a threshold where previously ignored data and data sources can now form some of the most useful and valuable agent context. Teams can build models and highly effective context layers for AI, applying the general principles we pioneered with analytics engineering as long as the techniques are applied with AI in mind. Analytics teams have always modeled quantitative data for BI. Context engineering is just modeling data for agents, and that data is far more than metrics: ## A dashboard needs metrics. Agents need context. **Untapped qualitative data is where AI truly pays off.** Instead of just serving structured data to dashboards, data teams can now use unstructured data like PDFs and JIRA tickets to engineer context. LLM advances make it possible to move, store, model, and process these files within the data warehouse and use them as agentic AI context. This is a new way of thinking for a lot of analysts because, historically, data organizations have been allergic to unstructured data like transcripts, and for good reason: our technology was not built to support it. We were forced to change because we were running out of Gong API calls. Everyone was asking very similar questions on very similar data, going right to Gong saying _give me all of my transcripts, now summarize each transcript_. The direct MCPs circumvented our entire traditional data modeling and warehousing world to repeatedly query raw data, which created redundant token costs and caused API constraints. This was when we realized _hey, you know, we actually already have this process whereby we model data into trusted, governed, more useful forms. Why don’t we apply that for data for AI?_ Following our merger with Fivetran, we realized [that moving and modeling unstructured data is already in our wheelhouse.](https://www.fivetran.com/blog/ai-requires-unstructured-data-to-unlock-its-full-potential) Platforms like Salesforce, Zendesk, or Gong now provide critical business context. Modeling this qualitative data and caching pre-built AI summaries in the warehouse delivers reliable, deterministic context while eliminating multiple token-heavy MCP calls, tool proliferation, and API constraints. Users seamlessly access this modeled data through a single context connector in the context layer via MCP. This drastically lowers token costs, simplifies the user experience to a simple chat, and opens up reusable, community-driven data patterns across the entire organization. ## How to do cost-effective context engineering There’s no magic in this approach. Data teams turn raw data, whether quantitative, qualitative, or semi-structured, into agent context through the same processes they already practice without materially changing the stack: - **Compress:** reduce a large corpus to the smallest forms that contains only what’s relevant - **Enrich:** join data to other relevant information so that it’s easier to access everything needed, like opportunity details and qualitative deal history. - **Describe:** provide information around what the data means and when it’s applicable - **Govern:** decide exactly what each agent is allowed to see to do its work ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/66f6a0a0ab2a94cd80765e8c857b7cf60ea679fb-1564x434.jpg) The data team models the context once, and every tool and agent reads from the same trusted context layer. The analytics engineer is now the context engineer. It’s the same skillset, except beyond modeling metrics you’re also mapping the entire data estate and every data source is now in play. All of this happens without materially changing your data stack. The data team builds a [structured shared context layer](https://www.getdbt.com/blog/bring-structured-context-to-conversational-analytics-with-dbt) for the entire organization, functionally layering AI over your original data stack in the same way that dbt is currently layered over your data warehouse. The whole company is now the consumer because every team and the agents they use can access this context layer, not just people who write SQL queries. Functionally, though, how do you move, transform, and manage structured, semistructured, and unstructured data into the context layer? 1. **Getting data, both traditional tables and unstructured files, into the data warehouse **(Snowflake, BigQuery, Databricks) is Fivetran's job. If tabular data needs further preparation (filtering, denormalizing, joining, and aggregating), then dbt allows the creation and execution of that transform logic in an open and portable way. 2. **Automating pipelines that utilize the SQL-based AI functions of your destination** is dbt’s job. For example, Snowflake Cortex provides SQL functions such as AI_EMBED, AI_COMPLETE and AI_PARSE_DOCUMENT that can be executed as part of, and orchestrated by, dbt models. BigQuery and Databricks offer similar functionality. Raw text fields, structured data from files like spreadsheets, and even replicated unstructured files like PDFs become queryable, retrievable context available to both human data users and AI agents in the shared context layer. The documents and data sources they seek information from are already in the same platform, moved and indexed automatically by Fivetran and processed as part of your dbt models. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/1543641da2fcc03b3228ed3497d1e432f7c8b4b8-1322x618.jpg) The stack beneath the context layer remains the same: Fivetran moves your data; dbt models it. Now, though, this includes the unstructured context that makes AI systems powerful. All the Salesforce email bodies, Jira ticket attachments, call transcripts, and contract PDFs that have always existed but never made it into a pipeline are now fully referenceable, trusted context. ## Best practices are beginning to emerge We are continuing to build the vocabulary, packages and open source projects that analytics engineers will need to use to do context engineering. We’ve come up with a few best practices we can share, with much more to come. Here’s a preview of some our team uses: - **Don’t aggregate context, generate it: **Use warehouse-native AI to create new context by parsing, chunking, embedding, transforming and joining related information into modeled data objects that agents can query. Aggregating can cause context to get lost, so maintaining the context quality through the whole pipeline is critical. - **Context layer architecture:** One shared, structured and governed, open-format canonical context layer that every engine can write to and every agent can read from. The same single source of truth the data industry has always focused on is now pointed at AI as the primary consumer. - **Read once, write many: **Reads are the most expensive part of context engineering. Do the expensive read with an LLM once, in batch processing where it’s cheaper; then serve the modeled form to every agent in the context layer - **Incremental context maintenance: **Agents do actions, and they need correct, current information to act on. Context can’t simply be snapshotted, it must be updated as new events and information come in. Incremental models that update context’s current state (with pipelines built to capture the state changes so that they are auditable) are critical. Until recently, the context path of least resistance was to just take the raw API and connect the MCP server. Now, data analysts can step in and say, "Use our existing data. We'll augment that with some of the AI capabilities you're asking for, process it once, and make it accessible to everybody to use an unlimited number of times.” Now, the data team owns AI and becomes the mission-critical team for the next era of businesses. [Fivetran + dbt Labs are building the data foundation for agents you trust. Join us at dbt Summit, where data practitioners and leaders come together to shape the future of data and AI.](https://www.getdbt.com/dbt-summit) --- --- title: "Retiring the dbt Snowflake Native App" description: "The dbt Snowflake Native App retires in November 2026. Here's what it means for you." url: "https://www.getdbt.com/blog/retiring-the-dbt-snowflake-native-app" date: "2026-07-24" authors: ["Kyle Dempsey"] categories: ["Product"] --- # Retiring the dbt Snowflake Native App We’ve made the decision to retire the **dbt Snowflake Native App** from the Snowflake Marketplace. The app will enter maintenance mode in **July 2026** and will be fully removed in **November 2026**. If you are one of the small number of customers currently using the Native App, your dbt Labs account team will reach out directly to support a smooth transition. For the vast majority of dbt + Snowflake users, this change has no impact on your workflows. ## What this means for you ### If you don't use the native app No action is required. This change does not affect dbt platform, dbt Core, the dbt Semantic Layer, or any other dbt product or integration. ### If you currently use the native app For the small number of customers with an active installation: - **Your app will continue to function through November 2026.** - **During the maintenance window** (from July 2026 to November 2026), we will provide critical bug fixes and security patches only. No new features will be released. - **Your dbt Labs account team will contact you directly** to walk through your specific situation and discuss your options. If you do not hear from your account team, please email [support@dbtlabs.com](mailto:support@dbtlabs.com) by **July 2026**. - **After November 2026**, the app will be fully delisted from the Snowflake Marketplace, and access will be discontinued. ## Why we're making this change When we originally launched the dbt Snowflake Native App, our goal was to bring dbt-powered AI capabilities directly into the Snowflake environment. Since then, the ways dbt and Snowflake work together have evolved significantly — and customers now have access to better-supported, more capable options. - **dbt’s Semantic Layer is available in Snowflake today — without the Native App.** Customers can model governed metrics in dbt and use them in Snowflake through Snowflake Native integrations (e.g., Semantic Views / Snowflake Intelligence) with dbt providing the trusted semantic definitions behind the experience. If you were using the Native App to bring dbt context into Snowflake, these options deliver a deeper, more reliable integration. - **The AI landscape has matured.** When the Native App launched, bringing a dbt-powered chatbot into Snowflake was a reasonable experiment. Today, Snowflake Intelligence and Cortex provide a far richer AI experience — with dbt's Semantic Layer powering the trusted context behind them. - **We want to focus investment where it benefits you most.** Rather than maintaining an app that hasn't kept pace, we're directing our partnership efforts toward the integrations that customers are actually adopting and finding valuable — like Semantic Views, OSI, and native dbt execution on Snowflake. dbt will continue to support Snowflake data teams with interoperable products and strong integration. We’ll also continue to help teams deliver transformations while providing observability into AI agentic data operations — so customers can build and operate trusted, governed data products that power analytics and AI on Snowflake. ## Support and feedback If you have questions about this change or need help transitioning: - **Email:** [support@dbtlabs.com](mailto:support@dbtlabs.com) - **Your account team:** Your CSM or account executive can help with your specific situation We welcome feedback on the timeline. While the decision to retire the Native App is final, we're open to adjusting dates if you need additional time. Please reach out if the current timeline creates a hardship for your team. --- --- title: "Fivetran + dbt Labs: The future of dbt Core v2.0" description: "AI has changed the game. dbt is changing along with it. Learn more about dbt Core v2.0 and the future of dbt itself." url: "https://www.getdbt.com/blog/fivetran-dbt-20-future" date: "2026-07-21" authors: ["Daniel Poppy"] categories: ["Product"] --- # Fivetran + dbt Labs: The future of dbt Core v2.0 Last year, we made two announcements that injected a bunch of uncertainty into the dbt community: [the dbt Fusion engine](https://www.getdbt.com/product/fusion) and [coming together with with Fivetran](https://www.getdbt.com/blog/fivetran-and-dbt-are-one-company-now-here-s-what-that-means). Change is scary, and the questions came fast. What was the future of dbt Core? One Reddit user went so far as to predict the "final nail in the coffin of OSS dbt," shorthand for open-source software. What has‌ turned out to be true over the past year is that we have wanted to ship more code in the open, not less. And we anticipate that being true for many years to come. The clearest proof is [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0&name=Fusion): the same fast, capable foundation that powers the dbt platform, now under an Apache 2.0 license. Alongside it, we've shipped [dbt State](https://docs.getdbt.com/docs/deploy/dbt-state-about), a caching layer that cuts customers' dbt-driven compute by 30%+, and dbt Wizard, a coding agent purpose-built for dbt. ## Why the data stack needs to be rebuilt for agents Taylor Brown, Co-founder and COO, Fivetran + dbt Labs, and Tristan Handy, Co-founder and President, Fivetran + dbt Labs, have been working for a long time to bring the two companies together. The two joined a webinar recently to discuss the rationale, and share what comes next. For the full conversation, including the live Q&A, [watch the webinar recording](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a). And if you want to see the speed for yourself, [create a free dbt account](https://www.getdbt.com/lp/dbt-free-account) and invoke dbt Core v2.0. Back in the 2010 to 2014 timeframe, everyone was trying to solve the problem of getting data out of siloed places and leveling up the overall analytics infrastructure to drive business value, largely for reporting. That's where Fivetran plus dbt originally helped coin and build the modern data stack. Over 100,000 teams adopted dbt as a standard for transformation. Over 8,000 customers adopted Fivetran for data replication. [The age of AI has changed the game](https://www.getdbt.com/blog/how-ai-will-disrupt-data-engineering). As Taylor puts it, the outcome that data stacks drive is no longer just for humans and reporting. AI and agents are now large consumers of data, and they're driving business operations and revenue applications. That shift exposes three structural breaks in the stack at agent scale: - **Trust breaks.** If you don't have complete data, consistent metrics, and lineage, you don't have trust. And trust is now infrastructural, because agents are taking action on this data, not just humans who understand its quirks. - **Scale stalls.** With ungoverned data, fragmented context, and locked-in formats, agents will amplify how bad the underlying data is and operationalize it in negative ways. - **Cost explodes.** Repeated retrievals and unoptimized modeling leave you with many more agents running on very expensive compute. When we hit the cloud era, the cloud was so much more useful than on-prem systems that we had to rip everything up. Fivetran and dbt are a result of what happened there. We're doing that type of work again in the agentic era because we're going to see at least an order-of-magnitude increase in data consumption. Our answer is an architecture we call [Open Data Infrastructure](https://www.getdbt.com/blog/what-is-open-data-infrastructure): a data stack built for agents and humans that's flexible by design, fresh and trusted by default, and efficient at scale. An open data infrastructure is: - Built on open standards, separates storage and compute. - Loads into modern data lakes in open formats like Iceberg and Delta. - Relies on strong data movement and transformation in the pipeline layer. - Provides rich metadata and context in the management and governance layer. Companies are already putting this together. - [**Zendesk**](https://www.zendesk.com/), the global customer experience platform, [used Fivetran plus dbt to scale out analytics agents and AI across the entire enterprise](https://www.getdbt.com/resources/coalesce-on-demand/coalesce-2025-how-zendesk-built-a-cross-domain-multi-platform-data-strategy) in a fraction of the time it would normally take. - [**Shutterstock**](https://shutterstock.com) built a trusted, more real-time analytics and AI platform for emerging AI workflows. - At [**Inova Health**](https://www.inova.org/), a leading nonprofit healthcare provider, [Jon McManus and his team compressed a four-year data modernization roadmap into six months](https://www.fivetran.com/case-studies/inova-health-compresses-4-year-roadmap-into-6-months-to-power-ai), with Fivetran and dbt as key parts of the new architecture. For healthcare, that pace is unheard of. ## dbt Core v2.0: One engine for all of dbt If you've been on this journey with us for any length of time, you've seen us trying to do two things at once. The first is to support a widely used piece of open source software infrastructure. The second is to make money so we can keep doing the maintenance. About 18 months ago, [dbt Labs acquired SDF Labs.](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs) Why? Simple: their engine was, in many ways, technically superior to long-term dbt Core. It's written in Rust and includes a bunch of other capabilities. We combined dbt and the SDF engine and shipped it as the dbt Fusion engine. But Fusion and dbt Core lived side by side, on different technical foundations (Rust and Python) and different licenses (Elastic and Apache 2.0). Over the past year, as we did the work of making Fusion ready for general availability (GA), we realized we wanted to bring these two engines together. That work culminated in the launch of [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0&name=Fusion). The alpha shipped on June 1, and it's still early, but the Apache-licensed dbt project is now based on the Rust implementation that Fusion created. The first thing you'll notice is that it's just really fast. The first time you invoke dbt Core v2.0, the responsiveness in your command-line interface (CLI) is dramatically improved. And there's a change we’re especially excited about in how we deal with metadata. dbt has published artifacts like manifest.json and run_results.json for a long time, but dbt now produces Parquet files as well, which form the foundation of a context layer. You can query them locally with DuckDB and immediately find out facts about your dbt project. Fusion isn't going anywhere. Fusion is a superset of dbt Core v2.0. Both are based on the same core technology, and pure OSS dbt Core is now a strict upgrade over v1. Installing the Fusion binary gets you everything in Core plus more, and the heart of that "more" is SQL comprehension. Fusion natively understands the SQL you write. That unlocks developer experience benefits such as easier refactoring, autocorrect, and column-level lineage in your docs. [You can see the full comparison of Fusion and dbt Core v2.0 in our docs](https://docs.getdbt.com/docs/fusion/fusion-availability). The same binary you can use free, without ever speaking to us, now also supports dbt login, which gives you access to proprietary features, some free and some paid. And if you're a dbt platform customer, upgrading to v2.0 or Fusion is straightforward: an auto-migrator tool with AI to fix any remaining issues. All of this lands just past a milestone that's honestly shocking. We recently celebrated the 10-year anniversary of the first commit to dbt Core. There are now over 100,000 teams using it in production every single week, and over a billion downloads. Many of the ideas dbt started with, like testing your data code and keeping data in version-controlled repos, were controversial a decade ago. Now they're defaults. ## dbt State: A caching layer that saves you money [dbt State](https://docs.getdbt.com/docs/deploy/dbt-state-about) is one of the most important features we've ever shipped. dbt State is a caching layer for dbt directed acyclic graphs (DAGs). If nothing has changed at the column level, the table level, or the code level, and you're inside the freshness window you've defined, dbt just skips that model. It doesn't need to build it. The funny thing about building a caching layer is that you don't know how much more efficient everything could be until you implement it. It turns out dbt State saves everybody a lot. That works out to a conservative 30%+ of a customer's dbt-driven compute bill. For us at dbt Labs, it was bigger, more like 64% of the compute bill. We've been able to cut about $400,000 annually from our budget for our underlying data platform. We added headcount as a result. The savings aren't just in production. Taylor shared that Fivetran had a large revenue model that took something like 40 minutes to run; once dbt State was turned on, it ran in about 40 seconds. The analyst team said they literally can't go back. They did see more queries as a result, which drives prices up slightly, but the experience was profoundly better. In development, dbt State makes your workflow really tight because you don't have to sit and wait to rebuild your entire development environment. In production, it saves you a tremendous amount of money. One thing to be clear about: dbt State is a paid feature, but it doesn't require the dbt platform. If you're using dbt Core plus [Airflow](https://airflow.apache.org/) today, you can still run dbt State, and we'd love it if you did. The lowest supported version is dbt v1.7 with a plugin; from v1.12 on, and certainly in v2.0, dbt State is an incorporated feature. ## dbt Wizard: A coding agent purpose-built for dbt There's one more feature we have that will supercharge your development times: [dbt Wizard](https://www.getdbt.com/product/dbt-wizard). dbt Wizard is a coding harness, similar to [Codex](https://openai.com/codex/) or [Claude Code](https://claude.com/product/claude-code), but tuned for dbt tasks. It turns out vertical-specific harnesses have a ton of benefits relative to generic harnesses built for all software engineering tasks. The difference is that we control the harness from the ground up. So we can tune it to perform better and be more token-efficient for the kinds of problems [analytics engineers](https://www.getdbt.com/blog/what-is-analytics-engineering) are solving all day, every day. dbt Wizard uses whatever model you run internally. You plug in your API key and pull from the same token budget, but it will frequently do a better and more efficient job than a generic agent. Try it out, and let us know in the [dbt Community Slack](https://www.getdbt.com/community/join-the-community) what your experience is. ## A foundation for the next decade A decade in, the ideas dbt introduced are now just how data work gets done. The next decade is about making that same trusted foundation work for agents as well as humans. dbt Core v2.0, dbt State, and dbt Wizard are the first big steps. For the full conversation, including the live Q&A, [watch the webinar recording](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a). And if you want to see the speed for yourself, [create a free dbt account](https://www.getdbt.com/lp/dbt-free-account) and invoke dbt Core v2.0. That wow moment is the whole point. ## Frequently asked questions ### Will dbt Core continue to be open source? Yes. ### Are there plans to make dbt packages private? No. As Taylor put it, we think it's important that they remain open source and that anyone can access them. We'd still like you to run them on Fivetran plus dbt, but you don't have to. ### Are dbt Core v2.0 and the dbt Fusion engine the same thing? They're built on the same engine, and Fusion is a superset of Core. If you're already running dbt Core v2.0, there is no upgrade to Fusion; it's just a question of whether you want to flip on additional features. Many of them are free, so it's an exclusively better experience to run with them turned on. The [Fusion availability page](https://docs.getdbt.com/docs/fusion/fusion-availability) breaks down exactly what's included where. ### What advantages remain in Fusion if Core v2.0 uses the same engine? The core capability Fusion has that dbt Core doesn't ship is SQL parsing. Fusion natively understands the SQL you write, which powers improved editor capabilities and column-level lineage in your documentation, and gives us a foundation for future functionality like personally identifiable information (PII) classifiers that flow through the DAG. ### Is dbt Core v2.0 still a Python package, even though it's implemented in Rust? Yes. If dbt had started in the Rust community, it never would have been distributed via PyPI in the first place. But dbt grew out of the Python community, and there are over 100,000 deployments that start with pip install dbt. So, while much of the code is, in fact, Rust, we're continuing to make it available via PyPI. You can also install the binary directly. ### Is dbt moving to usage-based pricing? We do anticipate monetizing dbt State on usage. But dbt State is a brand new feature, and we often make independent decisions about how to monetize brand new things. There's no specific plan to change how the products and services that existed before dbt State are priced. Over time, we think of dbt more and more as infrastructure and want to monetize more in line with consumption. But that's a long-term arc, and we'll take every step carefully in consultation with customers. ### How do the dbt State costs actually work out? We haven't released final pricing yet. The entire point is that every dollar you spend on dbt State saves you more than that. As Taylor put it, we'll give you a dollar for 50 cents. With dbt State, you spend a lot less money on your compute provider, we charge you a little more, and your total cost of ownership goes down. When we deploy dbt State with a customer, we'll frequently run a proof of value in production for a couple of weeks. The before/after delta makes it a straightforward conversation. Over a hundred customers have adopted dbt State already. Everyone's environment is different, so it's worth testing to see how much you'd save. ### What's changing in the semantic layer? We released a new spec for the semantic layer in v2.0 and Fusion. Previously, you had to write a top-level semantic model block as a separate thing and map it to the model, which was clunky. Now, natively within the model YAML, you can declare a semantic model, mark dimensions in the same syntax you use to describe a column, and define metrics right in the YAML. We're also investing in making it more expressive, because folks coming from LookML sometimes hit patterns it can't support natively. That’ll take a good six to 12 months for the semantic layer, so keep an eye out for upcoming changes. ### If AI agents become the primary data consumers, do SQL files become obsolete? Some folks have said that in the future, the spec is the code, and everything else is an intermediate compilation artifact. There are versions of the world where that's true. But for the coming several years, we don't anticipate it. Intelligence is really expensive. Executing software is dramatically more resource-efficient than starting from natural language text every time. So we don't think SQL files are going anywhere, and neither are Python or Rust files. The ways we author them will change, and it's becoming easier to build the surrounding infrastructure of tests and documentation, but they aren't losing relevance. The previous data stack was about building analytics for humans. The emerging stack is about building context for agents. ontext engineering is a real, broad role, and analytics engineers will play a huge role in it. It's a tailwind for the career of every analytics engineer. ### Will Fivetran and the dbt platform become a single app? Today we haven't spent a ton of time on a single app experience, partially because the integration between the two products is already very tight. We'd rather invest the energy in innovation like dbt State, dbt v2.0, and dbt Wizard. Looking five years out, it's hard to imagine two separate auth systems for transformations and integrations, especially in a world where more of this gets built by agents. But our number one priority is integrating the teams and keeping up the velocity of shipping things that matter to customers. For the full conversation, including the live Q&A, [watch the webinar recording](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a). And if you want to see the speed for yourself, [create a free dbt account](https://www.getdbt.com/lp/dbt-free-account) and invoke dbt Core v2.0. --- --- title: "Your next level starts here: A preview of dbt Summit sessions, by role" description: "Preview dbt Summit 2026 sessions by role: hands-on labs and breakouts for analytics engineers, data leaders, and execs." url: "https://www.getdbt.com/blog/dbt-summit-2026-sessions-by-role" date: "2026-07-20" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Your next level starts here: A preview of dbt Summit sessions, by role Ten years ago, this community rewrote how the world works with data. Analytics engineers turned data teams into strategic drivers of the business. We're at another one of those moments. AI changes what it means to work with data. The consumers of your models are shifting from analysts in a BI tool to agents acting on their own, at machine speed and scale. Four out of five data leaders say their data isn't ready for enterprise AI. The teams who win this era are the ones who build the foundation everything else runs on. That's you. dbt Summit 2026 is where this community levels up together. September 15-18 at The Cosmopolitan in Las Vegas, with 100+ sessions across keynotes, breakouts, hands-on labs, and peer exchanges. We've pulled out just a sample of the sessions worth planning your week around and organized them by role, so you can jump straight to what fits your work. [Register here](https://www.getdbt.com/dbt-summit/registration). ## For analytics and data engineers You're the one shipping models. These sessions are about getting hands-on with what's new and hearing from peers who have already put agents to work. **Hands-on learning** [**Accelerating your deployment with dbt v2**](https://www.getdbt.com/dbt-summit/agenda/accelerating-your-deployment-with-dbt-v2). Get hands-on with the next-generation dbt engine and work through realistic analytics engineering problems in a dbt v2 environment. You'll leave able to predict how a model or source change will influence a run. **[Standardizing insights with the dbt Semantic Layer](https://www.getdbt.com/dbt-summit/sessions/standardizing-insights-with-the-dbt-semantic-layer).** Start from a dbt project, define semantic models and metrics, validate the logic and the grain, then confirm the metric holds up when it's consumed downstream. It's the difference between writing a definition and actually having one. [**Accelerating analytics with AI**](https://www.getdbt.com/dbt-summit/sessions/accelerating-analytics-with-ai). Implement a dbt feature end-to-end with AI assistance, reviewing it against a structured checklist for correctness, performance, and maintainability, then turning those standards into a reusable custom dbt agent skill. **[Migrate stored procs with dbt Wizard](https://www.getdbt.com/dbt-summit/sessions/migrate-stored-procs-with-dbt-wizard).** Everyone has the stored procedure nobody wants to touch. Turn a legacy proc into a dbt project you can maintain and test, using dbt Wizard to generate the model layers and `audit_helper` to validate parity so you can prove the migration is correct rather than hope. You'll finish by defining a Semantic Layer metric on top of the migrated models. [**Beyond the basics: Running dbt at scale on Microsoft Fabric**](https://www.getdbt.com/dbt-summit/agenda/beyond-the-basics-running-dbt-at-scale-on-microsoft-fabric). Build a production-grade medallion architecture on Fabric, orchestrated end-to-end through a dbt job. Incremental models, contract testing at layer boundaries, CI/CD, and metadata-driven orchestration with Fabric Data Pipelines so new sources onboard by config, not by hand. **Breakout sessions** **[YAML doesn't know why: Building the business context your agents are missing](https://www.getdbt.com/dbt-summit/sessions/yaml-doesnt-know-why-building-the-business-context-your-agents-are-missing).** The dbt MCP server, dbt skills, and dbt Semantic Layer give agents technical fluency. The context agents keep missing is organizational. Pedro Heyerdahl of Kilo Code shows how to build a living context layer that multiple agents can read from and contribute to. **[From AI experiment to production: How Okta governs context for agents at scale](https://www.getdbt.com/dbt-summit/sessions/from-ai-experiment-to-production-how-okta-governs-context-for-agents-at-scale).** Okta found that production AI depended less on a better model and more on a governed, discoverable semantic layer any agent could reason over from day one, built on dbt as the source of truth. Pooja Crahen shares how they got there. **[How to build governed agentic analytics on Amazon Redshift with the dbt Semantic Layer](https://www.getdbt.com/dbt-summit/agenda/how-to-build-governed-agentic-analytics-on-amazon-redshift-with-the-dbt-semantic-layer).** Amazon wires the dbt MCP server into the Amazon Redshift MCP Server, so an agent never authors a query, it resolves one compiled from a governed metric definition. The live demo traces one answer back to the dbt model that produced it and the tests that passed on it. **[No drift allowed: LangChain's context playbook with Hex](https://www.getdbt.com/dbt-summit/agenda/no-drift-allowed-langchains-context-playbook-with-hex).** Emily Hawkins and Logan Cochran of LangChain layered dbt, semantic models, and workspace guides into a context stack that turned their Hex agent into a source of truth the whole org trusts. They'll cover how endorsements guard against ungoverned data and how Context Studio closes the feedback loop. **[How Sigma manages the semantic layer with dbt and Dagster](https://www.getdbt.com/dbt-summit/agenda/how-sigma-manages-the-semantic-layer-with-dbt-and-dagster).** Matt Senick walks through the Dagster job that deploys every semantic asset at Sigma, Snowflake semantic views, Cortex Agents, Cortex Search Services, and Sigma Data Models, from a single dbt project each time something changes. **[Sweetwater's proactive data observability playbook with Datadog](https://www.getdbt.com/dbt-summit/agenda/sweetwaters-proactive-data-observability-playbook-with-datadog).** Daniel Gonzalez and Derek Andres on pairing dbt with Datadog to catch a long-running script delaying order updates before anyone noticed, and how dbt Mesh eased the shift to departmental ownership. **[SQL-first AI: bringing BigQuery AI into your dbt project](https://www.getdbt.com/dbt-summit/agenda/sql-first-ai-bringing-bigquery-ai-into-your-dbt-project).** Google's Alicia Williams and Jobin George on using AI.GENERATE, AI.CLASSIFY, and AI.SCORE to turn raw logs into structured insights without leaving your SQL workflow, plus the cost and latency tradeoffs of running AI at scale in dbt. **Peer exchanges** Peer exchanges are small-group, discussion-first sessions. You bring your experience and your notepad. **[Agents, MCPs, and buzzword fatigue: What AI actually changes for analytics engineers](https://www.getdbt.com/dbt-summit/sessions/agents-mcps-and-buzzword-fatigue-what-ai-actually-changes-for-analytics-engineers).** New AI tooling launches weekly and the terminology multiplies faster than the problems it solves. XiaoHan Li of Xebia hosts a hype-free conversation about which tools actually stuck, who owns the logic when AI writes your models, and the skills worth investing in as more of the boilerplate gets automated. **[How to build a successful data career](https://www.getdbt.com/dbt-summit/sessions/how-to-build-a-successful-data-career).** Three practitioners, Millie Symns of Justworks, Silja Märdla of Bolt, and Bruno Lima of phData, trade practical patterns for building career momentum as AI shifts what's expected of the role. Expect honest talk about durable skills, cross-industry moves, and making your impact visible without defaulting to "just become a manager." ## For data team leaders You're deciding how your team scales, standardizes, and stays ahead. These sessions are about rollout patterns, governance, and positioning your team for what's next. **Hands-on labs** **[Scaling trusted self-service for dbt stakeholders](https://www.getdbt.com/dbt-summit/sessions/scaling-trusted-self-service-for-dbt-stakeholders).** Scale dbt beyond the build team by helping stakeholders find, understand, and reuse trusted data products without turning everyone into a developer. Walk you through documentation patterns, ownership, and a stakeholder access model that expands governed consumption while protecting your development workflow. **Breakout sessions** **[From selection to scale: How ING is operationalizing dbt across a global bank](https://www.getdbt.com/dbt-summit/sessions/from-selection-to-scale-how-ing-is-operationalizing-dbt-across-a-global-bank).** Jarno Boeijink shares how ING drives governed enterprise adoption inside a regulated bank, and how the dbt Semantic Layer and the dbt MCP server are opening new ground for natural-language analytics and AI-assisted development. **[Governed by default: How data teams at Nordstrom turn dbt governance into an AI advantage](https://www.getdbt.com/dbt-summit/sessions/governed-by-default-how-data-teams-at-nordstrom-turning-dbt-governance-into-an-ai-advantage).** Nadine Bruxel makes the case for dbt as the control surface for safe AI: freshness as an AI SLA, tests and contracts and lineage as guardrails, and a conversational agent built on top of all of it. Governance-first is how Nordstrom gets to AI readiness. **[Multi-agent dbt orchestration at Riot Games: Redefining the analytics engineering SDLC](https://www.getdbt.com/dbt-summit/sessions/multi-agent-dbt-orchestration-at-riot-games-redefining-the-analytics-engineering-sdlc).** Jessica Zhang shows how Riot Games safely coordinates multiple agents to read metadata, translate legacy logic, and generate pull requests, all behind read-only guardrails. If you want to know how far multi-agent workflows can go in production, this is the session. **[Scaling dbt on Amazon Redshift: how KOHO cut transformation runtime 70% without rewriting a single model](https://www.getdbt.com/dbt-summit/agenda/scaling-dbt-on-amazon-redshift-how-koho-cut-transformation-runtime-70percent-without-rewriting-a-single-model).** KOHO re-architected from a single Redshift cluster to a Hub and Spoke model and a data mesh, cutting its nightly dbt runtime from 6.5 hours to 2 in a config-only migration, no model rewrites, no downtime. **[From Data to AI: Why Microsoft Fabric and dbt are better together](https://www.getdbt.com/dbt-summit/agenda/from-data-to-ai-why-microsoft-fabric-and-dbt-are-better-together).** Roy Hasson, Pradeep Srikakolapu, and Abhishek Narain of Microsoft on how OneLake, unified security, and AI-powered experiences pair with dbt's testing, lineage, and documentation to cut fragmentation and build a trusted foundation for analytics, AI, and agentic workloads. **[From prototype to production: How DoorDash built a scalable analytics SDLC with dbt and ThoughtSpot](https://www.getdbt.com/dbt-summit/agenda/from-prototype-to-production-how-doordash-built-a-scalable-analytics-sdlc-with-dbt-and-thoughtspot).** Harsha Reddy on leading the modernization of DoorDash's analytics stack onto dbt and ThoughtSpot, consolidating a large, multi-team org onto a platform that now serves more than 10,000 users. **Peer exchanges** **[Beyond the bottleneck: Position your analytics engineering team as a strategic force](https://www.getdbt.com/dbt-summit/sessions/beyond-the-bottleneck-position-your-analytics-engineering-team-as-a-strategic-force).** Kasey Mazza of HubSpot leads a discussion on moving your analytics engineering team from a service desk to a strategic driver, with the framing and language to make that shift stick with leadership. **[How to build a successful data career](https://www.getdbt.com/dbt-summit/sessions/how-to-build-a-successful-data-career).** Worth the crossover for leaders too. Millie Symns, Silja Märdla, and Bruno Lima swap patterns on durable skills and career growth, useful for anyone coaching a team through the AI shift. ## For business leaders and executives You're weighing where data investment turns into measurable outcomes. These sessions lead with results, governance, and the business case for a strong data foundation. **Breakout sessions** **[Real-time analytics at Bilt: Architecture and approach](https://www.getdbt.com/dbt-summit/sessions/real-time-analytics-at-bilt-architecture-and-approach).** James Dorado and the Bilt team walk through the architecture behind their real-time analytics, powering audience targeting and offer execution on fresh data. **[An AlphaSense case study: Scaling AI on enterprise data with dbt-first governance and context from Euno](https://www.getdbt.com/dbt-summit/sessions/an-alphasense-case-study-scaling-ai-on-enterprise-data-with-dbt-first-governance-and-context-from-euno).** Sarah Levy of Euno and Brad Levy of AlphaSense show how dbt-first governance, paired with automated context, keeps AI decisions explainable and traceable back to governed source data. **[Governed by default: How data teams at Nordstrom turn dbt governance into an AI advantage](https://www.getdbt.com/dbt-summit/sessions/governed-by-default-how-data-teams-at-nordstrom-turning-dbt-governance-into-an-ai-advantage).** The executive read on the Nordstrom story: governance is the thing that makes AI on your data safe to trust and safe to scale. Nadine Bruxel shows how a governance-first foundation turns into a real advantage in the AI era. **Peer exchanges** **[Empowering stakeholders in the age of AI](https://www.getdbt.com/dbt-summit/sessions/empowering-stakeholders-in-the-age-of-ai).** Lexi Galantino of Zipline hosts a conversation on what "talk to your data" actually takes: which models make it work, how you keep the answers correct, and what the role of the data team becomes when stakeholders can propose their own changes. ## Build your week around it These are a fraction of the 100+ sessions on the agenda. Anchor your schedule around the two keynotes, Level Up on Wednesday and the Community Keynote on Thursday, then fill in the breakouts, labs, and peer exchanges that map to what you're building next. Registration is open now, and the $1,695 registration includes a free training and certification while spots last. Bring home the templates, playbooks, and patterns you can put to work immediately. Your next level starts here. [Register for dbt Summit](https://www.getdbt.com/dbt-summit/registration). --- --- title: "OSI is now Apache Ossie (Incubating)" description: "Apache Ossie is currently undergoing incubation at The Apache Software Foundation (ASF)." url: "https://www.getdbt.com/blog/osi-is-now-apache-ossie" date: "2026-07-13" authors: ["Quigley Malcolm"] categories: ["Partnerships"] --- # OSI is now Apache Ossie (Incubating) If you've been following the Open Semantic Interchange (OSI) project, the open specification for semantic layer and ontology, there's an important update. The project has been accepted into the Apache Incubator. Along with this transition the name is changing to Apache Ossie (Incubating). The spec, the community, and the mission haven't changed, but the name, governance home, and long-term trajectory have. ## Why the new name? When work on this initiative was started, it was called the Open Semantic Interchange. This quickly got shortened to OSI, which became the GitHub repo name. Unfortunately this has caused some confusion along the way as the acronym OSI is used frequently to refer to Open Source Initiative. Now although it might be fun to say OSI OSI (Open Source Initiative Open Semantic Interchange), the community decided it was best to rename the project. Through discussion the community decided on Ossie, and with acceptance into the Apache incubator, it is now Apache Ossie (Incubating). In addition to the rename, a mascot has been chosen, a kangaroo. The Ossie Kangaroo is dedicated to carrying semantic metadata in its pouch from platform to platform. That is, it’s making your data hop. In short: - The project is Apache Ossie (Incubating) - Any reference to "OSI" in the project are historical (and will slowly be removed) - There is a kangaroo logo If you've been building on Open Semantic Interchange, nothing breaks. The name changed, but the spec didn't. ## What is Ossie? Ossie is an open specification for both semantic layer and ontology. It defines a vendor-neutral format for expressing business metrics, dimensions, relationships, as well as broader business concepts and rules. It allows any tool or platform in your semantic layer stack to produce and consume semantic definitions without loss of meaning. The problem it solves is important: it ensures that a given business concept (say, "Monthly Active Users") can be defined, interpreted, and resolved consistently across an organization's CRM, data warehouse, and BI tools. When a human analyst or an AI agent runs a query, they shouldn't have to guess which definition is correct. Ossie provides the shared, machine-readable format that encodes not just the data but the intent and business meaning behind it. ## Why the Apache Software Foundation (ASF)? Incubating Ossie with the Apache Software Foundation ensures that it remains an open standard with no single controlling entity. The goal of Ossie is to provide industry-wide standardization of semantic data, and to that end ensuring that it has a vendor neutral ground to operate in is imperative. Under incubation, Ossie operates with public mailing lists, GitHub-based development, a formal discussion-and-vote process for spec changes, and committership earned through contribution rather than employer affiliation. Note that as part of this transition, all mailing lists referring to Open Semantic Interchange will be retired; community members should use the ASF-provided project resources that are linked below instead. ## Importance of the Ossie community Ossie didn't start as a single-company project, it has been a community effort. Since the repository opened in November 2025: - More than 100 commits and 35 merged pull requests have landed from contributors at Snowflake, Salesforce, Databricks, dbt Labs, RelationalAI, GoodData, and Honeydew - The participating coalition has grown from 17 launch partners to [more than 50 organizations](https://www.snowflake.com/en/blog/open-semantic-interchanges-specs-finalized/) - Three working groups (Metric Language, Catalog, and Ontology) operate with dedicated leads, meetings and public channels - Implementations including the Ossie-to-dbt Semantic Layer converters and an Apache Polaris™ converter are already merged ## What's next dbt Labs was one of the founding organizations behind Open Semantic Interchange, and we'll continue as an active contributor to the project as it grows under ASF governance. As with any Apache project, the community will decide the direction together. That said, there are a few areas we're excited about and hope to work with the community to contribute proposals for: - Deepening the spec's expressiveness to accommodate what real enterprise models demand, including an expression language spec, advanced metric logic, windowing functions and complex relationships - Building converters for additional platforms and frameworks so that adopting Ossie doesn't require ripping out what you already have - A standardized semantic query specification that any engine can support - Integration with Apache Polaris so that semantic models are discoverable directly from the catalog None of this is predetermined. It will go through the same open discussion-and-vote process as everything else in the project. ## Get involved Ossie is transitioning to ASF infrastructure as part of incubation. Watch for updates on the new [project website](http://ossie.apache.org), join the [development mailing list](mailto:dev@ossie.apache.org), collaborate on [GitHub](https://github.com/apache/ossie) and join the [Ossie Slack workspace](https://join.slack.com/t/apache-ossie/shared_invite/zt-42i1xkgy8-7YQtKEDq7v~mceFmdiLhkA). Whether you're building an AI agent, BI tool, or a query engine that needs to understand business context, Ossie is the community working to make sure you don't have to tackle semantic interoperability alone. --- --- title: "The productivity gains hiding in your data infrastructure" description: "Budgets aren't growing, but the work is. See how dbt customers recouped 58.7 FTEs in capacity, worth $1.75M a year." url: "https://www.getdbt.com/blog/data-infrastructure-productivity-gains" date: "2026-07-08" authors: ["Daniel Poppy"] categories: ["Product"] --- # The productivity gains hiding in your data infrastructure Demand for data is exploding thanks to AI. Some experts estimate that global spending on data centers capable of handling advanced AI workloads [could hit $7T by 2030](https://www.mckinsey.com/industries/technology-media-and-telecommunications/our-insights/the-7-trillion-dollar-data-center-build-out-how-industrials-can-capture-their-share). Data teams aren't necessarily getting more resources to deal with it, though. When we talked to companies, we found that only 36% of data teams reported increasing budgets. That means data teams have to shoulder more work with existing capacity. AI itself helps, of course. Tools like [data copilots](https://docs.getdbt.com/docs/dbt-ai/copilot-overview) reduce the burden of generating code that produces clean, high-quality data for AI agents. But AI can only do so much. The hard problems still need humans to solve them. And currently, those humans are mired in maintenance, working assiduously to keep the data house of cards from falling down around them. The good news is that data teams can free up significant capacity by taking a modular, reusable, automated approach to managing AI and analytics data workloads. This isn't just speculation. [A recent IDC report quantifies](https://www.getdbt.com/resources/the-business-value-of-dbt-idc-report) exactly how much using a platform like dbt can save at every step of the data lifecycle. ## The maintenance nightmare No data team handles data from a single source. Everyone is constantly wrangling data from a variety of data storage platforms and formats. The advent of AI has made this even more of a challenge. Data teams aren't just dealing with structured relational data and semi-structured data sources any longer. They're also mining PDFs, emails, and social media posts for insights. All this means that most teams end up taking a scattershot approach to managing data pipelines. Most are written on the fly as quick and dirty one-offs meant to get the job done. The result? A maintenance nightmare. - Pipelines are brittle and prone to breaking. Respondents in our annual [State of Analytics Engineering Report](https://www.getdbt.com/resources/state-of-analytics-engineering-2026) reported that they spend a significant amount of their time maintaining data sets, platforms, and infrastructure. - Most work isn't reusable across data pipelines, forcing teams to rebuild what they need from scratch each time. - Testing, deployment, and review are slow, manual processes, if they exist at all. - New data contributors face a long ramp-up time learning how to navigate heterogeneous data systems. Many systems end up being too complex for non-technical contributors to use efficiently. Each of these factors eats up precious time that data team members could instead be spending on more strategic work, such as streamlining data intake, optimizing storage, improving overall scalability, and automating key processes for faster delivery and more consistent quality. ## dbt: Unlocking capacity without hiring For years, dbt has served as the [data control plane](https://www.getdbt.com/blog/data-control-plane-why) for companies worldwide. By taking a single, vendor-agnostic approach to modeling data, testing changes, and orchestrating data pipelines, dbt reduces complexity, boosts reusability, and improves quality across all analytics and AI data workloads. The results, as summarized by the [IDC Business Value report](https://www.getdbt.com/resources/the-business-value-of-dbt-idc-report), are real and measurable. IDC interviewed eight enterprise companies that use dbt, with an average of 20,000 employees and $16.75B in annual revenue. In headcount numbers, across various data functions, companies reported recouping the equivalent of **58.7 FTEs' worth of capacity**. The businesses reported doing more with less headcount across all data functions: - Analytics and data teams: +25 FTEs - Developers: +11.2 - Business analysts: +17.6 - Data governance: +2.5 - Platform management: +2.4 The analytics/data team savings alone are the equivalent of $1.75M of additional business value delivered every year just by using dbt. In each case, functionality provided out of the box by dbt enabled teams to achieve these gains: ## A speed boost across the data stack These savings come from reducing the cycles required to perform basic data tasks. At every step of the data lifecycle, dbt enables teams to deliver more in less time: - Report delivery drops from 16.3 days to 8.4 days, a **49%** acceleration - Testing is **44% faster** for new apps, and **46% faster** for pipeline updates - Development cycles are **41%** and **37% faster** for new features and updates, respectively - Teams reported delivering new solutions to market **34%** faster, and scaling **33%** faster "dbt has significantly improved developer collaboration and increased our development velocity," one customer told IDC. "It introduced a structured deployment process through its integration with Git." ### Faster onboarding and reuse Teams also reported significantly faster onboarding times. Onboarding dropped from an average of 3.3 weeks to 1.7 weeks, a **47% acceleration**. The driver? The availability of existing models. With data modeled consistently and available for self-service discovery via [dbt Catalog](https://www.getdbt.com/product/dbt-catalog), new hires have access on average to 101,970 reusable models from day one. This means the productivity gains offered by dbt aren't a one-time bump. Each new model is an investment that compounds into the future as new hires plug existing data transformation code into their own pipelines. "dbt platform is very easy to use for new employees," reported one customer. "It's all templates and SQL, so people are comfortable with it." ### Reducing the data quality tax Testing, automated deployment, and clear documentation have a measurable impact on data quality. Companies reported **33% fewer data quality issues** due to dbt. They also reported **35% fewer late-data** and **13% fewer incomplete-data** instances. This is important because data teams historically spend much of their time responding to and fixing data quality issues retroactively, after they make it to production. Studies in software engineering have shown that fixing bugs in production is significantly more time-consuming and expensive than ensuring those defects never ship in the first place. A bug that costs $100 to fix in the design stage of a project [may become a $10,000 problem](https://www.forbes.com/councils/forbestechcouncil/2023/12/26/costly-code-the-price-of-software-errors/) if it makes it to production. "We use dbt to run tests such as checking for duplicates, missing values, and incorrect data," another customer said. "dbt helps identify issues and traces them back to their source, allowing us to fix problems at the system level and deliver more reliable data." ## The capacity was never missing None of these gains come from a single feature. They come from applying a single discipline, the [Analytics Development Lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle), to every stage of data work: - **Build**: Reusable models replace one-off scripts. Teams solve each problem once, not repeatedly. - **Test**: Automated checks catch errors before they ship, not after they've reached a dashboard or an AI model and led to a costly business decision. - **Deploy**: Continuous Integration and Continuous Delivery (CI/CD) move changes automatically through a rigorous change management process, not via manual pushes and after-the-fact firefighting. - **Operate and discover**: Data lineage makes it easier for engineers to identify the root cause of data issues. Data cataloging and health reports enable self-service so that everyone can leverage what already exists instead of re-inventing the wheel. This capacity wasn't missing. It was buried in maintenance and the errors caused by manual work. dbt frees up this capacity so that teams can ship more and ship faster using the resources they already have. To learn more about how dbt saves companies time and money, [download the full IDC report](https://www.getdbt.com/resources/the-business-value-of-dbt-idc-report). To get started enhancing your data productivity, [talk to us today](https://www.getdbt.com/contact) about how dbt fits into your existing data stack. --- --- title: "Solving dashboard errors in minutes: How Integral Ad Science used MCP to connect agents to dbt and Databricks" description: "Integral Ad Science used MCP to connect AI agents to dbt and Databricks, turning hours of dashboard debugging into minutes." url: "https://www.getdbt.com/blog/mcp-dbt-databricks" date: "2026-07-07" authors: ["Daniel Poppy"] categories: ["Product"] --- # Solving dashboard errors in minutes: How Integral Ad Science used MCP to connect agents to dbt and Databricks Stop us if you've heard this one. A user goes into a report on a BI dashboard and starts asking questions of the embedded chatbot. They're drilling, filtering, checking the numbers… And all of a sudden, things don't add up. This isn't an uncommon problem for analysts in the AI data age. Mars Dauer, Senior Director of Enterprise Data and AI at Integral Ad Science (IAS), and his team ran into the same issue. They knew they had the information that their AI agent needed; it was already there in dbt and Databricks. The challenge was how to supply it to the agent so that it could solve what Dauer and his team call the "BI why" problem. At Databricks Data + AI Summit 2026, Dauer shared how teams can supply this missing context via model context protocol (MCP) servers that connect AI chatbot agents to dbt and Databricks. In this article, we'll dig into the architecture he used to solve the problem, the decisions made, and the honest lessons about what worked and what the team learned. ## The BI "why" problem: What copilots can't answer BI chatbots have what Dauer calls a "BI why" problem. They can tell you what the numbers say. However, they can't explain the context and reasoning behind them. Let's assume, for example, that the company is a supermarket chain. An analyst is digging into a revenue report and breaks it down by category — and Food jumps out as unreasonably high. But why? Can the AI agent tell them where the disconnect is? This is where the analyst realizes (to no one's surprise) that the agent can't. It can summarize the chart; it can tell them what filters are applied. But it has no real understanding of the business logic or the transformations that produced those numbers. So, the analyst digs deeper. They open [Looker](https://cloud.google.com/looker) and check the LookML. That points to a dbt model, so they jump into the model and read the SQL. That model leads to three upstream marts, so they open more SQL in more tabs. They eventually run a few validation queries in [Databricks](https://databricks.com) and discover that the models have miscategorized a popular drink as food. Looking at the architecture of a typical AI chatbot system in BI tooling, it's easy to see why the bot couldn't unearth this issue. The BI copilot can see only the topmost layer of the data stack: - The BI tooling itself - The semantic layer that defines a unified layer for key metrics and governance - The mart models containing the gold standard data [Slide 5] However, it can't see everything _under_ that—in this case, an errant CASE statement. The agent can't see the business logic built into the dbt intermediate models, or the sources in Databricks it could use to validate and correct this error. ## MCP: Your USB-C for LLM tools Dauer knew that data was available, which meant the answer was somewhere. The issue was supplying it to the BI copilot so that it could solve the issue in minutes, not hours. The answer his team hit on: MCP. [MCP defines a standardized framework](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2.0&name=Fusion) that enables AI applications, such as agents, to talk to tools and understand their capabilities without needing to know their internal APIs. Dauer refers to MCP as the USB-C for large language models (LLMs): the agent has one kind of port, and any tool that exposes an MCP server with the right shape can plug into it. Before MCP, every integration meant a new bespoke client. With MCP, tools advertise their capabilities, and any MCP-compatible agent can discover and utilize them. For the project, Dauer's team needed three MCP servers to solve the issues they'd detected: - The **dbt MCP server** (the structured context layer), which connected the agent to tested, version-controlled, and documented model definitions — including model metadata, [data lineage](https://www.getdbt.com/blog/what-is-data-lineage), SQL, and field/table descriptions; - A **Databricks SQL MCP server** (governed data validation layer), which enabled read-only query execution and [Unity Catalog](https://www.databricks.com/product/unity-catalog) governance; and - A **Looker MCP server** (the BI on-ramp), providing access to SQL parsing for [explore URLs](https://docs.cloud.google.com/looker/docs/viewing-and-interacting-with-explores) and dashboard tiles. The team had been using all these MCPs for a while in coding agents like [Cursor](https://cursor.com/), [Claude Code](https://www.anthropic.com/product/claude-code), and [Visual Studio Code](https://code.visualstudio.com/). Using them in conversational/analytics agents felt like the logical next step. ## How the analytics agent works The agent has five core capabilities. Dauer thinks about them in three buckets: **Lineage and logic.** Given a column or metric, it traces that object from the mart back to the source. It reads the compiled SQL at each hop, and explains in plain language, what each transformation is doing. **Validation.** Once it has the lineage, the agent can validate against the actual data in Databricks using read-only SQL. It can also perform freshness checks and monitor pipeline health. **Impact analysis. **If you change a column upstream, the agent tells you what breaks downstream, and because it reads the logic in each downstream model, it can reason about whether your change matters semantically, not just structurally. The agent has two personas depending on who's using it. In **analyst mode**, it returns the answers that data producers and analysts need: full technical details, including model names, compiled SQL, column catalogs, directed acyclic graph (DAG)-aware lineage, and raw validation query results. In **business mode**, it gives only the information that matters to decision-makers: plain English answers (without jargon or model names), business context, and summarized findings. As Dauer emphasizes: "One agent, two voices: analyst mode shows model names and compiled SQL; business mode strips both for plain English. Same investigation, different answer for a CMO vs a data engineer." ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/09d66b3d44fb7c377a210cc235d98eeb9c2c3d72-2046x1068.jpg) For Integral Ad Science, Looker was the primary point of entry. The agent was in the BI layer so that users didn't have to context-switch. The agent runs as a window to the right in the main Explore to provide the best user experience. The backend is a FastAPI server running on Cloud Run. It handles auth through Google OAuth and an authorized group, manages sessions, and streams responses back over Server-Sent Events. Then there's the agentic runtime, which uses Google's Agent Development Kit. It's composed of a single root agent and four sub-agents. The root agent takes the user's question, decides which specialists to call, fans out the work, collects the results, and formats the final answer. The LLM layer here is deliberately provider-agnostic via the LiteLLM library. This means each subagent can target Databricks Foundation Model APIs, [Gemini Enterprise Agent Platform](https://cloud.google.com/products/gemini-enterprise-agent-platform), Anthropic, [OpenAI](https://openai.com/) - whatever best fits the job. This means subagents can use smaller, less expensive models for simpler jobs, such as routing or analyzing metadata. Finally, the MCP layer connects to dbt platform, Databricks SQL, Looker, and a few internal tools. All of the company's source systems are dbt projects built upon Unity Catalog in Databricks. ## Why MCP, and other design decisions A key question is: why MCP? Why build this using a universal adapter as opposed to creating custom integrations? A few reasons: **Speed.** Adding a new tool source meant pointing at a new MCP server, not writing a custom API client. This enabled dbt and Databricks to come online in days, not weeks. That speed compounds when you're still experimenting. **Decoupling.** Using MCP means the agent doesn't need to know how dbt, Databricks, or Looker work. Dauer's team could swap out any of these tools, and the agent's code wouldn't have to change. **Governance.** This is the one that matters most to security. The MCP server controls what's exposed. Integral can scope tool access, audit at the server boundary, and avoid handing raw API credentials directly to the agent. Databricks provides governed query execution; Looker provides the BI on-ramp. dbt provides the semantic and lineage layer that makes the whole investigation possible. It's the reason the agent can trace a metric from a dashboard tile all the way back to a source table — and trust what it finds there. Why rely on dbt MCP for lineage? Because dbt treats lineage as a first-class concept. The agent doesn't reconstruct the graph from scratch or hope documentation keeps pace with reality — it reads dbt's tested, version-controlled graph of how data actually flows. Alongside the graph, the agent reads the compiled SQL itself — and that's the ground truth. LookML is a model of a model. Documentation is what people intended (or hoped) the model did. Compiled SQL is what actually runs in the warehouse. The lineage tells the agent where to look; the SQL tells it what's happening there. ### Trade-offs and lessons learned Every architectural choice was a trade-off. Some of the key trade-offs that Dauer's team made: **Multi-agent vs. single-agent** The team started with a single agent. As tool complexity grew and prompts got longer, they decided to go multi-agent — specialists under a main routing agent. They landed at four. Every additional sub-agent adds a routing failure mode. The orchestrator has to choose the right specialist, and when it chooses wrong, they wouldn't always get a clear error. That led Dauer to make a rule: decompose only when the specialist has a genuinely different job and needs a meaningfully different system prompt. If two agents both query Databricks, and one is called "validation" while the other is called "exploration," they probably should be one agent. **Embedded vs. standalone chat** A standalone chat is tempting because you can ship it quickly. But it forces the user to copy context, meaning that every investigation starts with thirty seconds of context-pasting. The embedded Looker extension avoids that: it captures the dashboard/explore context automatically, so the question and the data live on the same surface. The trade-off here is real. Supporting another BI tool means redoing some of this work each time. But the savings per investigation compound. Bottom line: if your BI tool has an extension SDK, use it. **Per-sub-agent model tiering vs. one model for everything** The easy default is to pick one frontier model, use it for every sub-agent, and move on. IAS doesn't, Dauer said. This means simple agent tasks, such as the Looker resolver, aren't forced to use the most expensive model. They can save the more expensive model for operations such as the dbt analyst, where lineage reasoning over compiled SQL greatly benefits from the extra firepower. The rule: match the model to the work. The user doesn't feel a quality hit because the critical reasoning step still gets the strongest model. ## Only scratching the surface What used to take an analyst an afternoon, Dauer said, now takes minutes. And yet he feels like his team is only scratching the surface. The dbt platform MCP server exposes ~37 tools. IAS's solution only uses seven of them. That's actually a testament to how teams can get started narrow and expand without rebuilding anything — the breadth of the MCP surface area means the integration grows with your use case. Some of the tools the team is looking forward to including are: - Text-to-SQL with project context to generate SQL that knows the project's model names, joins, and conventions. - Code generation to create staging model YAML and sources. Leveraging this, agents could not only investigate issues but scaffold a fix and submit it for human review. - [Column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage?version=2.0&name=Fusion) via the Fusion engine — the agent already traces column-level lineage today by reading compiled SQL and inferring how columns flow. The Fusion-powered tool would give the agent that same lineage as a deterministic graph straight from dbt's compiler — replacing LLM inference with explicit column-to-column edges. Faster, cheaper, more precise. "The most interesting agents in this space probably haven't been built yet," Dauer said. "And, honestly, that's the part I'm most excited about." --- --- title: "A guide to implementing AI data pipelines" description: "AI coding took off. AI pipeline management didn't. Here's a practical playbook for closing that gap." url: "https://www.getdbt.com/blog/a-guide-to-implementing-ai-data-pipelines" date: "2026-07-06" authors: ["Stephen Thibeault"] categories: ["Insights"] --- # A guide to implementing AI data pipelines There’s a startling pair of statistics in dbt’s newly released [2026 State of Analytics Engineering](https://www.getdbt.com/resources/state-of-analytics-engineering-2026) report: 72% of teams prioritize AI-assisted agentic coding for data work, while only 24% prioritize AI-assisted data pipeline management. Here are my thoughts on why this is happening and, if this is the case in your own org, how to go about building AI into your everyday workflows. ## Why AI pipeline management is lagging behind AI coding Looking at these numbers in relation to each other really shows the difference between **AI adoption** in terms of giving developers AI tools to work with, and **AI integration** by working AI into the data infrastructure itself. AI-assisted coding is very self-driven. You can be on your individual system, working on your own, using an AI coding tool on whatever dbt project you personally have pulled down. But using AI for pipeline management, like looking for job errors and feeding that data into an agent, has to happen at the team level. I can absolutely spin up an agent in a five-minute Claude Code session, but actually building the bots that would monitor a data pipeline takes time, dedication, and collaboration from basically the whole team. One of the largest clients I'm currently working with has a team of 80 engineers throughout their entire company that touch their dbt project. But when it comes to actually orchestrating that project, it's a team of maybe five engineers managing the orchestration out of the dbt platform. I think that's why we're seeing a slower uptake on using AI to assist data pipeline management, because it has to go through different approvals and the team has to be aligned on what exactly they want that to look like. Ultimately, the difference in those numbers are really about accessibility: AI coding is individual, personal, and easy for developers to adopt, versus actually building and managing AI pipelines in production as a team. ## Your data stack is already AI-ready The typical data architecture for many orgs today follows the ELT pattern: data is extracted from source systems, loaded into a data warehouse, and transformed for downstream consumption by business users. **A lot of bringing AI into your daily workflows does not necessarily change that stack. **You’re just layering AI over it, in the same way that dbt is currently layered over your data warehouse. The original data stack structure is still very solid, so we're just keeping that; AI now exists on top and covers the whole thing (or, alternatively, most of the original stack, because you absolutely can add this AI layer in pieces). ### Where to start The first step to implementing this new AI layer is to start with the highest-value layers. Here are places to consider adding in AI. **Code review** is where you can get the most out of AI by integrating it for things like reviewing PRs, CI checks, flagging issues like _'Hey, this is missing. We need to take a look at it.' _This is a natural starting point for introducing AI into data pipelines and workflows because it’s where developers are already using AI for their individual work. **Error triage** is a logical place to layer in AI: making AI part of your code review process, using it for error tracking in your data transformations, and even in your data loading workflows. In practice this would look like setting up AI to do that initial triage work and then setting up alerting, be it through email, Slack, PagerDuty, or whatever system you use for error communications. **Extract/load** is also a small but very targeted opportunity. Here, AI could be useful for evaluating what's coming in, making sure nothing coming in is messed up, evaluating errors, etc. Your** ticketing intake system** can be another great place to start. One dbt customer created a ticket intake chatbot to make sure they had access to all relevant context around requests. This chat interface gathered information about the problem that they're trying to solve, rather than simply intaking a verbatim request that essentially dictates a solution. Business users often don't know what the data team has access to or even really grasp the whole breadth of what that team can do. Applying AI here to create collaboration can produce solutions that give more than the person requesting it even realizes is possible, which transforms a data team from order takers into internal consultants. **Ultimately, each layer of the stack from intake to data loading to data warehousing to consumption have use cases for AI. **It's all about identifying those use cases and implementing them into that workflow, and you find them by looking for places where the addition of an AI tool would be helpful for you and your team overall. ## A playbook for getting started For an applied example, let’s say that you are currently the leader of a small data team of 10 people and you provide analytics to the sales team. You have a current modern data stack that consists of five ingestions, Snowflake as your data warehouse, [dbt to do the transformations](https://www.getdbt.com/blog/data-transformation), and Tableau for your data visualizations. A typical day-to-day for this team right now looks like this: business users put requests in, your data teams take those requests and create what's needed for the business, and finish with going through the review process. Where does AI enter into this picture? ### Step 1: Find the first, low-stakes pain point The first step in determining what you should do with AI is finding your use case. You do this by asking,_ 'Hey, where are our biggest pain points currently? Do we not have enough time to do extremely thorough code reviews? Or do we get incredibly vague information from business requests? Do we frequently get failed loads from our loading tool into the database due to changing APIs or other reasons?’_ **Code review **is the example we will use here as it is what a lot of people start with to really get their feet under them because, if something goes wrong, there is typically already an assigned human in the loop. To do this, you would: 1. Start by deciding orchestration, figuring out what automations exist and how you want to run those automations. Are you going to just use something like GitHub Actions? Do you have an external orchestrator that will run this process whenever a pull request is created in your change control tool of choice? 2. Next, determine which AI company you want to use (which is often determined by whichever one(s) your company already has enterprise agreements with) and which of their AI models you want to use. Obviously, because it's your code, you want to make sure that there are agreements for security, for not using your data for training, etc. All this should be handled in tandem with the IT team and the security team. 3. Build the behavior as skills: You would put together a skill that says, 'Hey, you're a code reviewer. These are our coding standards. We don't use leading commas. We always use CTEs in this specific format, we're using the dbt standards.' Whatever is important to you, it all goes here. _Pro tip: dbt has a [curated collection of agent skills](https://github.com/dbt-labs/dbt-agent-skills) that help AI agents understand and execute dbt workflows more effectively._ 4. Add the business context into the skill: You would also put information about the business, like the sales team expects this certain format, here's the basic structure of our data. You would build this all in markdown, use it as a skill that whatever AI tool you selected to automate this process can actually follow those instructions. ### Step 2: Set up a pilot Now, it’s pilot time for your AI-powered automated code review: roll it out, then let it run, monitor it closely, make sure you're still doing human code reviews. You’ll want to: - **Test and iterate before going wide**. It all comes down to testing that automation. So either putting it on a small test repo or putting it out there, but only running it kind of ad hoc to make sure that it works. So, doing a couple of code reviews, giving it some code, refining what that skill looks like, refining what you're sending to the AI. - **Pilot small and expand outward.** Your pilot is what builds that trust over time. Just releasing it to 500 people all at once will just create a mess trying to keep up with managing it. - The final goal is up to you, but for most teams this will be putting the AI code review agent on your main repo so that whenever somebody pushes to that repo, they're going to get this automated code review. ### Step 3: Gather feedback, create evals Let your people who are excited do the actual work of setting the pilot itself. When it comes to evaluation, however, you need to have people from multiple walks of life do that review so that you can make sure you're getting a well-rounded understanding of what's working and what's not working. The best feedback is going to come from a standard senior member on your team that is perhaps indifferent or maybe even a little combative when it comes to AI adoption. They'll be able to spot things that someone who's very excited about it may not spot. Have more senior members of the team who do a lot of code reviews also review the output and see if they can spot things that the AI is not doing correctly. This feedback can then be used to refine the Skill you are utilizing. Once those more senior members of the team feel comfortable with the results that you’re getting, now you work with them to create the eval for keeping things accurate over time (we’ll talk more about evals shortly). Finally: remember that this is all brand new, so don't feel like your pilot has to be perfect out of the gate. Making mistakes is how you learn not to make those mistakes; that's the IT way of life. The rite of passage for every data engineer out there is accidentally dropping a production table or breaking something in a transaction. But I predict that pretty soon another rite of passage for data engineers will be having an AI pilot that just didn't work out. Because, after all, how can you truly understand something unless you've broken it and then had to fix what you broke? ## Maintenance and evals are part of the planning process The crucial work in integrating AI into your data pipelines isn't actually building the pipelines, but making sure they stay reliable once they’re up and running. The most important message I can give you is that **integrating AI isn't a one time effort**. These systems need to be regularly evaluated by the team because LLM performance can degrade over time. Models change, something's being throttled, context window limits change, and sometimes the cost to performance ratio no longer makes sense. **Maintenance has to be a large part of the planning process, not an afterthought. **If you don't think about maintaining an AI workflow from the moment you first decide that you're going to implement it, you'll end up with either something that nobody uses or something that degrades so much it erodes trust in the system. Going back to our AI code reviewer example, most people set one up within probably a few days through GitHub Actions, a couple of API keys, a couple of calls, it's good, it does what you need. You’re not done, though. Where you're going to win or lose is how you maintain that system and how you make sure that that model is giving consistent results, otherwise all you're doing is creating noise that people won't use. ### Evaluation loops are how you keep the system honest An essential piece of integrated AI workflows is having evaluation loops where you can make sure that the LLM is still performing as expected in each one of these processes. **Implementing these evaluation loops is one of the biggest pieces in productionizing AI.** In practice, evals look like different things depending on the workflow you’re attempting to automate. For an AI ticket intake chatbot, this could be having a generic ticketing workflow that you can automate to pass through the LLM, say once a week. For an automated error handling system, you’d send out a generic error message through your error workflow at a preset cadence so that it can be reviewed and you can make sure that the performance of the AI system you're using stays on point. If you try to skip this part of the process, what can happen is people don't notice the degrading performance at first, until suddenly you're getting skewed data. Again, faulty data is really going to poison any kind of adoption. ### Evals are OG ML Evals aren’t something we’ve just added due to agents. People who've been in the traditional AI frameworks space largely understand that evals are an integral part of the machine learning flow. Before I came to dbt I worked on a very large project doing predictive modeling for a manufacturing company; this was long before agents, but we still did evaluations as part of those workflows. One thing we would do for example is generate a month-over-month capture of basically,_ ‘Hey, we were doing good at predicting X last year. We're not doing so good now. Do we need to retrain or use a different model?’_ A lot of people are coming into AI now because it's the new thing, and often they skip looking into things like evaluating and making sure it's a system that can stand the test of time in production. But it’s always been the case that, whenever you're dealing with a machine doing things on its own, you have to give it guardrails and give it controls, to make sure that you know what's going on under the hood. And I feel like that's missed by a lot of people in the AI race nowadays because everyone's just scrambling to even keep up with what's going on. ### Evals run on humans, too Ultimately, evals are only partly automatable. You could have automated systems that compare the output of your evaluation month over month, but you need a human in the loop for review. If there is any sign of degradation in performance you would want to have a human reviewer apply judgement: is this true degradation, or is the model just giving slightly different answers? Other areas where humans still need to be involved in AI workflows include making decisions around whether to switch models or change something about the skills that you're using in your markdown files or changing what tools the agent can access through MCP. Changes like these need a human to review the results you’re currently getting, and make those determinations around any changes needed, and then rerun those evaluations to make sure that it's still giving the outputs that they would expect My baseline message to builders here is, don't get lazy with it. Don't just accept whatever the AI says, or even use a majority of what the AI says unless it's very boilerplate. Make sure that these outputs are being reviewed, they're being vetted and holding up to the same standards that you've already set for yourself or your team. The advent of AI doesn't mean that we can now just relax on the work that we've been doing for years. ## AI-ready data starts with leadership Right now, AI is very fractured at companies, with different levels of adoption in different areas. Every company wants to adopt AI, and I think most have moved a little too quickly. They give everyone some kind of AI tool and tell them to use it, but what's missing is actually working AI into their _way_ of working. You think you've got AI in engineering because everybody's using Claude Code or Codex, but that's just the top level. That's like saying everybody uses Google at this point. **Making your organization’s data AI-ready means building AI into your workflows.** Not just having your data engineers use AI tools, but also integrating them into your triage flow and your pull request flow and everything else your team does every day. The more a data leader can work AI into the daily workloads of an organization, the more that will actually drive adoption in a meaningful way. This takes buy-in from leadership to say, “When someone submits a pull request, we want to have an automatic agent review. When there's a pipeline failure, we want to be able to see an automatic triage done by the AI system.” For those of you in charge of actually building this, especially when you get those top-down “adopt AI now!” edicts from the C-level, VP level, people who may not be working with this technically — your challenge is communicating what it’s really going to be like. ### The things your org’s leaders need to understand about AI When it comes to fully integrating AI into your organization’s workflows, there are two important pieces of information that need to be communicated up the chain: First, this is extremely exciting technology. But, because AI is non-deterministic by nature, **making AI practical in the enterprise space means making it more deterministic. **You achieve this through evals, and with aligning skills and giving clear instructions and guardrails to make it incrementally more deterministic. As a result, leadership needs to understand that, "Hey, we can't implement this tomorrow." It's like any other major software implementation; actually, it’s even more complex because you're dealing with non-deterministic LLMs. Second, setting realistic expectations is key. Leadership needs to understand that **velocity is not going to 10x overnight.** Velocity will happen, but first the outputs need to be reviewed, the system needs to be trusted, evaluations need to be solid. Developers on the ground floor that are using these AI tools and seeing these outputs for things like error triage, code review, tickets coming in, and even text to SQL, will still need to review the outputs of the LLM. Realistically, you can get efficiency gains with AI, but in these early days, I would say a lot of the efficiency gains you'll get with AI are going to be at least partially eaten up by the amount of work you need to put into making AI reliable and trustable. ## Make AI organizational, not just individual Using AI tools personally is great. It's fantastic, it's super cool, you're spending all your Claude credits, that's awesome. But **if you really want to get the most out of your AI spend and the most out of your AI tooling, you need to work on productionizing it and building it into your daily tasks over time.** That's the muscle we need to be building if we want AI to‌ make sense in our organizations going forward, rather than just as a personal tool. --- --- title: "Data platforms were built to store. Intelligence platforms are built to reason." description: "Data platforms store information. Intelligence platforms govern meaning so AI can reason on it reliably." url: "https://www.getdbt.com/blog/data-platforms-were-built-to-store-intelligence-platforms-are-built-to-reason" date: "2026-07-02" authors: ["Dustin Dorsey"] categories: ["Insights"] --- # Data platforms were built to store. Intelligence platforms are built to reason. _This is a guest post from Dustin Dorsey, Senior Director of Data Engineering at [phData](https://www.phdata.io/) and co-author of Unlocking dbt. He works with enterprise teams to build the data foundations that make AI reliable at scale._ For a long time, the data platform was the destination. You invested in the infrastructure. You built the pipelines. You centralized the data, cleaned it up, made it available. You stood up a warehouse, connected a BI tool, trained your analysts, and delivered dashboards. And for a season, that felt like winning. Because compared to the chaos of disconnected data marts and spreadsheet-driven reporting, it genuinely was. But here is the thing nobody says out loud at the end of that journey: the platform was never the point. The point was better decisions. Faster decisions. Decisions made by people who actually understood what was happening in the business rather than guessing at it. The data platform was supposed to be the foundation for that. In most organizations, it became the ceiling. ## The problem was never infrastructure The data platform era solved the infrastructure problem. Data is more accessible, performant, and widely available than it has ever been. Cloud warehouses eliminated the constraints that used to make analytics painful. Modern ingestion tools reduced the friction of getting data in. Transformation frameworks made it easier to shape data into forms that could be queried and understood. By any technical measure, the infrastructure problem is mostly solved. What was never fully solved is the reasoning problem. The gap between having data and making good decisions from it was always bridged by people. Experienced analysts who understood the business. Data scientists who knew which model outputs to trust and which to question. Executives who had the pattern recognition to know when a number felt wrong. The platform gave these people better tools and faster access. It did not replace the judgment they were applying. For years, that was fine. The platform supported the humans who were doing the reasoning. That was enough because the humans were the bottleneck, and the platform reduced that bottleneck meaningfully. AI changes the equation not because it eliminates human judgment (it does not, and the organizations that believe otherwise are going to discover that the hard way) but because it now sits in the layer where that judgment was always applied. And the platform underneath was never built to support it. An AI system reasoning over your data is not a smarter dashboard. It is a different kind of consumer entirely. Dashboards were built to show humans pre-defined answers that humans then interpreted. AI is expected to form interpretations autonomously. To do that reliably, the platform needs to communicate business meaning, not just store data. And most platforms were never designed to do that. ## The category shift that is actually about philosophy There is a temptation to frame the evolution from data platform to intelligence platform as a technology upgrade. New tools, new capabilities, a few additional layers in the stack. That framing is wrong, and it leads organizations to make the wrong investments. The shift is not about technology. It is about purpose. A data platform is organized around the question: how do we make data available? Every architectural decision flows from that question. Where does data land? How fast can we query it? How do we manage access? How do we ensure it is clean and current? These are the right questions for a platform built to store and serve. An intelligence platform is organized around a different question: how do we make meaning available? The architectural decisions look different when you start from there. It is not enough for the data to be technically accessible. It needs to be interpretable by systems that have no institutional knowledge, no context, no ability to pick up the phone and ask someone what this number actually means. Making meaning available requires decisions that most data platforms deferred. What is the authoritative definition of this metric? What business process does this fact table represent? What does one row in this dataset represent? What relationships between entities are semantically intentional, not just technically possible? These are not new questions. They are questions that were always the right ones to ask. They just did not feel urgent when humans were mediating between the data and the decisions. They are urgent now. ## Overinvesting in intelligence, underinvesting in knowledge Look at where most organizations are putting their AI investment. A large share of it goes into the intelligence layer: the models, the inference infrastructure, the fine-tuning, the orchestration frameworks, the agent architectures. This is where the visible innovation is happening, and it is genuinely impressive. The layer that is chronically underinvested is the knowledge layer. The place where meaning lives. The semantic models, the metric definitions, the ontologies, the shared vocabulary that tells the intelligence layer how to interpret what the data layer contains. This imbalance is understandable. The intelligence layer has vendors, conferences, press releases, and benchmark comparisons. The knowledge layer has unglamorous conversations about what "customer" means and whether that definition should include trial accounts. Organizations gravitate toward work that feels like progress. Getting a model to generate a coherent summary is immediately satisfying. Spending three weeks aligning two teams on a metric definition is not. But the intelligence layer is only as good as the foundation it reasons over. An AI agent that can navigate complex workflows and synthesize information across domains will still produce unreliable outputs if the data it is reasoning over contains five competing definitions of the same concept. The capability of the reasoning layer is bounded by the quality of the knowledge layer beneath it. This is the central mistake of the current wave of AI investment in enterprise organizations: treating the intelligence layer as the place where reliability is established, rather than recognizing that reliability is established in the layers below it and the intelligence layer simply expresses it. ## Meaning has to become a product The organizational shift required here goes beyond architecture. It requires treating enterprise semantics the way mature organizations treat data products. With ownership, investment, stewardship, and a recognition that the value compounds over time. Most organizations do not currently have anyone who owns the definition of their core business metrics. They have people who maintain the pipelines that produce them, people who use them in dashboards, and people who argue about them in meetings. Ownership of the definition, the authoritative, governed, consistently enforced representation of what a metric means and how it should be calculated, is typically nobody's job. For an intelligence platform to function reliably, that has to change. Meaning needs an owner in the same way that a data product needs an owner. Someone who is accountable for whether the definition is current, whether it is consistently applied, whether changes to it are evaluated for downstream impact before they are made. Someone who treats the semantic layer as infrastructure worth maintaining rather than documentation worth archiving. This is not just an engineering problem. It is a product and governance problem. The organizations that recognize this early will build the organizational muscle to maintain their knowledge layer the same way they maintain their data infrastructure. The ones that treat it as a one-time modeling project will find themselves rebuilding it repeatedly as the business changes and the accumulated semantic drift undoes the previous round of work. ## What an intelligence platform actually looks like Calling something an intelligence platform is easy. The organizations doing it seriously are making a specific philosophical commitment, and it is worth naming clearly. A data platform asks: what do we have and where does it live? An intelligence platform asks: what does it mean and how should it be interpreted? That sounds like a semantic distinction until you see how differently the two platforms behave when an AI system is reasoning over them. The data platform produces a technically accessible answer. The intelligence platform produces the right one. The difference lives in a layer that most data platforms treated as an afterthought: the place where business meaning is governed. Not catalogued since catalogues tell you where data lives and what columns it contains. Governed. Where the definition of a metric is owned, maintained, and enforced in the transformation layer rather than documented in a wiki and applied inconsistently. Where the question "what does one row in this table represent" has an answer baked into the structure itself rather than living in an analyst's institutional memory. Organizations that try to build AI capabilities without that layer will keep encountering the same problem in different forms. They will tune the model and get better outputs for a week. They will refine the prompts and get consistency on the queries they thought to write prompts for. But the underlying ambiguity will surface somewhere else, because you cannot reason reliably over a foundation that was never designed to be reasoned over. The layer that makes AI reliable is not the AI. It is what the AI is built on top of. ## In five years, "data platform" will sound like a storage description Predictions about technology timelines are usually wrong in the specifics and right in the direction, so take this with appropriate skepticism. But the direction feels clear. The organizations building intelligence platforms today are not building something entirely new. They are building what data platforms were always supposed to be. The original promise of the data platform was that better data access would lead to better decisions. That promise was partially fulfilled. The AI era is forcing the other half to be taken seriously. The half about making meaning accessible, not just data. As that shift happens, "data platform" will increasingly feel like a description of what something stores rather than what it does. The platforms that persist and grow will be the ones organized around meaning and reasoning, not just storage and query. The investments that compound will be the ones made in the knowledge layer. In canonical definitions, in intentional data models, in the semantic infrastructure that makes it possible for both humans and AI systems to reason over enterprise data consistently. The organizations that are already making those investments do not look dramatically different from the outside today. Their AI demos are not necessarily more impressive than anyone else's. But when they move those demos into production, the outputs hold. The answers are consistent. The metrics mean the same thing to the finance team and the sales team and the AI system querying the data at 2 in the morning. That consistency is not a technical artifact. It is the result of someone having made deliberate choices about meaning and encoded them into the platform itself. That is what an intelligence platform actually is. Not a new set of tools layered on top of an existing data stack, but a fundamentally different philosophy about what the stack is for. And a recognition that the data engineering discipline has always been building toward this, even when the immediate deliverable was just a faster pipeline. If this series has prompted you to rethink what your own data foundation needs to look like for AI to operate reliably on top of it, [Building the Foundational Layer for Reliable AI on Structured Data ](https://www.phdata.io/offers/ai-data-foundation-whitepaper/)is the place to go deeper. It covers the structural conditions that separate AI environments that scale reliably from the ones that keep surfacing the same unresolved problems and why dimensional modeling is the prerequisite that most organizations are skipping. ## Where dbt and phData fits dbt sits at exactly the inflection point this shift requires. As organizations move from data platforms to intelligence platforms, the transformation layer becomes the place where meaning gets encoded and enforced, and dbt provides the model structure, semantic governance, and integration surface to make that possible. phData helps teams operationalize that shift by translating the philosophy into an executable data foundation: designing the knowledge layer, implementing the supporting models and semantics, and helping organizations move from ideas about reasoning to systems that support it in practice. --- --- title: "What Fivetran + dbt Labs brought to Databricks Data + AI Summit (and what you can take home)" description: "See what Fivetran + dbt Labs shared at Databricks Data + AI Summit: dbt Wizard, dbt State, and dbt Core v2.0, plus what's next." url: "https://www.getdbt.com/blog/what-fivetran-dbt-labs-brought-to-databricks-data-ai-summit-and-what-you-can-take-home" date: "2026-06-23" authors: ["Daniel Poppy"] categories: ["Partnerships"] --- # What Fivetran + dbt Labs brought to Databricks Data + AI Summit (and what you can take home) If agents are becoming a primary consumer of your data, what does that mean for your data work? It’s a lot to think about, and a lot to talk about, and ask questions about, which is exactly what we heard at the 2026 Databricks Data + AI Summit. Agents are only as reliable as the governed, tested, traceable data beneath them, which is why we showed up in San Francisco last week to show what’s important to us: data infrastructure for agents you trust. [We recently announced that dbt Labs and Fivetran are one company](https://www.getdbt.com/blog/fivetran-and-dbt-are-one-company-now-here-s-what-that-means). It makes Databricks Data + AI Summit all the more meaningful to us to show up with our shared mission to build Open Data Infrastructure for the agentic era. ## What we shared at Databricks Data + AI Summit ![photograph of dbt staff at Databricks Data + AI Summit](https://cdn.sanity.io/images/wl0ndo6t/main/00b2de5e789d2a58e99a256743abea59a46e45c2-4284x5712.jpg) We had a great time at Databricks Data + AI Summit, and if you were there, we hope you had a chance to chat with us. Based on the number of folks that attended our sessions, visited our booth, or attended our events,‌ a lot of you tried. And if you didn’t have a chance, we’d love to talk with you about how you’re approaching the data and agents landscape. [Talk to a dbt expert now](https://www.getdbt.com/contact). Kicking things off, Fivetran + dbt Labs CEO George Fraser and Fivetran + dbt Labs President Tristan Handy took the stage to share their vision of how AI is reshaping data infrastructure for the next decade. As agents become a primary consumer of data, they query the warehouse continuously, which raises the stakes on cost, observability, and above all, context. For George and Tristan, the answer to this is Open Data Infrastructure. The context your agents need lives in your data platform. For data teams, that means providing trusted, governed context for AI becomes a core part of the job. [To get direct access to Tristan and Fivetran + dbt Labs COO Taylor Brown, join our upcoming webinar and ask any questions you have about the new company.](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a) ## dbt and Fivetran customers sharing their journey ### The “BI why” problem Our customers provide an on-the-ground experience of what it’s like for data teams right now. Mars Dauer, Senior Director of Enterprise Data and AI at Integral Ad Science, spoke about what he calls the “BI why” problem: when a BI copilot can’t tell you why a number is wrong. It happens because the copilot only sees the top of the stack, and not the business logic underneath that produced the number. His team’s solution is an AI agent that connects Databricks and dbt through the model context protocol (MCP). The agent draws on three MCP servers, the dbt MCP server for model metadata, lineage, and compiled SQL, a Databricks SQL MCP server for read-only validation governed by Unity Catalog, and a Looker MCP server as the BI on-ramp. What used to take an analyst an afternoon, now takes minutes, Dauer says, and they are only scratching the surface. We recently announced [dbt Wizard, an agent built for analytics engineers](https://www.getdbt.com/product/dbt-wizard). It works alongside you with deep domain knowledge and a complete understanding of your project and workflows. [Join us for a deep dive to demo dbt Wizard in two surfaces: the CLI for engineers who live in the terminal, and the dbt platform for teams that want a shared interface](https://www.getdbt.com/resources/webinars/dbt-wizard-an-agent-purpose-built-for-analytics-engineering). ### Scale data and AI with Fivetran, dbt, and Databricks Inova Health Chief Data + AI Officer Jon McManus shared how the healthcare provider compressed a 4-year modernization roadmap into 6 months with Fivetran, dbt, and Databricks. [His team cut data movement costs by up to 8x with Fivetran](https://www.fivetran.com/case-studies/inova-health-compresses-4-year-roadmap-into-6-months-to-power-ai). dbt introduced a shared layer for transformation and [governance](https://www.fivetran.com/governance), which enables the team to define consistent metrics and improve trust in data across the enterprise. And Databricks provided a scalable platform to support advanced analytics and AI workloads. ## dbt on display at Databricks Data + AI Summit We’ve announced a lot recently, and it was pure joy giving folks their first look at everything in our booth and sessions and hands-on labs. With Fivetran and dbt Labs now as one company, we’re building the open data foundation for the agentic era so that your business logic and governance travel with you. We loved talking with folks new to dbt as they discovered the power of the dbt platform. And for current dbt users, we got to show off what’s new: [**dbt Wizard**](https://www.getdbt.com/product/dbt-wizard) is an AI agent purpose-built for analytics engineering to understand your lineage, tests, contracts, and metric definitions. [**dbt State**](https://www.getdbt.com/product/dbt-state) skips unchanged models across production and local runs for 30% average warehouse compute savings. And we announced the [**first alpha release of dbt Core v2.0**](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0&name=Fusion), now built on the same foundations as the dbt Fusion engine. dbt Core v2.0’s key feature developments include: - **Significant parse time improvements** - **A tightly-defined language spec** that gives anyone integrating with the dbt ecosystem a stable interface to build against. - **New Parquet artifacts** as a high-performance alternative to large JSON files that can be directly queried through your [agent of choice](https://www.getdbt.com/product/dbt-wizard). - **A completely revamped local documentation experience,** powered by those new artifacts and capable of scaling. - **A more streamlined way to build new adapters**, powered by Arrow Database Connectivity (ADBC) and the Arrow ecosystem. - **A simplified installation process** that removes the need to fight with Python's virtual environments. ## The dbt community brings the energy Whether you stopped by our booth, attended our talks, or joined our Happy Hour by the Bay or executive dinners, we’re grateful to connect with you. The dbt community is doing the most exciting work in data. If you want to see how the dbt community REALLY shows up, [**join us at dbt Summit Sept. 15-18 in Las Vegas**](https://www.getdbt.com/dbt-summit). dbt Summit is the world’s largest gathering of dbt users. Level up your data and AI work with dbt community members to shape the future of analytics engineering. Whether it’s at [dbt Summit](https://www.getdbt.com/dbt-summit), our [webinars](https://www.getdbt.com/resources/webinars), a [dbt meetup](https://www.getdbt.com/events), or [dbt Slack](https://www.getdbt.com/community/join-the-community), we’ll see you soon. ```json { "_key": "acdecba82c76", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/PcLd5bDRanQ" } ``` --- --- title: "The semantic debt crisis no one is talking about" description: "Two teams. Same metric. Different numbers. That's semantic debt, and AI will make it impossible to ignore." url: "https://www.getdbt.com/blog/the-semantic-debt-crisis-no-one-is-talking-about" date: "2026-06-22" authors: ["Dustin Dorsey"] categories: ["Insights"] --- # The semantic debt crisis no one is talking about _This is a guest post from Dustin Dorsey, Senior Director of Data Engineering at phData and co-author of Unlocking dbt. He works with enterprise teams to build the data foundations that make AI reliable at scale._ Picture the scene. Two teams are presenting to the same leadership group. Same company. Same time period. Same metric. Different numbers. Both teams defend their number with confidence. Both can show their work. Both are, by their own definition of the metric, correct. The meeting derails. Someone from finance says the sales team's number is wrong. The sales team says finance is using the wrong date field. Someone points out that marketing has a third version that does not match either of them. An hour passes. Nothing gets decided. The executive sponsor asks someone to "get aligned on the numbers" before the next meeting, and everyone leaves knowing that conversation will take months if it happens at all. This is not a data quality problem. The data is accurate. The pipelines ran correctly. The numbers are precisely what they claim to be. It is a design problem. Business meaning was never encoded into the data itself, so every team built their own version of the truth, and now there are three of them. Each defensible, none authoritative, all quietly costing the organization more than anyone has ever calculated. __ _Dustin goes deeper on the data design decisions that make AI reliable in production: [Your AI isn't broken. Your data model is.](https://www.getdbt.com/blog/your-ai-isn-t-broken-your-data-model-is)_ ## This is what semantic debt looks like Technical debt is a concept most engineers know well. You take a shortcut to ship faster, and you pay for it later in maintenance, refactoring, and fragility. Semantic debt works the same way, except the thing being deferred is not code quality. It is shared meaning. Every time a team builds a metric without encoding a canonical definition, that is a withdrawal from the semantic account. Every time a dashboard hardcodes a filter that belongs in the data layer, that is a withdrawal. Every time a business concept gets defined slightly differently in two separate pipelines, both of which work fine in isolation, that is a withdrawal. No individual shortcut looks catastrophic. The accumulation is. In most organizations, semantic debt has been building for years, sometimes decades. It was built quietly, because humans are remarkably good at compensating for it. A finance analyst knows their "revenue" is not the same as the sales team's "revenue" and calibrates accordingly. A data engineer knows which version of a customer record is authoritative for which use case. A business intelligence developer knows to add that one filter that nobody can explain but everyone agrees must be there. These workarounds are so ingrained they stop feeling like workarounds. They feel like just how things work. What they actually are is institutional knowledge plugging structural gaps that should have been designed out years ago. ## The meeting AI cannot call In a human-driven analytics environment, semantic debt gets resolved through a specific, recognizable process. Two teams disagree on a number. Someone calls a meeting. In the meeting, each team explains their logic. The group agrees (usually informally) on which definition is correct for which purpose. Someone writes a note in a Confluence page that nobody will find in two years. The organization moves on, with a slightly better shared understanding that lives, once again, in people rather than systems. This process is inefficient. It is slow. It only surfaces the conflicts visible enough to generate a meeting. But it more or less works, because the scale of human-driven analytics is limited enough that the gaps can be managed. AI cannot call that meeting. When an AI system is asked a question that depends on a metric your organization has defined inconsistently, it does not pause, identify the conflict, schedule a working session, and wait for consensus. It picks an interpretation and produces an answer. Confidently. Without flagging which version of the definition it used. Without knowing that a different version exists. Without any signal to the downstream consumer that the answer may depend on a definitional assumption that three teams would argue about if they knew it was being made. The conflict that your organization has been managing through meetings and tribal knowledge for years does not disappear when AI arrives. It gets made at machine speed, at scale, without anyone in the room to catch it. ## Where meaning lives in most organizations If you were to trace the authoritative definition of your organization's most important metrics, you would find them in some combination of the following places: a Confluence page last updated eighteen months ago, a comment in a SQL file that three engineers have each interpreted slightly differently, the institutional memory of a senior analyst who has been with the company for eight years, a Slack thread from 2022 that was pinned for a while and then wasn't, and a spreadsheet that lives on someone's desktop that nobody else knows about. This is not an exaggeration, and it is not a sign of dysfunction. It is the default state of organizations that built their data stacks to move fast and deliver value quickly within specific domains. Meaning accumulated in people because people were always there to apply it. Encoding it into the data structure itself was extra work that did not obviously improve the metric that mattered, which was whether the dashboard loaded. The problem with meaning living in people is not that people are unreliable. It is that meaning needs to travel to places people cannot go. It needs to travel to the AI system that a business user is querying at 11pm trying to understand why their numbers look different from last week. It needs to travel to the automated pipeline that is making decisions based on customer behavior without a human in the loop. It needs to travel to the new analyst who joined the company six months ago and has never heard the story of why that filter is always set to exclude refunds. It needs to travel to every downstream consumer of the data, reliably and consistently, every single time. People-stored meaning cannot do that. Only data-encoded meaning can. ## The stakes are higher than they appear Most organizations experience semantic debt as an inconvenience. Meetings that take longer than they should. Reports that require manual reconciliation before they can be shared. Analysts who spend a third of their time validating numbers before they trust them enough to use. These costs are real, but they are diffuse and largely invisible on any financial statement. AI changes the cost structure entirely. When AI is operating on data where meaning is inconsistently defined, the outputs are not just inconvenient, they are unreliable at a scale no human team could produce. An analyst who does not know which revenue definition to use asks a clarifying question. An AI that does not know which revenue definition to use picks one and answers three hundred queries with it before anyone notices the problem. The other shift is who is affected. Traditional analytics tools were used primarily by trained practitioners who understood the limitations of the data they worked with and knew, from experience, where the landmines were. AI-powered interfaces lower the barrier dramatically. Business users who have never seen the underlying data model, executives who expect answers to just be right, external stakeholders in some cases, all of them are now interacting with data whose meaning was never designed to be self-evident. Semantic debt that was manageable when experts were the only users becomes genuinely dangerous when anyone with a natural language prompt can query your data directly. ## You have a preview, not a problem Here is the argument I want to make directly: if your most important business metrics mean different things to different teams today, you do not have an AI problem yet. You have a preview of one. The meeting where two teams present different numbers? That is semantic debt becoming visible. It is uncomfortable and inefficient, but it is still being caught. Someone is in the room. The conflict surfaces. It gets resolved, however imperfectly. The question is whether your organization addresses the underlying cause intentionally (before AI amplifies it) or reactively, after AI has already made it impossible to ignore. Reactive is expensive. Not just in technical remediation time, but in organizational trust. The erosion of confidence in AI outputs is one of the hardest things to reverse once it begins. Business users who receive inconsistent answers stop trusting the system. Executives who got burned by a bad AI-generated number become skeptical of every AI-generated number. The technology gets blamed for a problem that the technology did not create, and the actual cause (the accumulated semantic debt in the data layer) stays unaddressed because the room full of people who saw the symptom did not get far enough upstream to see the disease. The organizations that act now have a significant advantage. Not because they will finish building perfect semantic foundations before anyone else because that is not how this works, and the organizations claiming they will are deceiving themselves. The advantage is that they will start building them as a deliberate strategic investment rather than as an emergency response to a crisis that has already damaged trust. ## Making meaning a system property, not a people property Think about what it would actually take for a new analyst to answer the question "what is our monthly active user count" correctly in their first week. They would need to know which table is authoritative, which definition of "active" the business uses, which date field governs the period, which user types to exclude, and probably the history of why each of those decisions was made. None of that is in the data. It lives in people. The new analyst calls a meeting. Someone explains it. They write it down in a place only they will ever look. Two years later, a different analyst asks the same question and starts over. This is the loop that data-encoded meaning breaks. When business logic is enforced in the structure itself (in how grain is declared, in how relationships are modeled, in what the transformation layer enforces) a new analyst and an AI system get the same correct answer without needing to know the history. Not because someone documented it well. Because the structure made the wrong answer impossible to produce. The distinction matters because meaning needs to travel. It needs to travel to the new analyst who joined six months ago, to the automated pipeline running without human supervision, to the AI system answering questions at 11pm. None of those consumers can call a meeting when they encounter ambiguity. They either get a consistent answer because the structure provides one, or they get a defensible guess because it does not. This is also not a one-time project. It is ongoing stewardship. Meaning drifts as the business changes. New products launch. Definitions evolve. Mergers happen. The organizations that stay ahead of semantic debt are not the ones that did a modeling exercise in 2022 and called it done. They are the ones that treat their semantic foundation the same way they treat their data pipeline. As infrastructure that requires ownership, investment, and deliberate evolution. What most organizations are missing is not the technical ability to do this. It is the organizational will to treat it as infrastructure rather than overhead. That shift (from viewing semantic work as a documentation task to viewing it as a platform investment) is what separates the organizations that will navigate the AI era well from the ones that will spend the next several years explaining why the answers keep changing. ## What you should be asking right now The diagnostic for semantic debt is not a technical audit. It is a business conversation. Pick your organization's three most important metrics. Ask five people from different teams to define them. Not to query them, but to define them. To explain what counts, what does not, what edge cases apply, and how they should be calculated. If you get the same answer from all five people, you are in better shape than most. If you get five different answers, each defensible, you are looking at your AI risk surface. The follow-up question is where the decision lives: do you want to address this before the AI initiatives your organization is investing in expose it at scale, or after? There is no neutral choice here. Semantic debt compounds. Every month it goes unaddressed is another month of diverging definitions, additional teams building their own versions of the truth, and growing distance between where your data is and where it needs to be for AI to operate reliably on top of it. Addressing semantic debt is not just about fixing individual models, though. It requires rethinking the fundamental purpose of the data platform itself. Most organizations are building platforms designed to store. What the AI era requires is platforms designed to reason. That is the argument in the next post and it is a bigger shift than it sounds. If the patterns in this post feel familiar, [Building the Foundational Layer for Reliable AI on Structured Data](https://www.phdata.io/offers/ai-data-foundation-whitepaper/) goes deeper on the structural conditions that separate data environments where AI operates reliably from the ones where it keeps surfacing the same unresolved ambiguity. Worth reading before your next AI planning conversation. ## Where dbt and phData fit dbt is one of the few tools in the modern data stack where semantic debt can be paid down incrementally and sustainably. Its combination of model-layer definitions, the dbt Semantic Layer for centralized metric governance, and documentation gives teams a practical path to encoding meaning as part of how transformations are written and maintained. phData helps teams put that into motion by working through the harder operational pieces alongside dbt: clarifying definitions, reshaping models, and building the delivery discipline needed to make semantic consistency hold up beyond a single project or team. --- --- title: "Start fresh, don't lift and shift: a dbt migration guide" description: "dbt migrations underdeliver when teams rebuild legacy patterns in a new tool. Here's how to start fresh instead." url: "https://www.getdbt.com/blog/start-fresh-don-t-lift-and-shift-a-dbt-migration-guide" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # Start fresh, don't lift and shift: a dbt migration guide We've seen this pattern often enough to name it. A team migrates to dbt, spends six months, and ends up with a dbt project that looks exactly like their old workflow, just with Jinja templating instead of drag-and-drop. The data model still has the same problems. It's only newer. The migration project gets checked off as complete. The downstream problems show up six months later: reports that contradict each other, engineers who can't explain what a model does without reading the SQL, a semantic layer that makes data trust worse instead of better. This is the lift-and-shift problem, and dbt isn't the cause of it. ## Signs your dbt project is a legacy migration The symptoms are recognizable once you know what to look for. Here are the six most common. **Every stored procedure has a 1:1 dbt model equivalent.** The migration team mapped source code to dbt models one-for-one. The logic lives in SQL now instead of PL/SQL, but the structure is identical, and no one asked whether the original structure was worth keeping. **Models are named after source systems, not business entities.** You see `salesforce_accounts_cleaned` and `hubspot_contacts_deduped` instead of `customers` and `leads`. The names describe where data came from rather than what it means to the business. **No staging models exist.** Everything jumps directly from raw source to reporting table. The [dbt project structure guide](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview) is explicit: [staging models](https://docs.getdbt.com/best-practices/how-we-structure/2-staging) are the atomic building blocks of a dbt project, and skipping them skips the part that makes the architecture maintainable. **Tests are bolt-ons rather than design principles.** The migration got done, then someone added [`not_null` tests](https://docs.getdbt.com/docs/build/data-tests) to the most important columns as an afterthought. **Documentation is empty or copied from source-system field descriptions.** Column descriptions say things like "from MDM system." That records provenance, which is useful, but it leaves the actual documentation undone. **No `ref()` lineage.** Models depend on hardcoded schema names instead of [`ref()`](https://docs.getdbt.com/reference/dbt-jinja-functions/ref). When something breaks, teams find out from a failed report rather than a CI check. None of these is hard to fix individually. The problem is what they signal together: the team migrated code, and left the thinking behind. ## The rewrite feels risky and the migration feels safe. Look closer. Here's how teams end up here. Migration projects get measured by completion, not quality. There's a deadline, there's a checklist, and the checklist says "migrate 200 stored procedures." Whether those procedures represent good data modeling never makes the checklist. The path of least resistance is to replicate what exists. Translate the stored procedures, get to green on the migration tracker, and ship. Quality gets deferred, and deferred quality hardens into permanent technical debt. A recent consulting engagement shows the pattern. An insurance company undertook a major migration from a legacy ETL platform to dbt. The migration ran on schedule and on budget. A year later, the data team was spending more time debugging model failures than shipping new analytics. The architecture had carried over all the original fragility, in a different tool. The symptoms: 300-plus models with no staging layer, no tests on roughly 40% of models, reports built directly on raw source joins, and a semantic layer that three different teams maintained independently. The migration was complete. The architecture was broken. Here's the framing that helps. A lift-and-shift migration is really a replatform: the tool changed, and the system stayed the same. If the data model was wrong before dbt, it's still wrong after. ## The triage decision: which legacy models deserve to exist Before writing a line of dbt SQL, teams should ask a harder question than "how do we migrate this?" The better question: "Does this deserve to exist?" Here's a practical triage framework. **Eliminate.** Reporting tables built for a deprecated BI tool, summary tables that exist only because the old database couldn't handle the underlying query, and "just in case" tables nobody queries. Leave these behind. **Rewrite.** Any logic that joins raw source tables with no staging layer. Any model that does transformations and aggregations in the same step. Any model named after source systems rather than business entities. Redesign these rather than translating them. **Translate with care.** Validated, tested business logic the organization depends on. Bring it over with explicit [tests](https://docs.getdbt.com/docs/build/data-tests), document every column, and have someone who understands the business domain review the logic, not just the SQL. **Build fresh.** The [semantic layer](https://www.getdbt.com/product/semantic-layer). This should always be built from business requirements: what is a customer, how do we define revenue? Migrating legacy SQL into the semantic layer inherits every ambiguity and compromise baked into the original definitions. The migration health conversation with stakeholders usually centers on timeline and scope. The more useful conversation is about triage: what are we keeping, what are we rebuilding, and what are we cutting? ## Seven signs of a healthy dbt migration Use this as a diagnostic. If you're mid-migration, run through it this week. 1. **Every source model has a [staging model](https://docs.getdbt.com/best-practices/how-we-structure/2-staging).** The staging model cleans and standardizes the data: type casting, column renaming, basic validation. Nothing downstream touches raw source tables directly. 2. **Business entities are represented as marts.** Final models are named for business concepts: a `customers` mart, an `orders` mart, a `revenue` mart. They represent what the business cares about, not what the source systems happen to contain. 3. **Every model has at least [`not_null` and `unique` tests](https://docs.getdbt.com/docs/build/data-tests) on primary keys.** These are the minimum, and they catch the most common failure modes: duplicate rows and unexpected nulls. Without them, you have data hope, not data quality. 4. **Documentation coverage is tracked and improving.** Not every column needs a long description, but every model should have one that tells a new team member what it is and why it exists. 5. **[`ref()`](https://docs.getdbt.com/reference/dbt-jinja-functions/ref) is used everywhere.** No hardcoded schema names. If the project can't run in a fresh schema without breaking, it isn't production-ready. 6. **CI runs on every PR.** Every pull request runs tests before merge, so nobody merges broken SQL. 7. **A team member who didn't write a model can explain what it does from the documentation.** If someone can't understand a model without reading the underlying SQL, the documentation has failed, and eventually the model will too. ## Why migration quality matters even more in 2026 The pressure to migrate from legacy systems to dbt is real, and so is the pressure to do it fast. In 2026, a third pressure has arrived: the AI systems being built on top of data infrastructure are only as good as that infrastructure. A lift-and-shift migration produces exactly the kind of foundation that makes AI unreliable. Models without semantics, tests, or documentation mean AI systems inherit every gap. The semantic layer doesn't know what "customer" means. Tests don't exist to catch when something breaks. Documentation can't help a model trace back to its source. The data shows how wide the gap is. Gartner [predicted in early 2025](https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk) that through 2026, organizations would abandon 60% of AI projects built on data that isn't AI-ready. And according to a March 2026 report from Cloudera and Harvard Business Review Analytic Services, only 7% of enterprises say their data is completely ready for AI, while 73% say their organization should prioritize AI data quality more than it currently does. The teams building reliable AI systems are doing it on data infrastructure they actually trust. That starts with the migration, well before the AI project. ## Closing Two practical moves from here. If you're mid-migration right now, run the seven-sign checklist as a diagnostic this week. Don't wait until the migration is "done." The checklist will tell you how far the project has drifted from good architecture and how much rework is piling up. If you're planning a migration, the triage framework is the architecture conversation to have before writing a line of code. Get the team in a room, look at the models in scope, and answer one question: does this deserve to exist in the new system? --- --- title: "The analytics engineer in 2026: system designer, governance owner, AI context provider" description: "AI is reshaping the analytics engineer role. Here's what system design, governance, and AI context look like in 2026." url: "https://www.getdbt.com/blog/the-analytics-engineer-in-2026-system-designer-governance-owner-ai-context-provider" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # The analytics engineer in 2026: system designer, governance owner, AI context provider ## What analytics engineering looked like in 2023 In 2023, the core of an analytics engineer's job was model development. You wrote SQL, organized it into dbt models, wrote tests, and built pipelines that turned raw data into something stakeholders could use. Documentation was a best practice you aspired to. [Column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage) was a nice-to-have. The bottleneck was your capacity to write and review code. In 2026, that bottleneck has eased. AI can write dbt model scaffolding faster than any human. It can generate first-draft documentation from lineage metadata. It can write the boilerplate tests most models need. AI-assisted coding is now part of how most analytics engineers work: 72% of them, according to the [2026 State of Analytics Engineering report](https://www.getdbt.com/resources/state-of-analytics-engineering-2026). The repetitive parts of model production are increasingly automated. That clarifies the role rather than shrinking it. With the repetitive work no longer the bottleneck, what's left is the work analytics engineers were always most valuable for. ## The three new responsibilities of the analytics engineer in 2026 **System design** The analytics engineer in 2026 focuses less on individual model implementation and more on how the system of models works. Which models are the source of truth for which metrics? Where are the boundaries between domains? How should the [semantic layer](https://www.getdbt.com/product/semantic-layer) be structured so downstream AI queries return consistent answers? These are architecture decisions that require business judgment and an understanding of how the data gets used, not just how it gets built. AI can scaffold a model. It can't decide whether revenue should be defined at the order line level or the order level, or which grain is correct for a retention metric. That judgment requires understanding the business, which remains a human capability. (For a real-world look at the tradeoffs, see [who should own the semantic layer](https://www.getdbt.com/blog/semantic-layer-ownership).) **Governance ownership** As AI-assisted development accelerates data production, the governance layer becomes more important. Tests, [contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts), column-level lineage, and ownership assignment are now the outputs that separate a trustworthy data system from a fast but unreliable one. The analytics engineer owns those outputs. This changes how the role gets evaluated. In 2023, an analytics engineer's output was models. In 2026, it's also the [contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) that protect those models, the [tests](https://docs.getdbt.com/docs/build/data-tests) that validate them, and the semantic definitions that make them machine-readable. Governance has become a primary deliverable. (More on this in [semantic layer for data governance and security](https://www.getdbt.com/blog/semantic-layer-data-governance-security).) **AI context provision** This responsibility has emerged most visibly in the past eighteen months, and it's the one analytics engineers are often not trained for explicitly. AI agents need context to reason reliably, and [that context has to come from somewhere](https://www.getdbt.com/blog/how-a-semantic-layer-prevents-ai-hallucinations-in-analytics). In a well-structured dbt project, it comes from [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metric definitions, column-level lineage, model documentation, and schema contracts. The analytics engineer who understands how to structure that context, what to name things, how to document them, which definitions to make machine-readable, is directly improving the reliability of every AI agent that runs against the data stack. The data community has spent a lot of 2026 debating "context" as a buzzword, and the joke is fair. The underlying problem it names is real: organizations are failing at AI not because their models are wrong, but because AI agents lack the context to reason about what the data means. Analytics engineers build that context. That's a significant expansion of the role's leverage. ## Analytics engineer skills that matter in 2026 SQL fluency still matters, and it isn't going away. But the premium on raw SQL productivity is lower than it was two years ago, because AI can produce syntactically correct SQL faster than most humans. Three things have gotten more valuable: business judgment, semantic precision, and system thinking. Business judgment means understanding what the data represents well enough to know when an AI-generated model is wrong, even when it looks syntactically correct. It means knowing that a metric definition that works for one use case will mislead in another. That judgment isn't automatable. Semantic precision means writing metric definitions and documentation precise enough to be unambiguous, both to a human reading them and to a model reasoning over them. This is a new skill, and the analytics engineers who develop it are more valuable in an AI-native data stack. System thinking means understanding how models relate, where the dependencies are, and how architectural decisions propagate through the stack. As AI takes over individual model implementation, the analytics engineer's comparative advantage lies in the decisions that span models. ## The career case for analytics engineers in 2026 Analytics engineers are more valuable in 2026 than they were in 2024, and AI is the reason why. In 2024, some analytics engineering work created value and some was maintenance. AI is eliminating the maintenance, and what remains is disproportionately the value-creating work: semantic definition, governance ownership, architecture decisions, and the business judgment that determines whether fast data is also accurate data. An analytics engineer who spends 2026 competing with AI on code production will find the role shrinking. One who focuses on what AI can't do, business judgment, context design, and governance ownership, will find it expanding. Treat that as a specific description of what to build toward. ## How dbt supports the evolving analytics engineer role dbt is well-positioned for this shift because of what it has always stored: semantic context in code. The metric definitions, tests, contracts, and documentation in a dbt project are exactly the context AI agents need to reason about data reliably. The analytics engineer who maintains that context is the person making AI-assisted data work trustworthy. The tooling supports this directly. [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) gives analytics engineers an AI agent grounded in their project's lineage, contracts, tests, and metrics. The [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) makes the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) queryable in natural language. [Column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage) gives AI agents the provenance they need to cite their answers. None of that replaces the analytics engineer. It amplifies what the analytics engineer already does: define what data means, govern how it's used, and build the context AI depends on. That work is the most valuable in the data stack right now, and the question is whether analytics engineers recognize it as such. --- --- title: "How dbt makes agentic data pipelines trustworthy: the transformation layer's role in autonomous data systems" description: "AI agents can run your pipelines. But who decides what \"correct\" looks like? The transformation layer does, and that layer is dbt." url: "https://www.getdbt.com/blog/how-dbt-makes-agentic-data-pipelines-trustworthy-the-transformation-layer-s-role-in-autonomous" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # How dbt makes agentic data pipelines trustworthy: the transformation layer's role in autonomous data systems ## What's missing from every agentic data pipeline diagram Agentic, self-healing pipelines are having a moment. Dagster has published its [AI-driven data engineering vision](https://dagster.io/blog/announcing-ai-driven-data-engineering). The Airflow community is discussing [agentic workloads in Airflow 3](https://airflow.apache.org/blog/agentic-workloads-airflow-3/). Datafold's [2026 predictions](https://www.datafold.com/blog/data-engineering-in-2026-predictions/) put autonomous data engineering on the near-term roadmap. The architecture diagrams all show the same thing: sources feeding agents that run tasks that produce results. None of them shows the layer that determines whether those results are correct. That layer is the transformation layer. And the question nobody seems to be asking yet: when an AI agent builds and runs a data pipeline autonomously, who defines what correct looks like? Without an answer, a self-healing pipeline is just a fast pipeline. It heals quickly, and it propagates wrong answers quickly. Speed isn't the value here. Correctness is, and correctness requires a governed transformation layer. ## What the transformation layer does in an autonomous system In a human-operated pipeline, the transformation layer is where raw data gets shaped into something meaningful. Models define how tables relate. Tests assert that specific conditions hold. Contracts enforce that the shape of data at a boundary can't change without an explicit decision. In an agent-operated pipeline, all of that still has to happen. The difference is that the agent making the changes doesn't inherently know your business rules. It knows syntax. It knows patterns from training data. It doesn't know that your revenue metric must exclude refunds, or that a customer is only active if they've logged in within 30 days, or that a null in this column means something different than a null in that one. That knowledge has to be encoded somewhere. In dbt, it lives in models, tests, [contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts), and [semantic layer](https://www.getdbt.com/product/semantic-layer) metric definitions. Encoding it is the core of what dbt does. (For more on this, see [what agentic AI requires from your data](https://www.getdbt.com/blog/agentic-ai-data-requirements).) When an AI agent operates a pipeline with a governed dbt transformation layer, it isn't making autonomous decisions about what the data means. It executes transformations whose semantics were defined by humans, validated by tests, and protected by contracts. The agent gets speed. The business gets correctness. That's the value of governed agentic workflows. ## Why governance has to come before autonomy Teams that skip the governance step and connect AI agents directly to their transformation layer are automating the wrong thing. An AI agent running against a dbt project with no [contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) can change a column type in a core model and break every downstream metric silently. An agent running against a project with no [semantic layer](https://www.getdbt.com/product/semantic-layer) definitions interprets metric names however seems reasonable from the table structure. Sometimes it's right. Often it's confidently wrong in ways that are hard to debug. The pipeline self-heals. The numbers are still wrong. The CFO doesn't care that the pipeline ran without errors. Governance at the transformation layer is the prerequisite for AI autonomy, not a constraint on it. An agent that can trust the semantic definitions it works with operates faster and with more autonomy, because the boundaries are what make autonomy safe. ## How data contracts and live project context create a trusted layer dbt gives agentic workflows two things that matter most: trusted boundaries and current context. [Model contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) define the shape of data at a model boundary: which columns must exist, what types they must be, which constraints must hold. When a contract is defined, breaking changes are caught at compile time, before they run. An AI agent that tries to remove a column a downstream contract depends on produces a compilation error, not a silent pipeline failure. Context matters just as much. An agent working from stale manifests reasons about a project that may be hours out of date. dbt's metadata layer and [column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage) in dbt Explorer give agents accurate column types, current dependencies, and real lineage, so an agent writing code is working from the actual state of the project. Together, these give an agent a context it can trust. It knows what the columns mean. It knows where the boundaries are. It knows that crossing one without an explicit decision means the pipeline won't run. That structure is what makes real autonomy possible, rather than uncontrolled automation. ## MetricFlow: defining what "correct" means for AI agents Contracts protect structure. The semantic layer defines meaning. [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metric definitions encode the business logic that turns data into answers. Revenue isn't just a `SUM(amount)` column. It's a `SUM(amount)` with specific filters, from specific models, under specific conditions, for specific purposes. That definition, written once in MetricFlow and version-controlled in the project, is the canonical answer to "what is revenue?" When an AI agent queries the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) through the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp), it isn't making up an answer. It queries a definition a human wrote, reviewed, and committed. The answer holds whether the question comes from a BI tool, a Slack bot, or an autonomous agent running a scheduled pipeline. That's what "correct" looks like in an autonomous data system: a model that returns the answer your CFO would agree with, from a definition that's version-controlled, tested, and auditable. ## What this means for analytics engineers In an agentic world, the analytics engineer's job shifts from competing with AI on code production to defining the semantic layer that AI agents need to be trustworthy. That means metric definitions, contracts, governance ownership, and lineage documentation. It means making the implicit knowledge about what data means explicit enough that a model can reason about it reliably. It means being the person who decides what correct looks like before the agents start running. This is the higher-leverage version of the job. The definitions you write are used by every AI agent that runs against your data stack, not just your team. The governance you put in place is the foundation that makes autonomous pipelines trustworthy. Every architecture diagram circulating today shows AI agents running data pipelines. The ones that work in production will have a transformation layer in the middle, owned by analytics engineers, that defines what the data means and enforces what correct looks like. That layer is dbt. The people who build it matter more now, not less. --- --- title: "Context engineering is the new analytics engineering skill: a practical guide for dbt users" description: "Analytics engineers are already doing context engineering. Here's how your dbt project becomes context for AI." url: "https://www.getdbt.com/blog/context-engineering-is-the-new-analytics-engineering-skill-a-practical-guide-for-dbt-users" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # Context engineering is the new analytics engineering skill: a practical guide for dbt users ## What is context engineering? When the idea that 2026 is "the year of context" started making the rounds, Joe Reis answered with a [send-up of context lakes, context products, and the analyst singularity](https://joereis.substack.com/p/gartner-declares-2026-the-year-of). Underneath the satire is a real problem: most organizations are bad at giving AI systems the context they need to reason reliably. Context engineering is the practice of structuring information so that AI models and agents can use it accurately. It works at a different level than prompt engineering, which operates on instructions, and it's broader than retrieval-augmented generation, which is one specific implementation. Context engineering decides what information an AI system gets, in what format, and at what level of specificity, so it can reason about a domain without hallucinating or guessing. For AI systems working with data, that discipline has a specific shape. The AI needs to know: what does this metric mean? What is the grain of this table? What does a null value in this column represent? What business rules make this metric different from that one? Those are questions about how knowledge is encoded in the data layer. And analytics engineers have been encoding that knowledge in dbt projects for years. ## Why analytics engineers already have the advantage Here's the insight this piece is built around: analytics engineers are already doing context engineering. They just don't call it that. When you write a [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metric definition, you're encoding context. You're making explicit the business logic that turns a column in a table into an answer a stakeholder can trust. When you write a model description that explains why a model exists and what it represents, you're encoding context. When you define a [schema contract](https://docs.getdbt.com/docs/mesh/govern/model-contracts), you're encoding context about what can and cannot change at this model's interface. The difference between an AI agent that reasons reliably about your data and one that hallucinates is almost entirely a function of how well-structured that context is. Agents that query raw tables with no semantic definitions make reasonable guesses and are frequently wrong. Agents that query a governed [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) return answers that are consistent and defensible. The skill analytics engineers already have, making implicit business knowledge explicit, structured, and machine-readable, is exactly the skill that makes AI systems trustworthy. Context engineering is analytics engineering pointed at a new consumer. ## How dbt structures context for AI agents A well-structured dbt project provides context to AI systems through four mechanisms. **Model descriptions.** A description field in a model's YAML does more than document for humans. When an AI agent reads the lineage graph through the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp), model descriptions are part of the context it receives. A sparse or absent description means the agent has to guess what the model does. A precise one means it doesn't. **MetricFlow metric definitions.** These are the most powerful context mechanism in the dbt stack. A [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) definition encodes not just what to compute, but how: the entity, the measure, the filters, the grain. When an AI agent queries the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), it isn't writing SQL. It's querying against definitions that already encode the business rules. The accuracy gain is large: in an [early dbt experiment](https://www.getdbt.com/blog/semantic-layer-as-the-data-interface-for-llms), LLMs answered natural language questions correctly about 83% of the time when grounded in governed MetricFlow definitions, far above raw text-to-SQL on real-world schemas. **Schema contracts.** [Contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) define the interface at a model boundary. They're context about what can be relied upon: which columns will always exist, what types they'll be, what constraints hold. For AI agents writing code that depends on downstream models, contracts prevent a class of errors that would otherwise need human debugging. **Column-level lineage.** Lineage tells an AI agent where data comes from. When an agent generates documentation or debugs an anomaly, [column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage) in dbt Explorer is the context that lets it reason about causality rather than just correlation. ## Three context engineering patterns to build this week Three patterns analytics engineers can implement using dbt this week: **Pattern 1: Metric-as-contract documentation** For each of your core business metrics, write a [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) definition that includes the measure plus explicit documentation of what the metric excludes, what the grain means, and what edge cases apply. This is the work that keeps your AI analytics tool from returning a plausible-but-wrong answer to a CFO query. Estimated time: 2-4 hours for 5-10 core metrics. The test: can a new team member, or an AI agent, read the definition and understand not just how to compute the metric but why it's defined this way? If not, the definition isn't doing its job as context. **Pattern 2: Model-level context headers** Add a structured description block to your 10 most important dbt models. Include what the model is (the grain, the purpose), what it isn't (common misuses or misinterpretations), and what changes downstream if its definition changes. This is the context an AI agent needs to reason about the model correctly without running the data. Keep the description complete in 3-5 sentences. Shorter is better if it's precise. The goal is to write everything an AI system needs to reason about the model accurately, no more. **Pattern 3: Freshness and quality metadata as context signals** Configure [source freshness](https://docs.getdbt.com/docs/deploy/source-freshness) checks and make the freshness metadata accessible. An AI agent that knows a source was last loaded 4 hours ago can make a better decision about whether to run a downstream pipeline. An agent with no freshness context runs regardless. Freshness metadata is context engineering at the pipeline-operations level: it gives autonomous systems the information they need to decide when to run. ## The vocabulary shift that matters Context engineering is a better name for work analytics engineers have always done: making implicit business knowledge explicit, structured, and reliable. The consumers of that work have changed. The work itself has held steady. What's changed is the leverage. In 2023, a well-documented dbt project made a human analyst's job easier. In 2026, a well-documented dbt project makes every AI agent that runs against your data stack more accurate. The context an analytics engineer builds is used by a system. That's a meaningful expansion of impact. The year of context, whatever the analysts mean by it, is really the year analytics engineers discover that the work they've always done is worth more than they realized. --- --- title: "Building the agentic data stack: A practical dbt guide for the AI era" description: "AI builds data infrastructure fast. Here's how to make your dbt project ready to support agents without it falling over." url: "https://www.getdbt.com/blog/building-the-agentic-data-stack-a-practical-dbt-guide-for-the-ai-era" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # Building the agentic data stack: A practical dbt guide for the AI era ## The AI readiness gap is bigger than it looks Ali Ghodsi [said earlier this year](https://www.cnbc.com/2026/02/09/under-the-hood-of-the-ai-economy-with-databricks-ceo-ali-ghodsi.html) that over 80% of databases on Databricks' platform are now built by AI agents. AI-driven data engineering has arrived. The headline left out the harder truth. Most of those databases run on foundations built for human-operated workflows, not for agents. Only 7% of enterprises say their data is completely ready for AI, according to the [2026 Data Readiness Index from Cloudera and Harvard Business Review Analytic Services](https://www.cloudera.com/about/news-and-blogs/press-releases/2026-03-05-only-7-percent-of-enterprises-say-their-data-is-completely-ready-for-ai-according-to-new-report-from-cloudera-and-harvard-business-review-analytic-services-reveals.html). The tests, contracts, semantic definitions, and lineage that make AI-generated data trustworthy are still the exception. Both facts describe the same gap. AI builds fast, and without a governed foundation underneath it, fast can also mean unreliable. Accuracy is where it shows: agents that look strong in a pilot slip the moment they meet real users and real edge cases. This gap is exactly what the June 1 launches set out to close. Fivetran and dbt Labs are [one company now](https://www.getdbt.com/blog/what-we-announced-at-snowflake-summit-and-why-it-matters), focused on open data infrastructure for trusted agents. dbt State, dbt Wizard, and dbt Core v2.0 all point at the same goal: a data foundation that's open, portable, and ready for what agents need. That leaves you with one practical question. What does your dbt project need to support AI agents safely? Four things, working together. ## Trust foundation: tests and contracts Start here. [Model contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) enforce schema constraints at the model level, so breaking changes get caught before they reach production. [Data tests](https://docs.getdbt.com/docs/build/data-tests) validate the data itself. For AI-generated models, both matter more. An agent does not know which downstream model breaks when a column type changes. Contracts and tests catch what the agent never thought to check. ## Query reliability: the dbt Semantic Layer Without a semantic layer, agents write raw SQL against undecorated tables, and accuracy drops sharply. In an [early dbt experiment](https://www.getdbt.com/blog/semantic-layer-as-the-data-interface-for-llms), LLMs answered natural language questions correctly about 83% of the time when grounded in governed [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) definitions, far above the accuracy of raw text-to-SQL on real-world schemas. The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) turns a natural language question into a consistent, defensible answer. Without it, your agent guesses what your metrics mean, and it guesses with confidence. ## The agent built for the work: dbt Wizard General coding agents write plausible code, then hand you an hour of verification because they do not know your project. [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) is built for the way analytics engineers actually work: investigating, building, validating, and shipping. It is grounded in your dbt project, so it knows your lineage, contracts, tests, and metric definitions before it writes a line. It knows which tool to call, validates its own work, and shows you what changed and why. You get it inside the dbt platform and as a terminal-native CLI, whether your team runs on the dbt platform or self-hosted. Under the hood, dbt Wizard draws on [dbt Agent Skills](https://docs.getdbt.com/blog/dbt-agent-skills), the open-source skills maintained by dbt Labs and the dbt community. ## Efficiency at AI speed: dbt State Point an agent at your pipelines and it will run them at machine speed. Rebuild everything on every run and your compute bill climbs fast. [dbt State](https://www.getdbt.com/product/dbt-state) checks your metadata and model SQL on each run, builds what's changed, and skips what hasn't, for an average of 30% reduction in warehouse compute. It runs in the dbt platform and in the orchestrator you already use. Build what's changed, skip what hasn't. For agentic workflows, that is the difference between automation that scales and a compute bill that punishes you for using it. ## Audit your own project Before you invest in new tooling, run this audit on the project you have. Each question maps to one of the four pillars. **Trust (contracts and tests)** - Are [model contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) enforced on your most important models? If your revenue and customer models have no contract, an AI-generated change upstream can break them quietly. - Do your critical models have [data tests](https://docs.getdbt.com/docs/build/data-tests), and do breaking changes fail in CI? If a schema change passes CI without failing, your governance is misconfigured. **Query reliability (Semantic Layer)** - Are your [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metrics defined in the project? If your core metrics live only in a BI tool, agents cannot query them in a governed way. - Can an external system reach your semantic layer through the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp)? If not, natural language queries are hitting raw tables. **The agent (dbt Wizard)** - Is your team using a dbt-native agent like [dbt Wizard](https://www.getdbt.com/product/dbt-wizard), or general coding agents that do not understand your project? Generic agents do not respect your contracts or see what breaks three models downstream. - Do agents reach your project through a governed, auditable interface, dbt Wizard or the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp), rather than raw warehouse credentials? **Efficiency (dbt State)** - Does every run rebuild everything, including models that have not changed? [dbt State](https://www.getdbt.com/product/dbt-state) skips the work that hasn't changed. - Before you connect agents that trigger runs, do you have cost controls in place? Machine speed and a per-query bill are a costly surprise to discover later. ## What an agentic dbt workflow looks like in practice A prompt or an automated trigger starts a task: write a new model, optimize an existing one, generate documentation. [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) reads the lineage graph, works out what the affected models do, and drafts the change, validating its own work against your [contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) and [tests](https://docs.getdbt.com/docs/build/data-tests) before you see a diff. The PR runs through CI with contract checks. If it breaks a downstream contract, CI fails, and the agent either fixes the issue or routes it to a human. When the pipeline runs, [dbt State](https://www.getdbt.com/product/dbt-state) builds only what changed, so the run stays cheap. Merged work feeds updated [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metrics and documentation. The next time an analyst asks a question in plain language, the [semantic layer](https://www.getdbt.com/product/semantic-layer) answers from the new model and cites its lineage. Every step depends on at least one pillar. Remove one and the workflow breaks in a predictable place. ## Common failure modes in agentic dbt workflows Teams that build agentic workflows without these foundations hit the same failure modes repeatedly. **AI-generated models with no tests.** The model runs. The numbers look plausible. Three months later a column gets renamed upstream, the model starts returning nulls, and nobody notices until finance flags a wrong number. [Data tests](https://docs.getdbt.com/docs/build/data-tests) catch this. Untested AI-generated models are technical debt at AI speed. **Plain-language queries against raw tables.** An agent queries a table called `orders`. It cannot tell placed orders from fulfilled or cancelled ones, so it guesses, and the answer is wrong. A governed [semantic layer](https://www.getdbt.com/product/semantic-layer) removes the entire category of failure. **Agents with direct warehouse access.** An agent holding database credentials is a liability. [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) and the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) give agents scoped, auditable, governed access to your project. Use them. **Rebuilding everything on every run.** Connect an agent that triggers pipelines and your compute bill climbs while you are looking elsewhere. [dbt State](https://www.getdbt.com/product/dbt-state) pays only for new work. **Treating AI-generated documentation as final.** AI-generated documentation is fast and usually close, and it is not authoritative. Build a human review step into your workflow before AI-generated descriptions publish. ## Build trust before you build automation Sequence matters. Contracts and tests first, the trust foundation. Then the semantic layer, for query reliability. Then the agent and the cost controls, dbt Wizard and dbt State. Teams that skip the foundation and jump straight to agents automate on ground they cannot trust, and all that earns them is faster data debt. The companies reaching production-grade agentic data workflows are the ones that build the right infrastructure before they connect AI to it. --- --- title: "The trust-speed paradox: Governing AI-accelerated data work" description: "72% of data teams use AI to write code. Only 24% invest in checking what it produces. Here's how to close that gap" url: "https://www.getdbt.com/blog/the-trust-speed-paradox-governing-ai-accelerated-data-work" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # The trust-speed paradox: Governing AI-accelerated data work If you lead a data team or build inside one, the [2026 State of Analytics Engineering report](https://www.getdbt.com/resources/state-of-analytics-engineering-2026) has a finding worth sitting with. Data trust is now the top priority for 83% of data teams, up from 66% a year ago. At the same time, 72% say AI-assisted coding is part of how they work. Only 24% say the same about AI-assisted observability. Read those three numbers together. Teams are using AI to produce more data, faster. They are not keeping pace on the systems that test, validate, and govern what AI produces. That gap is what happens when you adopt AI acceleration without investing in AI-level quality control. Gartner [predicted in early 2025](https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk) that through 2026, organizations would abandon 60% of AI projects built on data that wasn't AI-ready. That's tracking. The failure mode is always the same. AI gets deployed on infrastructure built for dashboards, not for reasoning. The code ships fast. The data it produces is inconsistent, untested, and undocumented. Stakeholders stop trusting it. The project stalls. That's the trust-speed paradox. AI makes data work faster, and faster data work without governance destroys the trust that made the work worth doing. ## Why the gap is worse than it looks AI doesn't just speed up good data practices. It amplifies whatever practices you already have. A team that runs tests, enforces contracts, and documents its models gets faster and sharper with AI. A team that doesn't gets faster at shipping undocumented, untested data. A distribution shift that used to cause a minor dashboard error becomes a hallucination once an agent is reasoning over the data. A schema change that used to break one report now breaks an entire agentic workflow, silently. The data debt you've carried for a decade turns into a different kind of liability the moment an agent has to trust the answers it's getting. You can't govern your way out of that after the fact. The governance has to live in the infrastructure before the AI runs on top of it. ## What a governed AI workflow actually looks like For practitioners, this isn't a new process. It's the dbt workflow you already know, applied consistently, before you add AI acceleration. Tests before merges. Every model gets [`not_null`, `unique`, and `accepted_values` tests](https://docs.getdbt.com/docs/build/data-tests) at a minimum. [Model contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) are enforced on your critical models. Breaking changes get caught in CI, not by someone eyeballing a dashboard after the fact. Contracts at ingestion. Source freshness checks run. Sources and exposures have owners. When something shifts upstream, your dbt project knows, and the failure surfaces where it should. The semantic layer as the source of truth. Your [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metric definitions live in code, not in a BI tool's proprietary layer. A natural language query from an agent returns the same number your CFO uses. The definition is version-controlled, auditable, and consistent everywhere it's queried. That's the whole point of the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl). None of this is a new architecture. It's the disciplined use of what dbt was built for, extended to cover the failure modes AI introduces. ## Five investments that close the trust-speed gap If you're a data leader staring at the gap between how fast your team adopts AI and how ready your governance is, the path forward is a sequence of investments, not a transformation program. **Enforce contracts on your most important models.** Not every model needs one today. Start with the models feeding revenue metrics, customer-facing dashboards, and any AI pipeline. [Contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) catch breaking changes automatically, and they document the interface other teams depend on. **Turn on column-level lineage.** If you can't trace a number back to its source, you can't debug an agent that returns the wrong answer. [Column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage) gives you that trace automatically in dbt. **Define your core metrics in MetricFlow.** Revenue, active users, churn. If those definitions live in a BI tool's proprietary layer, an agent can't reach them in a governed way. Moving them into the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) makes them machine-readable and consistent. **Add AI-assisted observability to match your AI-assisted coding.** The 72% to 24% gap is the most fixable part of this. If your team uses AI to write models, invest the same way in monitoring what those models produce. **Assign ownership before something breaks.** Ownership is cheap to set and expensive to reconstruct after an incident. Give every source and exposure a named owner. It's the governance investment that pays back fastest when an AI-generated model breaks and nobody can diagnose it. ## Why dbt is both the accelerator and the guardrail Databricks CEO Ali Ghodsi [told CNBC in February 2026](https://www.cnbc.com/2026/02/09/under-the-hood-of-the-ai-economy-with-databricks-ceo-ali-ghodsi.html) that over 80% of the databases on Databricks' platform are now built by AI agents. AI is building data infrastructure faster than people can. dbt sits in an unusual spot here. It's the tool that makes AI-assisted development faster, through [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) and the [dbt MCP server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) that ground agents in your dbt context. It's also the tool that makes AI-generated data trustworthy, through tests, contracts, column-level lineage, and a semantic layer that grounds agent queries in governed definitions. Most tools in the stack do one or the other. dbt does both. It's a function of what dbt stores: semantic context in code, version-controlled and machine-readable, at the layer where AI needs to understand what the data means. The teams that close this gap first will be the ones that build the foundation before AI runs on top of it. --- --- title: "Your AI isn't broken. Your data model is." description: "AI proof of concepts work. Production doesn't. The gap between those two things lives in your data model, not your model." url: "https://www.getdbt.com/blog/your-ai-isn-t-broken-your-data-model-is" date: "2026-06-08" authors: ["Dustin Dorsey"] categories: ["Insights"] --- # Your AI isn't broken. Your data model is. _This is a guest post from Dustin Dorsey, Senior Director of Data Engineering at phData. Dustin is a co-author of Unlocking dbt. He works with enterprise teams to build the data foundations that make AI reliable at scale._ There is a pattern showing up in organizations everywhere right now, and it is so consistent it almost feels scripted. A team runs a proof of concept. The AI performs brilliantly. Questions get answered in seconds that used to take days. Executives are impressed. Someone uses the word "transformative." Budget gets approved, the rollout begins, and within a few months the same executives are quietly wondering why the thing keeps giving different answers to the same questions. The data team gets called into a meeting. Someone suggests they need a better model. Someone else suggests better prompts. A third person wonders aloud if they should evaluate a different vendor Almost nobody asks the question that actually matters: Is the problem in the AI, or is it in the data the AI is trying to reason over? In most cases, it is the data. And more specifically, it is not a data quality problem. It is a data design problem. ## Why the POC always works The proof of concept works because it was designed to work. This is not cynicism. It is just how POCs operate. When you stand up an AI pilot, you pick a domain you understand well. You choose datasets that have been curated and used repeatedly. You scope the questions narrowly enough that the answers are unambiguous. You have subject-matter experts in the room who course-correct when something looks off. And critically, you are working within a slice of your data where meaning has already been implicitly enforced through months or years of human use. The AI is not discovering meaning in those scenarios. It is operating inside a perimeter where meaning was already established, and it is doing a very good job of navigating that perimeter quickly. The problem is that this creates the impression that your organization is ready for AI at scale. It is not. It is ready for AI in the narrow, well-maintained corner of your data estate that you chose for the demo. The rest of your data is a different story. When AI moves into production, the perimeter disappears. Real users ask questions that span domains. They phrase things differently. They ask about concepts that exist in three tables with slightly different definitions. They want to compare metrics that two separate teams calculate in two separate ways, and both of them call it the same thing. The AI, without a human expert in the room to catch the ambiguity, picks an interpretation and runs with it. Sometimes it picks correctly. Often it does not. And the frustrating part is that you cannot always tell which is which from looking at the output. This is where confidence erodes. Not because the technology failed. Because the expectations were built on a foundation that was never as solid as the demo made it appear. ## The human buffer no one talks about Here is the uncomfortable truth that most AI discussions sidestep: your analysts have always been compensating for this problem. Every time a business user asks "what was our revenue last quarter," an experienced analyst does not just run a query and send back a number. They instinctively clarify intent. They know that "revenue" means three different things depending on who is asking. They know the CFO wants recognized revenue on the accrual schedule; the sales team wants booked orders net of cancellations; and the operations team wants billed invoices for the period. They know which dataset is authoritative for each use case. They know which edge cases to handle and which filters to apply. They run the query, sanity-check the result against a number they already have a rough expectation for, and only then do they send it. That entire process is invisible to the business. It looks like "the analyst ran a query." What it actually is, is a highly experienced person serving as a translation layer between a messy, ambiguous data environment and a decision that needs to be made. AI does not have that translation layer. It cannot. It sees the raw structure of your data, reasons over what it finds, and produces an answer. If the structure is ambiguous, the answer will be inconsistent. Not randomly inconsistent, which would at least be easy to catch, but defensibly inconsistent. Inconsistent in ways where every answer it gives is technically justifiable based on what the data says. That is the hardest kind of wrong to catch, because nothing looks broken. The query ran. The numbers came back. The dashboard loaded. The output just happens to be answering a slightly different version of the question than the one the business was asking. ## One question, five defensible answers Let me make this concrete. "What was our revenue last quarter?" In a typical enterprise data environment, this question has multiple technically valid answers. Revenue might exist at the transaction level, the invoice level, or the recognition level. It might include or exclude returns, internal transfers, or credits depending on who configured the pipeline and when. There might be a table in the CRM that tracks bookings, a separate table in the ERP that tracks invoices, and a reconciliation table in the finance system that is the authoritative source for period-close reporting. All of them have a revenue column. All of them have a date field. All of them will give you a number. A human analyst knows which one to use. They know it because they were told, or because they learned it the hard way, or because they asked and someone explained it in a meeting two years ago that was never documented. That knowledge lives in their head, not in the data. Now ask an AI to answer the same question across all of those tables. It will pick one interpretation based on the structure it can see, the column names it recognizes, and whatever contextual signals exist in the prompt. Ask the same question worded slightly differently and it may pick a different interpretation. Ask it twice on different days and you may get two numbers that cannot be reconciled without knowing exactly which path each query took through your schema. The AI is not making mistakes. It is doing exactly what you would do if you were handed a schema with no documentation and asked to answer a business question as fast as possible. It is making reasonable inferences. The problem is that reasonable inferences are not the same as business-defined answers, and at scale, the gap between those two things becomes very expensive. ## Centralized storage is not the same as centralized meaning Most organizations spent the last decade centralizing data. They moved from distributed data marts to cloud warehouses. They built pipelines. They established governance frameworks. They invested in tooling. By most measures of data infrastructure maturity, they are in a strong position. What they did not centralize is meaning. Centralizing storage answers the question of where data lives. Centralizing meaning answers the question of what that data represents and how it should be used. These are completely different problems, and solving the first one does not automatically solve the second. When you bring data from multiple source systems into a single warehouse without establishing a shared interpretation of that data, you have not created clarity. You have created a larger surface area for ambiguity. You have more tables, more join paths, more definitions of the same concept, and more ways for a system trying to reason over that data to arrive at a different answer than the one you expected. In a human-driven analytics environment, this ambiguity gets resolved through people. It gets resolved in the meeting where two teams present conflicting numbers and someone explains which calculation is correct for this context. It gets resolved through tribal knowledge that experienced analysts carry around and apply every time they touch a dataset. It gets resolved through the dashboard filter that is always set to "exclude refunds" even though there is nothing in the underlying table that enforces that rule. AI cannot attend those meetings. It cannot acquire that tribal knowledge. It cannot apply that filter unless someone has encoded it into the data structure itself. The ambiguity that humans have learned to work around for years does not disappear when AI arrives. It becomes visible. It becomes consequential. And it becomes your most important data problem, regardless of how good your LLM is. ## What actually needs to change There is a version of this problem that gets solved by better prompts. If the ambiguity is narrow and well-understood, you can often describe it in the prompt and get consistent outputs. But that approach has a ceiling. You cannot prompt your way to consistency across a data estate where meaning is systematically implicit. At some point, the only real fix is to encode the meaning into the data itself. This is what dimensional modeling is actually for, and it is why people who have been doing data engineering for a long time have been saying for years that the fundamentals still matter. Dimensional models are not a legacy pattern for old-school BI tools. They are the structural mechanism for making business meaning explicit. They organize data around business processes rather than source systems. They separate what happened (facts) from the context required to understand it (dimensions). They declare grain. They make relationships intentional rather than inferred. The reason structure works where documentation and prompts cannot is worth stating directly. A documented definition of revenue gets ignored by the analyst who never found the Confluence page, and it is invisible to the AI that never reads documentation at all. A prompt can describe which table to use, but only for the one query you thought to write the prompt for. A modeled definition of revenue is enforced at query time, for every query, automatically, whether or not anyone remembers the rule exists. AI cannot read intent. It can only navigate structure. That is why the fix lives in the data layer, not in the prompt layer. When data is modeled this way, the questions AI can reliably answer expand dramatically. Not because the model becomes smarter, but because it has fewer opportunities to be wrong. The structure communicates intent. The interpretation is constrained. The answer space is bounded in ways that align with how the business actually defines its processes. This is not about going back to a rigid schema that cannot accommodate modern analytical needs. It is about recognizing that flexibility without structure is not an asset when AI is doing the reasoning. AI needs guardrails in the data layer that tell it what things mean and how they relate to each other. Dimensional modeling provides those guardrails. ## A question worth asking before you buy anything else Before you evaluate a new model, hire a prompt engineer, or stand up another AI platform, ask your team one question: Can you point to a single authoritative dataset for each of your core business processes? Not a general answer. A specific one. If someone asks "which table is the source of truth for customer lifetime value," is there an answer that everyone agrees on? If someone asks "how is an active customer defined," does the data enforce that definition, or does it live in someone's head and get applied inconsistently? If you cannot answer those questions with confidence, the problem is not your AI. It is the foundation the AI is trying to reason over. And no amount of model tuning or prompt engineering is going to fix a foundation that was never designed to communicate business meaning in the first place. The good news is this is a solvable problem. Organizations that invest in getting the foundation right do not just get better AI results. They get better analytics, better reporting, and better alignment across teams. The AI becomes an accelerant rather than an amplifier of existing confusion. The organizations that skip this step and keep tuning the model instead of the data will find themselves in the same meeting six months from now, still trying to explain why the numbers do not match. And that meeting (the one where two teams present the same metric with different values and neither of them is technically wrong) is worth understanding in its own right. Because that scenario is not an edge case. It is the default state of most enterprise data environments, and AI is about to make it impossible to ignore. If you want to go deeper on the structural conditions that need to exist before AI can operate reliably on your data, I've written a full white paper on this topic. [Building the Foundational Layer for Reliable AI on Structured Data ](https://www.phdata.io/offers/ai-data-foundation-whitepaper/)covers why dimensional modeling functions as trust infrastructure, what it actually means for data to be AI-ready, and why ‌organizations that skip this foundation keep struggling in the same ways. ## Where dbt and phData fit This is exactly why phData and dbt fit so naturally together. dbt provides the implementation home for the kind of intentional, process-oriented data modeling this argument is built on, with model-layer structure, testing, documentation, and the dbt Semantic Layer giving teams a practical way to encode business meaning directly into the transformation layer. phData’s role is to help teams operationalize that in practice: aligning on definitions, designing models around real business processes, and turning the principles described here into systems that can actually be built, governed, and trusted in production. --- --- title: "What is enterprise data infrastructure?" description: "Why your GenAI projects need a robust enterprise data infrastructure that acts as a single control plane for your data." url: "https://www.getdbt.com/blog/enterprise-data-infrastructure" date: "2026-06-04" authors: ["Joey Gault"] categories: ["Pulse"] --- # What is enterprise data infrastructure? Data infrastructure isn't anything new. Companies across various industries have striven to create reliable data pipelines that run at the optimal cost. It's one thing to run a couple of pipelines, however. It's another to process data at the scale, speed, and quality required of modern data-hungry applications. This is especially true in our new [Generative AI (GenAI)](https://www.techtarget.com/searchenterpriseai/definition/generative-AI) era. GenAI use cases require a large volume of highly accurate data from across the organization to succeed. A lack of such data is what will keep [up to 30% of GenAI prototypes from making it into production](https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025). Supplying the data needed for GenAI use cases requires scalability and a robust enterprise data infrastructure. In this article, we'll examine what an enterprise data infrastructure is, where companies struggle when it comes to implementing it, and how to build out an enterprise data infrastructure framework that minimizes the heavy lifting. ## What is enterprise data infrastructure? [Data infrastructure](https://www.getdbt.com/blog/data-infrastructure) is the set of systems and processes that businesses use for data management. In the past, many businesses didn't approach data infrastructure in a uniform manner. Instead, they built one-off data pipelines that were hard to manage and maintain. Data infrastructure establishes a technological and business framework that makes data processing uniform, repeatable, and reliable. Enterprise data infrastructure turns the volume up to 11, enabling a scalable framework for data ingestion, transformation, and data analytics that can feed a company's most data-hungry scenarios. Enterprise data infrastructure also prioritizes data quality and compliance. It ensures that data is only accessible by authorized personnel and processed in accordance with all applicable regulatory requirements and industry standards, including regulations such as GDPR and HIPAA. ### Common components of enterprise data infrastructure Enterprise data infrastructure consists traditionally of multiple tiers or components. The key ones include: Data ingestion. Data can come in many types of data and formats: unstructured, structured, and semi-structured. It can arrive from many data sources. The ingestion layer uses systems such as [Fivetran](https://www.fivetran.com/) to import data from sources and stores it in a consistent format for consumption by other downstream applications. This layer often relies on [Extract, Transform, and Load (ETL) or Extract, Load, and Transform (ELT) pipelines](https://www.getdbt.com/blog/etl-vs-elt) to standardize and load data into downstream systems. Data storage. Cloud computing brought us scalable data warehouse systems like [Snowflake](https://snowflake.com/) and [Amazon Redshift](https://aws.amazon.com/redshift/) that can easily store terabytes of structured data on demand. Other cloud-based data storage architectures, such as [data lakes](https://www.getdbt.com/discover/understanding-data-lakes) and data lakehouses, enable storing datasets in unstructured, semi-structured, or hybrid formats using low-cost cloud storage solutions. Data transformation and modeling. A data transformation layer takes data in its raw state and makes it available to derive insights. Data is generally stored in its raw format and then transformed multiple times to serve specific use cases. This layer typically involves tools to model data as well as [to test data models](https://docs.getdbt.com/docs/build/data-tests) to ensure validity and data quality. Data serving. The data serving layer makes data assets available for discovery and use in business intelligence and analytics, where they take their final form and can be used to drive valuable business insights and data-driven decision-making. This layer consists of tools such as data catalogs, visual dashboards, and automated report generation services to deliver a regular stream of insights in a timely manner. Data governance. [Data governance](https://www.getdbt.com/blog/data-governance) is a set of pillars that include data quality, data stewardship, data security, and data management. Rather than being a specific tool, it's a set of processes built into the fabric of the system, supporting access controls and compliance at every stage of the [analytics development lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle). ## How enterprise data infrastructure can hobble or help your GenAI projects As the AI age accelerates, enterprise data infrastructure is more important than ever. Gone are the days when companies could "just get by" with hard-coded, brittle data pipelines that require constant manual intervention. Successful GenAI projects rest upon a data infrastructure that supports automation, is robust, and can scale to manage terabytes and petabytes of data. What's stopping companies from getting there? Recently, [Fivetran created a global benchmark of enterprise data infrastructure spend](https://www.fivetran.com/blog/the-enterprise-data-infrastructure-benchmark-report-2026) to find out where the gaps are between AI expectations and reality. The report found three glaring issues: Operational efficiency is killing AI data scalability. Data teams dedicate 53% of their time to maintenance. The costs involved in data pipeline downtime and repair may cost companies more than $36 million a year. ROI on data integration is low despite massive spend. This manual, high-intervention approach to data pipelines is killing ROI. On average, enterprises allocate 14% of their data budgets to integration. However, only 27% say that ROI exceeded their expectations. Pipeline modernization delivers higher ROI. The good news is that these problems are avoidable. Fivetran found that 47% of organizations running fully managed and standardized data pipelines exceeded their ROI expectations. The decrease in maintenance work enabled their data teams to address more pressing issues in analytics, AI, and governance, thereby increasing data quality and velocity. ## Why enterprise data infrastructure remains a hard problem The thing is, companies don't end up where they are by accident. There's a reason why organizations end up with a hodgepodge of data pipelines written in different languages and running on their own infrastructure, whether cloud-based or on-premises: Data is fragmented. Without a centralized data catalog to track what's available, [a lot of data ends up in silos](https://www.ibm.com/think/topics/data-silos). These data silos use their own formatting and storage conventions, making them difficult to integrate with other corporate data systems. They often go ungoverned and unmonitored, creating a data security risk for the company, including potential data breaches involving sensitive data. Data exists in heterogeneous systems. Data storage systems have evolved over the past several decades. New architectures, such as data lakehouses, and new storage formats, such as [Apache Iceberg](https://iceberg.apache.org/), have brought needed improvements in quality, performance, and governance. This evolution means data is naturally scattered across disparate, heterogeneous data warehouses, data lakes, and other systems. Often, it's not economically feasible or technically advisable to move that data into newer systems. The time, cost, and risk entailed are too high. Building an enterprise data infrastructure takes time and money. With data engineers spending half of their waking hours fighting fires, there's little time left in a month to tackle the bigger issues. Many companies remain stuck in their current data infrastructure because building a replacement from the ground up requires time and computing resources they don't have. ## Creating a single data control plane for analytics and AI Companies looking to support GenAI don't need to centralize all of their data. What they need is an enterprise data infrastructure that supports delivering high-quality data at scale, no matter where it lives. They need adaptability and workflows that can respond to rapidly changing business needs. What they need, in other words, is a data control plane. A [data control plane](https://www.getdbt.com/blog/ai-ready-platform-generative-ai) is an architectural layer that sits over all of your end-to-end data activities. It enables data integration, access controls, data governance, and protection so you can manage the behavior of people and processes in a distributed and dynamic data environment. At dbt Labs, we've long worked to enable this vision of a vendor-agnostic, flexible, and trustworthy data control plane. dbt enables working with your data where it lives, transforming raw inputs into the high-quality data needed for both analytics and AI use cases. This approach gives stakeholders and data teams a real-time view of data health and streamlines the path from raw source to trusted insight. With dbt, you're not starting from zero. dbt provides a rich framework out of the box for enabling a consistent approach to modeling, testing, and publishing data. That means you can focus less on your enterprise data infrastructure and more on your use cases. dbt provides all of the essential tools you need to curate data for analytics and AI: Authoring and testing data transformations locally. Using the [dbt Fusion engine](https://docs.getdbt.com/docs/fusion/about-fusion), data producers can model and test data transformation code locally, using standard tools like Visual Studio Code. Fusion implements a SQL compiler that automatically validates code as engineers write, enabling data producers to work faster and deploy changes more quickly than ever. Version, test, and monitor continuously. All data analytics and AI code is version-controlled and reviewed before going live. Data developers can easily craft [data tests](http://docs.getdbt.com/docs/build/data-tests) to validate their workloads locally. Tests can also be run as part of a [Continuous Integration (CI) pipeline](https://docs.getdbt.com/guides/custom-cicd-pipelines) to verify any changes before deploying to production. This optimizes the development lifecycle and reduces the risk of errors reaching production environments. Document data for users and AI. dbt's built-in support for rich documentation means users can understand where data comes from and what it means. Using the [dbt Model Context Protocol (MCP)](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) server, you can also feed your docs to your GenAI systems as invaluable context for understanding your data. Define global metrics for GenAI apps. Metrics definitions often differ from group to group, causing confusion when the data is fed to GenAI systems. The dbt Semantic Layer provides a single location for defining and accessing key business metrics. You can supply these "gold" metrics to LLMs and GenAI apps to eliminate confusion between competing definitions and improve machine learning model performance. Drive governance across systems. [dbt simplifies data governance](https://www.getdbt.com/product/governance) as well. By standardizing on dbt as a single data control plane, you give everyone in the company a single source of truth for data and metrics. [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) provides a one-stop shop for global data discovery, with data access enforced via role-based access controls (RBAC). Delivering data for AI requires a solid enterprise data infrastructure that regulates access via a centralized control plane. dbt enables building that foundation on top of your existing, heterogeneous data architecture and gives your organization a competitive advantage in a data-driven world. [Contact us today](https://www.getdbt.com/contact) to learn more about how dbt can modernize your approach to data pipelines and set you up for GenAI success. --- --- title: "Building a data stack for trusted AI" description: "Trusted AI requires governed, consistent, and contextual data. Here’s how to build it without tying yourself down." url: "https://www.getdbt.com/blog/data-stack-trusted-ai" date: "2026-06-03" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Building a data stack for trusted AI AI adoption is stalling. Not for lack of ambition, and not because the models aren't capable. The data foundation underneath isn't ready. The numbers back this up. About 16 percent of organizations surveyed have deployed AI agents, according to Gartner. Meanwhile, 80 percent of IT leaders don't believe their data is ready for AI, and over 70 percent worry about governance in a world increasingly run by agents. That last number should be higher. Way higher. Trusted AI requires trusted data: governed, consistent, and contextual. Without it, every AI initiative follows the same arc: a promising pilot, a scaling problem, a stall. The model isn't to blame. The foundation is. ## The shape of the consumer has changed The data stack was built for people. Every layer, from how we model data to how we serve it, assumes a human is sitting at the end of the pipeline: someone with judgment who can pause when a number looks off and decide whether to trust it. That assumption is being overturned. Agents are querying your data continuously, taking action autonomously, and they don't stop to reconcile. They act. The analyst who runs 50 queries a day has been joined by agents that might run 50,000. Where a human analyst can tolerate lag and ambiguity, an agent making inventory, pricing, or routing decisions cannot. The core issue is **context**. An analyst carries context in their head: they know that "customer" in the revenue model means paying subscribers, not trial users, because someone told them that in onboarding. Agents don't have that. They take what they're given. What a column means, what's authoritative versus stale, what's a trusted curated model versus a raw staging table: that context now has to live in the data itself, in metadata, in contracts, in governed definitions that any system can read. ## Why AI is guessing in the dark Here's what actually happens when you point a generative AI model at your databases, warehouses, or BI datasets without a governed context layer: - The AI scans whatever tables and columns it can see. - It guesses which ones to use based on their names. - It pulls data straight from the warehouse, mixing raw, staging, and curated tables alongside siloed SQL in views and notebooks, with no reliable way to know which represents the actual source of truth. The results are consistent. Without governed context, agents: - Generate unreliable SQL because it can't identify the right models - Invent or misapply your metric definitions - Create governance and trust issues with no clear audit trail - Drive up costs as unreliable queries burn tokens and compute All of this happens because the agent is guessing in the dark. The context it needs to behave reliably simply isn't there. ## Two pillars: governance and structured context Trusted AI rests on two pillars. The first is **governance**: control, [data lineage](https://www.getdbt.com/blog/what-is-data-lineage), and quality. The second is **structured context**: the semantics that agents can actually reason over. Without both, every agent reinvents the truth. Over 50,000 companies use dbt in production. The governed structured context layer it provides fixes the context gap. It tells your AI or agent how your data is defined, how it connects, and what it actually means. It exposes the rich metadata that already lives in your dbt project: your models, lineage, metrics, freshness definitions, and documentation. Then, it surfaces all of that through open standards, including the m[odel context protocol (MCP)](https://docs.getdbt.com/docs/dbt-ai/about-mcp) and the [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) using [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow), so any AI system can rely on a single governed source of truth rather than reconstructing context from scratch every time. **[dbt Wizard](https://www.getdbt.com/product/dbt-wizard) **complements this by packaging proven dbt workflows - including [testing](https://docs.getdbt.com/docs/build/data-tests), debugging, migrations, and metrics definition - into reusable patterns. Agents not only know the context; they know how to follow a consistent, proven process when acting on it. **[dbt Wizard CLI](https://docs.getdbt.com/docs/dbt-ai/about-dbt-wizard-cli?version=2.0&name=Fusion) **surfaces that same context for fast, cost-efficient local development. The benefits compound quickly, with structured models and tests giving you AI that generates reliable, reusable SQL. A governed, query-optimized semantic layer gives AI the right definitions and logic, while centralized lineage surfaces clear ownership and makes every change auditable through Git and pull request (PR) workflows. The result is both increased accuracy and greater cost efficiency. When AI works from curated context rather than your entire warehouse, your token, compute, and review costs stay manageable. [**Check out dbt Wizard, your personal dbt agent, available wherever you work.**](https://www.getdbt.com/product/dbt-wizard) ## The semantic layer is not optional The semantic layer is the component that makes AI actually reliable. When metrics and business entities are defined once and reused everywhere, whether the consumer is a dashboard, an operational workflow, or an agent, they all work from the same trusted logic. That consistency is what separates AI that amplifies good decisions from AI that amplifies errors, with no human in the loop to catch the difference. Gartner estimates that by 2027, enterprises without a semantic layer will spend 40 percent more on AI rework and remediation than those with one. Don't be those enterprises. At dbt Labs, we approach this through open standards. [The Open Semantic Interchange (OSI)](https://www.getdbt.com/blog/dbt-labs-affirms-commitment-to-open-semantic-interchange-by-open-sourcing-metricflow) is an open standard for how semantics move between tools. MetricFlow gives you one definition of every metric, consumable anywhere. Define once, use anywhere, and it works with whatever agent or application comes next. The practical upshot: dbt is agent agnostic. We work with whatever LLM framework or vendor you choose. Your governed context travels with the data. As your AI stack evolves with new models, tools, and interfaces, the structured foundation stays the same. You centralize logic in one layer, reduce redundant queries, and get more reliable AI outputs across the board. ## What this looks like at ACV Auctions [ACV Auctions](https://www.acvauctions.com/) runs a wholesale dealer-to-dealer automotive marketplace. Their analytics manager, Darren Peters, has been building through this challenge firsthand. The starting point was familiar. ACV's embedded reporting system couldn't keep pace: simple report tweaks for dealers could take weeks or months. Their internal BI tool had accumulated over 1,500 dashboards. A new employee searching a common business term would see 400 different results, each a variation on the same concept, with no reliable way to identify ground truth. Darren's instinct, well before AI chat was part of anyone's roadmap, was rigorous data governance at the dbt layer. He talks about writing column descriptions the way you'd write a board game rule book: specific enough that no one argues over what a rule means. The goal was traditional data governance: a clean data dictionary, automated quality tests, and descriptions piped from [BigQuery](https://cloud.google.com/bigquery) into [Confluence](https://www.atlassian.com/software/confluence). Precise, unambiguous definitions for every column. That discipline turned out to be the foundation for everything that came after. When ACV brought in [Omni Analytics](https://omni.co/) for their semantic layer, Omni could pull column-level metadata directly out of their data warehouse automatically. Darren built the governance layer for traditional reasons. But it also set up ACV perfectly for AI. A few months into using Omni's AI chat, the team's workflow has shifted. When a [Slack](https://slack.com) message arrives from a stakeholder with a data question, Darren pastes it straight into the chat. If the response looks right, he sends back a shareable query URL: a live, governed analysis the recipient can view or iterate on themselves, without it being formally published and adding to the content pile. No more ad hoc dashboards that bloat the system just to answer a one-off question. Product managers who used to wait days or more for ROI analyses on feature work are now running those analyses themselves. Darren's team validates and adjusts, but a backlog that used to go untouched is actually getting cleared. Customer support teams are answering dealer-specific account questions in real time, without ticket queues. Darren describes himself as a former AI skeptic, specifically in the context of self-serve analytics. Two things shifted his view: the maturity of the semantic layer itself, and access to more capable reasoning models. When both came together, it wasn't a gradual improvement. "It was a light switch," he said. One question in the chat, one response, and he knew it was real. The shift ACV is living now: less time building content, more time engineering context. When the core analytics loop is about defining a dimension precisely and describing it in plain language, the team has to stay close to the business and keep sharpening its understanding. That rigor pays off across every consumer, whether human or agent. ## Shipping AI with confidence Solving for the context gap is what finally lets you ship AI with confidence. It leads to fewer hallucinations, better decisions, lower security and governance risk, reduced token and compute spend, and faster data development. Most importantly, context enables AI initiatives that actually scale beyond the pilot stage. They scale because teams are willing to adopt and trust them. That's the part people underestimate. Governance isn't the enemy of speed. It's the condition for adoption. When data semantics are open and portable, your stack stops being a series of vendor decisions and becomes a foundation you actually own. Every layer, whether ingestion, storage, or semantics, stays interoperable and yours to evolve. That's what lets enterprises scale AI without being held hostage to architectural choices made three years ago. [**For a hands-on look at all of this in practice, watch the full session recording.**](https://www.getdbt.com/resources/webinars/the-future-is-open-building-a-flexible-ai-ready-data-stack-without-lock-in) Build the governed foundation your AI initiatives need: [Talk to the team at dbt today](https://www.getdbt.com/contact). --- --- title: "dbt Labs Named Snowflake Data Integration Product Partner of the Year" description: "dbt Labs wins two Snowflake Partner honors: Data Integration Product Partner of the Year and Snowflake’s CoCo Adoption Award" url: "https://www.getdbt.com/blog/dbt-labs-named-snowflake-data-integration-product-partner-of-the-year" date: "2026-06-02" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Named Snowflake Data Integration Product Partner of the Year _Fourth consecutive annual award honors commitment to exceptional outcomes for joint customers_ **PHILADELPHIA – June 2, 2026: **[dbt Labs](https://www.getdbt.com/), a leader in standards for AI-ready structured data, announced today at [Snowflake Summit 2026](https://www.snowflake.com/summit/) that it has been named the 2026 Data Integration Snowflake Product Partner of the Year, in addition to being recognized for Snowflake’s CoCo Adoption Award for leading adoption and delivering customers transformative results via Snowflake’s coding agent and control plane for builders. dbt Labs is being recognized for its achievements as part of the Snowflake AI Data Cloud, helping joint customers unlock production-grade workflows that are built on a reliable, governed and trusted data foundation, ready to run AI agents reliably, at scale. dbt has become a preferred transformation and context engine for customers’ AI and analytics use cases, with over 75% of customers with Snowflake accounts using dbt and powerful tools like the dbt Semantic Layer, the dbt Fusion engine, and dbt MCP server to unlock a faster developer experience and execute new and complex AI use cases. With 90% of joint customers actively using Snowflake Cortex AI, dbt is an integral part of their AI journey, delivering the reliable data foundation that enables these organizations to take full advantage of Snowflake's AI functionality. “Trust in data is the most widely prioritized organizational objective, and Snowflake Marketplace is a powerful resource to connect enterprises to dbt and its latest features that bring structure, governance, and velocity to what data teams are building in the AI era," said Shawn Toldo, Vice President, WW Partner Organization at dbt Labs. "This award underscores our mutual dedication to supporting our joint customers and delivering remarkable results, allowing for data-driven innovation at scale to expand the reach of data and AI capabilities." This is the fourth consecutive year that Snowflake selected dbt Labs as a Snowflake partner award winner, which is a testament to the depth and durability of their collaboration and the impact of this longstanding partnership. The companies are united in their mission to help customers cost-effectively build AI-powered insights and data assets, ultimately driving deeper organizational trust in data and the teams that power it. With transactions on the Snowflake Marketplace growing by approximately 30% year over year, dbt Labs is helping customers utilize their full spend from Snowflake contract commitments to ensure optimal ROI. "dbt Labs has been an incredible partner to us over the years and we're proud to name them as Snowflake's 2026 Data Integration Partner of the Year," said Amy Kodl, SVP, Worldwide Alliances & Channels. "The work their team is doing with the AI Data Cloud ecosystem continues to deliver strong results for our joint customers." Joint customer [WHOOP](https://www.getdbt.com/case-studies/whoop) is a testament to this continued collaboration. As the WHOOP team grew, they used the dbt platform as a scalable solution to ensure clean and well-governed data was being migrated into Snowflake. dbt provided the foundation that gave all stakeholders at WHOOP access to reliable data, which allowed the team to [use Snowpark](https://www.snowflake.com/en/customers/all-customers/case-study/whoop/) to build the WHOOP AI/ML financial forecasting model. What once was a major roadblock is now streamlined, saving time for the WHOOP analyst and data engineering teams to focus on more strategic initiatives. Learn more about dbt Labs and Snowflake [here](https://www.getdbt.com/data-platforms/snowflake), and visit the dbt Labs booth (#2112) during this week's Snowflake Summit to explore the [latest dbt innovations](https://www.getdbt.com/blog/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents). Stay on top of the latest news and announcements from dbt Labs on [LinkedIn](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=82576339&u=https%3A%2F%2Fwww.linkedin.com%2Fcompany%2Fdbtlabs%2Fmycompany%2F&a=LinkedIn), [X](https://x.com/dbt_labs), [Instagram](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=255235389&u=https%3A%2F%2Fwww.instagram.com%2Fdbt_labs%2F&a=Instagram), and [YouTube](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=508887317&u=https%3A%2F%2Fwww.youtube.com%2Fc%2Fdbt-labs&a=YouTube). **###** **About Fivetran + dbt Labs** Fivetran + dbt Labs deliver the data infrastructure layer that makes agents trustworthy — from the moment data moves, through every transformation, to the context an agent reasons from. The Fivetran platform moves, manages, and transforms data from every system a business runs on into a secure, reliable foundation engineered to evolve, with the flexibility to work across clouds, engines, and tools. With Fivetran, analytics, operations, and AI run on data you trust and control. Thousands of organizations worldwide, including OpenAI, LVMH, Pfizer, and Verizon, rely on Fivetran to turn data into a competitive advantage. Learn more at [Fivetran.com](http://fivetran.com), or follow Fivetran on [LinkedIn](http://linkedin.com/company/fivetran). Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 100,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at getdbt.com, and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "What we announced at Snowflake Summit and why it matters" description: "dbt State, dbt Wizard, dbt Core v2.0, and the Fivetran merger" url: "https://www.getdbt.com/blog/what-we-announced-at-snowflake-summit-and-why-it-matters" date: "2026-06-01" authors: ["Corinne Hallander"] categories: ["Product"] --- # What we announced at Snowflake Summit and why it matters We showed up at Snowflake Summit differently this year. In case you missed it: [Fivetran and dbt Labs are now a single company](https://www.getdbt.com/blog/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents), with a shared vision of building a foundation for trusted agents using open data infrastructure principles. There's a lot to cover among all the announcements, so let's get into it. ## Fivetran and dbt Labs have merged If you've worked in data for the last decade, you know both names. Fivetran is the standard for getting data into your warehouse reliably and automatically, with connectors for virtually everything. dbt is the standard for data transformation. For years, teams have run Fivetran and dbt together. The combination was already the backbone of the modern data stack for thousands of organizations. What changes now is that the Fivetran and dbt teams are working toward the same thing: a data foundation that’s open, portable, and ready for what agents need. Customers of both Fivetran and dbt are already leading successful, trusted AI and agent initiatives, backed by our products. The value of this combined approach was emphasized by Piyush Bhargava, Sr. Director Global Data & Analytics at DocuSign: "AI and agents are only as strong as the data behind them. By investing in Fivetran and dbt, we've built the reusable, trusted data assets that are central to how we scale AI and drive innovation." [You can read a lot more about the merger thesis, dbt's continued commitment to open source, and what's next in Tristan's blog](https://www.getdbt.com/blog/fivetran-and-dbt-are-one-company-now-here-s-what-that-means). ## dbt Core v2.0: The open-source foundation, rebuilt on Fusion dbt Core started in 2016 as a way for data practitioners to work like software engineers with modular SQL, version control, testing, and documentation in an open-source framework. It quickly became the industry standard, and now over 100,000 organizations run it today. But the original dbt engine—built in Python—has its limits. Parse times grew with project size. The execution layer had no understanding of the SQL it was running. Errors surfaced only after hitting the warehouse. A ground-up rewrite was needed. That rewrite became the dbt Fusion engine: built in Rust, with native SQL comprehension and parse times up to 10x faster than the original engine. We launched Fusion last year and today, over 4,500 projects run on Fusion. But maintaining two engines meant confusing licensing, an all-or-nothing-feeling migration path, and real uncertainty for contributors, adapter maintainers, and partners about where to build. So this week, we changed that. We open-sourced the Fusion runtime and released it as dbt Core v2.0 under the Apache 2.0 license. This means that the two-engine era is coming to an end, and every practitioner can enjoy the dbt experience they know on a faster, more scalable foundation. Consolidating to one engine also means that our commercial investments will directly drive improvements in the open-source distribution of dbt. This is the biggest open-source expansion that dbt has made in years. Fusion extends dbt Core v2.0 in a new proprietary distribution with richer capabilities, such as SQL comprehension, column-level lineage, and instant feedback while you work. dbt Core v2.0 is in alpha release, while the proprietary distribution is in preview. The most used adapters are supported at launch: Snowflake, BigQuery, Databricks, and Redshift. The path forward is simple: `pip install dbt==2.0.0-preview.x` for a single binary that brings together the open runtime, Fusion-powered capabilities, and access to platform-connected workflows. [**Read more: dbt Core v2.0 is here**](https://docs.getdbt.com/blog/dbt-core-v2-is-here) ## Announcing dbt State: Build what’s changed, skip what hasn’t Every time a pipeline runs, it rebuilds everything, even tables that haven’t changed. You pay for every one of those compute cycles. Over time, this becomes a significant cost, not just in warehouse spend, but in time engineers spend managing job schedules, building sub-selectors, and orchestration around the inefficiency. Custom, manual workflows mean slower data and slower answers. **dbt State changes this**. On every run, dbt State checks your metadata and model SQL to see what’s changed. When upstream data or code has changed, it builds the model. Otherwise, it skips the build by reusing existing state, cloning existing state, or auto-deferring (in development) to production state. In production, this results in an average of 30% reduction in warehouse compute and prevents breakage. In development, this means faster iterations without fear of costly mistakes. CarGurus Vice President of Data Parag Shah put it simply: > "Before dbt State, every job rebuilt every model in the lineage. Every. Single. Time. Now, with dbt State, dbt checks if source data changed. If it didn't, the model is skipped. For us, that resulted in a 9% compute reduction, 35% fewer models built, and a 15% reduction in Snowflake backfill costs." ```json { "_key": "033b51c767a3", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/Wb-b4BjoVUE" } ``` The impact goes beyond compute savings. When you know you're only paying for new work, you stop rationing runs. Teams that used to schedule refreshes carefully, because more frequent runs felt wasteful, can now run as often as the business needs fresh data, without the fear of an expensive bill at the end of the month. And with orchestration logic living in the project rather than a spreadsheet of job definitions, engineers spend less time on maintenance and more time on what matters. State-aware orchestration has been delivering cost savings to dbt platform users on Fusion during preview. dbt State opens that capability to _every_ dbt user: running locally on dbt Core (1.7+) or the dbt platform, in your orchestrator of choice, true to open data infrastructure principles. [**Available now in Preview: get started today**](https://www.getdbt.com/product/dbt-state). _Want to learn more?_ [Join us for a live virtual event on July 15th to see dbt State in action](https://www.getdbt.com/resources/webinars/dbt-state-build-what-s-changed-skip-what-hasn-t/?utm_medium=internal&utm_source=blog&utm_campaign=q2-2027_dbt-state-deep-dive-product_aw&utm_content=themed-webinar____&utm_term=all_all__) with a real customer story and answers to the questions practitioners always ask before they deploy. Save your seat now. ## Announcing dbt Wizard: an AI agent built for analytics engineering Coding agents are good at writing decent code. They're less good at writing code that knows your dbt project: code that respects your upstream models, doesn't break downstream dependencies, and handles the governance requirements your team has spent years building. The agent generates something plausible, but you spend the next hour verifying it. Analytics engineering is not the same as software engineering. You're not just editing files. You're reasoning about lineage. You're aware of what breaks three models later. You're accountable when something goes wrong in production. **Generic tools weren't built for that context.** **dbt Wizard is.** dbt Wizard is an AI agent built specifically for the analytics engineering workflow: investigating, building, validating, and shipping. It's grounded in your dbt project natively. That means it knows your lineage, your contracts, your tests, and your metric definitions before it writes a single line. It knows which tool to call. It validates its own work. It shows you what changed and why. The difference this makes is concrete. > “Before dbt Wizard, our engineers were spending more time correcting AI output than they were writing models. Now the agent actually knows our project. It gets the joins right, it respects our contracts, and it doesn't break things downstream. We've seen a 15-20% reduction in production incidents since we rolled it out." - Erion Krasniqi, Junior Data Scientist, Endress+Hauser InfoServ ```json { "_key": "cce9e4d9c7d5", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/sQ8VQbOimng" } ``` dbt Wizard is available in two surfaces: inside the dbt platform and from the terminal via dbt Wizard CLI for teams developing locally. It works across Snowflake, BigQuery, Databricks, Redshift, and every other dbt-supported warehouse. [Get started today](https://www.getdbt.com/product/dbt-wizard). _Want to learn more?_ [Join us for a live virtual event on July 22nd to see dbt Wizard in action across the CLI and dbt platform](https://www.getdbt.com/resources/webinars/dbt-wizard-an-agent-purpose-built-for-analytics-engineering/?utm_medium=internal&utm_source=blog&utm_campaign=q2-2027_dbt-wizard-deep-dive_aw&utm_content=themed-webinar____&utm_term=all_all__)—real use cases, real analytics teams using it today, and answers to the questions you'll have after watching the demo. Save your seat now. ## Recognized as Snowflake's Data Integration Partner of the Year The week brought more than product news. At Snowflake Summit, dbt Labs was named the 2026 Data Integration Snowflake Product Partner of the Year, along with Snowflake's CoCo Adoption Award for leading adoption of Cortex Code and delivering customers transformative results. It's the fourth consecutive year Snowflake has recognized dbt Labs with a partner award, a reflection of how deep the collaboration has become. The reason is simple: trusted data is the foundation for trusted AI. Over 75% of customers with Snowflake accounts use dbt, and 90% of joint customers actively use Snowflake Cortex AI. dbt is the preferred transformation and context engine for those AI and analytics use cases. [Read the full announcement](https://www.getdbt.com/blog/dbt-labs-named-snowflake-data-integration-product-partner-of-the-year). ## What’s next? For data teams, the question has always been: how do we do more with what we have? Better pipelines, smarter infrastructure, less time managing things that should manage themselves. That's what we're building. Still have questions? On June 25th, Tristan Handy (President and co-founder, Fivetran & dbt Labs) and Taylor Brown (COO and co-founder, Fivetran & dbt Labs) are hosting a [live virtual event](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a) to walk through the merger, what's shipping in dbt, and what it means for your stack, and then taking your questions live. This is the most direct access you'll have to the people making these decisions. [Save your seat today](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a). --- --- title: "Fivetran and dbt are one company now. Here's what that means." description: "Fivetran and dbt Labs are officially one company to deliver data infrastructure for agents you trust." url: "https://www.getdbt.com/blog/fivetran-and-dbt-are-one-company-now-here-s-what-that-means" date: "2026-06-01" authors: ["Tristan Handy"] categories: ["Company"] --- # Fivetran and dbt are one company now. Here's what that means. We [announced](https://www.getdbt.com/blog/dbt-labs-and-fivetran-merge-announcement) the merger in October. Today it's official: [Fivetran and dbt Labs are one company](https://www.getdbt.com/blog/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents), with a shared mission to build open data infrastructure for the agentic era. The timing couldn't be better. In the months since October, the shift toward agentic AI has gone from a curiosity to a force of nature. What started as a strategic combination of two organizations to deliver best-in-class data pipelines has become something even more urgent: an absolute imperative to prepare the world for agents. We've spent years building complementary halves of the same stack. Now we get to build them together…at exactly the moment it matters most. ## The thesis Agents are rapidly becoming the primary consumers of your enterprise data. Agents are no longer just in production at frontier companies, but in enterprises globally. Agents approve insurance claims, deflect support tickets, and write code. And they do it at increasingly massive volumes: up at least 1-2 orders of magnitude YoY. Many of these agents will need access to your organizational data. Unfortunately, most data infrastructure deployed today was not built for agents, but for humans: analysts running queries, dashboards refreshing on a cadence, data consumers clicking around in a BI tool. Humans are patient. Humans bring their own context. Humans do one task at a time. Agents don’t. Agents operate at machine-speed. They’re running 24/7. They can issue queries at a volume human traffic never approached. And crucially: they don't have the institutional knowledge that a human analyst carries in their head. They can't fill in the gaps. When they encounter undefined metrics, ungoverned data, or fragmented business logic, they don’t shoulder-tap their colleagues...they often just (confidently!) produce the wrong answer. And the more autonomous they are, the harder that wrong answer is to identify and troubleshoot. 80% of IT leaders believe their enterprise data is not ready for agentic AI. 70% worry about AI governance with the proliferation of agents, according to Gartner. These aren't anxieties about the future. They're descriptions of the present. The bottleneck has shifted. For the past decade, the bottleneck in data was infrastructure, and we, collectively, largely solved it. The modern data stack worked, although sometimes it could be a bit of a pain. **The new bottleneck is trust**: trust in the data, the context, the governance. That is precisely the problem we built to solve. Together. ## What Fivetran and dbt each bring, and why coming together matters **Reliable, governed, and trusted data for agents.** Fivetran ensures data is complete, fresh, and reliably moved. When an agent reaches for your data, it's current and it reflects the actual state of your business across many sources. Agents querying stale, siloed data don't just produce wrong answers—they take wrong actions. Freshness matters more when the consumer is autonomous. But fresh data alone isn't enough. Agents don't bring context; they inherit it. dbt's role is making sure that context is governed, defined, and trustworthy: business logic versioned in code, metrics defined once and tested, lineage traced from raw source to downstream consumer. This matters well beyond conversational analytics. The agent approving a loan, deflecting a support ticket, or triggering a supply chain reorder needs the same thing an analyst needs: reliable data. When an agent asks, "what’s the shipment status of this order?" or “is this customer in good standing?” it has to know exactly how to get that answer, without guessing. **Flexible and portable. Any engine, cloud, or model. No lock-in.** Here's the thing about that business context: it needs to travel. The AI landscape is shifting fast: new models, compute platforms and tools are coming in and out of favor so fast that organizations can’t fully adopt one framework before the tide turns and it’s onto the next. The platform-native approach to data infrastructure—where your semantic definitions and pipelines are tightly coupled with a single compute vendor—breaks down the moment you need to either migrate or go multi-engine. Fivetran moves data across any source without owning your storage. dbt keeps your business logic in code you control, portable across any engine or cloud. Deep integrations across the AI and analytics ecosystem means this infrastructure is pluggable at every layer, even as you migrate or go multi-engine. Together, you get a flexible, portable foundation that evolves with your architecture instead of constraining it. That's not a feature: it's a core design principle. **Scalable and optimized for production AI demands**. Agents will issue queries at a volume human analysts never approached, and without an architecture built for it, costs grow incredibly fast. The data supports this: Despite a[ 95% drop in token costs](https://venturebeat.com/orchestration/cheaper-tokens-bigger-bills-the-new-math-of-ai-infrastructure), enterprise AI spend has[ exploded from $1.7B in 2023 to $37B in 2025](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/). Cheaper tokens, yet bigger bills. Agents repeatedly querying fragmented systems for missing context is expensive in compute and tokens. Managed data connections reduce operational overhead. Governed context minimizes unnecessary retrieval hops. Decoupled storage and compute means you can route workloads to the most efficient engine for each job. The goal is for per-query cost to fall as volume scales, not rise with it. **Together, Fivetran and dbt deliver open data infrastructure for agents you trust, at scale:** ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/58d77b8611ec19dd44538560287d8192b86599e5-2048x1645.jpg) This is just the beginning. Today, we announced [two new exciting dbt features](http://www.fivetran.com/fivetran-dbt-labs-merger) available to users wherever they write code. ## Our commitment to open source, and what we're doing to reinforce it Our commitment to open source isn’t changing. Open collaboration and community contributions have defined dbt's evolution from the start. It’s the very thing that’s made dbt the de facto standard for transformations on structured data, and that same ethos is what’s going to propel our users to succeed in this AI wave. Today, [we announced the alpha of **dbt Core v2.0**](https://docs.getdbt.com/blog/dbt-core-v2-is-here): the next major version of dbt Core. As it has been since the first commit, **dbt Core remains licensed under Apache 2.0 **with this release. What changes in v2.0 is what's under the hood: the kernel of the dbt Fusion engine, a full Rust rewrite of the dbt runtime, is being open-sourced under Apache 2.0. The Rust foundation powering Fusion—faster parsing, more consistent execution, dramatically better performance in agentic workflows—is now available to the entire community under the same license you’ve been using for a decade. Instead of two codebases, two licenses, and ambiguity about what’s available on dbt Core versus Fusion, where to build or how to migrate, **we will simply have one engine: dbt**. Not only are we excited to deliver all of these benefits to every dbt user, having a single engine to invest in allows us and the dbt community to innovate faster in the AI age where speed is everything. This is the single biggest drop of new Apache 2.0-licensed code we've shipped in years…potentially ever. ## What’s next Personally, I’m excited about the future of the data ecosystem, of the role of the data practitioner, of what Fivetran and dbt Labs can build together. After several years of fascinating progress in consumer AI but a constant refrain of it-doesn’t-quite-work-yet in the enterprise, _AI and agents are truly here_. Which has caused the pace of change for everything inside of the data ecosystem to absolutely skyrocket. This is great. This is what we should all want. Data practitioners—from engineers to analysts—are critical to the AI future, but have been sitting on the sidelines for the past few years saying “put me in, coach!” Guess what: **it’s time**. While none of us know _exactly_ how the next several years will play out, things will move quickly, and data practitioners everywhere will be central to the story. What you can see us doing—with the merger, with our product launches, with our investments in open data infrastructure—is to answer the question “How do we position ourselves to elevate and advocate for data practitioners into the next decade?” The “how” is changing, but the mission remains constant. **We want to be your partners in building data infrastructure for the age of AI and agents.** If you have questions about our shared vision, Core v2.0, and new products we announced at Snowflake Summit, [join me on June 25th for a live Q&A](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a/?utm_medium=internal&utm_source=blog&utm_campaign=q2-2027_fivetran-dbt-merger_aw&utm_content=themed-webinar____&utm_term=all_all__). --- --- title: "Fivetran + dbt Labs Complete Merger to Create the Data Infrastructure for Trusted AI Agents" description: "Fivetran and dbt Labs are now one company, focused on building the data foundation for the agentic AI era." url: "https://www.getdbt.com/blog/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents" date: "2026-06-01" authors: ["Elaine Green"] categories: ["Press"] --- # Fivetran + dbt Labs Complete Merger to Create the Data Infrastructure for Trusted AI Agents **OAKLAND, Calif. — June 1, 2026 — **[Fivetran](http://www.fivetran.com), the data foundation for AI, today announced the completion of its merger with [dbt Labs](http://getdbt.com), the creator of dbt and the leader in standards for AI-ready structured data. Initially operating as Fivetran + dbt Labs, the all-stock transaction, originally [announced](https://www.getdbt.com/blog/dbt-labs-and-fivetran-sign-definitive-agreement-to-merge) on October 13, 2025, brings together two category-defining platforms to advance a new era of trusted, Open Data Infrastructure for AI at scale. George Fraser will continue serving as CEO, and Tristan Handy will serve as President. Together, Fivetran + dbt Labs support a global community of more than 100,000 data teams across analytics, data engineering, and AI initiatives, including some of the world’s most recognized brands such as OpenAI, Zendesk, Coupa, and HubSpot, as well as leading enterprises across financial services, retail, manufacturing, and healthcare. **A new foundation for agentic AI** AI agents are quickly becoming the primary consumers of enterprise data, and they behave differently from the human analysts that today's data stack was built to serve. Agents operate continuously, in parallel, and at machine speed. And many organizations want to move into a world where most agents are autonomous — no human in the loop. This shift raises the bar for the data they run on, requiring it to be reliable, fresh, governed, and accessible across every system in the enterprise. Fivetran + dbt Labs are building the data foundation for the agentic AI era. Together, these companies deliver the data infrastructure layer that makes agents trustworthy, from data movement and transformation to the governed context needed for reasoning and action. Fivetran ensures agents operate on complete, continuously synced, and reliable data. dbt ensures that data is defined, tested, and trusted through governed business logic, shared semantic context, and software engineering best practices embedded throughout the data lifecycle. Built on open standards, this foundation works across any cloud, engine, and tool, giving organizations the freedom to evolve their architecture without lock-in while maintaining portable business logic, cost efficiency, and control. “The next generation of enterprise AI will be defined by the quality and trustworthiness of the underlying data,” said George Fraser, CEO and Co-Founder of Fivetran + dbt Labs. “Together, Fivetran and dbt Labs are creating the infrastructure layer that helps organizations deliver governed, high-quality, and semantically rich data to power trusted AI agents at scale.” "The companies that deploy AI successfully over the next decade will be the ones whose agents can be trusted to act," said Tristan Handy, President and Co-Founder of Fivetran + dbt Labs. "Trust is built at the infrastructure layer, on high-quality tooling and on open standards. That's the bet we're making together." **Joint product innovations** The merger also marks the first major milestone in a shared innovation roadmap, with the first combined innovations from Fivetran + dbt Labs debuting today. The announcements span agentic development workflows, intelligent orchestration, and continued investment in open source innovation, including extending powerful dbt capabilities to the dbt Core open source user base. Key innovations announced today include: - **dbt Core v2.0 (alpha): **The open sourcing of the dbt Fusion engine runtime, released as dbt Core v2.0 under an Apache 2.0 license, giving every practitioner the dbt experience they know on a faster, more capable foundation. Additionally, the locally installable distribution of dbt gives developers free access to the full breadth of Fusion’s capabilities — core language features and warehouse adapters — with the ability to seamlessly unlock additional platform features by logging in directly from the terminal. - **dbt State (preview):** dbt State acts as a caching layer for data pipelines. It only builds what’s changed and skips what hasn’t, helping companies reduce underlying infrastructure costs by 30% or more. - **dbt Wizard (beta): **dbt Wizard brings autonomous assistance for model authoring, refactoring, and debugging, grounded in full dbt project context, including lineage, tests, contracts, and defined metrics. The result is governed recommendations and trusted SQL generation that reflect how enterprise data is actually structured and defined. - **Agents Schema:** An open source standard for agentic context that designates a single schema in the warehouse or lake as the shared context layer for AI agents. Metric definitions, semantic models, dbt lineage, and business documentation are stored in plain SQL tables and can be published from existing systems through tools such as GitHub Actions, metadata connectors, or custom integrations. Compatible with any warehouse, lake, ingestion tool, or SQL-capable agent, Agents Schema gives organizations a customer-owned context layer that works within existing security and governance policies, improves token efficiency through richer context, and eliminates the need for new infrastructure or vendor-locked agent systems. **Customers building with Fivetran + dbt Labs** “With Fivetran and dbt, what used to take months now happens in weeks,” said Akshay Agrawal, Director of Data Engineering at Zendesk. “It gives the business faster access to trusted data and creates the foundation we need to scale analytics, agents, and AI across the enterprise.” “Our focus now is on how we operationalize AI across Inova. With Fivetran and dbt, we’re creating the foundation for AI agents and applications that can act on trusted, governed data — not just generate insights, but drive action,” said Jon McManus, Chief Data and AI Officer, Inova Health. “The combination of Fivetran and dbt isn't just about efficiency today,” said Lakshmi Ramesh, VP of Data Services at Tinuiti. “It's about being ready for what's next. Analytics, AI, and agentic workflows all run on trusted data, and together, Fivetran and dbt are the data infrastructure that makes it all possible.” “By building an AI-ready data foundation with Fivetran and dbt, we’re improving how teams across Shutterstock access and operationalize trusted, real-time data for analytics and emerging AI-driven workflows,” said Jitesh Kumar, Senior Software Development Manager at Shutterstock. "AI and agents are only as strong as the data behind them,” said Piyush Bhargava, Sr. Director, Data Architecture and Engineering at DocuSign. “By investing in Fivetran and dbt, we've built the reusable, trusted data assets that are central to how we scale AI and drive innovation.” The combined innovations are now available and will be showcased throughout Snowflake Summit 2026. Summit attendees can visit the Fivetran booth (booth #2313) or the dbt Labs booth (booth #2112) to connect with product experts, experience demonstrations of the combined innovations, and explore customer use cases powering trusted AI agents at scale. [Learn how](http://www.fivetran.com/blog/fivetran-dbt-an-open-agent-ready-future-for-data-teams) Fivetran + dbt Labs are building the data foundation for trusted AI agents. **About Fivetran + dbt Labs** Fivetran and dbt Labs deliver the data infrastructure layer that makes agents trustworthy — from the moment data moves, through every transformation, to the context an agent reasons from. The Fivetran platform moves, manages, and transforms data from every system a business runs on into a secure, reliable foundation engineered to evolve, with the flexibility to work across clouds, engines, and tools. With Fivetran, analytics, operations, and AI run on data you trust and control. Thousands of organizations worldwide, including OpenAI, LVMH, Pfizer, and Verizon, rely on Fivetran to turn data into a competitive advantage. Learn more at [Fivetran.com](http://fivetran.com), or follow Fivetran on [LinkedIn](http://linkedin.com/company/fivetran). Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 100,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at getdbt.com, and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "What data infrastructure do agents need?" description: "Discover the data infrastructure architecture AI agents need to eliminate hallucinations, handle sub-second freshness, and query" url: "https://www.getdbt.com/blog/what-data-infrastructure-do-agents-need" date: "2026-05-28" authors: ["Joey Gault"] categories: ["Pulse"] --- # What data infrastructure do agents need? AI agents do not fail because of a weak model. Rather, they fail because of using infrastructure that was never meant for autonomous decision-making. As enterprises shift from simple chatbots to autonomous agents executing multi-step business processes, data infrastructure requires a fundamental change. Traditional AI infrastructure built for batch model training and human analytics fails in agentic workflows. It becomes a bottleneck for systems whose success depends on freshness, low latency, and semantic consistency. Production-ready [AI agents](https://www.getdbt.com/product/dbt-wizard) need an adaptive [data infrastructure](https://www.getdbt.com/blog/data-infrastructure) that combines real-time ingestion with a centralized semantic layer and standardized interfaces for task execution. In this article, we'll discuss what makes agent data infrastructure different and how to build it. We'll also look at the tools you can use to build a foundation that agents can trust. ## What makes agent data infrastructure different? Traditional AI infrastructure supports passive, chat-based retrieval, where systems tolerate high latency and rely on loosely structured documents to generate text. Autonomous agents operate as active software applications that execute multi-step workflows, invoke tools, trigger APIs, and retrieve specific records to drive real-world business actions. This shift from passive text generation to active execution introduces three key differences from standard AI infrastructure. - Freshness requirements. Agents demand sub-minute data freshness. A fraud detection agent cannot rely on hours-old data to approve live transactions. Without instant access to current status and support tickets, agents risk making incorrect decisions based on conditions that no longer exist. - Access patterns. Agents require diverse, concurrent access patterns. While humans typically run single analytical queries, agents use multimodal retrieval and combine vector searches, SQL queries, and key-value lookups in one workflow. Infrastructure must support these complex patterns instantly without performance degradation. - Semantic grounding. Agents need rigorous semantic grounding to avoid hallucinating definitions. Data must include metadata and governed definitions so agents can use exact programmatic logic without guessing relationship keys. ## Key infrastructure gaps that cause agent failure in production Deploying autonomous agents on legacy infrastructure leads to numerous failure modes that impair reasoning and cause systemic logic failures. - The context vacuum. It occurs when raw database tables lack metadata or documentation, leaving agents unable to interpret field relationships. For example, without machine-readable logic to define status codes, agents must infer meanings from generalized training data. This results in operational failures when AI assumptions conflict with enterprise realities. - Lineage blind spots. When agents lack end-to-end lineage, they cannot trace where a data field originated, which transformations it passed through, or whether an upstream pipeline change has corrupted it. This makes it impossible to evaluate whether the data they are acting on is reliable. - Metric drift. Siloed metric definitions across business units produce conflicting agent reasoning. When the finance team defines "active customer" differently from the product team, an agent pulling from both systems will generate outputs that contradict each other. Those contradictions surface as unpredictable operational outcomes that are difficult to trace back to a root cause. - Integration overhead. Building, securing, and maintaining custom API endpoints for every new agent use case slows deployment momentum. Teams spend more time on integration plumbing than on agent logic, and each custom endpoint introduces a new failure point that requires monitoring and maintenance. ## Building the five-layer agent data architecture: A step-by-step blueprint To move AI initiatives from pilot to production, [data engineering](https://www.getdbt.com/blog/what-is-data-engineering) teams must implement a five-layer data architecture optimized for machine consumption. ### Step 1: The ingestion layer (transition from batch to log-based CDC) Implement real-time data pipelines to capture changes in source data the moment they occur. This removes the processing lag of [batch ETL pipelines](https://www.getdbt.com/blog/etl-pipeline-best-practices) and prevents agents from acting on outdated warehouse snapshots. Log-based [Change Data Capture (CDC)](https://docs.getdbt.com/blog/change-data-capture) is the optimal approach for the ingestion layer. It reads database transaction logs to identify modifications such as inserts, updates, and deletes. Each change in the log is linked to an ordered log sequence number (LSN) that helps the CDC system determine the sequence of modifications. The CDC approach tracks changes within milliseconds of the database committing the transaction and streams these changes as discrete events to downstream destinations. This sub-second latency ensures agent data remains aligned with operational reality in near real-time. ### Step 2: The processing and context store layer (reshape streams for rapid retrieval) Raw CDC streams are often too technical for direct use. The processing layer applies the transformations to filter Personally Identifiable Information (PII), manage schema changes, and consolidate transactional events into business entities. Processed data then feeds low-latency environments such as Redis and vector databases. This multi-store strategy lets agents select the retrieval method that best matches their current task. For example, if an agent needs to recall a session history, it queries a fast NoSQL key-value store. Similarly, if it needs to match a natural language query to a technical manual, it performs a similarity search in the vector database. The processing layer provides high-velocity data through real-time data pipelines. This ensures information is optimized for live model queries. It also reduces delays associated with batch-style reporting. ### Step 3: The semantic layer (centralize meaning in a machine-readable store) Most traditional [data pipelines](https://www.getdbt.com/blog/data-pipelines) stop after loading data into a warehouse. While they govern the [data transformation](https://www.getdbt.com/blog/data-transformation), they leave the data retrieval logic ungoverned. Without [governance](https://www.getdbt.com/product/governance), AI agents may misinterpret the necessary join logic and aggregation methods. The [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) mitigates the issue by centralizing meaning in a machine-readable format. It standardizes logical concepts so agents do not misinterpret data grains or confuse conflicting column names. Semantic layer uses entities, dimensions, and measures to create a semantic graph. When an agent requests a business metric, the semantic layer automatically references the graph to find the best path between tables and generates optimized SQL for accurate results. This method avoids fan-out queries, chasm joins, and agents from retrieving inaccurate results. ### Step 4: The validation layer (instrument pipelines with automated quality checks) Since AI agents process data as probabilistic engines, poor data leads to confident but corrupted outputs. The validation layer acts as a circuit breaker, preventing anomalous data from reaching inference engines and causing hallucinations. The validation layer runs testing queries during the data build phase. These checks use [SQL SELECT statements](https://docs.getdbt.com/sql-reference/select) to verify schema integrity, check for null constraints, and monitor distribution shifts. If the test returns zero failing rows, the data passes validation. But if failures are detected, the pipeline halts and quarantines data to protect the agent's context window. ### Step 5: The interface layer (standardize access control via open protocols) Connecting AI systems to enterprise data has long suffered from an integration explosion known as the N × M problem. If an enterprise uses four different AI models and needs to connect them to four different internal data sources, teams must build and maintain sixteen separate integrations. Every custom integration requires unique OAuth flows, distinct message parsing logic, isolated error handling routines, and custom rate limiting. The interface layer replaces this fragile, hard-coded custom API integration with open connectivity frameworks like the [Model Context Protocol (MCP)](https://www.getdbt.com/blog/mcp). MCP acts as a universal transport protocol for AI and enables teams to build a single server for a specific enterprise tool or data source. Any MCP-compliant AI model connects to that server instantly using standard JSON-RPC 2.0 messages. This standard allows autonomous agents to discover, reason about, and query corporate data assets through a unified interface. It ensures that the organization enforces strict governance and access policies before transmitting any data to the agent. ## How dbt helps you build a data infrastructure that agents can trust Implementing the five layers is difficult through custom tools. dbt provides that foundation out of the box. ### MCP Server The [dbt MCP Server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2.0&name=Fusion) uses the Model Context Protocol to expose [dbt models](https://docs.getdbt.com/docs/build/models), [metrics](https://docs.getdbt.com/docs/build/build-metrics-intro), and semantic definitions to AI agents without custom APIs. This open standard establishes a secure, restricted interface for LLMs to access governed corporate data through standardized host, client, and server roles. Data teams can deploy the [dbt MCP Server](https://github.com/dbt-labs/dbt-mcp) using two architectures: - Local Server: It runs on developer machines via uvx dbt-mcp and integrates with [dbt Core](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0&name=Fusion) for local testing and documentation. - Remote Server: The remote server connects to [dbt platform](https://www.getdbt.com/product/dbt) via HTTP for high-scale production use, enabling multi-user metadata discovery and lineage tracing. ### Semantic Layer The[ dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) centralizes business metric definitions using [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow). Every agent queries the same governed version of business logic regardless of which tool or workflow it operates from. It removes the metric drift and context vacuum gaps that cause the most persistent agent failures in production. ### Lineage and metadata dbt maps projects as [Directed Acyclic Graphs (DAGs)](https://www.getdbt.com/blog/dag-use-cases-and-best-practices), providing agents with model relationships and source metadata. This metadata transfers to the agent through the [Discovery API](https://docs.getdbt.com/docs/dbt-apis/discovery-api?version=2.0&name=Fusion), giving the agent a logical map to discover datasets, trace column-level lineage, and audit the origin of its inputs. ### Fast, low-cost development with the dbt Fusion engine The[ dbt Fusion engine](https://docs.getdbt.com/docs/fusion) validates SQL models against your data definitions locally as you write them. It parses projects up to 30x faster than dbt Core and cuts warehouse compute costs by up to 30%, so data teams can iterate quickly and ship agent-ready data without the build/break/fix cycles that slow production deployments. ### Built-in data quality dbt implements [data quality checks](https://docs.getdbt.com/docs/build/data-tests) at transformation time, so that agents never act on corrupted records. It supports singular data tests, like custom one-off SQL queries that check highly specific business logic, and generic data tests, which function as parameterized, reusable macros applied across multiple columns via YAML configurations. ## Conclusion Traditional batch data architectures cannot support autonomous agents due to a lack of data freshness, semantic grounding, and standard interfaces. Using raw data without metadata causes agent model hallucinations and broken pipelines. An adaptive data infrastructure architecture offers a reliable, machine-readable environment for agents. It centralizes semantics, automates validation, and standardizes interfaces, giving agents the reliable foundation they need to succeed in production environments. The dbt platform provides a governed, semantically consistent data foundation that agents need to operate reliably from day one.[Get started with dbt](https://www.getdbt.com/) today. --- --- title: "Are you ready for the dbt Fusion engine?" description: "A practical look at how Brooklyn Data’s Fusion Readiness Assessment helps teams plan a migration with confidence." url: "https://www.getdbt.com/blog/are-you-ready-for-the-dbt-fusion-engine" date: "2026-05-20" authors: ["Michael Carlone"] categories: ["Insights"] --- # Are you ready for the dbt Fusion engine? _This guest post comes from Michael Carlone at Brooklyn Data._ The dbt Fusion engine raises the ceiling on what data teams can build, and how fast they can build it. Fusion delivers faster, more responsive dbt development. It catches SQL errors in your IDE before anything hits the warehouse, and gives AI coding tools the project context they need to generate accurate code. For most organizations, the question is not whether Fusion is valuable. It is whether their current foundation can support the move successfully. That is why Brooklyn Data built a **Fusion Readiness Assessment**. It helps teams evaluate what has to be true across their project, their delivery practices, and their team structure to migrate with confidence instead of guesswork. ## Why readiness matters now Without the dbt Fusion engine, teams working with dbt often rely on warehouse compute just to catch issues that are relatively small: syntax errors, broken references, and type mismatches. That slows development down and adds cost to work that is often just part of normal iteration. Fusion improves that loop by giving teams faster feedback as they work. The result is a development experience that feels quicker, more responsive, and better suited to the way modern analytics teams build, and is designed for an agentic development experience. That upside is exactly why readiness matters. A move to Fusion does not only test the engine. It also tests the broader system around it: project design, development and deployment habits, governance, ownership, and technical depth. ## What “Fusion ready”‌ means Brooklyn Data defines Fusion readiness in a simple way: you are Fusion ready when you have confidence that your data foundation can support the move smoothly, securely, and at scale. That confidence should not come from instinct alone. It should come from a clear view of how your organization works today and where migration friction is most likely to appear. That is also why readiness is not a universal checklist. Every data team operates with a different level of maturity, a different tolerance for change, and a different mix of technical and organizational complexity. A good assessment should reflect these realities. ## How the assessment works Our assessment looks at readiness across three pillars: **Project, Process, and People**. We use these three because, in practice, they shape the outcome of nearly every dbt migration. Looking at all three pillars helps teams avoid blind spots. Migration risk rarely lives in one place, and focusing only on the dbt project itself can cause teams to miss process gaps, unclear ownership, or concentrated expertise that could slow the migration once technical work begins. Even when the technical foundation is ready, the transition can still become harder to manage if deployment practices are inconsistent, governance expectations are unclear, or key knowledge sits with only a few people. Assessing Project, Process, and People together gives teams a more complete view of what needs attention before they move forward. ### Project Within **Project**, we look at the technical patterns that affect migration complexity. That includes how standardized the project is, how much custom logic is in play, how heavily the team relies on advanced features, and whether the broader foundation reflects dbt best practices. ### Process Within **Process**, we look at governance, delivery, and repeatability. Can the team develop, review, deploy, and govern work in a way that reduces risk? Are the workflows consistent enough to support change without introducing avoidable instability? ### People Within **People**, we look at organizational clarity and depth. Does the team know who owns what? Do they have the hands-on knowledge required to support a migration? Is the organization positioned to learn, adapt, and handle issues without over-relying on a single person? Each pillar is assessed, quantified, and then rolled into an overall Fusion Readiness score. The goal is to make the output easy to understand and useful in conversation, not to create a false sense of precision. In practice, the assessment is meant to answer a short list of planning questions: - How ready are we today? - Where is migration friction most likely to show up? - What should we focus on first? - What should improve before we reassess? ## Why the pillars need to be read together The assessment looks at Project, Process, and People separately, but the real value comes from reading them together. A dbt project may be technically mature, but if delivery practices are inconsistent, migration can still become harder to manage. A team may have established governance and release patterns, but if ownership is unclear or hands-on expertise is uneven, execution can still slow down. And a capable team should not have to compensate for project issues that could be surfaced and addressed earlier in the process. That is why we do not treat any one pillar as the answer. Fusion readiness is shaped by how the technical foundation, the operating model, and the team support one another. Looking at them together gives a more realistic view of where migration will be smooth, where friction is likely to appear, and where focused improvement will have the biggest impact. ## What the output helps teams do The value of the assessment is not just the score. It is the interpretation behind it. When the results come back, teams can see where readiness is more established, where there are gaps to address, and what should be prioritized before moving forward. In some cases, that may point to project cleanup. In others, it may highlight process refinement or the need for clearer ownership and broader hands-on expertise. That makes the output useful at multiple levels. Practitioners get a clearer view of the technical and operational work that will reduce migration friction. Leaders get a shared language for planning, prioritization, and risk management. Instead of relying on instinct or general enthusiasm, teams can make decisions based on a more complete view of their starting point. A good readiness assessment does not just tell you where you stand. It helps you decide what to do next. Fusion is the destination. Readiness is the roadmap. Fusion represents a meaningful step forward for dbt teams. The speed improvements are real. The developer experience is materially better. The long-term upside is clear. But the teams that will benefit most are not the ones that move first. They are the ones that understand what they are moving from, what needs attention before the transition, and how to make that move with confidence. That is the role of readiness. Our Fusion Readiness Assessment helps teams take that first look inward, quantify what matters across Project, Process, and People, and turn that picture into a more practical migration plan. If your team is thinking seriously about Fusion, the smartest first step is not to assume readiness. It is to assess it. ## Want to know your Fusion Readiness score? We can walk through the assessment with your team, quantify readiness across the three pillars, and identify the next steps that will make a migration smoother and lower risk. [Measure your readiness here](https://www.brooklyndata.co/campaigns/dbt-readiness-assessment) --- --- title: "Get dbt certified. Stay certified. Stay ahead." description: "Get dbt certified -- and stay that way. Here's why certification matters for your career and how to earn it." url: "https://www.getdbt.com/blog/get-dbt-certified-stay-certified-stay-ahead" date: "2026-05-20" authors: ["Laurent Goldsztejn"] categories: ["Learn"] --- # Get dbt certified. Stay certified. Stay ahead. dbt keeps moving. The practitioners who build with it should move with it. Analytics engineering has matured quickly. Employers expect more, projects are more complex, and the bar for what good looks like keeps rising. If you work with dbt, getting certified isn't just a nice-to-have. It's how you prove you're operating at the standard the industry now demands. ## Why get certified? Because credibility matters, and proof beats claims every time. The [**dbt Labs certification program**](https://www.getdbt.com/dbt-certification) offers two exams designed for the practitioners actually building data pipelines, writing models, and owning data quality in production. Whether you're an analytics engineer, a data analyst moving into engineering, or a dbt platform power user, there's a path built for where you are. Passing the exam tells employers, clients, and collaborators something a resume line can't: you've been tested against a real standard, and you passed. It helps you: - Stand out as an expert in a crowded field of dbt practitioners - Build trust faster with new teams and clients - Move into senior roles with a credential that backs you up - Show your employer that their investment in you is paying off as you continue learning ## Two exams. One standard. The [**dbt Analytics Engineering Certification**](https://www.getdbt.com/certifications/analytics-engineer-certification-exam-version-1-11) tests your ability to apply dbt's core framework—modeling, testing, documentation, and deployment—the way it's actually done in production environments. The [**dbt Architect Certification**](https://www.getdbt.com/certifications/dbt-architect-certification-exam) tests your ability to design secure, scalable dbt implementations, with a focus on environment orchestration, role-based access control, integrations with other tools, and collaborative development workflows aligned with best practices. Together, they cover the full range of what it means to work with dbt at a high level. By getting certified, you’ll join 3,100+ other dbt developers and architects who have earned the credential. ## Get certified on dbt Because relevance matters as much as experience. dbt evolves. Best practices shift. An active certification tells the world you're keeping pace, and not coasting on knowledge from two years ago. If you let your certification lapse, it doesn't erase what you know. But it weakens the signal at exactly the moment someone is deciding whether to trust you with their data stack. Stay certified to show you're current, not just experienced. Committed, not complacent. ## This is your moment. This isn't starting over. With the [**dbt Analytics Engineering Exam**](https://www.getdbt.com/certifications/analytics-engineer-certification-exam-version-1-11) updated to align with [dbt Core 1.11](https://www.getdbt.com/blog/dbt-core-v1-11-is-ga), it's your chance to close any gaps, get hands-on with what's changed, and prove your skills against the standard that matters today. To sharpen your preparation, we’re including 10 sample questions on the web page and in the study guide so you know what to expect before you sit for the exam. And when you pass the exam again to regain your credential? It'll mean even more. ## Keep your edge Certification isn't a one-and-done achievement. It's a standard you maintain, and a signal to everyone you work with that you take this craft seriously. So whether you're earning it for the first time or earning it again: Stay certified. Stay relevant. Stay ahead. [Stand apart with dbt Certification.](https://www.getdbt.com/dbt-certification) --- --- title: "AI-ready data in practice: What dbt Semantic Layer and dbt's MCP server and agent skills do for your team" description: "Clean data is just the start. See how dbt's semantic layer, MCP, and agent skills give AI the business context it needs." url: "https://www.getdbt.com/blog/ai-ready-data-in-practice-what-dbt-semantic-layer-and-dbt-s-mcp-server-and-agent-skills-do-for" date: "2026-05-19" authors: ["Stephen Thibeault"] categories: ["Insights"] --- # AI-ready data in practice: What dbt Semantic Layer and dbt's MCP server and agent skills do for your team When it comes to getting their data AI-ready, many organizations start with cleaning and structuring their data and then simply stop. This is an important first step, but it’s not the last step, because AI-ready data relies heavily on context: the layer of meaning that explains what your data‌ actually represents. You need to gather as much information as you‌ can about that data: Where are data points coming from? Which team defines the metric? Which team owns inputting this data into a system? Without answers to questions like these, even clean, well-structured data can lead AI astray. One way to think about AI is as a great teammate that knows SQL and analytics really, really well but knows zero about your organization. An agent doesn't know the different acronyms used in your industry, for example, and it doesn’t understand your business goals. For AI to work effectively and efficiently, you need to give it all that important context to make the data meaningful. In practice, teams use dbt’s AI capabilities to make data meaningful to AI agents. dbt lives on top of tools like Snowflake, BigQuery, and Databricks to transform data without having to use stored procedures or other data transformation techniques, and there are three key pieces to dbt’s AI stack: the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=1.12), and [dbt agent skills](https://docs.getdbt.com/blog/dbt-agent-skills?version=1.12). Here’s what they are, how they work together, and how to use them to ensure high-quality, AI-ready data. ## The semantic data layer is your lens The semantic data layer provides all of the context that the AI will need to understand your data: the structure of the data, how you work with the data, and what exists in the data. I think of it like this: I have very bad eyesight. When I take my glasses off, I can still see things, but they are far from in focus. There will be some things that I miss and other things that are incomplete in my vision because I can't fully see everything. When I put my glasses on, I'm able to see clearly and completely. This is essentially what a semantic layer does for your data. A generic semantic layer is like buying plain, off-the-rack reading glasses. It makes things somewhat clearer; you will get answers some, but not all, of the time, and you’re not getting the most detailed vision possible. A governed, [dbt-backed semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) gives you prescription lenses that are custom-focused for your business's vision, signed off by someone trusted, and updated through scheduled exams as your vision (your data, your definitions, your business) change. AI wearing drugstore readers might see something somewhat clearly, but it'll squint and need to occasionally guess. AI wearing your prescription sees exactly what your business means by "revenue," "active customer," or "churn" and keeps seeing correctly as those definitions evolve. So when we talk about gathering context around data, most of that context is typically handled within the semantic layer. This is especially true when it comes to what certain columns mean, what certain metrics are, and how different values or properties are to be calculated. ### You don't need a perfect semantic layer to start You can get a lot of use out of dbt’s AI tooling even without a semantic data layer in place. The semantic layer is mainly used for conversational AI, letting agents query your actual data and return reliable AI outputs. But if you want to use dbt's AI tooling for development workflows, you don't need it. There are still things that you can do with dbt's AI tools outside of it, like diagnosing job failures, finding column-level lineage, and other things that really speed up your workflow. Don't let not having your data fully cleaned up, or not yet having your data fully defined in the semantic layer, be what stops you from using dbt’s AI tools. You can absolutely start using them now, and you can even use some of them to help build your semantic layer as you go. ## Three pieces of the dbt AI stack Terms like "agent skills" and "MCP server" can be‌ intimidating when you first hear them. Let's demystify these. **MCP server: the tools.** An MCP server is a set of tools like API calls that can be used to communicate with applications on the backend. Its function is to give the agent instructions on how to make those calls and how to use what it gets back. For example, there's a tool called **list_metrics** used to pull data from the semantic layer, and another one called **get_job_run_error** for diagnosing failures available as functions in the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=1.12). The dbt MCP server grounds those interactions in structured, dbt-native context, so agents are working from what your data actually means, not guessing from static documentation. **Agent skills: the instructions.** [dbt’s agent skills](https://docs.getdbt.com/blog/dbt-agent-skills?version=1.12) are workflow instructions that give your agent proven, opinion guidance for common dbt tasks like writing tests, debugging failures, defining metrics, handling migrations. They load on demand and only when relevant. An agent skill gives the agent a set of clear instructions needed to complete a specific task. Skills also provide the agent with rules and guardrails: never do this; here are common pitfalls you may run into; here are things that you need to look out for. **How the semantic layer, MCP, and agent skills fit together:** Each piece has a distinct role, and together they cover everything an agent needs to work effectively with your data. The semantic layer provides the context, MCP provides governed access, and agent skills provide the proven workflows agents need to query the data or to get the tools they need out of the MCP server. ## dbt’s AI tools in production to speed up data development The best way to understand how these three pieces work together is to see them in action. One of our clients, a very large technology company, used them to feed structured data into a Slack channel where dbt errors are automatically sent. They hooked dbt's MCP server, along with Claude, into that error triage channel to look at those job failures and actually diagnose them. The integration uses the **get_job_failure** function in the dbt MCP server, looks at the error, and then has the agent analyze what happened and why. By the time a developer actually gets to that error they're able to see a quick triage that was already done, along with some possible solutions. This integration is not fully set up for self-healing just yet. There are definitely controls around the AI, and it doesn't get everything right all of the time, but it's a huge time save. Instead of having to go into the dbt platform and dig through the logs to find the specific problem, you have it all laid out there by your agent. That same team is also working on a GitHub action: if somebody creates a model and doesn't include a semantic layer definition, the agent will try to create one and send it back to the developer with a note: _here's what I created, add on to it to make your semantic layer._ The goal is to encourage that hygiene of getting that context as a natural part of the workflow, rather than an afterthought. And, notably, both of these are use cases that don't require a semantic layer at all. ## Where to start: pilot small and smart If you're ready to include AI in your data pipelines, the most important advice I can give is to do it in steps. Really hone in on one business unit that is willing to work with you on a pilot program for AI readiness, and focus on gathering semantics around the data for that small subset. (Pilots within the data team itself, like the error triage example above, are a great place to start. They can be very useful, and they don't require a well-crafted semantic layer to work effectively. So there's no reason to wait!) Gathering that semantic information, though, will really allow you to get your feet under you when it comes to building a semantic layer, and it will allow you to iterate very quickly. When you collaborate with one team in a pilot project, you're able to break things and learn from your mistakes before bringing it out to more business units. So: start small, really focus in on what you're able to do (and what you reasonably _can_ do), and then apply what you learned. Then you can use the momentum you gain by providing something great to that particular team or business unit to expand the semantic layer to more teams across your org. ## Why semantic standards matter: Open Semantic Interchange Once you’re ready to build out your semantic layer it’s important to understand that, right now, basically every data tool implements semantic definitions in its own proprietary format.. Power BI has one, Omni has one, Databricks has one, Snowflake has one, and of course [we have one](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl?version=1.12). That fragmentation creates a portability problem: if your semantic layer definitions (metric names, calculations, business logic) are expressed in a format that’s specific to one tool, you can't move them to another tool without rebuilding everything from scratch. So, for example, if you define "monthly recurring revenue" in dbt's Semantic Layer and then want to also expose that definition in Power BI or Snowflake, you'd have to redefine it natively in each system. Besides redundancy it also creates inconsistency risk and a lot of maintenance overhead. This is why dbt, along with Snowflake, Databricks, and a large number of other major organizations in data, have joined an initiative called the [Open Semantic Interchange](https://open-semantic-interchange.org/). The v0.1 a vendor-neutral spec is already live and open source. It’s an industry-wide specification that standardizes how we exchange semantic metadata across analytics, AI and BI platforms. The OSI spec serves as a common language for metrics, dimensions, and relationships, so metrics can be interpreted consistently across tools (e.g., Snowflake, Tableau, dbt) while minimizing vendor lock-in. The dbt Semantic Layer complements the spec by making those definitions operational: you define and govern your metrics in the dbt Semantic Layer using MetricFlow, and OSI provides the interchange format to move those definitions across other tools like Snowflake and Tableau. Author once, use everywhere. Think of it‌ like the same reason we needed MCP in the first place: when there's no common standard, every tool reinvents the wheel and nothing moves cleanly between systems. A shared standard changes that. --- --- title: "What's shipped in dbt — May 2026" description: "A roundup of everything we've shipped since January—across agents, Fusion, security, developer experience, dbt Core, and more." url: "https://www.getdbt.com/blog/what-s-shipped-in-dbt-may-2026" date: "2026-05-19" authors: ["Corinne Hallander"] categories: ["Product"] --- # What's shipped in dbt — May 2026 It's been a big few months of shipping at dbt. We've got a lot to cover — from the dbt Developer Agent going into preview, to making the upgrade to the dbt Fusion engine self-serve, to new ways to lock down your account security, to quality-of-life improvements for practitioners who live in the IDE. Here's everything that's landed since January. ## AI that works with your data, not around it ### dbt gets an AI-native developer: the dbt Developer Agent (Preview) General-purpose coding agents are now everywhere, ready to help anyone code. But the question we kept hearing from teams this year was some version of: can we get an agent that actually works like an analytics engineer? One that‌ understands my whole dbt project? One that can read the graphs, knows the lineage, validates before it touches anything, and helps me build dbt models without breaking anything? This is why we’ve built the dbt Developer Agent, which is now available in Preview for dbt platform customers with dbt Copilot enabled. Simply describe the change you want to make — rename a model, add a metric, migrate a stored procedure, fix a failing build — and the agent reads your graph, understands what's upstream and downstream, and drafts the edits across every file that needs to move. SQL, YAML configs, tests, documentation: coordinated changes in one pass. That means less time context-switching between files, fewer broken builds, and data work that ships faster. → [Read our full announcement blog to learn more](https://www.getdbt.com/blog/the-dbt-developer-agent-is-now-in-preview) ### dbt Agent Skills - GA Earlier this year we released [dbt Agent Skills](https://github.com/dbt-labs/dbt-agent-skills) — an open-source repository of best practices that teach generalist coding agents how to think like an analytics engineer that actually understands how to work with dbt projects. Skills are structured knowledge files that agents load on demand. They encode things like: when to preview data before writing tests, how to structure a semantic model, how to debug a job failure without chasing the wrong root cause. Check out our growing repository of skills by clicking below: → [dbt Agent Skills on GitHub](https://github.com/dbt-labs/dbt-agent-skills) ### Securely connect dbt to your favorite AI tools (Beta) The dbt MCP server now supports OAuth, so you can now connect OAuth-enabled AI tools — Claude, ChatGPT, Glean, and others — to dbt using your existing dbt login. No token management, no configuration hand-off to an admin. Your identity, properly permissioned and secure, in a few clicks. [→ OAuth integrations docs](https://docs.getdbt.com/docs/platform/manage-access/connect-apps-oauth) ### Remote MCP Server: Admin API support + product docs tools Two new sets of tools landed in the dbt Remote MCP Server. First, the MCP server now supports Admin API calls — which means AI assistants (Claude, Cursor, etc.) can help troubleshoot job errors directly, not just write queries. Second, the MCP server now includes search_product_docs and get_product_doc_pages tools that pull from docs.getdbt.com in real time, so you get answers grounded in the actual docs rather than training data. → [dbt MCP repo](https://github.com/dbt-labs/dbt-mcp) ### Bring your own Anthropic key dbt Copilot now supports BYOK (bring your own key) for Anthropic, so teams can power their AI workflows in the dbt platform using their own Anthropic API key — with the usage, cost, and data handling that comes with it. BYOK is also available for OpenAI and Azure OpenAI, giving teams flexibility to build with the model provider that fits their security, compliance, and cost requirements. → [Read the docs to learn more](https://docs.getdbt.com/docs/platform/enable-dbt-copilot#configure-your-ai-provider) ## Getting to Fusion just got a lot easier The big headline on the Fusion side this cycle is that adoption is now self-serve in dbt platform. By accelerating your upgrade to Fusion, you can take advantage of 30x faster parsing time, richer metadata for AI, realtime feedback on SQL as you type, and more. But upgrading your projects manually one-by-one could take hours or days…why not let dbt do the hard parts for you? ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/713a573e9671b3ed3b7351cf89a66b8499c12b92-1478x926.jpg) ### Upgrade to Fusion project by project If you're a dbt platform customer, you can now see which of your projects are eligible for Fusion and move them one at a time, directly from the platform UI. Pick a project, follow the prompts – no ticket, no wait, no overhead. ### Fusion migration skill (Beta) Upgrading to Fusion shouldn't mean fixing conformance errors manually. The Fusion migration skill in the dbt Developer Agent brings an automated approach to getting your projects Fusion-ready, faster: - It classifies every conformance failure, - Applies only validated high-confidence fixes automatically, and - Walks you through medium-confidence changes with clear diffs and your approval. Blocked issues– those caused by Fusion bugs or framework limitations– are surfaced immediately, with context and a path forward. No wasted effort chasing unfixable errors. The skill re-validates after every fix to handle cascading errors correctly and ends every session with a transparent report. This gives you faster triage, safer fixes, and trust in your upgrade. **How to get started:** 1. **In dbt Studio: **Find a job or project that’s ineligible for Fusion. Attempt the Fusion run so you can see the build conformance errors. **Studio will surface a new entry point directly in that conformance error experience** so you don’t have to dig through error logs. From there, launch the conformance skill and enjoy! 2. **Via VS Code: **The dbt VS Code extension now makes Fusion setup and upgrade significantly easier. When you're ready to upgrade your project, you can run the CLI onboarding flow in the terminal or let an AI agent handle it via the dbt Agent Developer or Cursor, no command line required. Start your seamless upgrade to the dbt Fusion engine: → [Learn more](https://docs.getdbt.com/guides/upgrade-to-fusion?step=3#step-1-start-the-upgrade-assistant) ### More from Fusion this cycle: Beyond easier adoption, we've invested in making the engine faster and more capable. - **UDF-aware deferral.** When you run with --defer and --state, dbt now resolves function() calls from the state manifest — so models that depend on UDFs don't require you to rebuild those functions in your current target first. - **Python UDFs** are now supported on Snowflake and BigQuery in the Fusion engine CLI. - **DuckDB support** **(Beta)**. Run local dbt projects without a warehouse account. Useful for testing, exploration, and CI scenarios where warehouse costs matter. - **Apache Spark 3.0 (Beta)**. Fusion engine CLI support for Spark means faster compilation and execution for Spark-based dbt projects – no Python runtime, no subprocess overhead. For dbt platform customers: - **dbt compare** **from local dev to CI**. You can now compare changes at every stage of your workflow. In local development, the dbt VS Code extension previews how your edits affect your data (added/removed rows, join verification) before you open a PR. Then at the CI stage, dbt compare runs in orchestration on Fusion, giving you model-level diffs as part of your pipeline gate automatically. - [**Fusion release tracks**](https://docs.getdbt.com/docs/dbt-versions/cloud-release-tracks?version=2.0#fusion-release-tracks) give you control over your update cadence: Nightly, Stable, Extended, and Fallback. Choose the release track that matches your team’s stability requirements, risk tolerance and change management processes. - **New projects default to Fusion Stable**. New environments in Developer, Starter, and Enterprise accounts now provision on the “Fusion Stable” release track by default – for any supported adapter (Snowflake, Redshift, BigQuery, Databricks). Want to fast-track your migration to Fusion? Use our quickstart guide. [–> Quickstart guide for Fusion](https://docs.getdbt.com/guides/fusion?step=1) ## For dbt builders: Developer experience improvements This cycle we focused on the things practitioners have been asking for: faster navigation in the IDE, more context at a glance, broader warehouse support for query history, and a meaningfully simplified semantic layer spec. ### Studio IDE: search, replace, and command palette The Studio IDE now has search and replace across your project, a command palette, and the ability to jump to symbols and run IDE configuration commands. These capabilities have been long-requested, and now they're here. ### Studio IDE: Better status bar The status bar now surfaces deferral settings, dbt version, and project status with quicker access to change them. ### Model query history: Databricks and Redshift — Bet**a** ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/2f97a4089ef8dbe3d16d7510f05d251cb2bed244-2048x1117.jpg) Model query history now supports Databricks and Redshift in addition to Snowflake and BigQuery. If you're on either of those warehouses and want to understand query patterns at the model level, this is now available in beta. →[Read the docs](https://docs.getdbt.com/docs/explore/model-query-history) ### New semantic layer YAML spec ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0391c15f8ba49aa763bf214fe2e62de15fc028ac-1698x1276.jpg) The new [semantic layer YAML specification](https://docs.getdbt.com/blog/modernizing-the-semantic-layer-spec?version=2.0) introduces several key changes: semantic models are now embedded within model YAML entries (no more managing entries across multiple files), measures are now simple metrics, and frequently-used options are promoted to top-level keys. This is a meaningful spec simplification making it easier for anyone maintaining a semantic layer, and a lower barrier to adoption for those who haven't yet. The new specification is live in dbt Core v1.12 and on the dbt platform “Latest” release track. →[ Migrate to the latest YAML spec](https://docs.getdbt.com/docs/build/latest-metrics-spec?version=2.0) ## ## Access to dbt that’s secure, governed and self-serve We shipped several updates this cycle to make security configuration simpler — and in most cases, self-serve. ### Global login — GA ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/baf87657a055672e2bb41584ad2f6d2faa415aed-2048x1183.jpg) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/2fd33a5b1272ef7ac83d2653a6d3110eda2b2cdc-1910x1080.jpg) There's now a universal login URL that shows all the accounts you have access to across regions and tenancies, in one place. This is available now for multi-tenant accounts with an account-specific domain; single-tenant support is coming soon. →[ Log in to dbt platform](https://login.dbt.com/) ### Self-serve private endpoints — Beta You can now configure Snowflake PrivateLink endpoints directly in the dbt platform without filing a support ticket. Go to **Account settings → Integrations → Private endpoints** to request and manage Snowflake PrivateLink endpoints on AWS. If establishing secure connectivity for your dbt setup has been a multi-week support ticket process, that changes now. →[Read the docs](https://docs.getdbt.com/docs/platform/secure/private-connectivity/aws/aws-snowflake) ### Connection profiles — GA ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/2d593ec401c5b3c18b3e50281e4cd60a0ed92db4-1428x1482.jpg) Profiles let you define and manage connections, credentials, and attributes for deployment environments at the project level. dbt automatically creates profiles for your existing projects and environments, so there's nothing to migrate. Useful for teams that want more structured control over how credentials and connections are organized across environments. → [About profiles](https://docs.getdbt.com/docs/platform/about-profiles) ### Account-level Slack and Microsoft Teams notifications — GA ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/aee48940b197cb070aaef24e21ebabd7ed477232-1872x542.jpg) Job notifications can now be sent to Slack and Teams channels configured at the account level, not just per-job. This makes it easier to set up centralized alerting without touching every job's configuration. Both Slack and Teams notifications are now generally available. → [Slack notifications](https://docs.getdbt.com/docs/deploy/job-notifications#slack-notifications-account) · [Teams notifications](https://docs.getdbt.com/docs/deploy/job-notifications#microsoft-teams-notifications) ## dbt Core v1.12 is here in Beta The dbt language is continuing to evolve, and dbt Core v1.12 reflects that momentum. The beta release includes contributions from across the community. ### What's in v1.12: - New on_error config to control whether downstream models run when an upstream model fails. Set on_error: continue on a model to allow downstream nodes to still attempt to execute even when it errors. - Define project variables in root-level vars.yml to reference them within dbt_project.yml or to keep dbt_project.yml slim. - New selector method (selector:my_selector) to reference a named selector from selectors.yml inside --select or --exclude to combine with other selectors, graph operators, and set operators. - Support for the new semantic layer spec simplifies how you define metrics and dimensions by embedding semantic annotations directly alongside each model. - Expansions of user-defined functions (UDFs) - Use public third-party PyPI packages in your Python UDFs with the new packages config. - Write UDF logic in javascript. - Overloaded UDFs - define multiple functions with the same name but different argument signatures. - Execute ad hoc database statements (no macro needed) with dbt run-operation --sql - Improvements to exception handling so error messages are clearer and stack traces are easier to interpret. - and more coming soon! → [Learn more in the v1.12 upgrade guide](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.12?version=2.0) ## What’s next There's always more coming. Stay tuned on our blog for the latest announcements. In the meantime, the features above are live. If you have questions, _find us in [#product-updates](https://www.getdbt.com/community/join-the-community/) in the dbt Community Slack._ Or [contact us](https://www.getdbt.com/contact) to see what dbt can do for your data team. ### See us in San Francisco this June We’re at [Snowflake Summit June 1-4 (Booth #2112)](https://www.getdbt.com/events/snowflake-summit-2026) and [Databricks Data+AI Summit June 15-18 (Booth #430)](https://www.getdbt.com/events/databricks-summit-2026). We'll have live demos, the team on site, and a lot to show you. --- --- title: "AI-assisted analytics engineering: Docusign’s framework for scaling dbt unit testing" description: "How Docusign reduced dbt unit test authoring from 5 hours to 30 minutes using a structured AI-assisted framework." url: "https://www.getdbt.com/blog/ai-assisted-analytics-engineering-docusign-s-framework-for-scaling-dbt-unit-testing" date: "2026-05-18" authors: ["Sundar Subramanyam"] categories: ["Insights"] --- # AI-assisted analytics engineering: Docusign’s framework for scaling dbt unit testing _This guest post comes from Sundar Subramanyam, Lead Data Engineer at Docusign._ At Docusign, we support millions of customers worldwide in managing critical agreement workflows. As our analytics platform scaled to support new product launches and features, ensuring data quality before production data existed became a key challenge. Traditional dbt data tests (e.g., `not_null, unique`) rely on existing datasets. However, for new features and evolving pipelines, we needed **dbt unit tests**—tests that validate logic by mocking input data and asserting expected outputs. While powerful in theory, unit testing in dbt introduced a practical problem - **the effort required to manually author tests did not scale with the complexity of our models.** To address this, we explored a focused question: **_Can AI systematically reduce the friction of dbt unit testing?_** This led to the development of a structured approach using **GitHub Copilot (GPT-4+)** that significantly improved both **testing efficiency and adoption**. We used our own AI tooling to speed up unit test drafting by ~90%, while dbt provided the structure to govern and enforce those tests in CI, making data quality reliable even before production data existed. ### **The unit testing bottleneck** Unit testing in analytics engineering is critical for validating: - Complex `CASE` logic - Join conditions and fan-out scenarios - Filtering rules and edge cases However, the real challenge lies in the setup: For each model, engineers must: - Analyze SQL logic - Create mock input datasets - Manually compute expected outputs - Debug YAML syntax In practice, this process took **up to 5 hours per complex model**, making comprehensive testing difficult to prioritize. ## **Introducing the AI-assisted dbt unit testing framework** ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/4568838730aa4732f55c55c3505344144c4c2211-510x255.png) To address this, I developed the **AI-Assisted dbt Unit Testing Framework** — a structured, human-in-the-loop methodology that leverages generative AI to automate the creation of dbt unit tests. Rather than treating AI as a replacement for engineers, this framework positions AI as an **accelerator for repetitive tasks**, while preserving human validation for correctness. ## **Framework workflow** The framework follows a repeatable, multi-step process: ### **1. Model input** Engineers provide a dbt model (e.g., `dim_customer`) containing SQL transformations. ### **2. AI interpretation** A custom AI workflow parses: - Column-level transformations - Joins and filters - Logical branches (e.g., CASE conditions) ### **3. Logic summarization** The system generates a structured understanding of: - Source tables and references - Output columns - Transformation rules ### **4. Human validation** Engineers review and confirm the interpretation before proceeding, ensuring correctness and trust. ### **5. Test case generation** The framework generates: - Positive test cases - Negative test cases - Edge-case scenarios The AI focuses heavily on generating **synthetic mock data**, including: - Null handling - Boundary conditions - Join anomalies - Temporal edge cases ### **6. YAML output** The system produces a valid dbt unit test file (`*_unit_test.yml`) with: - Mock input datasets - Expected outputs - dbt-compliant structure ### **7. Iterative refinement** Engineers refine the generated tests and commit them into the dbt CI/CD pipeline. ### ### **Prompt pattern behind the framework** The core of this workflow was a structured prompt pattern rather than a one-off AI request. The prompt guided the AI through a repeatable sequence: - Interpret the dbt model logic. - Identify source references and output columns. - Summarize the logic and ask the engineer to validate the understanding. - Generate positive and negative unit test scenarios. - Create mock input data and expected outputs. - Ensure the `expect` section matches the model output columns. - Output the result as a dbt-compliant `_unit_test.yml` file. - Allow the engineer to refine the test cases through feedback. This structure helped make the workflow repeatable and reviewable, while keeping the engineer responsible for validating business logic and expected outcomes. ### **From SQL to test case** The real power is seeing how the AI handles mocking data. **Model SQL (Snippet):** SQL `CASE` `WHEN subscription_status = 'Active' AND renewal_date < current_date THEN 'Overdue'` `WHEN subscription_status = 'Active' THEN 'Current'` `ELSE 'Inactive'` `END as derived_status` AI-generated unit test: The framework generates test scenarios such as: YAML `unit_tests:` `- name: test_derived_status_logic` `model: dim_subscription` `given:` `- input: ref('stg_salesforce')` `rows:` `- {subscription_status: 'Active', renewal_date: '2023-01-01'} # Scenario 1: Overdue` `- {subscription_status: 'Active', renewal_date: '2025-01-01'} # Scenario 2: Current` `- {subscription_status: 'Pending', renewal_date: '2025-01-01'} # Scenario 3: Inactive` `expect:` `- rows:` `- {derived_status: 'Overdue'}` `- {derived_status: 'Current'}` `- {derived_status: 'Inactive'}` _The key advantage - The AI identifies logic branches and automatically generates test data to validate each scenario._ ### **Impact: 10x productivity and test coverage** The results of this small experiment were immediate and measurable: - **90% Reduction in cycle time**: Writing a comprehensive Unit test suite dropped from 5 hours to roughly 30 minutes. Engineers no longer start from a blank file - but they start with a working draft. - **Increased test coverage**: Because testing became easier, engineers tested more. We closed the gaps on edge cases that used to slip through manual review. - **Shift-left quality**: We caught complex logic bugs (mismatched joins, bad filters) locally, long before they reached the production dashboards. Implementing unit tests helped us catch at least 5–10 data defects that would otherwise have gone unnoticed. - **Scalable Trust**: Whether refactoring legacy code or building net-new models for new product and feature launches, we established a consistent baseline of quality without burning out the team. ### **What worked and what didn’t** **Where AI excelled:** - Parsing Jinja and SQL syntax to map logic branches. - Generating tedious mock data (rows of CSVs) in valid YAML format. - Identifying edge cases a human might overlook (e.g., "What if this date is null?"). **Where humans remain essential:** - Validating the _business intent_ of the logic. - Ensuring the "expected output" aligns with domain knowledge, not just code patterns. ## **Industry relevance and adoption potential ** The challenges addressed by this framework are not unique to a single organization. Many data teams struggle with: - Low adoption of unit testing - High manual effort - Inconsistent data validation practices This framework provides a **reusable and scalable approach** that can be applied across dbt projects and analytics engineering teams. ### **Looking ahead: From a win to a workflow** This initiative began as a focused experiment but has evolved into a repeatable pattern for integrating AI into analytics engineering workflows. Future directions include: - Integrating test generation into CI/CD pipelines - Generating tests from business requirements (e.g., Jira tickets) - Expanding the framework to other areas of data engineering ## **Conclusion** AI does not need to be complex to be impactful. By addressing a specific bottleneck—unit** test creation in dbt**—this framework demonstrates how targeted AI applications can deliver measurable improvements in productivity, reliability, and scalability. The broader takeaway: “Identify one friction point in your workflow—and use AI to systematically eliminate it.” --- --- title: "How Nasdaq built a governed intelligence layer with dbt and Databricks" description: "Nasdaq processes up to a trillion messages a day across 26 business lines. Here's why they use dbt and Databricks to do it." url: "https://www.getdbt.com/blog/how-nasdaq-built-a-governed-intelligence-layer-with-dbt-and-databricks" date: "2026-05-18" authors: ["Daniel Poppy"] categories: ["Insights"] --- # How Nasdaq built a governed intelligence layer with dbt and Databricks The stakes in financial services data are different from almost any other industry. In most business cases, the cost of a data error comes down to wasted time or a slightly wrong metric. In financial markets, a data break doesn't surface as an error message. It surfaces as a wrong regulatory filing, an incorrect client bill, or a risk desk making calls from numbers that were already stale. Jamie Nemeroff leads value engineering at dbt, where he quantifies what good data infrastructure is worth. In financial services, that calculation is easier to make than elsewhere because the cost of getting it wrong is so traceable. It's also what makes the story of what Nasdaq has built worth examining closely. Michael Weiss, AVP of Product at Nasdaq, leads Nasdaq Eqlipse Intelligence, an end-to-end data platform built on dbt and Databricks that Nasdaq originally developed for its own markets and now delivers to its financial market infrastructure (FMI) customers: the exchanges, clearinghouses, and central securities depositories (CSDs) that operate on Nasdaq's Eqlipse product suite. [Jamie spoke recently with Michael and Andrea DeSosa, who leads go-to-market for capital markets at Databricks, to discuss what they built, how they built it, and what it makes possible.](https://www.getdbt.com/resources/webinars/how-nasdaq-productized-a-governed-intelligence-layer-with-dbt-databricks-for-financial-market&sa=D&source=docs&ust=1778887535313379&usg=AOvVaw0hj8jcJYElXzLm3FrVUBqN) Four themes came up that Jamie tracks across every large financial services engagement he works on: scale, removing engineering bottlenecks, regulatory governance, and AI readiness. Nasdaq's story connects all four. And it's increasingly a product story: the architecture they spent a decade proving internally is now available to their customers in six months. ## Nasdaq Eqlipse Intelligence: architecture and origins Most people know Nasdaq as a market operator. A less-visible part of the company is its financial technology arm, which delivers infrastructure software to financial services firms globally, covering anti-financial crime, regulatory technology, and capital markets operations. Nasdaq's Eqlipse product suite focuses specifically on FMI customers. Michael described the intelligence platform as a data layer spanning the full lifecycle: ingestion from financial market protocols (Nasdaq-provided and third-party), transformation and mapping, validation and reconciliation, and business applications for reporting, analytics, and billing. Databricks serves as the primary computation layer. dbt plays two roles: transformation engine and [semantic layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl). On top sit three customer-facing products: InsightsHQ for visual dashboards, ReportHQ for structured report outputs like CSVs and PDFs, and RevenueHQ for managing billing and fees. The platform didn't arrive fully formed. Nasdaq has been on this path since 2012, starting with a regulatory data product for US broker-dealers. In 2014, the company moved all of its market data from on-premises warehouses to AWS. The intelligence platform was formally launched in 2021, extended to external FMI customers in 2024, and is now in active delivery to its first customers. What's notable about that timeline is the direction of proof. Nasdaq didn't build a product and then go find buyers. They built it for themselves, ran it at scale for years, and then recognized that the FMI customers they spoke with had the same problems they had already worked through. ## Scale at a trillion messages per day The volume Nasdaq manages helps explain why every architecture choice at this level carries weight. Michael described ingesting hundreds of billions of messages per day across Nasdaq's US and Nordic businesses, from thousands of data sources, across dozens of business lines. Regulatory compliance makes precision essential. US programs like the Consolidated Audit Trail (CAT) run on roughly a three-day window for getting data out. Billing has to reconcile to the exact contract. As Michael put it: "When we go to generate the bill at the end of the month, we can't have an unaccounted for or an extra contract in the bill we're sending. We need to make sure that information is right, both in terms of the input and the output." The approach Nasdaq uses to maintain consistency at this scale is what Michael described as a [dbt Mesh](https://www.getdbt.com/blog/data-mesh-architecture-explained) structure: a family of dbt projects organized from foundational models upward. The foundational layer defines contracts for raw data. Every message on an order chain looks the same, regardless of whether it comes from a Nasdaq market, a Southeast Asian exchange, or a non-Nasdaq trading platform. That consistent contract makes it possible to define metrics and KPIs once and deploy them everywhere. It also shapes how the team responds to errors. "Our stance," Michael said, "is we'd rather have the pipeline break with an error on the test and be late on sending data than send wrong data." Andrea framed scale from the Databricks side as a trust problem as much as a volume problem. "Every time a team builds its own pipeline and defines its own metrics," she said, "you've created a future audit finding, a future reconciliation break, or a future AI model trained on the wrong truth." Databricks' [Unity Catalog](https://www.databricks.com/product/unity-catalog) addresses this at the infrastructure level. The same definition of settlement finality or net position applies across all of Nasdaq's internal teams and FMI customers. That’s not because people agreed to use the same spreadsheet, but because the platform enforces it. ## Getting data to the people who need it Engineering bottlenecks are the second theme Jamie consistently work through in financial services business cases. The data exists. The business teams need it. But access runs through ticket queues, and the lag compounds. Michael described what changed at Nasdaq when they put dbt in front of non-engineering teams. Because SQL is broadly understood across the business, governed modeling tools could go directly to analysts and business users. The result was a significant acceleration in time to market for new data products and insights. The example he gave was concrete. Nasdaq's options sales team, working alongside the options business team that owned the underlying dbt models, built client-specific visuals to distribute as part of their sales process. No engineering involvement required. Nasdaq is now looking to offer the same model to its FMI customers: access to Nasdaq's foundational dbt models alongside governed tooling that lets customers' business teams build on top. "We're looking to let them take our foundational models and rebuild things from a business point of view that complement or supplement the models they need to make their business go." Andrea described why this approach compounds rather than complicates. "The biggest tax on delivery isn't talent," she said. "It's starting from scratch every time." A consistent foundation changes the unit of work: teams inherit a proven architecture and configure it to their markets. The people at FMIs who understand the business best, the clearing workflow specialists and surveillance officers, can start contributing to model development rather than waiting on engineering queues. The condition that makes this safe is governance. As Andrea said: "You can only safely put those tools in non-engineering hands when you have a governance layer enforcing the boundaries underneath." Unity Catalog handles that at the data level. dbt handles it at the model level. Together, they let more people build without fragmenting the foundation. ## Governance that holds up to regulators "Governance" can become abstract fast. In financial services, it's specific. Michael walked through two concrete examples: - Rule 17A for broker-dealer compliance in the US establishes a write-once, read-many (WORM) obligation. Firms must be able to prove that data wasn't manipulated after the fact. - SOC 2 compliance for billing means being able to show an auditor the full path from source data through every transformation to the fee applied and the invoice sent. "I can't go to a customer and say, here's a non-SOC 2 compliant billing solution," Michael said. "Every other billing solution on the planet is SOC 2 compliant." When compliance gaps appear, the cost runs beyond fines. There are regulatory filings that have to be amended, legal overhead, and time spent explaining and correcting. For Nasdaq's FMI customers, these same obligations apply. That means the platform Nasdaq delivers has to make defensible answers available on demand. Andrea described Databricks' role here as infrastructure-level auditability. Unity Catalog tracks every transformation, model, and dashboard: where the data came from, who touched it, when it changed, and what depends on it, automatically, without anyone needing to document it manually. "When an auditor or regulator asks," she said, "you have the answer." dbt adds the logic layer. Michael described pulling [a data lineage snapshot from dbt](https://www.getdbt.com/blog/what-is-data-lineage) to show a regulator the flow of a model from source to output. If there's a question about a specific calculation, the answer is in the documentation or the code itself. "dbt makes that part pretty effective and pretty easy," he said. It also helps customers building on top of Nasdaq's foundational models understand what those models mean and how they're calculated before making any modifications. ## Trusted data as the path to AI The AI conversation in financial services carries a particular weight. As Andrea framed it: "In financial markets, a confident but wrong answer is a risk event. It's not just an inconvenience." If a generative AI model hallucinates a product recommendation in a consumer app, a user gets a bad suggestion. If it hallucinates a position, a margin requirement, or a regulatory classification at a clearinghouse, that's a compliance breach. This is the gap the semantic layer is built to close. Michael described the root cause: AI gets things wrong not because models are incapable but because they lack context about the data they're working with. [The dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), which Michael called Nasdaq's "context layer," provides structured meaning alongside the data. When an agent accesses a metric, it knows what the metric means and how to apply it because the definition is explicit and governed, not inferred. dbt Labs [recently published a study on text-to-SQL accuracy](https://docs.getdbt.com/blog/semantic-layer-vs-text-to-sql-2026), comparing results with and without a semantic layer in place. The difference was stark: close to or at 100% accuracy when using semantic layer with both ChatGPT and Claude. Without it, agents querying raw tables produce confident answers that can be wrong in ways that are difficult to detect. Nasdaq is now building what Michael called an "agentic surface layer," investing in [Model Context Protocol (MCP)](https://www.getdbt.com/blog/mcp) tooling and skills that let customers build their own agentic workflows on top of the intelligence platform. The longer-term vision is a marketplace where Nasdaq, its partners, and its customers can share those workflows in a governed, verifiable way. "How do you do everything in a controlled environment such that you always know the result is guaranteed?" Michael asked. The foundation they've built is designed to answer that question, whatever the use case. When Jamie works through a financial services business case, the four themes he tracks—scale, productivity, governance, AI readiness—keep pointing back to the same thing: the work that satisfies a regulator today is the same work that makes AI trustworthy tomorrow. Lineage, contracts, semantic definitions, etc., aren't AI features. They're data infrastructure decisions. But they're the infrastructure decisions that determine whether AI agents can be trusted with the calculations financial markets depend on. Michael put it directly near the end of our conversation: "Delaying these initiatives is just becoming a bigger issue. There's a lot of focus on just trying to get there as quickly as possible." For Nasdaq's FMI customers, the option now exists to skip the three-year build and get there in six months. The combination of dbt and Databricks underneath the intelligence platform is a significant part of why that acceleration is real. [Watch the full recording](https://www.getdbt.com/resources/webinars/how-nasdaq-productized-a-governed-intelligence-layer-with-dbt-databricks-for-financial-market) to hear more of what Michael and Andrea covered. And if you're working toward the same kind of trusted data foundation, [talk to our team](https://www.getdbt.com/contact). --- --- title: "Ship smarter agents in production with dbt Agent Skills" description: "Just like humans, autonomous agents need faster feedback loops. Here’s how to develop them." url: "https://www.getdbt.com/blog/ship-smarter-agents-in-production-with-dbt-agent-skills" date: "2026-05-18" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Ship smarter agents in production with dbt Agent Skills Coding agents are doing a tremendous amount of useful work today. Since Claude Code dropped last year, followed by Opus 4.5 and GPT 5.2, software engineering has very clearly passed a phase change. We've gone from copilot-style autocomplete to agents that can run end-to-end across the SDLC. But anyone who's tried to point one of these agents at a dbt project has hit the same wall I have. Ask a coding agent to build a new dbt model, and it'll happily make five or six changes across your DAG, then try to run the new model at the end. It breaks. The agent didn't know which columns existed, didn't iteratively run queries as it walked the DAG, and didn't think about lineage or contracts. The technology is marvelous. The agents simply haven't been taught how to do data work yet. Without a governed foundation—your actual models, lineage, contracts, and metrics—a coding agent is working from guesswork. It can write SQL that looks right and still return numbers no one can verify. That's what we're fixing with [**dbt agent skills**](https://docs.getdbt.com/blog/dbt-agent-skills). ## Coding agents are generalist agents, including for data We call them coding agents, but that framing undersells what they are. dbt agent skills can do all types of work. One of those types is data work, which is where most of you reading this probably want to put them. The catch is that the agents have been specialized for coding workflows. There's a long list of small tweaks and improvements that make them slot neatly into a software engineering loop. Data has its own additional bits that haven't been baked in by default: understanding [data lineage](https://www.getdbt.com/blog/what-is-data-lineage), respecting contracts, iteratively running queries to validate as you go, and knowing when to materialize what. A lot of teams have been layering those in by hand with AGENTS.md files. But there's a ceiling to how big an AGENTS.md can get before it becomes its own problem. ## What agent skills are, and what dbt's are doing Agent skills are [a protocol Anthropic released late last year](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) and donated to an open foundation. They're how you give an agent context into specific processes and workflows it needs to know about, packaged as markdown plus optional supporting scripts that the agent loads when relevant. We've taken everything dbt Labs has learned about analytics engineering and ported it into a series of agent skills, [in an open repo](https://github.com/dbt-labs/dbt-agent-skills) that works with any agent supporting the protocol. You can think of agent skills as **dbt best practices ported directly into your agent**. That includes: - Building a model iteratively rather than one-shotting it - Writing unit tests - Understanding your Directed Acyclic Graph (DAG) - Building your semantic layer - Debugging incremental models (which any longtime dbt user has spent their share of time on) The goal is to have the entire [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) captured within skills, so the combination of a generalist coding agent and the dbt skills gives you a powerful data agent out of the box. That's part of the way there. The rest is your custom context, the things only your organization knows. Naming conventions, materialization choices, source quirks, gotchas. Skills are designed for that too: anyone can write one for their own project, and the strongest setups combine the general dbt skills with org-specific custom skills layered on top. ## How Factory scaled up with dbt Nikhil Harithas from [Factory](https://factory.ai/) has been building this out from scratch. Factory is in the business of bringing autonomy to software engineering through [Droids](https://docs.factory.ai/cli/configuration/custom-droids), their generalist agents that work across the SDLC, including the terminal, web, CLI, GitHub, and Teams. Nikhil's a field engineer there, and over the last few months, he's been standing up Factory's data posture using dbt as the central piece. Factory's materialization strategy changed as the team scaled up. Factory is a younger company, so the default was to materialize as late as possible, with a blanket rule that worked fine in the early days. As certain queries took long enough that things started getting expensive, and as the team wanted fresher data, that rule needed to change. The math, Nikhil notes, is both money and time, and also how many people you have to maintain the pipeline. Up until a few weeks ago, Factory was rebuilding entire tables on every single run because it was fine. The shift to incremental builds came because that was the only way to run more often without blowing up cost. Incremental builds are finickier than entire builds, so testing started to matter a lot more. As Nikhil puts it: "I wasn't as familiar with the [dbt test suite](https://docs.getdbt.com/docs/build/data-tests) until a couple of weeks ago, when I was like, okay, it's time to shore up all these kind of implicit contracts in these table definitions. We realized that every row has to have a distinct ID of some kind, depending on the table." Tests went on the most important tables first. The principle behind it: "How can you increase individual leverage as far as you can by systematizing as much as you can, by giving Droid the same kind of feedback you or I would?" ## Lessons learned from the build-out Nikhil’s top-line claim from the build-out: months of work that would have taken five or six people was done by one and a half people, in a couple of months. That alone is worth taking seriously. A few of the patterns he ran into: **Build vs. operate are different motions, and you need to teach the agent both.** Nikhil's framing: "What is it like to develop in dbt, and what is it like to operationalize in dbt?" Those things are related but slightly different. The historical reason agents have struggled with data teams more than they've struggled with generalist software engineering is, ultimately, a context problem on both fronts. Data is a mix of writing code and running operational workflows. It's closer to SRE work than pure software engineering in places. **Skill creep is real.** Early on, Nikhil ran into trouble with too many skills, which led to inconsistent triggering and ambiguity about which skill applied. The fix: Be intentional about how many skills you have, and make each one denser. Skills you're confident the agent will discover on its own can stay as habit-forming background. The high-value skills are the ones that anchor behavior on the most important tasks. **Say it louder in AGENTS.md.** Skills get auto-invoked sometimes, but you shouldn't bet on it. AGENTS.md is what's guaranteed to be in context, so Nikhil's pattern is to point at the relevant skills there explicitly: "You're going to be using [BigQuery](https://cloud.google.com/bigquery) and dbt and a few other vendors that [Extract, Transform, and Load (ETL)](https://www.getdbt.com/blog/extract-transform-load) data to us. You have to pay attention to the skills of these particular frameworks, because there's going to be how you do anything at all." **Build the repo so agents bump into skills naturally.** Even when a skill isn't auto-invoked, an agent grepping around to get its bearings should run into it. The mere mention of "dbt" anywhere in the project should surface the relevant SKILL.md in the search results. **Skill golf.** Doug Bady at dbt coined this. The practice: go through your skills and try to rip out everything you can to make them as tight as possible. Distractions in context cost you. **Documentation as a hook.** Factory now has a check that fires when someone changes a column or table. It requires documentation to land somewhere in the project before the change can merge. The agent doesn't have to write the docs by hand, but the structural requirement creates a flywheel where good behavior produces more good behavior. [To see these principles in action, watch a full demo of Factory’s Droids and dbt agent skills in our webinar.](https://www.getdbt.com/resources/webinars/ship-smarter-agents-building-for-production-with-dbt-agent-skills) ## Building faster feedback loops "Look at the last 20 years of software engineering,” Nikhil says.. “We’ve gotten faster because of feedback loops. It's easier to write tests, easier to write integration tests, logs are easier to look at. Humans are getting faster feedback loops. We have to give the same thing to agents." Nikhil's framing of where coding agents are today is worth leaning on. We're at a point, he says, where "if you can imagine it and if you are determined to build it, you can build it." That's a pretty magical thing to be able to say. But it feels like not everyone is experiencing that reality. The reason, in his view, is that while a lot of things that used to be difficult are now easy, some of the things that were hard are still hard. Integrations. Human context and assumed knowledge. Naming things, famously. If you're going to put one thing on your list this week, Nikhil's call was to figure out how to give your agents access to the most tools you can and the most amount of direct feedback you can. Nothing is more powerful than read-only access to the database. If an agent can't query the database, everything goes slower. The reframe he offered is the one I'd lead with: _What would need to be true for you to give an agent access to your database?_ That's the question. Maybe the work this week isn't a data engineering task at all. Maybe it's setting up your environment so you'd feel safe handing an agent that access. Once they have it, they fly. This is a singular moment in technological progression. The combination of generalist coding agents, dbt agent skills, the[ dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp), your own custom skills, and platforms like Droid is making it possible for very small teams to do work that used to require very large ones. Get involved. The [dbt agent skills repo](https://github.com/dbt-labs/dbt-agent-skills) is open, and we want to see what you build with it. --- --- title: "The dbt Developer Agent is now in Preview: the coding agent for analytics engineering" description: "dbt Developer Agent is now available in Preview—grounded in your dbt project so you ship faster without breaking downstream." url: "https://www.getdbt.com/blog/the-dbt-developer-agent-is-now-in-preview" date: "2026-05-06" authors: ["Chakshu Mehta", "Sam Ferguson"] categories: ["Product"] --- # The dbt Developer Agent is now in Preview: the coding agent for analytics engineering Every few years, the way we work with data evolves. We’ve gone from SQL editors to BI dashboards to conversational chat, and each shift moves data work forward. Now we’re entering a new phase defined by agents that can reason through tasks, plan, and act on their own. But despite being remarkably good at general software engineering, today’s coding agents struggle on dbt projects, because the project itself is the context and the guardrails. Without that grounding, they write SQL that looks right but references a column that doesn't exist, breaks a dialect rule, or silently breaks tests, contracts, and governed definitions three models later. The result is SQL that's syntactically correct and semantically wrong. Analytics engineers need an agent built for the job, grounded in your dbt project from the start. With today's launch, that's finally possible. The [**dbt Developer Agent**](https://docs.getdbt.com/docs/dbt-ai/developer-agent?version=2.0) is now available in Preview for dbt platform customers. It’s the next evolution of dbt Copilot, built directly in the Studio IDE, and it works across every file a change touches. That matters in analytics engineering, where the risk isn’t just a syntax error—it’s a change that slips past your project’s guardrails and breaks tests, contracts, or governed definitions downstream. No new tools to install. No context switching. Nothing to set up. It ships with [dbt Agent Skills](https://github.com/dbt-labs/dbt-agent-skills) and [dbt's product docs toolset](https://docs.getdbt.com/docs/dbt-ai/mcp-available-tools?version=2.0#product-docs) built in, so best practices and canonical answers are there while you build. Open dbt Studio to try it, or [read the docs](https://docs.getdbt.com/docs/dbt-ai/developer-agent?version=2.0#prerequisites) to learn more. ```json { "_key": "61029b8610f4", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/T5vRS9XSZSY" } ``` ## Agents that can think across your dbt project When we launched [dbt Copilot](https://www.getdbt.com/blog/introducing-dbt-copilot) in 2024, we were solving a real but bounded problem. Writing boilerplate, drafting tests, documentation, and YAML configs by hand slows everyone down without making anything better. dbt Copilot made those workflows much faster. But the thing that actually slows teams down isn't writing a model. It's what happens _around_ the model. You rename a column and something three dashboards deep breaks. You add a metric and spend an hour chasing down every semantic definition and exposure that needs to move with it. You open a PR and realize the last person who touched this file left your company six months ago. None of that is a coding problem. It's a data problem, and addressing that problem requires an agent that understands your data. During our beta, we spoke with dozens of customers, and kept hearing the same thing: they wanted an agent that actually knew their whole dbt project. One that could read the graph, understand the lineage, pick up on the contracts and the semantic definitions, and then make the change—without breaking the guardrails their teams depend on to keep data trustworthy. When something broke, they wanted it to troubleshoot the way a senior analytics engineer would: trace the failure, find the root cause, fix what's actually wrong instead of patching the symptom. So we built it. ## Meet the dbt Developer Agent The dbt Developer Agent lives right in dbt Studio, so you see every change in context directly in the IDE, right alongside the code you’re changing. In the IDE you can describe the change you want to make—like renaming a model, updating a column, or adding a metric—and the agent will analyze your full dbt graph. Not just file dependencies, but lineage, contracts, semantic definitions, tests, and governance. Then the agent drafts edits across those files and shows them to you as a sequence of reviewable diffs. You approve or reject each step, so only the changes you accept are saved to your project. > _“What really sets the dbt Developer Agent apart is its precision in identifying and resolving bottlenecks. It doesn't just suggest code; it understands our existing tests and lineage well enough to troubleshoot issues almost instantly. This has significantly reduced our build times and allowed us to scale our dbt project with total confidence in our data's trustworthiness. It’s the perfect balance of AI speed and strict architectural governance.”_ _Vishal Kaviraj, AVP Data Architect_ ## Grounded in your dbt project for faster, safer changes What separates the dbt Developer Agent from coding agents built for software engineering is two things: context and governance. Context is what makes an agent good at analytics engineering. The more of your project an agent can see, the safer its changes. Governance is what makes those changes trustworthy at scale, the tests, contracts, semantic definitions, and review controls that your team already depends on. Other agents built for data have pieces of that context. Warehouse-native agents read your schemas and write SQL within the warehouse they're built for. IDE-native agents (like Cursor, Claude Code) refactor across files and know your repo. Both are useful, and both are getting better every month. Some can even pull in pieces of your dbt project, but none of them are grounded in it. A dbt project isn't just SQL. It's years of accumulated decisions about how your organization's data should be shaped, tested, owned, and understood. That's the context that decides whether a change is actually safe, and that context lives in dbt. ## How it works The dbt Developer Agent runs on a loop, not a single shot. You describe what you want. It drafts the edits. It proactively asks to run `dbt compile` or `dbt build` to validate its own work. You stay in control with the ability to approve or deny each command as it goes. It sees the result, adjusts if it needs to, and keeps going until the change is ready for your review. it’s designed to help you move fast, but still work inside the same validation and review guardrails your data teams already trust. It also works with [Fusion](https://docs.getdbt.com/docs/fusion/about-fusion) out of the box. This ‌gives the agent a fast, local, deterministic feedback loop so it can validate changes (like missing columns, dialect rules, and downstream breakage) without waiting on expensive warehouse round-trips. A few things that make it great for developer workflows: - **It’s grounded in your whole project, not just the file you have open**. The agent understands your full dbt graph. When it writes or refactors a model, it understands what’s upstream, what’s downstream, and what a change means for the rest of the project. - **It keeps related files in sync.** A single prompt can produce coordinated changes across models, YAML configs, and documentation. If you rename a model, the refs follow. When you change a column, the downstream tests update with it. - **It ships with dbt Agent Skills.** Earlier this year, we launched [Agent Skills](https://docs.getdbt.com/blog/dbt-agent-skills?version=2.0), built by dbt Labs and the dbt community, that encode a decade of analytics engineering best practices. It brings the kind of knowledge that usually only lives in a senior engineer’s head, now available out-of-the-box to the agent, without any configuration. It also offers support for **project-level context skills**, enabling the Developer Agent to seamlessly detect and prioritize custom markdown skills alongside dbt Labs-managed ones. - **It ships with dbt's product docs.** So you can verify documented behavior and recommended patterns without leaving dbt Studio. - **It shows its work.** You see the agent’s reasoning and tool calls as it works, not just the final output, so if it takes a wrong turn you can see exactly where and why. You can copy the suggestion or open it directly in the editor. - **It keeps a human in the loop**. “Ask-for-approval” mode (the default) surfaces every edit as an inline diff before anything saves, while “edit-automatically” mode applies its work as it goes, with reasoning you can follow. You pick the right level of autonomy for the task, and nothing lands without your say. - **It validates as it goes.** A built-in comparison loop catches problems before you see the final diff, so agent-generated changes meet a higher bar than "the code compiles." - **It runs commands on your behalf.** Beyond just file edits, the agent can execute dbt commands, open pull requests, and handle the workflow steps around the change, not just the change itself. You approve each command before it executes, with options to allow it once for the session or deny it. ## On trust and autonomy We spent a lot of time thinking about how much autonomy a native dbt agent should have. The answer we landed on is that "it depends," and that the tool should support that flexibility rather than forcing a single workflow mode. - **“Ask-for-approval” mode (the default)** surfaces everything for review before saving. You see inline diffs, approve what looks right, and push back on what doesn't. Nothing happens without your say. - **“Edit-files-automatically” mode** writes and saves as it works, which can be right for ‌tasks where your team has enough confidence in both the agent and the change to let it run. Most teams will probably live somewhere between the two depending on the model, the stakes, and how familiar the work is. ## What you can do with the Developer Agent today The Developer Agent lives in the dbt Copilot panel in dbt Studio, so the whole loop of intent, change, validation, and review happens next to the code and lineage you're already working in. No context switching, no separate tool to learn. Here are a few of the things it handles well today: - **Refactor across files.** Describe the change, the agent reads your full graph, coordinates every file that needs to move, and surfaces the diffs for review. Rename a model and the refs follow. Restructure a mart and the downstream tests update - **Fusion migration.** For teams with projects blocked on conformance failures between dbt Core and Fusion, the agent classifies what's fixable, applies high-confidence fixes automatically, and surfaces what needs your input or is blocked at the engine level. - **Update models.** Describe a change in natural language and let the agent write or refactor the SQL, tests, and docs together, not separately. - **Enhance your semantic layer.** Add or modify metrics and dimensions with full project context. The agent knows your existing definitions and builds consistently with them. - **Migrate stored procedures.** Describe the logic you're moving off of and the agent translates it into dbt models, tests, and docs, with full awareness of where it fits in your existing graph. - **Create tests and docs.** Generate test coverage and documentation for existing models in a single pass, grounded in how each model is actually used downstream. - **Work without breaking your flow.** The Developer Agent lives in the Copilot panel, so you can stay in whatever file you're already editing and let the agent handle the related changes in the background. No switching to a chat tab to kick off work, then switching back to review it. > _We went from about 60 conformance errors to 7, using the Fusion migration agent. That's the difference between too hard and actually doable._" - Michael Fridolfsson, Data Architect, Brighte ## What's next This is just the beginning, and here are a few things we're actively building toward: - **Built-in data previews and data diffs** to improve validation and help you review the outcomes of agent-generated changes before they hit production. If your team works primarily in Claude Code or VS Code, the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) brings the same structured project context to the agent you already use. ## Get started The Developer Agent is available in preview for dbt platform customers with dbt Copilot enabled. Simply open dbt Studio, find the dbt Copilot agent pane, and describe what you want to build or change. “Ask for approval” mode is a good place to start: review the diffs, approve what looks right, and let the agent handle the rest. Not on the dbt platform yet? [Talk to our team](https://www.getdbt.com/contact) to learn more. --- --- title: "5 dbt MCP server patterns that work in production" description: "Five dbt MCP server patterns from real production use, including one that doesn't work the way you'd expect." url: "https://www.getdbt.com/blog/5-dbt-mcp-server-patterns-that-work-in-production" date: "2026-05-01" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # 5 dbt MCP server patterns that work in production The Model Context Protocol (MCP) crossed 97 million monthly SDK downloads last December, the same month Anthropic handed it to the Linux Foundation. [The dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2.0&name=Fusion) is one of the more widely adopted data MCPs in that ecosystem. Here's what practitioners are doing with it. Five patterns. Two that work well, one that doesn't do what people expect, and two for when you want to go further. **1. Conversational analytics against governed metrics (works well)** Point a Claude or ChatGPT interface at the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) through the MCP server. The server exposes your MetricFlow metrics, and the model queries them in plain language. When the model is grounded in governed definitions instead of writing raw SQL against undecorated tables, accuracy improves. **Where it shines:** metric-level questions against well-defined models. W**here to be careful:** complex multi-join queries still want a human in the loop. **2. First-draft documentation (works well)** Use the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2.0&name=Fusion) to generate documentation from column-level lineage and test coverage. The model reads the lineage graph, drafts model, and column descriptions, and you review them in the PR. Teams running this go from almost no coverage to most of their models documented. Setup takes a few hours. Treat the output as a starting point, not the final word. The review step is the whole game. **3. Real-time queries against very large tables (not what you'd expect)** Here's the one that bites people. Pointing an agent at full-scan queries against large, unpartitioned tables runs up warehouse costs and timeouts that you won't see coming at setup. The fix is simple: pre-materialize the semantic layer views you need, or add query guards before the agent runs anything. If your warehouse bill jumped and you're not sure why, start here. **4. CI/CD integration (advanced)** Wire the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2.0&name=Fusion) into your PR review. A pull request (PR) opens, the agent reads the diff, runs the impacted downstream tests through MCP, and writes a review comment. Teams that have set this up report much faster review cycles. This needs GitHub Actions or equivalent, plus a clear model impact map. There's a learning curve. It pays off. **5. Cross-tool orchestration (advanced)** Let an orchestration agent decide which dbt models to run based on upstream data quality signals. Source freshness fails, the agent skips the dependent models and pages the on-call engineer instead of running on stale data. Works in n8n and LangGraph. Not trivial to build, but you get a pipeline that degrades gracefully instead of quietly shipping bad numbers downstream. [**Check out dbt Wizard, the AI agent built for the way analytics engineers work.**](https://www.getdbt.com/product/dbt-wizard) --- --- title: "How Obie cut compute costs by 30%, reclaimed engineering hours, and built stronger governance" description: "See how Obie used the dbt Fusion engine and state-aware orchestration to cut costs, speed up pipelines, and scale with confidence." url: "https://www.getdbt.com/blog/how-obie-cut-compute-costs-by-30-percent" date: "2026-04-24" authors: ["Elaine Green", "Aika Zikibayeva"] categories: ["Product"] --- # How Obie cut compute costs by 30%, reclaimed engineering hours, and built stronger governance **** At Obie, an embedded insurance platform serving real estate investors, the data reality looked a lot like a typical fast-growing startup. With a lean footprint of just 130 employees and a recent, high-stakes acquisition, the pressure on the company's data platform was immense. Behind the scenes, infrastructure costs were climbing and starting to create friction across the business. The team adopted the dbt Fusion engine as their transformation engine and deployed state-aware orchestration (SAO) to remove operational drag and ensure confidence in the numbers the business depended on. When the team started, Obie's data layer was built on a Data Vault 2.0 methodology. This is a pattern designed for large enterprises. For a smaller fast-moving insurance tech company, it slows development and results in overhead that doesn’t match a smaller tech company’s size or pace. Obie made the call to re-architect entirely on Fusion, converting legacy models while preserving the underlying business logic. That migration is now over 90% complete. ## SAO reduces warehouse costs by ~30% and makes more frequent refreshes possible Previously, large portions of the pipeline were rebuilt whether upstream data had changed or not. As data volumes increased, that approach became expensive. Moving to [SAO on Fusion](https://www.getdbt.com/blog/announcing-state-aware-orchestration) meant running only what was necessary. Compute usage dropped, and warehouse spend became more predictable. The team could refresh data more frequently, from daily to every two hours, [without worrying about runaway costs.](https://www.getdbt.com/blog/dbt-compute-cost-reduction-fusion-state-aware-orchestration?utm_source=chatgpt.com) "We're saving at least 30% on compute costs, just from reusing models with state-aware orchestration,” said Tyson Doberneck, senior data engineer. Matt Karan, senior data engineer, added, “Knowing we're being proactive about costs gives leadership more confidence that we can have our data volume grow and still operate in a lean, young-company environment.” ## Fusion's orchestration and CI workflows recover up to 5 engineering hours per week With Fusion's built-in orchestration, version-controlled models in GitHub, and CI-driven staging environments that let engineers compare production vs. staging data before merging, pipeline interruptions decreased. Engineering time was freed to focus on analytics and product-facing work. ## Consistent metric definitions ensure consistent data The data team implemented best practices of data transformation: consolidating models, adding testing, and documenting definitions all within their Fusion project, the business can move faster because of consistent data. "Everything about Fusion has sped up my workflows,” said Karan. “I feel like it's just going to keep going in that direction. Eventually, I'll never have to leave my coding environment and be able to work with all the data in one pane of glass." ## What's next: Building the "Middleware" for agentic workflows and self-service analytics with dbt Semantic Layer Data development now feels noticeably smoother, and engineers can now focus on delivering business value. Next on the roadmap: an internal Slack bot that queries dbt's Semantic Layer to answer business questions on the fly. Because the bot references governed metric definitions rather than querying the database directly, the risk of hallucination drops significantly. The team is also evaluating dbt Mesh and the dbt MCP server as part of a broader push toward self-service analytics across Obie's 130-person organization. Doberneck concluded, “The dbt Fusion engine is a non-negotiable for me. With anything else in our stack, we could make a change. I would never switch out dbt.” --- --- title: "Using dbt with Databricks: Architecture decisions that determine success" description: "The cost of skipping dbt on Databricks compounds quietly. A solution architect explains what to watch for and when to act." url: "https://www.getdbt.com/blog/using-dbt-with-databricks-architecture-decisions-that-determine-success" date: "2026-04-22" authors: ["Keith Ludeman"] categories: ["Insights"] --- # Using dbt with Databricks: Architecture decisions that determine success _This guest post comes from Keith Ludeman, a solution architect at [Analytics8](https://www.analytics8.com/)._ Databricks gives you the platform. dbt gives you the structure, speed, and operating layer required to make it work for your whole organization. But getting the combination right requires decisions most teams underestimate, and the cost of getting them wrong compounds over time. I work with organizations at every stage of their [Databricks journey](https://www.analytics8.com/technologies/databricks-partners/)—teams standing it up for the first time, and teams that have been running on it for years and are hitting walls they didn't expect. The question I hear most often: do we really need dbt? The answer, in almost every case, is yes. But when you introduce it, and how, depends on where you are. Below I’ll walk you through: - The cost of going Databricks-native without a transformation framework - What dbt adds to a Databricks environment - Starting fresh: Why dbt belongs early in a Databricks implementation - Signals it's time to introduce dbt in an existing Databricks environment - Common questions and objections About dbt with Databricks - Practical lessons from real implementations ## The cost of going Databricks-native without a transformation framework Databricks gives you a lot of power right out of the gate. You can land data, shape it, orchestrate pipelines, run models, and serve analytics, all in one place. It is flexible enough to support almost any approach, which is exactly why teams lean into building things their own way early on. That usually starts small. An engineer writes a few transformations in notebooks. Someone wires up orchestration to keep things moving. As more use cases come in, more logic gets added, often by different people, each solving for the problem in front of them. Nothing feels wrong in the moment. The system is working. The problems show up later. As the environment grows, it becomes harder to answer basic questions. What does this model do? Where is this logic defined? If I change this, what breaks? There is no single way of doing things, so every answer depends on who built it and when. What used to feel flexible now feels unpredictable. That is where the friction starts to build: - **You lose consistency across the environment**. Different engineers structure transformations in different ways. Over time, it becomes difficult to follow the logic, compare approaches, or confidently reuse anything. - **Technical debt builds quietly**. What began as a handful of workflows turns into a network of dependencies that no one fully documents or owns. Testing becomes an afterthought; teams end up with basic row counts to confirm something ran, rather than automated, granular validation. Even small changes require extra caution, and new engineers need time just to understand how things fit together. - **Your team spends time maintaining the system instead of improving it**. Instead of focusing on analytics, you end up managing orchestration, testing patterns, and documentation standards on your own. The platform starts to demand attention rather than enable progress. - **Issues surface late, and often without clear signals**. Without consistent testing, lineage, and dependency tracking, problems do not show up where they start. They show up downstream, in dashboards, reports, or decisions, where they are harder to trace and fix. Most teams do not decide to build their own transformation framework. It happens gradually, through a series of reasonable choices. By the time it becomes a problem, you are not just dealing with messy code. You are dealing with slower delivery, rising costs, and a growing lack of confidence in the data. ## What dbt adds to a Databricks environment Before getting into what dbt adds, it’s worth clearing up a common misconception. dbt does not replace Databricks, and it is not another engine doing the work somewhere else. Your data still lives in Databricks, and your transformations still run there. dbt is simply the layer that defines how those transformations are written, tested, and maintained over time. That distinction matters, because most of the issues that show up in a Databricks-native environment are not about the platform itself. They come from how transformation logic evolves as more people start contributing to it. Without structure, that logic tends to live wherever it was first created. A notebook here, a pipeline there, maybe a few different patterns depending on who built what. It works, but it does not give the team a consistent way to understand or extend what is already in place. dbt changes that by giving you a single, shared way to define transformations: - You can follow how data moves from one model to the next. - You can see dependencies instead of guessing what might break - You have testing, documentation, and version control built into how the work gets done The shift is subtle at first, but it changes how teams operate. Engineers spend less time figuring out what already exists and more time building on top of it. Analysts are not blocked by how the data was originally created. And when something looks wrong, they can trace the logic themselves, propose a fix, and have it reviewed before anything changes in production. New team members do not have to reverse-engineer the environment just to contribute. The work becomes easier to reason about, which makes it easier to scale across teams. That same structure becomes critical when you layer [AI or advanced analytics](https://www.analytics8.com/services/ai-data-analytics-consulting/) on top of your data. Weak points show up quickly, and if transformations are hard to trace or validate, those issues do not stay contained. They surface in downstream outputs, where they are harder to catch and more expensive to fix. A structured transformation layer gives AI systems the context they need to return trustworthy answers and makes it easier to identify sensitive fields before they surface somewhere they shouldn't. ## Where the dbt platform advances this further As teams grow, the challenge usually shifts from how transformations are written to how they are run. In many dbt Core setups, orchestration becomes something you have to solve alongside everything else. Jobs are scheduled outside of dbt, dependencies are managed across tools, and over time you end up with another layer of logic that someone on the team is responsible for maintaining. The dbt platform changes that dynamic by pulling orchestration into the same place where transformations are defined. More importantly, it adds awareness of the data itself. If nothing upstream has changed, there is no reason to run the same transformations again. Instead of executing full pipelines on a fixed schedule, you run only what needs to run based on the state of the data. At smaller scale, that might feel like an optimization. At larger scale, it directly affects both cost and reliability. It also changes the day-to-day experience in ways that are harder to quantify but easy to feel over time. When you can trace dependencies more precisely, make changes across multiple models with confidence, and catch issues before anything runs in Databricks, the work becomes less about managing risk and more about moving forward. ## Why dbt belongs early on a Databricks implementation If your organization is implementing Databricks for the first time, the instinct is to start simple. Get Databricks running, prove it out, and add tooling later. That instinct makes sense. It is also where the retrofit tax begins. Most engineering teams have deeper SQL expertise than Spark or Python expertise. dbt is SQL-native, which means your team can start contributing immediately without a steep learning curve or relying on a small group of Spark specialists. It also changes your hiring options. SQL talent is easier and more affordable to find, and that advantage grows as your team scales. Most organizations are not starting from scratch. They are migrating from something else, whether that is SQL Server and SSIS, Google Dataform, or another SQL-based transformation environment. In those cases, the transformation logic is already written in SQL, and it ports cleanly into dbt. We have agentic migration accelerators that automate 70–80% of that work, compressing what would otherwise take years into weeks. That advantage depends on staying in a SQL-native model. If you go Databricks-native, you are rewriting that logic in a different paradigm. If you are building the platform now, this is the least expensive moment to make the right architectural decisions. A few foundational choices made at the start will determine whether your platform is easy to scale or expensive to untangle: - **Adopt Unity Catalog from day one.** Don’t take shortcuts. It is far easier to implement correctly upfront than to retrofit after you have already built on top of a structure that does not support it. - **Separate compute by workload from the start.** Segregating compute for ingestion, transformation, and reporting gives you cost visibility and performance control. You can see where costs are coming from and tune each workload independently. Without that separation, costs become harder to interpret and performance becomes harder to improve. - **Don’t reinvent orchestration with dbt Core.** Starting with the free version and building orchestration, testing, and observability yourself shifts the cost into engineering time. In practice, that effort often exceeds the cost of using a managed solution. The decisions that are cheap to get right at the start are expensive to fix later. Starting with dbt isn’t adding complexity. It’s avoiding it. ## Signals it's time to introduce dbt in an existing Databricks environment Some organizations bring dbt in from day one. Others start Databricks-native, run a successful small team, production environment, or proof of concept, and then hit a wall. These are the signals that tell you it is time to introduce, or reinforce, structure. - **Your team has grown past two or three people.** A small team can move quickly in Databricks without much formal structure. That breaks down once the team grows or when business analysts start contributing as analytics engineers. At that point, you are dealing with concurrency and change control. Multiple people working in the same environment without proper CI/CD and version control leads to conflicts, overwritten work, and compounding risk. dbt provides the structure that makes that collaboration manageable. - **You’ve added a second data domain.** A proof of concept built on sales data works. Leadership pushes to bring in supply chain. The moment you move beyond a single domain, both your data structures and your team structures need to evolve. Without that shift, each new use case becomes harder to add, maintaining what already exists gets more expensive, and both technical and organizational debt start to build. It is far easier to address this before the second domain is fully in production. - **The data team is still a bottleneck.** If business users are still waiting on IT for analytics, or if analysts are exporting data and rebuilding logic on their own instead of working from shared, governed definitions, your operating model has not kept pace with your platform. “I don’t want to be the chokepoint. I want to activate the business.” That is what you hear from data leaders at this stage. dbt makes it possible to bring business analysts into the analytics engineering process. Business analysts already know SQL. Very few know Spark. A SQL-based transformation layer helps close the gap between IT and the business, both technically and organizationally. When teams make the switch, the first thing they feel is speed—not in terms of compute, but in how quickly someone can answer a question. When an analyst asks why a number looks off, the engineer is not launching a research project. The logic is traceable from the final output all the way back to the source, model by model, in a way that is easy to follow and fast to navigate. Engineers also feel the difference in how safely they can make changes. With CI/CD guardrails baked into how dbt is structured, a change gets tested and reviewed before it touches production. That security changes how the team works—less caution, more momentum. ## Common questions and objections about dbt with Databricks ### **Can't we just do this natively in Databricks?** You can. Databricks is flexible enough to support almost any pattern. The tradeoff shows up over time in development effort, maintenance, velocity, and risk. dbt establishes the many patterns that teams often try (but fail) to build for themselves. Testing, documentation, version control, and standardized transformations are not new problems. dbt gives you a consistent way to handle them without reinventing that layer inside notebooks and pipelines. ### **When does dbt become necessary?** It rarely comes down to a single moment. What changes is the cost of continuing without it. The longer a team operates without structure, the more expensive it becomes to introduce it later. Patterns that are easy to establish early require rework once pipelines, dependencies, and team processes are already in place. If you are seeing the signals from the previous section, the cost has already started to compound. And if your organization is starting to think seriously about AI, that is another clear signal. The foundation of clean data, traceable logic, documented definitions that AI requires, is the same foundation that dbt helps you build. ### **Isn't dbt expensive?** dbt Core is free and can even be run within a Databricks job. That makes it a practical starting point. The real question is where the surrounding work lives. Orchestration, testing, observability, and maintenance do not go away with Core. They move to your team. The dbt platform shifts responsibility off your engineers. In many cases, the total cost balances out, while the operational burden drops. ### **How do teams decide between dbt Core and the dbt platform?** Most teams start with dbt Core because it is free and easy to spin up. That is a reasonable starting point, and for smaller teams it may be all you need. The conversation shifts, however, when complexity grows. When orchestration requires a separate tool, when you need more sophisticated dependency management, when the surrounding work starts pulling your engineers away from the analytics itself. At that point, the dbt platform tends to make sense. The licensing cost is relatively modest, and what you get in return (orchestration, testing, observability, all in one place), typically offsets what you were spending in engineering time to maintain everything separately. ## Practical lessons from real implementations Based on our experience implementing dbt with Databricks, here are a few practical recommendations to consider: - **Use each tool for what it does best.** Databricks is the right place to land and organize raw data. dbt is the right place to transform it into something the business can use. Teams that blur that line end up with an environment that is harder to maintain and harder to explain. - **Keep all transformation logic in dbt.** This is the single most common pattern we see in struggling implementations: logic split between notebooks and dbt models. It starts with one exception and compounds from there. Once logic lives in multiple places, lineage breaks down, debugging becomes a research project, and onboarding new team members takes significantly longer. - **Don't let logic creep into orchestration.** Orchestration should trigger work, not define it. When business logic gets embedded in orchestration tools, it becomes invisible to the rest of the pipeline. Changes to source systems break things in ways no one anticipates, because the logic affecting the data isn't where anyone thinks to look. - **Adopt Unity Catalog from the start.** Teams that skip it consistently pay for it later. It is one of those decisions that is low cost to get right early and high cost to retrofit once an environment has been built on top of something else. When these things are in place, the team stops spending its time maintaining pipelines and debugging failures. It starts spending that time on what matters—expanding capabilities, supporting new use cases, and engaging with the AI and advanced analytics work the business is asking for. _[Keith Ludeman](https://www.linkedin.com/in/keithludeman/) is a Solution Architect at [Analytics8](https://www.analytics8.com/) (a [dbt Labs Visionary Consulting & Services Partner](https://www.analytics8.com/technologies/dbt-partners/)) who specializes in designing and recovering complex data platforms across Databricks, Snowflake, and dbt, from migrations to governed transformation layers. He focuses on uncovering the root causes behind delivery challenges and putting the right structure in place so teams can scale without rework or technical debt. Known for stepping into high-risk situations and bringing clarity, he approaches each engagement as a long-term partner rather than a tool-driven implementer._ --- --- title: "dbt Labs Wins a 2026 Google Cloud Partner of the Year Award" description: "dbt Labs recognized for empowering thousands of Google BigQuery users to deliver trusted analytics and AI at scale" url: "https://www.getdbt.com/blog/dbt-labs-wins-2026-google-cloud-partner-of-the-year-award" date: "2026-04-21" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Wins a 2026 Google Cloud Partner of the Year Award **PHILADELPHIA – April 21, 2026: **[dbt Labs](https://www.getdbt.com/), the leader in standards for AI-ready structured data, announced today that it has received the 2026 Google Cloud Partner of the Year award for Data and Analytics: Data Pipelines and Governance. dbt Labs works together with Google Cloud to provide the foundation for an organization's transition to AI leadership and innovation. The combination of rich data warehousing capabilities and the democratization of complex data transformation removes technical barriers, enabling analysts and business leaders to accelerate their time-to-value. dbt Labs is being recognized for its achievements in the Google Cloud ecosystem, helping joint customers manage data at scale on Google Cloud and turn it into trusted, actionable insights with speed and efficiency. Thousands of organizations run dbt on Google BigQuery globally, an integration designed to accelerate the delivery of trusted analytics and AI. By consolidating data transformation into a single, unified tool, joint customers quickly gain increased operational efficiency through advanced orchestration features. dbt Labs empowers customers to manage and trust results, ensuring high-quality data is ready to power analytics and AI initiatives both today and in the future. “Every AI strategy needs to be underpinned by a standardized foundation and process to control, govern and document progress for high-quality, trusted results,” said Shawn Toldo, Vice President, Worldwide Partner Ecosystem at dbt Labs. “Together, dbt Labs and Google Cloud enable organizations to build that foundation for an AI-ready future. We are excited for the recognition and growing partnership with Google.” This recognition is the latest example of dbt Labs’ momentum since [launching](https://www.getdbt.com/blog/dbt-labs-launches-on-google-cloud-and-google-cloud-marketplace) on Google Cloud Marketplace one year ago. The partnership’s trajectory is driven by extensive global adoption and usage across diverse industries and a rapidly expanding community of active practitioners. Additionally, dbt Labs’ partner team earned two Google Partner All Star awards, reinforcing the deep collaboration and commitment to driving mutual success. “The Google Cloud Partner Awards honor the strategic innovation and measurable value our partners bring to customers,” said Kevin Ichhpurani, President, Global Partner Ecosystem and Channels, Google Cloud. “We are proud to name dbt Labs a 2026 Google Cloud Partner Award winner, celebrating their role in driving customer success over the last year.” By bringing Google AI capabilities into dbt workflows, joint customers gain the trustworthy, well-documented, governed foundation that reliable analytics and AI demand. To learn more about how dbt Labs and Google Cloud are enabling AI-ready data pipelines, watch the on-demand webinar “Building dbt Models Faster with Google AI” at [https://www.getdbt.com/confirmation/building-dbt-models-faster-with-google-ai-recording](https://www.getdbt.com/confirmation/building-dbt-models-faster-with-google-ai-recording). **About dbt Labs **Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 80,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at getdbt.com, and follow dbt Labs on [LinkedIn](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=82576339&u=https%3A%2F%2Fwww.linkedin.com%2Fcompany%2Fdbtlabs%2Fmycompany%2F&a=LinkedIn), [X](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=2299196361&u=https%3A%2F%2Fx.com%2Fdbt_labs&a=X), [Instagram](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=255235389&u=https%3A%2F%2Fwww.instagram.com%2Fdbt_labs%2F&a=Instagram), and [YouTube](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=508887317&u=https%3A%2F%2Fwww.youtube.com%2Fc%2Fdbt-labs&a=YouTube). --- --- title: "Why metric definitions matter for reliable AI agents" description: "Learn how dbt's semantic foundation enables reliable, governed agentic analytics." url: "https://www.getdbt.com/blog/metric-definitions-ai-agents" date: "2026-04-21" authors: ["Joey Gault"] categories: ["Pulse"] --- # Why metric definitions matter for reliable AI agents ## The challenge of semantic ambiguity [dbt](https://www.getdbt.com/product/what-is-dbt) agents operate fundamentally differently than human analysts. When a human encounters ambiguity in a metric definition, they can apply context, ask clarifying questions, or make informed assumptions based on institutional knowledge. AI agents lack this intuitive understanding. Without precise definitions, they generate inconsistent results that undermine trust and create operational risk. Consider a seemingly straightforward metric like "monthly revenue." Across different departments, this could mean revenue recognized in a given month, revenue booked that month, revenue from contracts starting that month, or revenue adjusted for returns and refunds. A human analyst working with the finance team understands which definition applies in context. An AI agent querying data autonomously does not. When multiple agents operate across different teams—a sales agent analyzing pipeline performance, a finance agent generating forecasts, and a customer success agent evaluating retention—semantic inconsistency creates a compounding problem. Each agent might calculate "monthly revenue" differently, producing conflicting outputs that require manual reconciliation. This defeats the purpose of autonomous analytics and erodes confidence in agent-generated insights. The scale of this challenge becomes apparent in modern data environments. Organizations now work across an average of 400 data sources, with nearly one in five enterprises managing more than 1,000 sources. In these complex ecosystems, the same business concept might be represented dozens of different ways across systems. Without a shared semantic layer that provides consistent definitions, agents amplify rather than resolve this fragmentation. ## Structured context as the foundation for agency For AI agents to operate safely and effectively in enterprise environments, they require more than instructions and advanced language models. They need structured context: the schemas, semantics, relationships, permissions, and lineage that describe how data works within an organization. Structured context equips agents with three essential capabilities. First, it provides memory through metadata, enabling agents to understand what data assets exist and how they relate to one another. Second, it establishes boundaries through clear definitions, permissions, and rules that prevent agents from operating outside established guardrails. Third, it enables useful actions by providing validated tools and interfaces for reading and writing data safely. Metric definitions sit at the heart of this structured context. When an agent needs to answer a question about customer churn, it must know precisely how "churn" is defined, which data sources contain the authoritative calculation, what business rules apply, and who has permission to access the underlying data. Without this semantic foundation, agents resort to guessing or hallucinating definitions, producing unreliable results. dbt excels at creating this structured foundation through [data transformation](https://www.getdbt.com/product/develop) workflows that convert raw data into analytics-ready models with built-in testing, lineage tracking, and semantic meaning. By defining metrics consistently within dbt models and exposing them through the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), organizations create a single source of truth that both humans and agents can reliably query. ## The cost of poor definitions The consequences of weak metric definitions become severe when agents operate autonomously at scale. Poor data quality is cited as the primary reason AI projects fail to deliver expected value, and organizations lose an average of [$12.8 million annually](https://www.ibm.com/thought-leadership/institute-business-value/en-us/report/data-quality) due to data quality issues, with some companies losing as much as 6% of annual revenue from flawed AI outputs. High-profile failures illustrate the risk. Airlines have faced legal action when chatbots promised refunds based on hallucinated policies. Lawyers have submitted briefs citing fictional cases generated by AI systems. News organizations have published AI-generated travel content directing readers to unsafe destinations. Even the most advanced language models hallucinate at significant rates when operating without proper grounding in structured, validated data. In the context of business analytics, these failures manifest as agents that confidently report incorrect metrics, make recommendations based on flawed calculations, or trigger automated actions using inconsistent business logic. When an agent autonomously adjusts pricing based on a miscalculated margin metric or sends customer communications based on an incorrect churn definition, the operational and reputational damage can be substantial. Regulatory frameworks are making these risks explicit. Under the EU AI Act, particularly Articles 10 and 27, organizations deploying high-risk AI systems must demonstrate that their data is complete, accurate, representative, and error-free. This includes comprehensive documentation of data sources, quality checks, and bias mitigation measures. Metric definitions are a core component of this compliance obligation: organizations must be able to prove that their AI systems are calculating business-critical metrics correctly and consistently. ## Governance through definition Effective governance for AI agents cannot be bolted on after deployment. It must be embedded in the [data transformation](https://www.getdbt.com/product/develop) layer where metrics are defined and calculated. This approach ensures that governance policies flow automatically through dependent models and that agents inherit the correct definitions and access controls. When metrics are defined centrally in dbt, changes propagate consistently across all downstream uses. If the definition of "active user" changes to reflect new product features, that update flows automatically to every dashboard, report, and agent that references the metric. This eliminates the drift that occurs when definitions are scattered across multiple systems or hardcoded into individual queries. Column-level security and row-level access controls become particularly important for agents. Unlike human users who might access a dashboard showing aggregated metrics, agents often query underlying data directly. A conversational analytics agent responding to a sales manager's question about team performance should only access data for that manager's region and team members. These access boundaries must be defined at the metric level and enforced consistently regardless of how the data is accessed. The [dbt MCP (Model Context Protocol) server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) provides a standardized interface for exposing dbt models, lineage, and semantic context to AI systems while maintaining fine-grained, policy-aware access controls. This enables agents to discover available metrics, understand their definitions and lineage, and query them safely within established governance boundaries. ## Enabling multi-agent collaboration As organizations move beyond single-purpose agents to multi-agent architectures, consistent metric definitions become even more critical. In these systems, specialized agents handle specific functions and collaborate through orchestration layers to complete complex tasks. Consider a scenario where a discovery agent helps a business user identify relevant datasets, an analyst agent generates insights from those datasets, and a developer agent creates new data models based on the findings. For this workflow to function reliably, all three agents must share a common understanding of the metrics involved. If the discovery agent surfaces a "customer lifetime value" metric that the analyst agent calculates differently, the entire workflow breaks down. Event-driven architectures for multi-agent systems depend on semantic consistency. When one agent publishes an event indicating that a key metric has crossed a threshold, downstream agents must interpret that metric identically to respond appropriately. This requires metric definitions to be versioned, documented, and accessible through shared interfaces that all agents can query. Organizations implementing multi-agent systems should treat metric definitions as contracts between agents. Just as microservices rely on well-defined APIs, agents rely on well-defined metrics. Changes to metric definitions should follow the same rigorous change management processes as API changes, including versioning, deprecation notices, and backward compatibility considerations. ## Practical implementation for data engineering leaders Building a metric definition framework that supports reliable AI agents requires deliberate architectural choices. Data engineering leaders should focus on several key practices. Start by mapping all business-critical metrics and documenting their definitions comprehensively. This includes not just the calculation logic, but also the business context, data sources, refresh frequency, known limitations, and ownership. These definitions should live in version control alongside the dbt models that implement them, creating a single source of truth that evolves with the business. Implement comprehensive testing for metric calculations. [dbt's testing framework](https://docs.getdbt.com/docs/build/data-tests) enables data teams to validate that metrics are calculated correctly, that underlying data meets quality standards, and that changes don't introduce regressions. For AI agents, these tests serve as guardrails that prevent autonomous systems from operating on flawed data. Establish clear ownership and approval processes for metric changes. When an agent relies on a metric definition to make autonomous decisions, changes to that definition have operational implications. Metric owners should be identified, change requests should be reviewed by stakeholders, and impacts should be assessed before deployment. Expose metrics through a semantic layer that provides a consistent query interface for both humans and agents. The [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) enables organizations to define metrics once and query them consistently across tools, eliminating the proliferation of slightly different metric implementations that creates semantic drift. Monitor how agents use metrics in production. Observability for agentic systems should include tracking which metrics agents query, how they interpret results, and what actions they take based on those metrics. This visibility enables rapid intervention when agents misinterpret metrics and creates feedback loops for improving definitions. ## The path forward The shift toward agentic analytics is accelerating. According to the [IBM Institute for Business Value](https://www.ibm.com/thought-leadership/institute-business-value/en-us/report/agentai), 70% of executives consider agentic AI critical to their future, and 61% of CEOs report actively deploying or scaling AI agents. The agentic AI market, valued at $5 billion today, is projected to reach $50 billion by 2030. Organizations that establish rigorous metric definition practices now will be positioned to deploy autonomous analytics systems with confidence. Those that treat metric definitions as an afterthought will struggle with unreliable agents, inconsistent outputs, and erosion of trust in AI-generated insights. For data engineering leaders, the imperative is clear: invest in the semantic foundation that makes autonomous analytics possible. [dbt's semantic layer](https://www.getdbt.com/product/semantic-layer) provides the transformation framework for defining metrics consistently, testing them rigorously, and exposing them through interfaces that agents can query reliably. Explore [dbt's agent capabilities](https://docs.getdbt.com/blog/dbt-agent-skills) and the [dbt documentation](https://docs.getdbt.com/docs/introduction) to learn how to implement these practices at your organization. The organizations that will thrive in the era of agentic analytics are those that recognize metric definitions not as a documentation exercise but as critical infrastructure. By treating metrics as first-class data products with clear ownership, rigorous testing, and consistent governance, data engineering leaders create the foundation for AI agents that augment rather than undermine analytical capabilities. ## AI agent FAQs **What is the difference between AI agents and traditional business intelligence tools?** AI agents operate autonomously and lack the intuitive understanding that human analysts possess. When human analysts encounter ambiguity in metric definitions, they can apply context, ask clarifying questions, or make informed assumptions based on institutional knowledge. AI agents cannot do this: without precise definitions, they generate inconsistent results that undermine trust and create operational risk. Traditional BI tools are typically operated by humans who can interpret and contextualize data, while AI agents query and analyze data independently, making structured context and clear metric definitions essential for their reliable operation. **What analytics are available to assess AI agent performance against business objectives?** Monitoring how agents use metrics in production is essential for assessing their performance. Observability for agentic systems should include tracking which metrics agents query, how they interpret results, and what actions they take based on those metrics. This visibility enables rapid intervention when agents misinterpret metrics and creates feedback loops for improving definitions. Organizations should monitor agent outputs for consistency, validate that agents are calculating business-critical metrics correctly, and ensure that autonomous decisions align with established business logic and governance policies. **How do AI agents improve consistency and trust?** AI agents improve consistency and trust when they operate on a foundation of structured context with precise metric definitions. By defining metrics centrally in transformation workflows, changes propagate consistently across all downstream uses: every dashboard, report, and agent references the same definition. This eliminates drift that occurs when definitions are scattered across systems. When metrics are treated as contracts with clear ownership, rigorous testing, and consistent governance, agents can reliably query a single source of truth, producing consistent outputs across different teams and use cases rather than conflicting results that require manual reconciliation. --- --- title: "Meet Antigravity: Google’s agentic IDE enters the dbt orbit" description: "Google's agentic IDE is here, and paired with dbt, it might just give you your weekends back. Here's what you need to know." url: "https://www.getdbt.com/blog/meet-antigravity-google-s-agentic-ide-enters-the-dbt-orbit" date: "2026-04-17" authors: ["Stephen Robb"] categories: ["Learn"] --- # Meet Antigravity: Google’s agentic IDE enters the dbt orbit There’s a new player in the Agentic IDE space, and it’s coming in hot. Enter **Antigravity**, Google’s entrance into the world of AI-powered development environments. I spent some time with it this past weekend. Paired it with Gemini 3. Let’s just say…I did not expect to be that impressed. The power jump compared to traditional IDE workflows (and some other popular agentic IDEs) is significant, especially when you bring it into the dbt universe. Let’s break down what it is, how it fits with dbt, and a few pro tips to get the most out of it. ## **What is Antigravity?** At its core, Antigravity is a fork of Visual Studio Code. That’s great news because it means most of the extensions, tooling, and workflows you already love just work out of the box. But Antigravity isn’t “just VS Code with a new coat of paint.” It’s built for an agent-first experience. You’re not just coding, you're collaborating with AI agents that can reason across your project, propose plans, and execute tasks. Think less autocomplete.Think more “co-pilot who drank three espressos and read your entire repo.” ## **Getting started with dbt in Antigravity** If you’re working with dbt, step one is easy: install the official dbt extension. You can read about it here: [https://docs.getdbt.com/docs/about-dbt-extension](https://docs.getdbt.com/docs/about-dbt-extension) With the dbt extension installed, you immediately get: - Column-level lineage - Query preview - Rich dbt-aware IDE features - Improved navigation across models, sources, and tests In other words, your IDE‌ understands dbt instead of just politely pretending to. From there, you can immediately start using the built-in agent to help generate SQL and YAML files. The agent scaffolds models, adds tests, and writes the YAML documentation you keep forgetting. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/32f5a7dd0b35e93aa18f270d806507a13acdd77b-936x826.png) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/7e3f3380640d1cbf822a7451f2c31ff11b168be7-936x430.png) ## **Turning it up: Add the dbt MCP server** If you really want to unlock Antigravity's potential for dbt work, give it access to dbt's structured context. The dbt MCP server is how you do that: it surfaces your project graph, model definitions, lineage, test results, and semantic layer to the agent in a governed, queryable way. Inside the agent window: 1. Click the three dots in the top-right. 2. Select **MCP servers**. 3. Add the dbt MCP server. Eventually, the dbt MCP server will be available in the market. You can choose it from the dropdown menu, but for now, you can just change the mcp_config provided by Antigravity and it does the rest. Those configuration options can also be seen here: https://docs.getdbt.com/docs/dbt-ai/about-mcp This significantly expands what your local agent can do. And here’s where things get interesting. You’re no longer just asking for code snippets. You’re enabling deeper project-level awareness and workflows. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/c7639ef0030a5b2f49c444e04fc363dcf173187e-430x544.png) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/ada010f1bd8fc3206846babd02e29f95a74f3b11-936x412.png) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0d314eff6e132c3936a15d0640c4d74787197bb8-936x592.png) ## **MCP servers that pair beautifully** Beyond dbt, there are several MCP servers that elevate the experience: - GitHub - Google BigQuery - AlloyDB - Dataplex Now imagine this workflow: 1. A ticket is opened. 2. The agent reads it. 3. It reviews data classification tags in Dataplex. 4. It generates SQL, YAML, and tests. 5. It writes a pull request in GitHub. 6. CI kicks off dbt orchestration and validates everything. That’s not autocomplete. That’s workflow acceleration. We’re talking about reducing friction across development, governance, and deployment in one unified environment. ## **Pro tips for working with Antigravity + dbt** After a few days of experimenting, here are some practical lessons. ### **1. Pair Antigravity with the Gemini CLI** Use Antigravity for: - Multi-agent brainstorming - Large architectural work - Implementation planning Use the Gemini CLI for: - Focused terminal tasks - Deep maintenance - Headless execution - Specific, scoped operations Together, they create a powerful balance between high-level reasoning and low-level precision. ### **2. Define rules for your agent** Create global or workspace rules to guide how your agent behaves: - `~/.gemini/GEMINI.md` - `.agent/rule/` You can define: - Naming conventions - SQL style standards - Testing requirements - Documentation expectations Think of it as training your agent to be a senior analytics engineer instead of an enthusiastic intern. ### **3. Add skills for specific tasks** You can extend your agent with specialized skills depending on what you're building. dbt-specific skills are a great place to start: [https://docs.getdbt.com/blog/dbt-agent-skills](https://docs.getdbt.com/blog/dbt-agent-skills) Skills help tailor the agent’s behavior so it understands how to approach dbt models, testing strategies, documentation, and more. ### **4. Break large tasks into smaller ones** This one’s critical. Antigravity is very good at: - Creating implementation plans - Designing step-by-step execution strategies It is even better when you: - Break requests into smaller, precise tasks - Start at task 1 - Move sequentially With AI agents, smaller and more specific is almost always better. Think iterative, not monolithic. ## **Final thoughts** Google entering the agentic IDE space with Antigravity feels like a meaningful shift especially for analytics engineers living in the dbt ecosystem. Because it’s built on Visual Studio Code, adoption is frictionless.Because it supports MCP servers, it’s extensible.Because it integrates deeply with Gemini, it’s powerful. And because it can help you write SQL, YAML, tests, PRs, and documentation…it might just give you your weekends back. No promises. But it’s a strong start. --- --- title: "Exploring dbt and Google with AI agents" description: "What happens when you plug AI into a dbt project and let it do things? A practical guide to building your first dbt agent." url: "https://www.getdbt.com/blog/exploring-dbt-and-google-with-ai-agents" date: "2026-04-17" authors: ["Stephen Robb"] categories: ["Learn"] --- # Exploring dbt and Google with AI agents _This will be a longer blog.** **There will be some code if you want to follow along otherwise feel free to skim to see a glimpse into a simple guide for AI. **Technical requirements: Python, Git, dbt Core/dbt Fusion, or the dbt platform.**_ **TL;DR: **This post is about exploring what happens when you plug AI into a dbt project and‌ let it _do things._ By experimenting with Gemini, Google Agent Development Kit, the dbt MCP server, and the dbt Fusion engine, with dbt's structured context as the foundation, I built a working agent just to see how far it could go as a starter project. It’s less “here’s a perfect solution” and more “let’s see what’s possible.” And that’s what made it fun. I started my career as a software engineer. The kind who lived in an editor all day, shipping code, breaking things, and fixing them again. Over time, my day job shifted. I was still close to the code, but no longer _in_ it the way I used to be as I switched to different roles. Then AI tools showed up and quietly changed the rules of the game. Suddenly, working with code felt…fun again. I stopped asking, _“Do I still remember how to build this?”_ The question became, _“What can I build now?”_ At the same time, I switched to being more in the data world and started thinking about how I could bring that same energy here. Then dbt Labs released the dbt Fusion engine and I saw a true path forward to something exciting. With dbt’s rock-solid foundation for analytics engineering and Google’s AI tooling opening the door to agentic workflows, I decided to explore what it looks like to pair the two hands-on. The result of that exploration is a working **dbt agent**, powered by Google’s **ADK framework**. It’s not meant to be magic or perfect. It’s meant to be practical: a starting point for solving common dbt problems, poking at what’s available out of the box, and experimenting with what happens when you give dbt a little bit of autonomy. This post is a walkthrough of that journey. What I built, why I built it, and how AI changed the way I think about getting started with dbt again. ## Defining terms Before jumping into the fun stuff, here are a few terms worth knowing. **LLM (large language model): **An LLM is the “brain” behind modern AI tools. It’s what reads, writes, and reasons about text (and increasingly, code and data). It’s like a very fast reader who has seen _a lot_ of books and code. You ask it a question, and it predicts the best next words to respond with often in surprisingly useful ways. **MCP (model context protocol): **MCP is a standard way for AI models to safely interact with tools, systems, and data without hard-coding custom integrations everywhere. Think of MCP like a universal remote for AI. Instead of teaching the AI how to use every tool differently, MCP gives it a consistent set of buttons and rules so it doesn’t accidentally do something wild. **Agent: **An agent is an AI system that can reason, decide what to do next, and take actions using tools rather than just answering questions. A normal AI answers questions.An agent gets a goal, figures out steps, uses tools, and checks its own work. Think of it as a very junior, but very fast, teammate. **dbt MCP: **The dbt MCP server exposes dbt capabilities like metadata, models, tests, and commands as tools an AI agent can safely use. Instead of an AI guessing how dbt works, dbt MCP gives it a rulebook and a toolbox. The agent can ask things like “What models exist?” or “Run this dbt command” without breaking anything. **Gemini: **Gemini is Google’s family of AI models, designed to handle reasoning, code, and multi-step problem solving at scale**. **Gemini is the brain I’m plugging into this system. It’s the part doing the thinking, reading dbt projects, understanding context, and deciding what to try next. **Google ADK (agent development kit): **Google ADK is a framework for building AI agents defining how they think, what tools they can use, and how they interact with systems. If the agent is the worker and Gemini is the brain, ADK is the job description. It defines what the agent is allowed to do, how it calls tools, and how everything stays organized and safe. ## Why this matters for dbt All of these pieces: LLMs, agents, MCP, and Google’s ADK matter to dbt because they finally let AI move from _suggesting_ things to _safely doing_ things in analytics engineering. What really unlocked my excitement here was the **dbt Fusion engine.** Real-time parsing and a smart, deterministic compiler mean AI no longer has to “hope” its SQL or YAML is correct. Every output can be validated immediately against the warehouse, the project graph, and dbt’s rules. That’s a huge shift. Instead of treating AI like a clever autocomplete, Fusion makes it possible to treat AI like a junior analytics engineer: - It can propose models, tests, and metrics - Fusion can instantly tell us whether they compile, parse, and conform - Mistakes become feedback loops, not production risks When you combine: - **Agents** that can reason and act - **MCP** that enforces safe, intentional tool use - **Gemini** for multi-step reasoning - **Google ADK** to orchestrate everything - **dbt Fusion** as the guardrails and truth source You get something genuinely new: an AI workflow that can iterate on data logic in real time, with confidence. For someone who hasn’t been data engineering code-heavy ever, this felt like having a safety net that made building fun and that’s what pushed me to see how far this could go. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f33d3f6697498bb0ee2a9abbd32aeba3a1b3cf24-936x518.png) ### Part 1: A beginning Like most beginnings, let's start small. We are going to build an enterprise data scientist agent that would make Skynet jealous. Just kidding, we will get the dbt MCP to work. I won’t do an in-depth guide here but will provide some links: [Introducing the dbt MCP Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server): This is a fantastic blog by Jason Ganz explaining a few of the key operating principles of dbt’s MCP. There are two important takeaways from that blog: - dbt MCP can be deployed locally or connected remotely - It grants us access to most dbt functionality through tools There is a quickstart guide located here if you want to try it out: [https://docs.getdbt.com/docs/dbt-ai/setup-local-mcp](https://docs.getdbt.com/docs/dbt-ai/setup-local-mcp) Lastly, there is the repo itself, where you can find the tools diagram and even some agent examples that have been put together in the examples folder. If you want this to run locally, clone this repo: [https://github.com/dbt-labs/dbt-mcp](https://github.com/dbt-labs/dbt-mcp) Once your python environment is set up, requirements installed, and environment variables configured, you should be good to go. You can then use Claude, Cursor, Antigravity or other clients to connect and confirm it’s working. Here is what success looks like with Claude. If you have any problems up to this point, please revisit the MCP quickstart [guide](https://docs.getdbt.com/docs/dbt-ai/setup-local-mcp): Now if I prompt Claude to list my tools, it would provide a list across a few categories: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/186ba88892177365084503fb09e13d984821fd07-936x494.png) This is an important milestone. These tools let you explore your dbt project, query metrics, analyze lineage, monitor job runs, and troubleshoot issues. Most people never have to go any further. **But for those who do want to go deeper: **Try asking questions on your data, do some codegen in cursor/antigravity, translate syntax, break up stored procedures…it’s limited only by your creativity. ### Part 2: Let’s get agentic Alright, let’s add **Google’s Agent Development Kit (ADK)** to the mix. If you want to go deep, Google’s official documentation is available here: [https://docs.google.com/document/d/1yYaRUUJddrY5PZIJHLZD1a54PHvqCswRJu8TGXnjc_8/edit?tab=t.0](https://docs.google.com/document/d/1yYaRUUJddrY5PZIJHLZD1a54PHvqCswRJu8TGXnjc_8/edit?tab=t.0) It goes _far_ beyond what we’ll cover here and dives into the full breadth of functionality ADK brings to the table. At a high level, ADK provides: - Orchestration for adaptive agent behavior - Multi-agent architecture support - A rich, extensible tool ecosystem In short, it gives you the structure and tooling needed to build serious agent-based systems, and importantly, to take something from prototype to production. We won’t be publishing any agents today, but ADK absolutely provides the scaffolding to do exactly that when you’re ready. I chose Python as my language of choice and started with the official quickstart guide: [https://google.github.io/adk-docs/get-started/python/](https://google.github.io/adk-docs/get-started/python/) **That gives you a clean project structure right out of the gate:** `my_agent/` `│── agent.py # main agent code` `│── .env # API keys or project IDs` `│── __init__.py` From there, you’ve got two easy ways to run your agent: **Command-line access:** `adk run my_agent` **Or… the much nicer option, the web interface:** `adk web --port 8000` That spins up a beautiful local test interface where you can interact with your agent in real time. It’s fast, clean, and makes experimentation far more enjoyable than staring at raw terminal output. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/7d5c90497cfdaf416441a881c2e762f50ab3c3fb-936x460.png) ### Part 3: Customize the agent If you’re looking for the simplest possible agent experience, start with the example in the **dbt Labs** `dbt-mcp` repository: [https://github.com/dbt-labs/dbt-mcp/blob/main/examples/google_adk_agent/main.py](https://github.com/dbt-labs/dbt-mcp/blob/main/examples/google_adk_agent/main.py) That example does a great job of keeping things simple. It: - Reads your .env file to locate your dbt MCP endpoint - Registers the available MCP tools - Exposes those tools through an agent interface In just a few lines of code, you have an agent that can talk directly to your dbt environment. Clean. Practical. Effective. ### Part 4: Extend Of course, we couldn't stop there. I was just getting started. I created an extended example here: [https://github.com/StephenR-DBT/dbt-gemini-agent-starter](https://github.com/StephenR-DBT/dbt-gemini-agent-starter) The goal was to push beyond a single-tool agent and explore orchestration across multiple tools and subagents. Specifically, I built three components: ### **1. dbt_compile** A local dbt compilation tool with detailed JSON log analysis. This runs `dbt compile` (using the Fusion engine) to validate SQL before anything ever hits the warehouse. That means: - Syntax validation - Model resolution checks - Dependency validation - Structured log parsing for intelligent feedback In practice, this allows an agent to generate SQL, validate it locally, detect issues, and automatically iterate before shipping anything downstream. It’s like giving your agent a pre-flight checklist. ### **2. dbt_mcp_toolset** Cloud-based dbt platform operations via MCP. This exposes the full dbt MCP toolset to the agent, including access to: - Project metadata - Model definitions - Lineage information - Intelligent querying capabilities Instead of guessing about the structure of a project, the agent can inspect it directly. It can reason over metadata the same way an analytics engineer would. ### **3. dbt_model_analyzer** A specialized subagent focused on data modeling analysis. This was the fun part. Rather than giving one monolithic agent every responsibility, I created a purpose-built subagent that focuses purely on modeling logic: structure, best practices, and design patterns. It’s narrower in scope and offers a fun perspective. ## **Why this matters** What this experiment showed me is that agents don’t just “generate SQL.” They can: - Validate their own work before execution - Use metadata to reason about the broader system - Delegate to specialized subagents - Iteratively fix issues they create In other words, they can participate meaningfully in the development lifecycle and not just at the prompt layer, but at the systems layer. And‌ most importantly, it demonstrated something creative and genuinely new: a development loop where AI doesn’t just produce code, but critiques, validates, and improves it using the same tooling we rely on as engineers. That’s where this starts to feel less like a demo… and more like the beginning of a new workflow. Please reach out to me if you create anything fun. I’ll have demos, blogs, and things going forward, and I would love to see what the community creates with the dbt MCP server. --- --- title: "Tableau and dbt: structured context for reliable AI analytics" description: "Two MCPs, one config file. Learn how pairing dbt and Tableau unlocks impact analysis, metric reconciliation, and more." url: "https://www.getdbt.com/blog/tableau-and-dbt-mcps-together" date: "2026-04-16" authors: ["Stephen Robb"] categories: ["Learn"] --- # Tableau and dbt: structured context for reliable AI analytics We spend plenty of time talking about the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=1.12). And for good reason. It’s practical, reliable, and genuinely useful when you’re building analytics workflows around dbt. But there’s another player that deserves equal airtime: the MCP for Tableau. So let’s fix that. ## **Why pair dbt and Tableau?** dbt handles transformation logic and ensures every metric definition is versioned, tested, and governed. Tableau handles visualization and distribution. One shapes the data; the other tells the story. Individually, their MCPs are powerful. Together, they’re streamlined. When both are wired into your agentic environment, you’re no longer bouncing between tooling contexts. You can: - Inspect and adjust dbt models - Validate exposures and lineage - Explore Tableau metadata - Align dashboards with transformed models - Iterate with a single conversational thread No context loss. ## **The setup (it’s easy)** This isn’t a 17-step integration guide. It’s simple. Using your preferred agentic IDE or client whether that’s Cursor, Claude, Antigravity, or another MCP-capable tool add both the dbt MCP configuration and the Tableau MCP configuration to the required config file. That’s it. Save the file. Restart if needed. Wait for the green arrows. Once both MCPs are live, they’re available to any prompt you run in that environment. No special invocation rituals. No manual switching. Just ask. ### **Let’s explore a few ideas** #### **1. Impact analysis dashboard** Touching a critical model like `fct_revenue` has downstream consequences that are easy to miss. - Trace lineage of the model in dbt. - Pull performance metrics from the dbt Semantic Layer. - Search Tableau for every dashboard/workbook using that model. - Automatically generate an impact report showing: downstream dependencies, dashboard usage stats, and potential stakeholders impacted. You go from “Uh oh, did I break something?” to “Here’s exactly what I need to check” in minutes. #### **2. Data quality health monitor** Keep your analytics trustworthy with a unified view: - Check dbt model health (tests, source freshness). - Pull trend metrics from the dbt Semantic Layer. - Snap in Tableau views for the same metrics. - Generate a health report with recommendations, like how often each dataset should be refreshed. Think of it as a fitness tracker for your data pipelines. Green arrows = data is in shape; red = time for a pit stop. #### **3. Metric reconciliation detective** The classic “why don’t the numbers match?” problem? Solved. - Query a metric through dbt Semantic Layer (e.g., monthly revenue). - Query the same metric via Tableau’s published data source. - Retrieve the compiled SQL from both systems. - Compare and highlight discrepancies automatically. Your CFO will finally stop asking why the dashboard number differs from the report. Mystery solved. This works because the dbt Semantic Layer is the one place where 'monthly revenue' has a single definition. #### **4. Self-service analytics enablement** Empower your team without endless hand-holding: - User asks: “What revenue metrics can I analyze by region?” - List available metrics from dbt Semantic Layer. - Search Tableau for existing dashboards with those metrics. - Show screenshots if dashboards exist; otherwise, query dbt and suggest creating a new viz. It’s like having an analytics concierge always ready to point people to the right metric or dashboard. #### **5. Performance optimization finder** Stop slow queries before they slow you down: - Get dbt model performance metrics (execution time trends). - Identify Tableau dashboards querying those slow models. - Analyze Tableau query patterns and data retrieval efficiency. - Recommend optimizations, e.g., “This dbt model takes 10 min but Tableau only uses 3 columns trim it down.” It’s the intersection of observability, efficiency, and a tiny bit of magic. ## The takeaway Individually, dbt and Tableau MCPs are powerful. Together, they turn what used to be multi-step, context-switch-heavy tasks into single-threaded, agent-powered workflows. One config file, two green arrows, endless possibilities. --- --- title: "New dbt Labs Report Finds AI-driven Acceleration is Outpacing Trust and Governance" description: "AI is accelerating data workflows, but governance and trust aren't keeping pace, according to a new dbt Labs report." url: "https://www.getdbt.com/blog/new-dbt-labs-report-finds-ai-driven-acceleration-is-outpacing-trust-and-governance" date: "2026-04-14" authors: ["Elaine Green"] categories: ["Press"] --- # New dbt Labs Report Finds AI-driven Acceleration is Outpacing Trust and Governance _The 2026 State of Analytics Engineering Report reveals data leaders’ concerns over data quality amid rising need for reliability_ **Key findings:** - 72% of respondents now prioritize AI-assisted coding in their development workflows, while only 24% prioritize AI-assisted pipeline management, including testing and observability; this highlights an imbalance between acceleration and quality - Trust in data and data teams as an organizational priority surged from 66% to 83% year over year, the steepest single-year increase of any measured objective; speed followed, climbing from 50% to 71% - 71% of data professionals cite incorrect or hallucinated outputs reaching stakeholders as a top concern, which carries greater consequence as autonomous agents operate on top of organizational data at scale - Data infrastructure costs are outpacing budget growth, with 57% reporting increased warehouse and compute spend, compared to just 36% reporting increased team budgets **PHILADELPHIA – April 14, 2026: **[dbt Labs](https://www.getdbt.com/), a leader in standards for AI-ready structured data, today released its fourth annual State of Analytics Engineering Report, revealing a growing gap between the speed at which AI is transforming data work and the systems designed to ensure its reliability. As AI becomes embedded in analytics workflows, organizations are producing data faster than ever, but governance, validation, and trust mechanisms are not keeping pace. As a result, trust in data has emerged as the most widely prioritized organizational objective, rising to 83% year over year. In this environment, organizations that invest in governance, validation, and data quality as strategic priorities are best positioned to scale AI-driven outcomes reliably and turn acceleration into sustainable impact. **AI moves from experimental to embedded** According to the survey, AI is scaling across two key areas of analytics engineering: AI-assisted coding that increases productivity and AI-generated, stakeholder-facing insights. The majority (72%) of respondents now prioritize AI-assisted coding in their development workflows, and 77% of leaders report pushing teams to improve productivity with AI. "Two years ago, most analytics practitioners and leaders didn't expect to be generating the majority of their analytics code with AI. But today, that’s where we are," said Jason Ganz, dbt Labs Director, Community, Developer Experience and AI. "This signals a fundamental shift in the role of data practitioners, away from manually creating code and toward building the systems that enable agentic data workflows at scale, while providing the trusted infrastructure those agents need to operate reliably. Organizations that treat governance as infrastructure, not an afterthought, are the ones that will make the most of what AI can do." **Trust and governance as key enablers of AI at scale** Even though technical integration challenges have declined (from 35% to 27% year-over-year), governance issues like ambiguous data ownership (41%) and poor data quality remain persistent obstacles. Nearly three-quarters (71%) of data professionals are concerned about incorrect data reaching stakeholders. In parallel, trust and speed have emerged as the dominant priorities among respondents, clearly separating from cost reduction. The importance placed on increasing trust in data rose sharply from 66% in 2025 to 83% in 2026, while the priority of "shipping data products faster" climbed from 50% to 71%. An emphasis on cost reduction, however, increased by only 5% (from 48% to 53%). “There’s a real tension between moving fast and building trust, and you can’t optimize for both without intention,” said Pooja Crahen, senior manager of analytics engineering at Okta. “That’s where discipline in modeling, validation, and ownership becomes a requirement, not a best practice.” On April 29, a panel of industry experts from Hex, Ramp and dbt Labs will host the 2026 State of Analytics Engineering Virtual Event. The conversation will focus on the report findings, what the year-over-year changes signal, and how trust isn’t a constraint on AI-driven impact but the determining factor in how far it can scale. Register here: [https://www.getdbt.com/resources/webinars/2026-state-of-analytics-engineering-virtual-event](https://www.getdbt.com/resources/webinars/2026-state-of-analytics-engineering-virtual-event) Download the 2026 State of Analytics Engineering report at [https://www.getdbt.com/resources/state-of-analytics-engineering-2026](https://www.getdbt.com/resources/state-of-analytics-engineering-2026). **Methodology** dbt Labs collected survey responses in late 2025 and early 2026 from 363 data practitioners and leaders across industries and regions. Of the respondents, 73% identified as practitioners, and 27% as managers or executives overseeing data teams. **About dbt Labs** Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 80,000 data teams use dbt, including those at Siemens, Roche, and Condé Nast. Learn more at getdbt.com, and follow dbt Labs on [LinkedIn](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=82576339&u=https%3A%2F%2Fwww.linkedin.com%2Fcompany%2Fdbtlabs%2Fmycompany%2F&a=LinkedIn), [X](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=2299196361&u=https%3A%2F%2Fx.com%2Fdbt_labs&a=X), [Instagram](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=255235389&u=https%3A%2F%2Fwww.instagram.com%2Fdbt_labs%2F&a=Instagram), and [YouTube](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=508887317&u=https%3A%2F%2Fwww.youtube.com%2Fc%2Fdbt-labs&a=YouTube). --- --- title: "From raw data to trusted AI: What dbt is bringing to Google Cloud Next" description: "See how dbt + BigQuery powers trusted, AI-ready analytics. Visit Booth #6606 at Google Cloud Next, April 22–24 in Las Vegas." url: "https://www.getdbt.com/blog/what-dbt-is-bringing-to-google-cloud-next-2026" date: "2026-04-13" authors: ["Stephen Robb", "Emily Hart"] categories: ["Partnerships"] --- # From raw data to trusted AI: What dbt is bringing to Google Cloud Next If you’re heading to Google Cloud Next this year, we’d love to show you what we’ve been building at dbt. Think of this as a quick preview before you stop by the booth. ## Build a trusted data foundation Great analytics starts with data you can trust. That’s why dbt works directly with Google BigQuery, AlloyDB, and BigLake to organize SQL transformations where your data already lives. But we don’t stop at transformations. dbt adds the layers that turn raw pipelines into reliable systems. Documentation helps your team understand what’s happening. Version control keeps everything accountable. Testing ensures your data is actually correct. The result is a foundation your team can build on with confidence. ## One data platform, finally Data teams have been juggling too many tools for too long. With open table formats like Apache Iceberg, your data can flow seamlessly across BigQuery and Spark without being locked in. On top of that, dbt’s Semantic Layer creates a single source of truth for metrics. No more debating definitions across dashboards or teams. Everyone works from the same logic, whether they are writing SQL, building dashboards, or training models. It is a simpler way to think about your stack. One platform. Shared definitions. Less friction. ## AI-powered analytics development that actually works AI is only as good as the foundation behind it. dbt’s [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) brings structure to how teams build with AI so results are accurate and production-ready. At the booth, you’ll see how dbt fits directly into your workflow with a [dbt VS Code extension](https://docs.getdbt.com/docs/install-dbt-extension?version=2.0), powered by [dbt’s MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2.0) and integrated with Google’s Antigravity IDE. This means you can develop, test, and iterate without leaving your environment, while still taking advantage of everything dbt knows about your data. Here’s a few dbt Platform features which make that possible: - [**dbt Catalog**](https://www.getdbt.com/product/dbt-catalog) helps you visualize exactly what agents and LLMs are building, so nothing feels like a black box. - [**Semantic Layer**](https://www.getdbt.com/product/semantic-layer) gives AI access to well-defined metrics, improving accuracy and consistency. - [**Advanced CI** ](https://www.getdbt.com/blog/announcing-advanced-ci)lets you preview every change to your data before it ships, making review simple even when AI is involved. - [**State-aware orchestration with Fusion**](https://www.getdbt.com/blog/announcing-state-aware-orchestration) ensures all of this runs efficiently, so you get powerful results without unnecessary spend. Put together, these pieces make AI development not just possible, but practical for real enterprise use cases. We’re excited to show all of this in action at [Google Next.](https://www.getdbt.com/events/summit/google-cloud-next-2026) ## Find us at Booth #6606 We'll be at Booth #6606 at Google Cloud Next, April 22–24 in Las Vegas. Come find us to: - **See live demos** — If you are new to dbt view the standard on data transformation. There will be a demo dbt Platform alongside AI, including our VSCode extension, Semantic Layer, Catalog, Advanced CI, and Fusion orchestration. - **Book a meeting** — Schedule a 1:1 with a dbt expert to dig into how dbt can help with your BigQuery infrastructure initiatives - **Enter our raffle** — Stop by for a chance to win Apple AirPods Max (winners announced April 24 at 12:00 PM PT) ## Join our afterparty Kick off the week with us! dbt Labs, Fivetran, and 66degrees are hosting an exclusive afterparty at Eyecandy Lounge on the Mandalay Bay Casino floor on Wednesday, April 22, from 6:30–10:00 PM. Come for the music, cocktails, and elevated bites, and stay for the conversation. [Secure your spot](https://www.66degrees.com/events/next-26-afterparty?utm_source=dbtlabs&utm_medium=email&utm_campaign=googlenext2026). --- --- title: "Mistakes I made as the head of analytics (and what I’d do differently now)" description: "A former head of analytics on the 6 mistakes he made with dbt—and what he'd do differently now." url: "https://www.getdbt.com/blog/mistakes-i-made-as-the-head-of-analytics-and-what-i-d-do-differently-now" date: "2026-04-09" authors: ["Kyle Salomon"] categories: ["Insights"] --- # Mistakes I made as the head of analytics (and what I’d do differently now) Before I joined dbt as a solutions architect, I spent time leading an analytics and data engineering team. I managed the business analytics (BA) and data engineering (DE) teams, helped build out our data platform, and made _a lot_ of decisions I thought were right at the time. Spoiler alert: some of them were not. Now that I sit on the other side—helping customers adopt and scale dbt—I see the patterns everywhere. The same mistakes I made are the ones I watch teams make every single day. So consider this my confessional. These are the things I got wrong, what I wish I'd done instead, and what I'd tell any data leader who's willing to listen before they learn the hard way. A quick note: **Jack Kennedy** deserves a ton of credit here. Jack was the brains behind implementing dbt at our company. I was his manager; my job was to remove blockers and give him room to operate. A lot of the sharpest insights in this piece come directly from conversations with him about what we would've done differently. ## Mistake #1: I didn’t push my team to stay ahead of the platform This is the one that stings the most because it was entirely within my control. When we adopted dbt, we learned what we needed to learn to get the job done. And then… we stopped. New features would ship, new paradigms would emerge, and we'd collectively shrug and say _"we'll look at that when we need it."_ Classic mistake. ### What I should’ve done: - **Made continuous learning non-negotiable.** I should have pushed every person on my team to regularly dig into newly released features and the roadmap. Not just read the changelog—actually think about how those features could solve their day-to-day problems. - **Set up monthly knowledge-sharing sessions.** Each person owns a topic, learns it, and teaches the rest of the team. Simple, cheap, high-impact. - **Gotten the team through SA-level training.** The depth of understanding you get from that kind of enablement is a game-changer. It would have completely shifted how my team thought about the platform. ### And this isn’t just about dbt: I'm telling this story through the lens of dbt because that's where I live now, but this lesson applies to **every tool in your stack.** Your cloud data warehouse, your BI platform, your orchestration layer, your ingestion tools—all of them are shipping new capabilities constantly. I know for a fact we left a lot of meat on the bone with dbt, but I'd bet we were underutilizing our data warehouse and other tools just as badly. We simply never created a structured way for the team to investigate what each vendor was offering and learn from it. If I could do it over, I'd assign ownership of vendor relationships across the team. Someone owns staying current on dbt. Someone else owns the warehouse. Someone else owns the BI tool. Each person is responsible for knowing what's new, what's coming, and how it could help us. That kind of distributed awareness is how you stop leaving value on the table across your entire data stack—not just one platform. ### The “faster horse” problem: This scales beyond individual habits into how entire teams think about solutions. It's the classic Henry Ford analogy. If you asked people what they wanted, they would've said _"a faster horse."_ Ford didn't build a better horse—he built a car. A completely different solution to the same underlying problem. That's exactly what happens when a team is so locked into their current approach that they can't see a fundamentally better one sitting right in front of them. They keep breeding faster horses—optimizing workarounds, layering on complexity—instead of stepping back and asking whether the whole approach needs to change. When your team knows what's on the roadmap and what's newly available, the question shifts. They stop asking _"how do I make this workaround better?"_ and start asking _"do I even need this workaround at all?"_ That shift in thinking is everything. It's the difference between a team that's perpetually catching up and one that's building ahead of the curve. ### The mesh example: Our team operated as a classic **hub and spoke model**—a Center of Excellence made up of data engineers and analytics engineers (the hub) supporting domain teams across Finance, Product, Marketing, Sales, and others (the spokes). On paper, it's a solid structure. In practice, we made it harder than it needed to be. We ran everything out of **one monorepo and one project.** Every spoke team was working inside the same massive codebase as the hub. That meant they were exposed to things they didn't need to see, navigating complexity that wasn't theirs, and constantly dependent on the core team for changes that should've been within their own control. If I had taken the time to learn **dbt Mesh** and multi-project architecture when it was introduced, we could have fundamentally changed this dynamic: - **Simplified what the spoke teams saw.** Each domain team could have had their own project scoped to the models and data that mattered to them—not the entire estate. - **Given them more responsibility and autonomy.** Most of those spoke teams _wanted_ more ownership over their data. Mesh would have let us give it to them in a controlled, well-bounded way instead of the all-or-nothing access of a monorepo. - **Set clear boundaries between hub and spoke.** The core team could have published curated, contracted interfaces for the spokes to build on, instead of everyone reaching into the same tangled codebase. Instead, we stayed in our old patterns—one repo, one project, one bottleneck—and missed an opportunity to fundamentally improve how we collaborated across domains. ### Why this matters: The dbt platform is evolving faster than most teams realize, and the ones who invest in staying current are the ones who get outsized value. The teams that treat dbt like a static tool are missing out on massive capability. ## Mistake #2: We were too married to our original data estate We built our data estate, and then we treated it like it was sacred. Every layer, every naming convention, every model structure—it was all inherited from the original design, and we never seriously questioned whether it still made sense. **It didn't.** ### What went wrong: - We should have changed the shape of our layers and our data as our needs evolved. Instead we kept bolting things onto a structure that wasn't designed for where we were headed. - Data discoverability suffered. When your architecture doesn't reflect how people actually need to find and use data, everything gets harder. ### The lift-and-shift trap: When we moved to dbt, the question was: _do we lift-and-shift first and fix later, or do we rebuild from scratch?_ Here's what I'd say now, with the benefit of hindsight: - **If your team is coming from nothing**—lift and shift into a controlled environment. Don't stress about perfection. Get into dbt, get version control, get testing. _Then_ begin to fix. - **But the follow-up is absolutely necessary.** You _must_ reorganize into proper architecture after the initial migration. Too many teams treat lift-and-shift as the finish line. It's the starting line. - **Before you build anything net-new**, make sure your business objects are well-defined and you have at least a base layer of documentation in place. Otherwise you're building a house on sand. ### But first ask yourself: should we even move this? This is the lesson Jack and I regret the most from our migration. Before you lift and shift _anything_, you need to ask two questions: 1. **Do we actually need this?** 2. **Does someone own this?** If the answer to either question is no—_do not bring it into your new environment._ Full stop. Don't migrate garbage just because you can. We learned this the hard way. One of our engineers wrote a script to lift and shift _everything (it was pretty sweet)_—every model, every object, regardless of whether it was actively used or owned by anyone. We knew a lot of it was garbage. We did it anyway. And then we spent an enormous amount of time dealing with the consequences: maintaining things nobody needed, debugging things nobody understood, and cluttering an environment that was supposed to be a fresh start. The instinct during a migration is to say _"let's just move it all and sort it out later."_ Resist that instinct. Be ruthless about what gets to come along. If it doesn't have a clear owner and a clear purpose, leave it behind. Your future self will thank you. ### Why this matters : If you’re thinking about migrating to dbt platform or expanding your deployment, you need to hear this. The migration itself is just step one. The real value comes from what you do *after—*rethinking your architecture, improving discoverability, and being willing to break from the old design. ## Mistake #3: Our job orchestration was stuck in the past We ran our jobs the way everybody runs their jobs: daily schedules, hourly schedules, cron expressions, and a prayer that nothing breaks overnight. It worked. Until it didn't scale. ### What we should’ve done differently: - **Used triggers instead of schedules.** Instead of running everything on a timer, we should have leaned into triggering jobs off of completed upstream jobs. Take advantage of the API—let the work flow naturally instead of forcing it onto a clock. The tools to do this existed. We just defaulted to what was familiar: cron jobs and hope. - **Communicated better with stakeholders.** If my team of data analysts that were closer to the business had truly understood how orchestration worked under the hood, they could have set better expectations with the business. Instead, we were reactive—explaining delays after the fact instead of designing for reliability upfront. And here's the kicker: **just because Rapid Onboarding gets you started doesn't mean that's the entire window.** There's a lot of value outside of that initial onboarding phase that teams completely overlook. ### The SAO connection: Now, I want to be clear—I left my previous role before Fusion and state-aware orchestration were even announced. There's no way we could have known SAO was coming. But that's almost the point. If we had built our orchestration the _right_ way from the start—using triggers, thinking in terms of dependencies rather than schedules—the team I left behind would have been in a much stronger position to adopt SAO when it arrived. The mental model would've already been there. The architecture would've been aligned. Instead, they'd have to unwind years of schedule-based patterns before they could even begin to take advantage of it. The lesson isn't "you should've predicted the future." The lesson is: **build with the best patterns available today, and you'll be ready for whatever ships tomorrow.** ### Why this matters: Orchestration is where a lot of hidden pain lives—wasted compute, late data, frustrated stakeholders. Adopting trigger-based orchestration now isn't just about solving today's problems, it's about being ready for where the dbt platform is headed. ## Mistake #4: We never treated our dbt infrastructure as code This one is squarely Jack's domain; he owned this and would be the first to say we should have done it sooner. I'm not a Terraform expert, but the lesson is clear. ### What we missed: - We should have used **Terraform with the dbt provider** to manage our jobs, groups, and infrastructure as code from the start. - Instead of manually configuring jobs in the UI, we could have had our entire dbt platform infrastructure defined in a Terraform repo—**1:1 mapping** between what's in the repo and what's running in production. - This would have given us: - **Reproducibility**—spin up environments, replicate configurations, no more "who changed that job?" - **Control**—permissions, jobs, resource blocks, all managed in version-controlled code - **Auditability**—every change is a PR, every PR is reviewable ### Why we didn’t do it: Honestly? It was a capability that was introduced after we were already on the platform, and we fell into the same trap as Mistake #1—we didn't stay ahead of what was new and possible. By the time we realized Terraform could have transformed how we managed dbt, we had already accumulated a bunch of manual configuration debt. ### Why this matters: If you are scaling your dbt deployment or managing multiple projects and environments, Terraform is a game-changer. It's the kind of thing that sounds like overhead upfront but pays for itself ten times over as complexity grows. ## Mistake #5: We ignored governance until it was too late Nobody wakes up excited to talk about data governance. I get it. But ignoring it cost us more time and headaches than almost anything else on this list. ### What went wrong: - **We built data products without clear owners.** If nobody owns it, nobody maintains it. And if nobody maintains it, it rots. We should have had a hard rule: _don't build objects that don't have owners._ Only pull in data products that have clear, accountable ownership. - **We underutilized our resident architect.** We had access to an RA (during our early stages of dbt) and we used them for short-term tactical wins—firefighting, basically. We should have engaged them for **long-term strategic value**: architecture reviews, best-practice adoption, asking what’s the future of dbt, and building a foundation that would scale. - **We ignored the dbt style guide**, especially around marts. This sounds minor, but when your marts layer is a mess, everything downstream suffers. Naming conventions, model organization, documentation—all of it compounds. - **We had little to no documentation and context.** This is one I knew was a problem from the day I took over. Our models were poorly documented, and instead of fully understanding what dbt, MCP, and the catalog could do for us, we reached for outside tools—Confluence, Google Docs, our company intranet—to try to fill the gap. We were patching documentation together across many different platforms when the capability to do it properly was sitting inside the tools we were already paying for. I just never took the time to research the best solution and understand what our vendors actually offered. - **We let data testing slide.** Documentation wasn't the only thing we neglected—data quality suffered right alongside it. We had _some_ testing in place, but nowhere near enough. It was the same pattern: we knew it was a problem, we knew better solutions existed, and we still let it slide for far too long. Governance, documentation, and testing are a package deal—when you let one go, the others tend to follow. ### Why this matters: Governance isn't sexy, but it's the difference between a data platform that scales and one that collapses under its own weight. Ownership, style guides, and proper RA engagement aren't "nice to haves"—they're the foundation. And here's the part that makes this mistake feel even bigger in hindsight: **AI is coming for every data platform, and it needs context to be useful.** Everything I neglected—documentation, testing, clear ownership, well-defined processes—is exactly the context that AI needs to actually help. Without it, AI doesn't accelerate your team. It just confidently generates garbage faster. Think about the dbt MCP server. If you expose it to an LLM today, that LLM is going to interact with your models, your definitions, your metadata. If none of that is documented, tested, or governed—what exactly is the LLM working with? It can't be trusted. Not because the tooling is broken, but because the foundation underneath it is hollow. The AI is only as good as the context you've given it, and if that context is incomplete, inconsistent, or missing entirely, you've just handed an LLM the keys to a house with no blueprints. Of all the mistakes on this list, this might be the one that compounds the hardest going forward. I did no favors to the leaders who came after me. I left no context behind—no documentation to build on, no testing framework to trust, no governance structure to lean into. And now, in a world where AI is supposed to unlock the next wave of productivity for data teams, the teams that ignored this work are the ones who will be the least prepared to take advantage of it. The ones who invested in governance, documentation, and testing? They'll plug AI into a well-documented, well-tested estate and actually get value from it. Everyone else will be starting from scratch—or worse, trusting outputs they shouldn't. So don't be me. Get this right _now_—or better yet, yesterday. ## Mistake #6: I stopped failing fast This one is personal. Early in my career as a web developer and data analyst, I lived by a simple principle: **fail fast.** Try things. Figure out quickly whether something is going to work or not. If it doesn't, pivot. Iterate. Move on to the next solution. Don't fall in love with your first attempt; fall in love with finding the _right_ answer, however many attempts it takes to get there. Be AGILE! That mindset was my edge. It's how I learned, how I solved problems, and how I built confidence in tackling things I'd never seen before. And then I became the head of analytics, and I lost it. ### What happened: When I took over a larger team, I defaulted to **what was safe and "what worked."** I stopped taking swings. Instead of rapidly testing approaches and iterating toward the best outcome, I leaned on the familiar. If something was functioning—even if it wasn't great—I left it alone. I chose stability over experimentation, and I told myself that was the responsible thing to do as a leader. But that's not leadership. That's risk avoidance dressed up as pragmatism. ### What I should’ve done: - **Kept the fail-fast mentality, even at scale.** Leading a team doesn't mean you stop experimenting—it means you create an environment where _the team_ can experiment safely. Timeboxed spikes. Low-risk proof-of-concepts. Small bets with big learning potential. - **Modeled the behavior I valued.** If I wanted my team to be bold and iterate quickly, I needed to show them what that looked like—not retreat into safe, predictable decisions. - **Stayed true to who I am.** This is the real lesson. When you step into a bigger role, the instinct is to become someone else—someone more cautious, more "executive." But the traits that got you there are the ones you need to keep. You have to bring your whole self to the job, not a watered-down version that's afraid to break things. ### Why this matters: No one should lose their edge when they take on more responsibility. If anything, a bigger team needs that willingness to try, fail, learn, and pivot _more_ than a small one does. The problems are bigger, the stakes are higher, and the cost of staying stuck in "what works" only compounds over time. Be true to yourself. Create the culture. Take the chances you believe in. And if they don't work out—good. You just learned something faster than everyone who's still playing it safe. ## The meta-lesson Every single one of these mistakes boils down to one thing: **we didn't invest time in understanding what was possible before we needed it.** We were always reactive. Always catching up. Always learning about a feature or a pattern _after_ we'd already built around the old way of doing things. If I could go back and give myself one piece of advice, it would be this: **Carve out time—real, protected, recurring time—for your team to explore what the tools you have at your disposal can do for you. Not what it does today for you, but what it _could_ do tomorrow. That's where the compounding value lives.** _Thanks to Jack Kennedy for being the real architect behind all of this—both literally and figuratively. Any of the smart ideas in here are probably from conversations with him. The mistakes are all mine._ --- --- title: "Operationalize analytics agents: dbt AI updates + Mammoth’s AE agent in action" description: "Learn how to operationalize your analytics agents by building context for LLM models with dbt and MCP servers." url: "https://www.getdbt.com/blog/operationalize-analytics-agents-dbt-ai-updates-mammoths-ae-agent" date: "2026-04-03" authors: ["Sai Maddali"] categories: ["Learn"] --- # Operationalize analytics agents: dbt AI updates + Mammoth’s AE agent in action In conversations with customers of all sizes - from small startups to some of the largest enterprises in the world - three themes consistently surface when it comes to AI and analytics workflows. The first is conversational analytics. Business users want to ask questions of their data, but in a governed, reliable way. The answers they get need to be consistent, predictable, and above all, accurate. Accuracy itself exists on a spectrum. Sometimes, you need precision every single time. Sometimes, a well-reasoned ballpark is enough. The challenge is building agents and experiences that can navigate that spectrum for your end users. The second is development acceleration. This is one of the biggest use cases we see across the industry: How do you accelerate the speed at which you ship data products with agents? Modern LLMs are powerful and very good at generating code, But they don't necessarily understand the dependencies in your data. How do you make these agents effective? How do you help them understand what your data actually means, so they can generate code that makes sense for your codebase, your data architecture, and your schema? The third is cost control. How do you get all these benefits - the magic, the speed, the accuracy - without your costs ballooning? Token costs might be cheap in isolation. When there's a lot of usage, however, not all of it is equally productive. The challenge is balancing consistency, accuracy, and speed without letting costs spiral. **** ## Giving agents the context they need What we're trying to do at dbt is give AI tools, agents, and the experiences you build on top of LLMs the context they need about your data so they can make effective decisions. LLMs are good at generating code. But are they actually good at generating code that makes sense **for your organization**? Take impact analysis as an example. When you give an agent a prompt to make a change to a dbt model because you've added a new schema upstream, how do you ensure the agent understands the downstream consequences of that change? It needs to know: if I make a change to this model, I'm affecting five tables that depend on it - and those five tables are being used by fifteen dashboards downstream. That's impact analysis. It's one of the most important capabilities an agent needs to operate effectively in a real data environment. The same logic applies to migration scenarios. Whether you're moving from a legacy system, a stored procedure, or simply modernizing your stack into dbt, agents need to guarantee that the logic of the old system matches the dbt models they generate. The agent may be good at writing code. But can it guarantee the logic stays the same? For relatively simple cases, probably. For complex ones with many dependencies, probably not. That's the gap we're working to close. The same pattern applies when you look across your dbt models at scale. If you have hundreds or thousands of models - and some teams have tens of thousands - you will have redundant logic. That redundancy drives complexity and cost. The tools we're building help you understand where that duplication exists and address it proactively, not just reactively. ## Operationalizing is where the real work is Once you've built a data product, deploying it into production and keeping it running is where a lot of your time and cost actually go. Operationalizing your data is one of the hardest things you do. One of the most critical metrics here is mean time to resolution. When a dbt pipeline fails, for whatever reason, the goal is to troubleshoot it as fast as possible and get back into production. Agents can help dramatically - but only if they have the context to diagnose what actually went wrong. Schema evolution is another persistent pain point. A source schema changes and breaks things downstream. The challenge isn't just detecting schema evolution automatically - it's ensuring the agent has the context to understand exactly what changed, and then make the right fixes across your [dbt directed acyclic graph (DAG)](https://www.getdbt.com/blog/dag-use-cases-and-best-practices) so that the end dashboard isn't affected. Cost optimization follows the same pattern. Most of us have at least one dbt model that's large, slow, and expensive. If you give an agent a one-shot prompt to optimize it, it can do a decent job when the model is relatively simple. However, for complex models with many dependencies, you need the agent to have enough context to preserve the same logic while making the necessary performance and cost improvements. That's where the tooling matters. ## Injecting dbt context into any AI tool As your data architecture evolves and you adopt more AI tools, the goal should be a consistent layer of context that works regardless of which data platform you run on or which AI client you use. Whether your teams are using Claude, ChatGPT, a BI tool with an embedded agent, or something you've built internally, they should be getting the same consistent, accurate answers about your data. What will stay consistent - regardless of how the AI stack evolves - are your tables, your data logic, and the context you've built into that logic. One of the most effective ways we've seen customers achieve this is through the [dbt MCP server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server). The dbt MCP server takes core dbt concepts - the metadata API, the dbt CLI commands, the admin API - and exposes them to any agent as specific tools it can call. What we're seeing with customers who use it is striking: their agents are significantly more effective at completing specific tasks. They have better context to operate with, which means lower token usage and better cost efficiency. They're faster. And the results are more accurate and more consistent, especially for natural language queries. Bottom line: If you're using any AI tools with your dbt projects today, I'd strongly recommend checking out the dbt MCP server. ## Agents across the dbt lifecycle For dbt platform customers, we're also investing in AI-native experiences across the dbt product lifecycle in the form of [dbt agents](https://docs.getdbt.com/docs/dbt-ai/dbt-agents). One example is the developer agent we're building into [dbt Studio](https://docs.getdbt.com/docs/cloud/studio-ide/develop-in-studio) and [Canvas](https://docs.getdbt.com/docs/cloud/studio-ide/develop-in-studio). We want to bring the same capabilities you see in frontier coding agents like [Claude Code](https://code.claude.com/docs/en/overview) or [Cursor](https://cursor.com/) directly into dbt Studio. That way, you can develop and ship data products more productively without leaving your environment. We also have the [analyst agent](https://docs.getdbt.com/docs/dbt-ai/analyst-agent) and the catalog agent, both of which are currently in private beta. If you're already using [dbt Catalog](https://www.getdbt.com/product/dbt-catalog), you can ask natural language questions and get answers about your data. The agent acts on behalf of the user, writes SQL, and uses a set of dbt tools under the hood. It's fully secured based on the role-based access control (RBAC) privileges you've granted to that user in your warehouse and within dbt. Similarly, within dbt Catalog, you can ask questions about what data is available. This makes it a genuinely useful tool for data discovery, not just a documentation repository. A new data analyst who joins a team and needs to understand what's been built, what's available, and how models interconnect can now get that context directly through the catalog. It spares them from digging through code or waiting for someone on the data team to walk them through it. ## What production-grade agentic development actually looks like One pattern we're seeing from teams building with agents effectively is that the quality of the output is directly proportional to the quality of the context you provide. This is a make-or-break difference that determines whether your projects even make it out of development. Agents that operate without rich context - documentation, business logic, schema definitions, best practices - produce code that you'd never push to production. Agents with that context produce code that looks like it was written by your best engineer. The teams seeing the most success are investing in codifying their standards. That means using documented best practices, defined schemas, and clear technical specs. A big caveat: this requires having the discipline to build those artifacts before asking an agent to implement anything. The agent then follows those standards every single time. This includes writing the documentation that engineers often skip. The combination of rich context, the dbt MCP server for data access and model validation, and structured workflows for review and deployment is what makes the difference between AI that's impressive in a demo and AI that actually ships production-grade data products. ### Mammoth’s Analytics Engineering Agents Mammoth Growth is one example of a company that successfully leveraged dbt MCP servers to operationalize AI agents. Their attempts at AI integration before MCP always stumbled over dbt code quality because the agents were only using their baseline capabilities. This meant their code style, architecture, and accuracy were far below production standards. Bringing on an MCP server gave their agents the context necessary to recognize common coding tasks and column names, replicate repository styles, and provide correct formatting for outputs. This meant that code generated by the agent fits within the framework of Mammoth’s projects. Access to these agents lets Mammoth’s data teams match the speed of business, so they don’t have to compromise between keeping pace and project quality. [See how they did it in this demo](https://www.getdbt.com/resources/webinars/operationalize-analytics-agents-dbt-ai-updates-mammoth-s-ae-agent-in-action). The combination of rich context, the dbt MCP server for data access and model validation, and structured workflows for review and deployment is what makes the difference between AI that's impressive in a demo and AI that actually ships production-grade data products. ## What's ahead The near-term investment at dbt is in making it easier to inject dbt context into whatever AI tools you're using. That can be through the dbt MCP server, the analyst and developer agents we're building natively into the platform, or the integrations that enable self-serve analytics for business users. The longer-term vision is a world where agents can proactively surface issues, manage schema evolution, optimize costs, and respond to incidents. Ideally, they’ll do this, not just when prompted, but as persistent, always-on participants in your analytics workflow. We're early in that journey. But the foundation is being built now, and the teams that invest in getting the context layer right today will be the ones best positioned to take advantage of it. If you want to explore this further - the agents beta, a demo of the dbt platform, or just a conversation about how to get started - [reach out to the dbt team](https://www.getdbt.com/contact). --- --- title: "How AI is reshaping the way data practitioners work" description: "What happens to data work when AI changes everything? The hosts of The View on Data podcast share what's shifting and what isn't." url: "https://www.getdbt.com/blog/how-ai-is-reshaping-the-way-data-practitioners-work" date: "2026-04-03" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # How AI is reshaping the way data practitioners work In this episode of The View on Data, hosts Faith McKenna, Paige Berry, and Erica "Ric" Louie sit down with Sam Ferguson, staff product designer at dbt Labs, to talk about what it means to design for data practitioners and how AI is reshaping that work in real time. The conversation covers everything from embedded natural language in SQL, to the typist analogy you didn't know you needed, to why your code comments might be more valuable than you think. 🎧 Listen & subscribe: [Spotify](https://open.spotify.com/show/0Itxg3BoF6zaQdORfVuY6I?si=6133f88257624765) | [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-view-on-data/id1835755886) | [Amazon Music](https://music.amazon.com/podcasts/84c47dbe-adf7-4d33-a371-a1fe55255495/the-view-on-data) | [YouTube](https://www.youtube.com/playlist?list=PL0QYlrC86xQk2WktL3vbdB1FcWlpay7VQ) [Watch video](https://youtu.be/Dj_H3zcVV7Q) ## From Mode to dbt: designing for the people who write the queries Sam came to dbt Labs after nearly a decade at Mode Analytics, where the product was built around a core belief: analysts who write SQL deserve great tooling. That philosophical overlap with dbt made the transition feel natural. Both companies were thinking hard about code-first data practitioners long before it was a mainstream conversation. But the move also meant expanding her mental model. At Mode, the center of gravity was the analyst. At dbt, it shifted upstream to analytics engineers and data engineers, people whose job is less about ad hoc answers and more about building the infrastructure that makes those answers possible. One of the projects that bridged those two worlds was an AI assist feature she worked on at Mode, back when text-to-SQL was just starting to emerge. The idea: instead of switching between a chat window and your SQL editor, you could embed natural language directly inside your query. Write the SQL you know, leave a placeholder in plain English for the parts you don't, and let the model fill in the gaps. The closer the prompt was to the actual output, the better the result. That experiment is still shaping how Sam thinks about AI UX today, specifically whether inline context (like code comments) might outperform separate AI instruction directories. More on that in a second. ## Comment your code. Your future agent will thank you. Sam mentioned that the closer a user's natural language prompts were embedded to the actual code, the higher the output quality. Ric and Paige didn't need convincing. Ric talked about writing comments not just for other humans, but for the logic itself: why this filter exists, what this variable is actually doing, what the weird edge case is and why it matters. Paige echoed it. She's learned more from reading well-commented code than from almost any other source. The "feed two birds with one scone" framing stuck: good commenting practices already make your codebase easier for humans to navigate. They might also make it significantly easier for agents to work with. That's not a trade-off. That's just a bonus. ## The typist analogy: what happens when a constraint disappears Sam referenced a book called _Reshuffle_ that offers a useful frame for thinking about how roles evolve under technological pressure. The example: typists. Typing used to be a discrete, specialized job because editing was expensive. If you made a mistake on a typewriter, you started over. Word processors didn't just make typists faster. They removed the constraint that made the role distinct. Typing got absorbed into everything else. The parallel for product design: one of the biggest historical constraints has been the cost of building. Designers spent enormous time simulating products before they existed, through mockups, prototypes, and user tests, to de-risk development before a single line of production code was written. Now, anyone at dbt Labs can come to a meeting with a working concept. The simulation phase is collapsing. Which means design needs to reorganize around different constraints. Sam's candidate for the new constraint: **human attention**. In a world of infinite features and infinite answers, the scarce resource isn't information. It's cognitive bandwidth. What gets prioritized? What gets ignored? What does "good" look like when you can generate a hundred options before lunch? For data teams, the parallel lands in a similar place. If agents can answer any question, the hard part becomes figuring out which questions actually matter. ## Outcome-driven workflows: familiar territory, new stakes A big theme in the episode was the shift from output-driven to outcome-driven AI workflows. In the early GitHub Copilot days, the goal was code completion: predict the next line, write faster. Then it became code validation: review what the agent wrote, approve or reject. Now, with tools like Claude Code, you might not look at the files at all. You describe what you want the end result to look like, and you evaluate from there. That's a meaningful shift in how practitioners relate to their work. But as the hosts noted: this isn't a new conversation in data. "Do you really need to spend 2,000 years optimizing this table if you're not actually getting to the answer the business cares about?" Faith put it directly. The outcome vs. process tension has always been there. AI just raises the stakes and maybe forces teams to get better at navigating it. Ric made a similar point through a different analogy: state-aware orchestration. The idea of saying "I want this table updated at this frequency, with these freshness guarantees" is structurally the same as saying "I want this outcome; figure out what needs to happen upstream." The framing is different. The principle isn't. ## On hero mode, accountability, and bringing people along One of the more honest stretches of the conversation: just because you _can_ build something start to finish with AI assistance doesn't mean you should. The concern isn't capability. It's isolation. If one person on a team can move twice as fast as everyone else but isn't documenting, sharing, or collaborating, the team doesn't go faster. It goes uneven. Ric's analogy was a boat with mismatched rowers. The strong rower doesn't win races alone. Faith added an accountability angle that felt important: when AI produces bad output, "ChatGPT did that" isn't an acceptable response. You prompted it. You shipped it. You own it. That's true in training content, in data models, in dashboards, in anything that lands in front of a stakeholder who trusted you. The flip side is worth naming too: AI makes it easier to create nonsense at scale. Overly complex queries, undocumented logic, architectures nobody can explain. The principle Paige has internalized is asking, at every step, what are you actually trying to accomplish? That question doesn't get less important when AI is in the loop. It gets more important. ## Governance, context, and the onboarding problem Something Sam is thinking about a lot right now: context management for agents isn't that different from onboarding a new teammate. When someone joins your data team, you spend time helping them understand how you define things, where the data lives, which models are canonical and which ones are legacy noise. That knowledge transfer is slow and often informal. Now imagine having to codify all of that for an agent, clearly enough that it can reason about your stack, your terminology, and your business logic without constant correction. That exercise forces you to ask: are these practices actually right? Do they still hold? Have we documented them anywhere, or do they just live in Paige's head? The upshot: the work of building good AI context is also the work of building a better-documented, more navigable data environment for humans. Again, two birds. ## Advice for staying level-headed The episode closed with some grounded takeaways for data and tech professionals trying to navigate all of this without losing their minds. **Paige:** Try stuff. Write it down. Track what's working and what isn't. The record you build over time becomes its own source of encouragement, proof that you're learning, even when it doesn't feel like it. **Sam:** Experiment fast, because the switching cost has never been lower. And think about the constraints your role has historically worked around. Which ones are disappearing? Which ones are you glad to hand off? Start there. **Ric:** Think in principles. Before adding AI to a workflow, ask what you're actually trying to get out of it and what guardrails you need to make sure it doesn't create more work than it saves. Move quickly, but bring people with you. **Faith:** Use AI to enhance the experience, not to replace the thinking. The hard work of figuring out what learners actually need, what questions actually matter, what's just enough complexity and no more, that's still yours to do. AI can help you do the other parts faster so you have more time for it. And maybe most importantly: good leadership, in AI or otherwise, raises the floor beneath everyone rather than enabling a few to touch the ceiling. --- --- title: "What does Databricks Lakebase mean for analytics engineers?" description: "Learn how to connect dbt, when to migrate, and what the tradeoffs are for your data team." url: "https://www.getdbt.com/blog/databricks-lakebase-analytics-engineers" date: "2026-03-30" authors: ["Joey Gault"] categories: ["Pulse"] --- # What does Databricks Lakebase mean for analytics engineers? ## Understanding the Lakebase architecture Lakebase exposes a PostgreSQL wire protocol interface to your Databricks environment. Instead of routing all queries through Databricks' native SQL endpoints or Spark-based interfaces, teams can now connect through standard PostgreSQL tooling. Any tool that works with PostgreSQL can connect to your Databricks environment—that's the core unlock. The PostgreSQL compatibility layer doesn't replace the underlying [lakehouse architecture](https://docs.getdbt.com/guides/databricks?step=1). It adds an access pattern that coexists with your existing connection methods. ## Connecting dbt to Lakebase Connecting [dbt](https://www.getdbt.com/product/what-is-dbt) to Lakebase means using the [dbt-postgres adapter](https://docs.getdbt.com/docs/supported-data-platforms), not the dbt-databricks adapter your team may currently use. That distinction changes how you configure and manage your dbt projects. For [dbt Core](https://docs.getdbt.com/docs/introduction), installation follows the standard pattern: `python -m pip install dbt-core dbt-postgres` Your connection configuration mirrors a PostgreSQL setup. Find the hostname in your Databricks workspace under Compute > Database instances > Connect with PSQL—typically structured as `instance-123abcdef456.database.cloud.databricks.com`. The default database name is `databricks_postgres`. Authentication is where things get operationally significant. The dbt-postgres adapter supports username/password authentication, but this requires enabling Native Postgres Role Login in your Databricks environment. You'll use the role name as the username and work within [Databricks' Postgres role and privilege management](https://docs.databricks.com/en/database/privilege-management.html). OAuth tokens are available as an alternative, though they require hourly refresh—a constraint worth factoring into your CI/CD pipeline design. ## Architectural implications for data teams Lakebase creates a real decision point: migrate existing Databricks-based [dbt projects](https://docs.getdbt.com/docs/build/projects) to use the PostgreSQL interface, stay on your current Databricks connection, or run a hybrid approach. For teams already running dbt on [Databricks](https://docs.getdbt.com/guides/databricks?step=1) using the dbt-databricks adapter, migration to Lakebase isn't automatic—and it's often not the right call. Your existing workflows, [materializations](https://docs.getdbt.com/docs/build/materializations), and optimizations were built around Databricks' native capabilities. The PostgreSQL interface won't expose the same performance characteristics, particularly around Delta Lake-specific operations or Databricks SQL extensions. New projects face a different calculation. If your analytics engineering team has deep PostgreSQL expertise and limited Databricks experience, Lakebase lowers the learning curve. Familiar patterns apply, existing tooling works, and knowledge transfers across PostgreSQL-compatible platforms. ## Deployment and infrastructure considerations Lakebase connections through dbt require attention to network topology and security. If you're using [dbt](https://www.getdbt.com/product/dbt), ensure that dbt's IP addresses can reach your Lakebase instance. For environments where Lakebase isn't publicly accessible, SSH tunneling through a bastion host is the path forward. SSH tunnel configuration for Lakebase follows the same pattern as standard PostgreSQL connections in dbt. You'll configure bastion server details in your [connection settings](https://docs.getdbt.com/docs/deploy/deployments), and dbt generates a public key to add to the bastion server's `authorized_keys` file. Each new SSH tunnel connection generates a unique key pair, so build a key management process into your infrastructure operations. Connection timeout management matters here. The tunnel must stay active throughout dbt job execution. dbt sends keepalive checks every 30 seconds and terminates the connection after 300 seconds without a response. Your bastion host configuration needs to align with those parameters—misalignment is the most common cause of premature disconnections. ## Adapter differences and feature parity Using dbt-postgres with Lakebase means working within PostgreSQL's feature set, not Databricks' full capabilities. The dbt-postgres adapter doesn't include the Databricks-specific materializations, optimizations, or configurations that come with the dbt-databricks adapter. That affects [incremental models](https://docs.getdbt.com/docs/build/materializations), [snapshot strategies](https://docs.getdbt.com/docs/build/snapshots), and performance tuning. Techniques that work well with Delta Lake through the native Databricks adapter don't always translate to the PostgreSQL interface. Test your specific use cases—especially large-scale incremental processing or complex merge operations—before committing. Your dbt project configuration, profile setup, and model-level configs will look like a PostgreSQL project, not a traditional Databricks one. Plan for that difference when onboarding team members or updating documentation. ## Strategic considerations for analytics engineering teams If your organization runs multiple data platforms—Databricks for some workloads, PostgreSQL-compatible systems for others—Lakebase's compatibility layer could enable more standardized tooling and processes across platforms. But standardization has tradeoffs. You may give up platform-specific optimizations for cross-platform consistency. The right answer depends on your team's skill sets, existing infrastructure, performance requirements, and long-term platform strategy. Also worth evaluating: adapter maturity. The [dbt-databricks adapter](https://www.getdbt.com/product/how-dbt-platform-compares) was built specifically to optimize for Databricks' architecture. The dbt-postgres adapter is mature for PostgreSQL workloads, but it may not expose all of Lakebase's capabilities or optimizations. Understand what you gain and lose in that translation before committing to either path. If you're running dbt at scale across multiple platforms, [dbt](https://www.getdbt.com/signup) provides managed infrastructure, job scheduling, and environment management that reduces this operational overhead significantly—worth evaluating as you think through your long-term setup. ## What Lakebase means going forward Databricks Lakebase is an interesting convergence of lakehouse architecture and traditional database interfaces. For analytics engineers, it's not a mandatory migration or an automatic upgrade over existing Databricks workflows. It's an additional option that fits specific use cases, team compositions, or strategic directions. The deciding questions: Does PostgreSQL compatibility solve a real problem for your team—skill set alignment, tool compatibility, or multi-platform standardization? If yes, Lakebase deserves serious evaluation. If your current Databricks workflows perform well and your team knows the existing tooling, Lakebase isn't an immediate priority. As with any architectural decision in the data platform space, the right choice depends on your specific context, constraints, and objectives. Lakebase expands the option set for how analytics engineers can work with Databricks. Understanding those options is the first step toward making an informed decision about your data infrastructure. [Ready to start building? Sign up for dbt for free.](https://www.getdbt.com/signup) Or [talk to our team](https://www.getdbt.com/contact) about what the right setup looks like for your organization. ## Databricks Lakehouse FAQs **What is Lakebase and how does it differ from traditional Databricks deployments?** Lakebase is a Databricks offering that exposes a [PostgreSQL wire protocol](https://www.postgresql.org/docs/current/protocol.html) interface to the Databricks environment. Traditional Databricks deployments use native SQL endpoints or Spark-based interfaces. Lakebase adds support for standard PostgreSQL tooling and protocols alongside those existing methods—it doesn't replace the underlying lakehouse architecture. **How do you connect dbt to Lakebase?** Use the [dbt-postgres adapter](https://docs.getdbt.com/docs/supported-data-platforms), not the dbt-databricks adapter. Install with `python -m pip install dbt-core dbt-postgres` and configure it like a PostgreSQL connection. Find the hostname in your Databricks workspace under Compute > Database instances > Connect with PSQL. The default database is `databricks_postgres`. For authentication, use username/password (requires Native Postgres Role Login) or OAuth tokens (require hourly refresh). Non-public environments need SSH tunneling through a bastion host. **Should existing Databricks dbt projects migrate to Lakebase?** Probably not, unless there's a specific reason to. If you're running dbt with the dbt-databricks adapter today, your workflows and optimizations are built around Databricks' native capabilities. The PostgreSQL interface doesn't offer the same performance characteristics for [Delta Lake](https://delta.io/) operations. New projects—or teams with deep PostgreSQL expertise—have more to gain from Lakebase's familiar interface and lower learning curve. --- --- title: "Introducing the dbt Community Champions Program" description: "Building the future of analytics engineering, together." url: "https://www.getdbt.com/blog/introducing-the-dbt-community-champions-program" date: "2026-03-26" authors: ["Bolaji Oyejide"] categories: ["Community"] --- # Introducing the dbt Community Champions Program Today, we're thrilled to announce the launch of the [**dbt Community Champions Program**](https://www.getdbt.com/dbt-champions)—a new initiative to recognize and empower the remarkable practitioners who make our community what it is. The dbt Community is powered by practitioners who share knowledge, mentor newcomers, create content, and push the boundaries of what's possible with analytics engineering. The Champions program formalizes our commitment to support these community leaders and amplify their impact. But this isn’t just about recognizing a cohort; it’s about strengthening the experience for everyone in the dbt Community. By lifting up Champions as visible resources, role models, and trusted peers, we’re creating more opportunities to learn, connect, and build together. This inaugural cohort is invitation-only, and we plan to open applications for future cohorts later this year. ## Meet the founding Champions Learn more about the program and meet the Champions on the [dbt Champions page](https://www.getdbt.com/dbt-champions). We're proud to introduce a **founding cohort of dbt Champions** from across six continents, representing diverse industries, company sizes, and technical backgrounds. These practitioners have demonstrated exceptional leadership through content creation, community mentorship, speaking engagements, and contributions to the modern data stack ecosystem. Our Champions span financial services, healthcare, technology, retail, and beyond—bringing perspectives from startups to Fortune 500 enterprises. They are analytics engineers, data engineers, platform architects, and technical leaders who have chosen to invest their expertise in helping others succeed. ![Champions quote- Jenna Jordan](https://cdn.sanity.io/images/wl0ndo6t/main/a307e00d24a0a8d5a0a5f70758a36af42c601971-1200x1200.png) ## What Champions do dbt Champions are more than community members; they're co-creators of the dbt ecosystem. They: - **Create content** that educates and inspires: blog posts, tutorials, and technical deep-dives - **Mentor community members** through Discourse, Slack, meetups, and virtual events - **Share thought leadership** at conferences, webinars, and podcasts - **Provide product feedback** that shapes the future of dbt products - **Organize and participate** in meetups and community events ## How we support our Champions We’re committed to supporting Champions with meaningful benefits: - **Connection** with fellow Champions in dedicated spaces - **Early access** to product features and beta programs - **Professional development** opportunities - **Social amplification** of Champion content - **Champions swag** and merchandise - **VIP experiences** at dbt Summit ## Why this matters The Champions program is about more than recognition; it's about building a sustainable model for community-driven innovation. By investing in our most engaged practitioners, we're creating: - **Better learning resources** through Champion-created content - **Product improvements** informed by real-world practitioner insights - **Stronger peer support** through mentorship and community guidance - **A more connected community** across geographies, industries, and experience levels Our goal is ambitious: to foster a culture of knowledge sharing and continuous learning across the entire data ecosystem. ![Champtions quote- Alexander Antonison](https://cdn.sanity.io/images/wl0ndo6t/main/beb31b22c491b0d6d190c599f9913c2bcc55d3eb-1200x1200.png) ## What's next Over the coming months, you'll see our Champions in action—publishing content, leading discussions, speaking at events, and shaping the future of analytics engineering. We'll share Champion spotlights, content highlights, and success stories through our newsletter and community channels. While our inaugural cohort is invitation-only, we'll open formal applications for future cohorts later this year. If you're passionate about contributing to the dbt community, stay tuned for opportunities to join. In the meantime, you can find the latest program updates on the [dbt Champions page](https://www.getdbt.com/dbt-champions). To all our founding Champions, thank you for your leadership, your generosity, and your commitment to making analytics engineering better for everyone. Here's to building the future of data, together. [Meet the dbt Champions](https://www.getdbt.com/dbt-champions). ## Join the conversation Want to connect with our Champions? You'll find them active in: - [**dbt Summit**](https://www.getdbt.com/dbt-summit): Speaking, mentoring, and connecting with the community - [**Local meetups**](https://www.getdbt.com/events): Organizing and presenting at events near you - [**Discourse**](https://discourse.getdbt.com): Look for the Champion badge on community profiles - [**Slack**](https://www.getdbt.com/community/join-the-community): Participating in discussions across channels --- --- title: "Types of data transformations for machine learning" description: "Explore key data transformation types for ML, including cleaning, scaling, feature engineering, and validation." url: "https://www.getdbt.com/blog/data-transformations-for-machine-learning" date: "2026-03-19" authors: ["Joey Gault"] categories: ["Pulse"] --- # Types of data transformations for machine learning ## Understanding data transformation in the ML context Data transformation converts raw data from its original format into a standardized structure that meets the requirements of machine learning workflows. This process encompasses cleaning, normalizing, validating, and enriching data to ensure consistency and usability across the entire ML pipeline. The transformation process typically unfolds across several stages. Discovery and profiling assess data structure, quality, and characteristics to identify anomalies and inconsistencies. Cleansing follows, correcting inaccuracies, filling missing values, and removing duplicates. Data mapping then structures information according to model requirements, converting data types and reorganizing fields as needed. Finally, transformed data loads into a central data store where it becomes available for model training and inference. Modern data transformation commonly occurs within [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) pipelines, where data is transformed after loading into its destination. This approach has largely replaced traditional ETL methodologies because cloud computing makes it more cost-efficient to load data prior to transformation. Raw data becomes immediately available to everyone with warehouse access, and teams with different needs can transform it however they see fit. ## Core transformation types for machine learning Machine learning workflows require several distinct categories of transformation, each addressing specific data quality or structural requirements. ### Data cleaning transformations Data cleaning removes errors and inconsistencies that would otherwise compromise model performance. Missing fields, inaccurate entries, duplicated records, and formatting issues all fall within this category. For ML applications, cleaning extends beyond simple error correction to include outlier detection and handling, where statistical methods identify data points that fall outside expected ranges and require investigation or special treatment. The challenge with cleaning transformations lies in balancing thoroughness with performance. Overly aggressive cleaning rules might inadvertently remove valid data points or introduce bias into training datasets. Establishing clear data quality standards and validation rules helps ensure consistent cleaning approaches across different transformation projects. ### Normalization and scaling transformations Normalization transforms data into a standard range or format to ensure consistency and comparability across features. Many machine learning algorithms perform poorly when features exist at different scales. A feature measured in thousands will dominate one measured in single digits, even if both carry equal predictive value. Common normalization approaches include min-max scaling, which transforms values to a fixed range (typically 0 to 1), and standardization, which centers data around zero with unit variance. The choice between these methods depends on the distribution of your data and the requirements of your specific algorithms. Neural networks often benefit from standardization, while tree-based models may require less aggressive normalization. ### Aggregation transformations Aggregation rolls up granular data into summary statistics that capture patterns more efficiently than raw observations. For time-series ML applications, aggregation might convert second-by-second sensor readings into hourly averages, reducing noise while preserving meaningful trends. For customer behavior models, aggregation transforms individual transactions into summary metrics like purchase frequency, average order value, or customer lifetime value. The key consideration with aggregation is selecting the appropriate level of granularity. Too much aggregation loses important signal; too little creates computational burden and risks overfitting to noise in the training data. ### Feature engineering transformations Feature engineering creates new variables from existing data to make patterns more accessible to ML algorithms. This category encompasses a wide range of techniques, from simple mathematical combinations of existing features to sophisticated domain-specific transformations that encode expert knowledge. Temporal features represent one common pattern: extracting day of week, hour of day, or season from timestamp data to capture cyclical patterns. Categorical encoding transforms text labels into numerical representations that algorithms can process, using techniques like one-hot encoding or embeddings. Interaction features capture relationships between variables that might not be apparent when considered independently. ### Validations transformations Validation verifies that data adheres to specified criteria before it becomes eligible for model training or inference. This includes data format validation (ensuring phone numbers match expected patterns), unique constraint validation (verifying that identifiers aren't duplicated), completeness checks (confirming no critical fields are empty), and range validation (ensuring values fall within acceptable bounds). For ML systems, validation takes on additional importance because invalid data can silently degrade model performance rather than causing obvious failures. Implementing comprehensive validation as part of your transformation pipeline prevents these subtle quality issues from reaching production models. ### Enrichment transformations Enrichment enhances internal data with external sources to provide additional context for ML models. A fraud detection system might enrich transaction data with device fingerprinting information, geolocation data, or historical fraud patterns. A recommendation engine might enrich user profiles with demographic data or market segment classifications. The challenge with enrichment is managing dependencies on external data sources and ensuring that enrichment processes don't introduce unacceptable latency into your ML pipeline. Careful architecture is required to balance the value of additional features against the operational complexity they introduce. ## Architectural considerations for ML transformation pipelines Building transformation pipelines for machine learning requires attention to several architectural concerns that distinguish ML workloads from traditional analytics. ### Separation of training and inference transformations ML systems require transformations in two distinct contexts: during model training and during inference. Training transformations can be computationally expensive and may operate on large historical datasets. Inference transformations must execute with low latency on individual records or small batches. This distinction requires careful design to ensure that the same transformation logic applies in both contexts. Inconsistencies between training and inference transformations create train-serve skew, where models perform well in development but fail in production because they encounter differently transformed data. Tools like [dbt](https://www.getdbt.com/product/what-is-dbt) help address this challenge by allowing teams to define transformation logic once and apply it consistently across different contexts. By representing transformation pipelines as version-controlled code, dbt ensures that the same business logic applies whether you're preparing historical data for training or transforming real-time inputs for inference. ### Managing features stores As ML systems mature, organizations often implement feature stores: centralized repositories of transformed features that can be reused across multiple models. Feature stores solve several problems simultaneously. They eliminate redundant transformation logic across different ML projects, ensure consistency in how features are calculated, and provide a catalog of available features that accelerates new model development. Implementing a feature store requires thoughtful integration with your transformation infrastructure. The transformation layer becomes responsible for populating the feature store with up-to-date values, while maintaining the lineage information that connects features back to their source data. ### Handling temporal consistency ML models trained on historical data must make predictions about the future. This creates unique requirements for how transformations handle time. Point-in-time correctness ensures that features used for training reflect only information that would have been available at the time being predicted, avoiding data leakage that would artificially inflate model performance during development. Achieving point-in-time correctness requires careful attention to how transformations incorporate slowly changing dimensions and time-varying features. Your transformation architecture must support querying historical states of dimension tables and applying business logic as it existed at specific points in time. ## Operational best practices for ML transformation pipelines Successful ML transformation implementations require engineering discipline and operational rigor. ### Version control and reproducibility ML experiments must be reproducible to be scientifically valid. This means transformation logic must be version-controlled alongside model code, with clear lineage connecting trained models back to the specific transformation versions used to prepare their training data. Modern transformation platforms like dbt provide built-in [version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics?version=1.11) capabilities, allowing data teams to track changes to transformation logic and understand how those changes impact downstream models. This becomes particularly important when debugging model performance issues or conducting experiments to improve model accuracy. ### Automated testing Unlike traditional software applications, data transformations operate on datasets that change over time. Effective testing strategies must validate both transformation logic and data quality. Automated testing frameworks should verify that transformations produce expected outputs given known inputs, that data quality metrics remain within acceptable bounds, and that schema changes don't break downstream dependencies. dbt's testing capabilities allow teams to define assertions about their transformed data and automatically validate those assertions as part of CI/CD pipelines. This prevents data quality issues from reaching production models and provides early warning when source data characteristics change in ways that might impact model performance. ### Monitoring and observability Production ML systems require comprehensive monitoring of transformation pipelines. Data drift (changes in the statistical properties of input data) can silently degrade model performance over time. Monitoring transformation outputs helps detect drift early, before it impacts business outcomes. Effective monitoring tracks both technical metrics (transformation execution time, failure rates, resource consumption) and data quality metrics (completeness, validity, distribution statistics). When anomalies occur, detailed lineage information helps teams quickly identify root causes and assess impact on downstream models. ### Scalability and performance optimization ML transformation workloads often process large volumes of historical data during model training while also supporting low-latency transformations during inference. This dual requirement demands careful attention to performance optimization. Incremental processing strategies update only changed records rather than reprocessing entire datasets, dramatically reducing computational costs for large-scale transformations. Materialization strategies determine whether transformed data is stored as tables, views, or incremental models, balancing query performance against storage costs and data freshness requirements. ## Building toward production ML systems The transformation layer sits at the foundation of production ML infrastructure. Well-designed transformations ensure that models train on high-quality, consistent data and that inference systems apply the same logic reliably at scale. For data engineering leaders, investing in robust transformation infrastructure pays dividends across the entire ML lifecycle: faster experimentation, more reliable production systems, and clearer paths from prototype to production. Modern transformation tools like dbt enable teams to apply software engineering best practices to data transformation, treating transformation logic as code that can be tested, versioned, and deployed through rigorous CI/CD processes. This approach creates ML systems that are not just functional but truly scalable, enabling self-service feature development, supporting confident model deployment, and freeing data teams to focus on high-value work rather than repeatedly solving the same data quality problems. For organizations building ML capabilities, the question isn't whether to invest in transformation infrastructure, but how to build transformation systems that support both current requirements and future growth. The different types of transformations outlined here (cleaning, normalization, aggregation, feature engineering, validation, and enrichment) form the building blocks of that infrastructure. Understanding when and how to apply each type allows you to construct ML pipelines that deliver reliable predictions at scale. Learn more about building production-grade transformation pipelines in the [Understanding data transformation guide](https://www.getdbt.com/discover/understanding-data-transformation). ## Data transformation for ML FAQs **What is data transformation?** Data transformation converts raw data from its original format into a standardized structure that meets the requirements of machine learning workflows. This process encompasses cleaning, normalizing, validating, and enriching data to ensure consistency and usability across the entire ML pipeline. The transformation process typically includes discovery and profiling to assess data structure and quality, cleansing to correct inaccuracies and fill missing values, data mapping to structure information according to model requirements, and finally loading transformed data into a central data store where it becomes available for model training and inference. **Why and when do you normalize data?** Normalization transforms data into a standard range or format to ensure consistency and comparability across features. Many machine learning algorithms perform poorly when features exist at different scales. A feature measured in thousands will dominate one measured in single digits, even if both carry equal predictive value. Common normalization approaches include min-max scaling, which transforms values to a fixed range (typically 0 to 1), and standardization, which centers data around zero with unit variance. The choice between these methods depends on the distribution of your data and the requirements of your specific algorithms. Neural networks often benefit from standardization, while tree-based models may require less aggressive normalization. **What is feature engineering?** Feature engineering creates new variables from existing data to make patterns more accessible to ML algorithms. This encompasses a wide range of techniques, from simple mathematical combinations of existing features to sophisticated domain-specific transformations that encode expert knowledge. Common examples include temporal features that extract day of week, hour of day, or season from timestamp data to capture cyclical patterns; categorical encoding that transforms text labels into numerical representations using techniques like one-hot encoding or embeddings; and interaction features that capture relationships between variables that might not be apparent when considered independently. --- --- title: "What are the most common data pipeline architecture patterns?" description: "Explore common data pipeline architecture patterns—from ETL and ELT to batch, streaming, and semantic layers." url: "https://www.getdbt.com/blog/common-data-pipeline-architecture-patterns" date: "2026-03-18" authors: ["Joey Gault"] categories: ["Pulse"] --- # What are the most common data pipeline architecture patterns? ## The evolution from ETL to ELT The transition from [ETL (Extract-Transform-Load) to ELT (Extract-Load-Transform)](https://www.getdbt.com/blog/etl-vs-elt) represents more than a simple reordering of operations. This shift reflects a fundamental change in how organizations leverage compute resources and structure their data workflows. In traditional ETL architectures, transformation happens on dedicated middleware servers before data reaches the warehouse. This approach made sense when storage was expensive and compute resources were limited. Teams would extract data from source systems, apply transformations on separate ETL servers, and only then load the processed results into the warehouse. ELT inverts this model. Raw data lands directly in the warehouse, where transformations occur using the platform's native compute capabilities. This pattern aligns naturally with cloud data warehouses like [Snowflake](https://www.getdbt.com/blog/elt-best-practices-snowflake), [BigQuery](https://docs.getdbt.com/reference/resource-configs/bigquery-configs), and [Redshift](https://docs.getdbt.com/reference/resource-configs/redshift-configs), which offer elastic compute that scales with workload demands. The ELT approach delivers several advantages for modern teams. Transformations become more transparent and easier to iterate on, since all logic executes in SQL within the warehouse. Multiple teams can work with the same raw data, applying different transformation logic for their specific use cases. Version control and testing become straightforward when transformation code lives in repositories rather than proprietary ETL tools. [dbt](https://www.getdbt.com/product/what-is-dbt) has emerged as the standard transformation layer in ELT architectures, bringing software engineering practices to analytics workflows. Rather than building transformations in GUI-based tools, teams write modular SQL that can be tested, documented, and deployed through standard CI/CD pipelines. ## Batch hub-and-spoke architecture The batch hub-and-spoke pattern remains prevalent in organizations with on-premises systems or strict compliance requirements. In this architecture, a central database serves as the hub, receiving scheduled extracts from various source systems that act as spokes. The hub processes incoming data and distributes curated outputs to downstream systems like BI tools or reporting marts. This pattern offers predictable scheduling and tight control over data movement. Organizations can enforce specific processing windows, ensuring that transformations complete before business users need access to refreshed data. The centralized hub provides a clear point of governance and monitoring. However, batch hub-and-spoke architectures face inherent limitations. Rigid batch windows create latency between when data changes in source systems and when those changes become available for analysis. As data volumes grow, these batch windows can extend uncomfortably, sometimes pushing into business hours. The monolithic nature of batch processing makes it difficult to isolate and debug specific transformation failures. Scaling batch architectures often requires careful orchestration to manage dependencies between jobs. When one transformation fails, it can cascade through the entire pipeline, delaying all downstream processes. Teams must carefully balance the desire for comprehensive batch processing against the need for timely data availability. ## Cloud warehouse as the central hub Modern cloud data platforms have enabled a more flexible architecture where the warehouse itself serves as the central hub for all data operations. In this pattern, raw data flows directly into platforms like Snowflake, BigQuery, or [Databricks](https://www.getdbt.com/blog/data-transformation-dbt-databricks), where all subsequent transformations occur using native compute. This architecture aligns naturally with ELT workflows. Teams use ingestion tools like Airbyte or Fivetran to land raw data in the warehouse, then apply transformations using dbt. The warehouse's elastic compute scales to handle varying workloads, from lightweight development queries to production transformation jobs processing billions of rows. The warehouse-as-hub pattern supports diverse analytical use cases. Data scientists can access raw data for exploratory analysis. Analytics engineers build tested, documented models for business intelligence. Machine learning teams can train models on the same data foundation that powers dashboards. This convergence on a single platform reduces data movement and eliminates the synchronization challenges that plague multi-system architectures. Cost management becomes critical in warehouse-centric architectures. Without careful attention to compute usage, costs can escalate quickly. Teams need clear conventions for [incremental processing](https://docs.getdbt.com/docs/build/incremental-models), appropriate materialization strategies, and monitoring of query patterns. dbt helps manage these concerns through features like incremental models that process only new or changed data, reducing both compute time and warehouse costs. Access control requires explicit configuration in warehouse-centric architectures. Unlike traditional enterprise systems with built-in role-based security, cloud warehouses require teams to thoughtfully design permission structures. However, this explicit approach also provides flexibility to implement fine-grained access controls that align with organizational needs. ## The semantic layer pattern As self-service analytics has expanded and teams work across multiple tools, the semantic layer has emerged as an essential architectural component. Rather than defining metrics repeatedly in different dashboards or applications, teams codify business logic once in a centralized semantic layer that serves all downstream consumers. The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) exemplifies this pattern. Teams define metrics like "customer churn" or "monthly recurring revenue" as declarative assets within dbt. These definitions include the underlying SQL logic, dimensional attributes, and calculation rules. Once defined, these metrics become available through APIs that BI tools, notebooks, and applications can query. This architectural pattern solves a persistent challenge in analytics organizations: metric inconsistency. When different teams calculate the same metric using slightly different logic, stakeholders lose confidence in data-driven insights. The semantic layer ensures that everyone works from the same definitions, eliminating discrepancies between reports. Implementing a semantic layer requires upfront investment and cross-functional alignment. Teams must agree on canonical metric definitions before building models. Governance processes need to manage changes to these definitions over time. Performance optimization becomes important as the semantic layer serves queries from multiple consumers simultaneously. When executed well, the semantic layer accelerates decision-making and improves data governance. Analysts spend less time reconciling conflicting numbers and more time generating insights. New team members can quickly understand available metrics without reverse-engineering dashboard logic. The semantic layer becomes the single source of truth that powers consistent analytics across the organization. ## Streaming and real-time patterns For use cases requiring low-latency data availability, streaming architectures process data continuously rather than in discrete batches. [Change Data Capture (CDC)](https://www.getdbt.com/blog/data-movement-patterns) enables near-real-time pipelines by detecting and synchronizing changes as they occur in source systems. Log-based CDC reads directly from database transaction logs, capturing inserts, updates, and deletes with minimal overhead on source systems. This approach provides low-latency replication suitable for operational reporting and real-time personalization. Trigger-based CDC emits change events from within applications when direct log access isn't available. Streaming data typically lands in message queues or streaming platforms before being processed and loaded into the warehouse. Teams must design for exactly-once delivery semantics to prevent duplicate records. Schema evolution requires careful handling to ensure that changes in source systems don't break downstream processing. dbt supports streaming architectures through [incremental models](https://docs.getdbt.com/docs/build/incremental-models) that efficiently process new events. Rather than reprocessing entire datasets, incremental models append or update only changed records, maintaining performance as data volumes grow. This pattern works well for event streams, CDC feeds, and other continuously arriving data. The complexity of streaming architectures is justified when data freshness directly impacts business outcomes. Fraud detection systems need immediate access to transaction data. Personalization engines require current user behavior to make relevant recommendations. Operational dashboards must reflect the latest system state. For these use cases, the investment in streaming infrastructure delivers clear value. ## Hybrid and federated patterns Some organizations employ hybrid architectures that combine multiple patterns. Data virtualization and federation allow queries across systems without physically moving data, which can be useful for proof-of-concept work or when compliance requirements prevent data movement. However, virtualized approaches introduce performance challenges. Live queries across systems create latency and place load on source databases. Joins between federated sources are often slow and fragile. Permission models don't always translate cleanly across system boundaries, complicating governance. For data that requires regular querying, joining with other sources, or deep analysis, loading into a centralized warehouse typically provides better performance and reliability. Data virtualization works best as a tactical solution for specific edge cases rather than as a core architectural pattern. ## Choosing the right architecture No single architecture pattern fits every organization. The right choice depends on latency requirements, data volumes, governance needs, and team capabilities. Many organizations employ multiple patterns simultaneously, using batch processing for historical analysis, streaming for real-time use cases, and a semantic layer to ensure consistency across tools. The most successful architectures share common characteristics: they're modular and testable, with clear ownership and documentation. They incorporate automated quality checks and monitoring. They balance performance with cost efficiency. And they're designed to evolve as business needs change. [dbt](https://www.getdbt.com/product/dbt) provides the transformation backbone for these diverse architectural patterns. Whether teams are building batch pipelines, processing streaming data, or serving metrics through a semantic layer, dbt brings consistency to transformation logic. [Version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics?version=1.11), [automated testing](https://docs.getdbt.com/docs/build/data-tests), and [comprehensive documentation](https://docs.getdbt.com/docs/collaborate/documentation) ensure that pipelines remain maintainable as they scale. For data engineering leaders evaluating architecture patterns, the key is selecting approaches that align with organizational objectives while maintaining flexibility for future evolution. The modern data stack provides powerful building blocks; the challenge lies in assembling them into architectures that deliver reliable, timely insights at scale. ## Data pipeline architecture pattern FAQs **What are the key differences between batch, real-time, and hybrid data pipeline architectures, and when should each be used?** Batch architectures process data on scheduled intervals through a central hub that receives extracts from source systems and distributes curated outputs to downstream tools. This pattern offers predictable scheduling and tight governance control, making it suitable for organizations with on-premises systems or strict compliance requirements. However, batch processing creates inherent latency and can struggle with scaling as data volumes grow. Real-time streaming architectures process data continuously using technologies like Change Data Capture (CDC) to detect and synchronize changes as they occur. Streaming data flows through message queues before landing in warehouses, where incremental models efficiently process only new events. This approach is justified when data freshness directly impacts business outcomes, such as fraud detection, personalization engines, or operational dashboards requiring immediate insights. Hybrid architectures combine multiple patterns to address different needs simultaneously. Organizations might use batch processing for historical analysis while employing streaming for time-sensitive use cases and a semantic layer to ensure consistency across tools. The right choice depends on latency requirements, data volumes, governance needs, and team capabilities. **When should a team choose ETL versus ELT for a cloud-native data pipeline, and what trade-offs does each pattern entail?** **What criteria should guide the choice between ETL and ELT patterns in a data pipeline?** --- --- title: "How a semantic layer prevents AI hallucinations in analytics" description: "Learn how a semantic layer gives AI systems the consistent, governed foundation they need." url: "https://www.getdbt.com/blog/how-a-semantic-layer-prevents-ai-hallucinations-in-analytics" date: "2026-03-17" authors: ["Joey Gault"] categories: ["Pulse"] --- # How a semantic layer prevents AI hallucinations in analytics ## The root cause of AI hallucinations in data contexts AI hallucinations in data analytics typically stem from three interconnected problems. First, ambiguous or undefined metrics create confusion. When business terms lack clear, standardized definitions, AI systems must make assumptions about what users are asking for. A column labeled "RevAdj_2023" might mean different things to different teams, and without clear metadata or context, an AI system cannot reliably interpret it. Second, inconsistent data definitions across teams compound the problem. One department might define "revenue" as gross sales; another subtracts discounts and returns. When an AI system queries data from multiple sources with conflicting definitions, it produces outputs that appear authoritative but are fundamentally unreliable. Third, ungoverned data access lets AI systems query raw, unvalidated tables directly. Without guardrails, these systems might pull from outdated sources, apply incorrect business logic, or combine incompatible datasets — all while presenting results that look perfectly legitimate to end users. ## How a semantic layer addresses these challenges A semantic layer provides a centralized framework that defines key metrics and business logic, embedding the metadata and context AI systems need to function reliably. Rather than letting AI query raw database tables directly, a semantic layer acts as an intermediary that enforces consistency and accuracy. When properly implemented, a semantic layer transforms how AI systems interact with data. Instead of making assumptions about undefined terms, the system queries only pre-approved, governed metrics. If a user asks about "total adjusted revenue" and no such metric exists, the semantic layer flags the query as invalid and suggests valid alternatives — such as "revenue adjusted for discounts and returns" or "revenue of active accounts." This creates a hub-and-spoke architecture. Metrics are defined once in a central location (the hub) and queried by any number of downstream systems (the spokes) — whether BI tools, embedded applications, or AI interfaces. Every endpoint accesses the same centralized definitions, ensuring consistency across the organization. ## Consistency for reliable insights A semantic layer eliminates the metric inconsistencies that cause AI systems to produce unreliable outputs. By aligning all metrics to single, standardized definitions, organizations ensure that AI systems always query trustworthy data. When the finance team and marketing team both ask about last quarter's revenue, they receive identical answers — because both pull from the same governed metric definition, regardless of which tool or interface they use. This consistency extends beyond simple calculations. A semantic layer captures the complete business logic behind each metric: how it should be aggregated, what dimensions it can be sliced by, and what relationships exist between different data entities. This context allows AI systems to understand not just what data exists, but how it should be used. ## Governance to protect and standardize Effective governance is critical for AI success. A semantic layer enforces governance by restricting access to sensitive metrics, tracking changes to definitions with clear audit trails, and preventing unauthorized data access. Teams can be scoped to only the metrics relevant to their function — preventing a customer-facing AI agent from inadvertently exposing sensitive internal data, for example. Governance also ensures consistency when business definitions change. Imagine the executive team updates the definition of "adjusted revenue" to include a new discount category. Without a semantic layer, that change requires manual updates across every BI tool, dashboard, and AI system that references that metric — a tedious, error-prone process that inevitably creates inconsistencies. With a semantic layer, the definition is updated once centrally, and all connected systems automatically use the new logic. AI interfaces, LLMs, and human users always work with the latest approved definitions. ## Context for smarter decision-making AI systems need more than data — they need context. A semantic layer provides this by embedding metadata and explicitly defining relationships between data elements. It links tables together (connecting "Customer ID" in a customers table to transactions, for example) so AI systems understand how purchases relate to customers or revenue to products. Defining these relationships explicitly means joins between tables are always performed correctly. The semantic layer also standardizes business logic, embedding rules like "revenue = price − discounts − returns" to prevent mismatched definitions. Each metric includes comprehensive metadata: a clear name, a description of what it measures, the calculation logic, and guidelines for appropriate usage. This eliminates the ambiguity that leads to AI hallucinations. ## Real-world impact on AI accuracy The difference between AI systems with and without a semantic layer is substantial. When an AI system has access to well-defined metrics, clear business logic, and proper context about data relationships, it can provide reliable answers to complex questions — rather than interpolating from ambiguous source data. Consider the earlier example of a retail company. With a semantic layer in place, when someone asks about "adjusted revenue for Product X," the AI system doesn't guess. It recognizes that multiple valid metrics exist and prompts the user to clarify: "Did you mean revenue adjusted for discounts and returns, or revenue adjusted for currency fluctuations?" This guided approach ensures users get accurate answers while building confidence in the AI system. ## Speed and scalability for AI adoption Beyond accuracy, a semantic layer accelerates AI adoption by improving query performance and enabling reuse. Through smart caching and precomputed metrics, AI systems deliver results faster — pulling from validated metric stores rather than scanning raw tables for every query. When AI systems are slow, users abandon them and revert to manual processes or ad hoc data team requests. The semantic layer also streamlines scaling by letting teams reuse standardized metrics across projects. Instead of rebuilding logic for every new AI initiative, teams leverage existing governed definitions. As AI adoption grows, quality and consistency don't degrade. ## Building AI on the right foundation The potential of AI to transform how organizations work with data is real. But that potential is only realized when AI systems are built on a foundation of consistent, governed, well-contextualized data. A semantic layer isn't optional for AI projects — it's a prerequisite. It provides the guardrails that prevent hallucinations, the consistency that builds trust, and the context that enables sophisticated analysis. For data engineering leaders evaluating AI initiatives, the question isn't whether to implement a semantic layer — it's how quickly. Organizations using [dbt](https://www.getdbt.com/product/what-is-dbt) already have a significant advantage. The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) translates dbt models into well-defined business metrics, creating a foundation for clean, reliable, [AI-ready data](https://www.getdbt.com/product/ai). It integrates seamlessly with existing dbt workflows, ensuring data is accurate, governed, and aligned with business goals. By defining [semantic models](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) alongside data transformations, teams create a single source of truth that serves both human analysts and AI systems. Before your next AI initiative, verify that your data is ready: Are metric definitions clear and standardized? Is data access governed? Is business logic codified and version-controlled? If any of those are uncertain, a semantic layer is where to start. [Get started with dbt for free](https://www.getdbt.com/signup) and build the governed data foundation your AI strategy needs, or [talk to our team](https://www.getdbt.com/contact) about implementing the dbt Semantic Layer at scale. ## FAQs **How does a semantic layer address AI hallucinations in data analytics?** A semantic layer prevents AI hallucinations by providing a centralized framework that defines key metrics and business logic with clear metadata and context. Instead of letting AI query raw database tables and make assumptions about undefined terms, the semantic layer acts as an intermediary that enforces consistency and accuracy. AI systems can only query pre-approved, governed metrics — and when a requested metric doesn't exist, the system flags the query and suggests valid alternatives rather than guessing. **What governance capabilities does a semantic layer provide for AI systems?** A semantic layer enforces governance by restricting access to sensitive metrics, tracking definition changes with clear audit trails, and preventing unauthorized data access. When business definitions change, updates are made once centrally, and all connected systems automatically use the new logic — ensuring AI interfaces and users always work with the latest approved definitions. This eliminates the manual, error-prone process of updating every downstream tool individually. **Why does a semantic layer improve AI system performance and scalability?** A semantic layer accelerates AI adoption through smart caching and precomputed metrics, allowing AI systems to deliver results faster by pulling from validated metric stores rather than scanning raw tables for every query. It also streamlines scaling by enabling teams to reuse standardized metrics across projects instead of rebuilding logic for every new AI initiative — maintaining quality and consistency as AI adoption grows. --- --- title: "How ETL tools fit into modern data pipeline architecture" description: "Explore ETL vs ELT and how modern transformation tools power scalable data pipelines." url: "https://www.getdbt.com/blog/etl-tools-data-pipeline-architecture" date: "2026-03-16" authors: ["Joey Gault"] categories: ["Pulse"] --- # How ETL tools fit into modern data pipeline architecture ## The evolution from ETL to ELT Traditional [ETL (Extract, Transform, Load)](https://www.getdbt.com/blog/extract-transform-load) tools were designed for an era when compute resources were expensive and storage was limited. In this model, data is transformed before loading into the warehouse, typically on standalone ETL servers outside the data warehouse environment. This approach made sense for on-premises systems with constrained resources and worked well primarily with structured data. However, as data volumes have grown and cloud data warehouses have become the standard, the limitations of traditional ETL have become apparent. ETL pipelines slow down as data size increases because all transformation must complete before any data reaches the warehouse. This creates bottlenecks that delay insights and makes reprocessing or adding data later difficult. The architecture also struggles with the variety of data types modern organizations need to process, including semi-structured and unstructured data. The shift to [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) represents a fundamental architectural change. In ELT pipelines, raw data is loaded into cloud data warehouses first, then transformed using the processing power of platforms like Snowflake, BigQuery, Redshift, or Databricks. This approach enables loading and transformation to happen in parallel, leveraging affordable on-demand cloud computing. Raw data remains available in its original form, providing flexibility for iterative data reprocessing and making it easier to adapt transformations as business requirements evolve. [dbt is purpose-built for ELT workflows](https://www.getdbt.com/product/what-is-dbt) where data is already loaded into the data warehouse for processing. This architectural shift has made dbt the industry standard for data transformation at scale, as it sits on top of the data warehouse and enables anyone who can write SQL to deploy production-grade pipelines. **** ## Core components and the transformation layer Modern data pipeline architectures consist of several interconnected components that work together to move data from source to consumption. [Understanding how transformation](https://www.getdbt.com/discover/understanding-data-transformation) tools fit within this broader ecosystem is essential for data engineering leaders. The pipeline begins with **ingestion**, where data is selected and pulled from source systems. Data engineers evaluate data variety, volume, and velocity to ensure only valuable data enters the pipeline. The **loading** step then lands raw data in cloud data warehouse or lakehouse platforms, emphasizing the "L" in ELT that allows subsequent transformation within the data repository. **Transformation** is where ETL and ELT tools primarily operate, though in fundamentally different ways. This is where raw data is cleaned, modeled, and tested. The process includes filtering irrelevant data, normalizing data to standard formats, and aggregating data for broader insights. With dbt, these transformations become modular, version-controlled code, making data workflows more scalable, testable, and collaborative. Traditional ETL tools handle transformation outside the warehouse, often using proprietary transformation engines with graphical interfaces. While this can provide visual clarity for simple workflows, it creates challenges for version control, testing, and collaboration. Modern ELT approaches using dbt transform data inside the warehouse using SQL, bringing software engineering best practices like version control, automated testing, and modular design to analytics workflows. **Orchestration** schedules and manages pipeline execution, ensuring transformations run in the right order at the right time. **Observability and testing** components provide data quality checks, lineage tracking, and freshness monitoring; critical for building trust and catching issues before they impact downstream analytics. Finally, **storage** and **analysis** components ensure transformed data is accessible for business intelligence, machine learning, and operational use cases. ## Where traditional ETL tools still fit Despite the shift toward ELT, traditional ETL tools maintain relevance in specific scenarios. Legacy databases that cannot be easily migrated to cloud platforms may require ETL approaches for integration. Regulated industries with strict compliance requirements sometimes mandate that certain transformations occur before data reaches the warehouse. Organizations with significant investments in existing ETL infrastructure may continue using these tools while gradually transitioning to modern architectures. However, even in these scenarios, the trend is toward hybrid approaches. Many organizations use traditional ETL tools primarily for initial data extraction and basic cleansing, then leverage ELT tools like dbt for more complex transformations within the warehouse. This allows teams to take advantage of cloud warehouse processing power while maintaining compatibility with legacy systems. ## Modern transformation in practice The practical advantages of ELT transformation tools become clear when examining how they address common data pipeline challenges. Traditional ETL pipelines often struggle with scalability as large, monolithic scripts become difficult to debug and maintain. Pipeline bottlenecks slow data processing and delay insights, while manual processes create operational overhead that doesn't scale with business growth. [dbt](https://www.getdbt.com/product/dbt) addresses these challenges through modular, version-controlled transformations. Each model is self-contained, making it easier to isolate and fix errors without affecting the entire pipeline. Git-based version control tracks data changes as code, enabling teams to collaborate, audit, revert updates, and maintain a single source of truth for scalable pipeline management. Incremental model processing in dbt transforms only new or updated data, reducing costs and minimizing reprocessing while improving efficiency. This approach enhances query performance, lowers warehouse load, and accelerates transformations compared to full-refresh patterns common in traditional ETL. Parallel microbatch execution processes data in smaller, concurrent batches, further reducing processing time and improving efficiency. ## Integration with the broader data stack Modern transformation tools don't operate in isolation; they integrate with the broader data infrastructure to create end-to-end pipelines. Data ingestion tools like [Airbyte](https://airbyte.com/) or [Fivetran](https://www.fivetran.com/) handle the extraction and loading phases, moving data from source systems into the warehouse. dbt then transforms this raw data into analytics-ready models. Orchestration platforms like Airflow or Kestra coordinate the execution of these steps, ensuring dependencies are respected and failures are handled gracefully. The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) provides a critical bridge between transformation and consumption, centralizing metric definitions to ensure consistency across all pipelines and datasets. This prevents metric drift and accelerates the creation of reliable, reusable data products. Column-level lineage in dbt Catalog helps consumers understand the journey of individual columns from raw input to final analytical models, building trust by allowing users to trace data origins and transformations. Workflow governance capabilities enable teams to standardize on a single platform, ensuring version control, lineage tracking, and access management that makes data transformations auditable and reliable. Integration with industry-leading data quality and observability tools ensures that data entering pipelines won't cause downstream errors. ## Architectural patterns for modern pipelines The architecture surrounding transformation tools determines how well pipelines scale and how teams collaborate. Cloud warehouse or lakehouse architectures have become the central hub for modern data integration, with raw data ingested directly into scalable platforms where all transformations happen using native compute. This setup aligns well with ELT workflows and supports diverse use cases across analytics, machine learning, and real-time reporting. However, without clear conventions, ad-hoc transformations can diverge across teams, creating inconsistencies in metric definitions. dbt addresses this by providing a framework for standardized transformation logic that can be shared and reused across the organization. [State-aware orchestration through dbt optimizes workflows ](https://www.getdbt.com/blog/using-state-aware-orchestration-to-slash-your-data-costs)by running models only when upstream data changes, reducing redundant executions and improving efficiency. Hooks automate operational tasks like managing permissions and optimizing tables, while macros bundle logic into reusable functions that enable parameterized workflows. Integration with CI/CD workflows automatically tests modified models and their dependencies before merging to production, ensuring changes don't break existing functionality. ## The future of transformation in data pipelines The role of transformation tools continues to evolve as AI and automation reshape data engineering. dbt has integrated capabilities like [dbt Copilot](https://www.getdbt.com/product/dbt-copilot), which leverages large language models to generate code, documentation, tests, metrics, and semantic models based on natural-language descriptions. This greatly reduces the time spent writing models and accelerates pipeline deployment. The [acquisition of SDF by dbt Labs](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs) brings high-performance compilation and validation capabilities that can catch breaking changes during development before code is even checked in. This shift-left approach to quality means errors are caught as developers type, well before they run transformations or deploy to production. Cloud data warehouses continue to release features that complement modern transformation tools. Dynamic Tables in Snowflake, for example, can be leveraged through dbt to deploy models where the warehouse handles incremental updates automatically. Technologies like Snowflake Snowpipe Streaming and Databricks Lakeflow enable efficient ingestion and transformation by leveraging high-throughput, low-latency processing. ## Conclusion ETL tools fit into modern data pipeline architectures primarily as legacy components being gradually replaced by ELT approaches, or as specialized tools for specific use cases involving on-premises systems and compliance requirements. The architectural shift to cloud-native data warehouses has fundamentally changed where transformation should occur, moving it from external ETL servers into the warehouse itself. Modern transformation tools like dbt represent the current state of the art, bringing software engineering best practices to data transformation through modular SQL models, version control, automated testing, and comprehensive documentation. These tools integrate seamlessly with cloud data warehouses, orchestration platforms, and observability solutions to create end-to-end pipelines that are scalable, reliable, and maintainable. For data engineering leaders, the strategic question is not whether to use ETL or ELT tools, but how quickly to transition legacy ETL workflows to modern ELT architectures that unlock the full potential of cloud data platforms. Organizations that make this transition gain faster insights, better data quality, improved scalability, and more efficient use of both human and computational resources. *** **Related resources:** - [AI Data pipelines: Critical components and best practices](https://www.getdbt.com/blog/ai-data-pipelines) - [Building reliable data pipelines: a foundational approach](https://www.getdbt.com/blog/building-reliable-data-pipelines) - [What is data infrastructure and how to design it](https://www.getdbt.com/blog/data-infrastructure) - [Data integration: The 2025 guide for modern analytics teams](https://www.getdbt.com/blog/data-integration) - [dbt Documentation](https://docs.getdbt.com/) ## ETL tools FAQs **What are the trade-offs between ETL and ELT when architecting a data warehouse pipeline?** Traditional ETL transforms data before loading it into the warehouse, which made sense when compute resources were expensive and storage was limited. However, this approach creates bottlenecks as data volumes grow because all transformation must complete before any data reaches the warehouse, delaying insights and making reprocessing difficult. ETL also struggles with diverse data types and typically uses proprietary transformation engines with graphical interfaces that create challenges for version control, testing, and collaboration. ELT loads raw data into cloud data warehouses first, then transforms it using the warehouse's processing power. This enables loading and transformation to happen in parallel, leverages affordable on-demand cloud computing, and keeps raw data available in its original form for flexibility. ELT approaches using tools like dbt transform data inside the warehouse using SQL, bringing software engineering best practices like version control, automated testing, and modular design to analytics workflows. The trade-off is that ELT requires cloud data warehouse infrastructure, while ETL may still be necessary for legacy systems or strict compliance requirements. **How do modern data pipelines differ from traditional pipeline architectures?** **When is it appropriate to use batch processing versus stream processing in a data pipeline design?** Batch processing is appropriate when data consumers need refreshed data on an hourly or daily basis rather than in real-time. Incremental model processing transforms only new or updated data in batches, reducing costs and warehouse load while improving efficiency compared to full-refresh patterns. Parallel microbatch execution processes data in smaller, concurrent batches to reduce processing time. Stream processing or real-time ingestion is appropriate when consumers need data refreshed every second or minute. Technologies like Snowflake Snowpipe Streaming and Databricks Lakeflow enable efficient high-throughput, low-latency processing for real-time use cases. The choice depends on evaluating data velocity requirements and whether directionally accurate data suffices or if high data quality with immediate availability is necessary for operational or analytical needs. --- --- title: "The Iceberg ecosystem today" description: "Anders Swanson explains what data teams can realistically expect when attempting to run on top of Iceberg in production." url: "https://www.getdbt.com/blog/the-iceberg-ecosystem-today" date: "2026-03-15" authors: ["Daniel Poppy"] categories: ["Insights"] --- # The Iceberg ecosystem today The data industry is moving towards open standards. The migration towards open standards throughout the data ecosystem is happening rapidly despite all the oxygen getting sucked out of the room from the rapid progress of AI and agents. The dbt Labs data team is moving to an all Iceberg lake with a mix of compute engines to power transformation, analytics, and agentic experiences. The team has been able to move quickly towards this architecture because the entire ecosystem has been laying the groundwork for years. All of it’s coming together to make this new open world a reality fast. In this episode, Tristan discusses the reality on the ground for data practitioners. Where’s the Iceberg ecosystem today? What can practitioners realistically expect when attempting to run on top of Iceberg in production? Tristan is joined by Anders Swanson, a developer experience advocate at dbt Labs. Anders has spent a lot of time over the years navigating open-source data ecosystems and tracking their progress. They unpack the open standards shift, define the core building blocks (query engines, object stores, catalogs), and dig into why external catalogs have become a fourth namespace tier across platforms. Anders outlines a pragmatic, phased adoption model for Iceberg integrations, explains why metadata performance and resiliency are hard requirements, and clarifies why vended credentials exist and what they solve. _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [Youtube](https://www.youtube.com/playlist?list=PL0QYlrC86xQm83Q9deiy4euEnbw8ceu3I) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) [Watch video](https://youtu.be/K7PvwU5ulrA) ## Key takeaways ### Tristan Handy: I wanted to have you on because of work you’ve been doing internally to summarize the state of the Iceberg ecosystem. We’ve talked about Iceberg a bunch lately with folks deep in specific parts. Your work is more of an overview: where we’re at with platform integrations, what’s easier now than a year ago, and what’s still hard. Before we dive in, I want to define a few terms. When you say “query engine,” what do you mean? **Anders Swanson:** It’s the thing that does your work. When you issue a CREATE TABLE or a SELECT statement, it’s what returns data or stores it somewhere for later. ### Object store. It’s the cloud service where you can store an object. An object is anything: a blob. ### Catalog. In this context, a catalog knows what tables and views exist and where they are, and how you can fetch or write to them. ### Let’s talk internal versus external catalogs. An internal catalog is what you get by default in a system like Snowflake or SQL Server. An external catalog is more like another directory, often managed by a different system. As you connect more disparate platforms, you can’t assume one system controls everything. ### The complexity comes from duplication. How do you make namespaces unique? Can you plug in many external catalogs? Abstraction matters. A common pattern emerging is one‑to‑one mapping of an external catalog into a database. That pushes a move to a four‑part namespace: catalog, database, schema, identifier. Spark moved toward this; Databricks Unity Catalog and Snowflake‑style catalog link approaches are in this family. ### So the downside? The devil is in the details, especially metadata performance and resiliency. For example, information schema listing. Users expect listing tables to be fast and reliable. In a federated world, if listing tables takes five seconds, users blame the vendor they’re using—even if the external system is slow. DuckDB draws a line by not mixing external catalog tables into information schema listing today. Snowflake’s catalog link databases appear to cache or mirror metadata so it feels as performant as native tables. ### With catalog link databases, Snowflake is doing mirroring. Yes. Mirroring exists in different flavors across platforms. Delta is sometimes seen as “simpler” because metadata can live in object store, but as soon as you want multiple engines writing, you still need a real catalog. ### Sharing across multiple platforms adds another layer. What’s the state of platforms reading and writing to the same Iceberg catalog? There are phases of integration. Phase one is the naive approach: you have Parquet and JSON in object storage, and an engine reads it. Reading is easier than writing. You can get a toy example working. Then you run into versioning and “what’s latest.” The next phase is connecting to an Iceberg REST catalog so engines can ask for the latest table version without users thinking about paths. Phase three is schema‑scale: it’s never just one table. You need discovery of new tables, keeping schemas up to date, and eventually things like multi‑table transactions. ### This maps to dbt Mesh and cross‑platform mesh. Producer vs consumer. A consumer‑led model requires the downstream team to create pointers (DDL) to external tables. It’s operationally messy. Producer‑led is cleaner: the producer writes to the catalog and it’s just there, immediately queryable downstream. ### Are platforms there yet? Some support writing directly to external catalogs. When it works, it’s great, but there are still kinks. We’re retrofitting race cars designed for isolation to be interoperable without losing performance. ### Identity is one of the hairiest issues. Vended credentials. Vended credentials solve the “two keys” problem. You authenticate to the catalog, the catalog tells you where data lives, but then you need separate object store credentials to read files. Vended credentials means the catalog vends short‑lived credentials so you can access the object store location without managing separate keys. ### That doesn’t solve user identity and grants. Correct. Vended credentials isn’t global authorization. Identity and access across platforms is still hard. Ideally you grant access once and it works everywhere, but enterprises have different identity providers and platforms have different permission models. Today, admins often have to configure grants separately in each platform. ### Is this mission creep? The goal is to reduce how many people have to think about storage details. Big tech had whole data platform teams solving reliability problems in Hive‑era lakes. Iceberg reduces that toil dramatically, but the long tail is still auth, mirroring, and cross‑platform governance. ### How does this reshape data teams? Analytics engineering abstracted a lot of work. Data engineering has also been simplified by replication/orchestration vendors. What remains is the open ecosystem complexity: identity, object store policies, and cross‑platform connections. Many enterprises already have teams with these skills (infra as code, Terraform, Snowflake management), but others will need to grow into them. ### Are vendors embracing Iceberg in good faith? The goodwill and collaboration in the past 18 months feels unprecedented. We’re getting “more problems” because we solved prior ones. The industry aligning on standards feels like F1 teams standardizing components so they can innovate elsewhere. ### In your internal writeup about Iceberg, you quoted Wolf Hall: “The making of a treaty is the treaty. It doesn’t matter what the terms are, just that there are terms, it’s the goodwill that matters. When that runs out, the treaty is broken, whatever the terms say.” Explain the relevance here. When I joined dbt, it was taboo to mention one partner to another. Now vendors openly acknowledge mutual customers and invest in interoperability. On the Iceberg repo you see competitors collaborating on proposals. The goodwill is the standard. ### Wrap us up with three things you’re excited for next year. Push‑based catalog updates so platforms can subscribe to changes rather than repeatedly listing and polling. Progress on the small files problem so Iceberg works better for smaller data too. And more platforms supporting writing directly to external catalogs, unlocking producer‑led sharing and cross‑platform mesh. ## Chapters 00:00:00 — Intro: why open standards are accelerating 00:01:20 — What practitioners can expect from Iceberg in production 00:05:00 — Lightning round: query engine, object store, catalog 00:06:20 — Internal vs external catalogs 00:09:30 — The “four-part namespace” and catalog-link style abstractions 00:11:30 — The downside: metadata performance, resiliency, and caching 00:17:10 — Sharing across multiple platforms: reality and tradeoffs 00:19:10 — Iceberg integration phases (1: naive table, 2: REST catalog, 3: schema-scale) 00:24:10 — Producer vs consumer model and cross-platform mesh 00:29:10 — Identity and “vended credentials”: what it is and what it isn’t 00:33:30 — The hard unsolved part: grants and global identity across platforms 00:37:00 — Is this mission creep? What Iceberg is optimizing for 00:39:50 — How roles on data teams evolve in an open ecosystem 00:43:40 — Are vendors genuinely aligned? Why Anders is optimistic 00:46:50 — “The making of a treaty is the treaty”: goodwill as the standard 00:51:50 — Three things Anders is excited for next year --- --- title: "Why metadata management is critical for modern data teams" description: "Metadata management improves discovery, governance, performance, and trust in modern data systems." url: "https://www.getdbt.com/blog/why-metadata-management-is-important" date: "2026-03-13" authors: ["Joey Gault"] categories: ["Pulse"] --- # Why metadata management is critical for modern data teams ## The foundation of data understanding Metadata management refers to the processes, tools, and practices organizations use to organize, control, and leverage information about their data assets. While data represents the actual values stored in tables and files, metadata describes the characteristics, structure, lineage, and context of that data. This distinction matters because modern data environments generate metadata continuously across disconnected systems: data warehouses track query performance, BI tools monitor dashboard usage, and transformation pipelines capture execution statistics. Without deliberate management, this metadata remains scattered and inaccessible, making it difficult to answer fundamental questions about data assets. The metadata layer encompasses multiple dimensions of information. Structural metadata describes technical characteristics like table names, column names, data types, and storage locations. Operational metadata captures information about data processes, including when tables were last updated, how long transformations take to run, and which jobs have succeeded or failed. Lineage metadata tracks how data flows through systems, showing which upstream sources feed into downstream models and reports. Business metadata adds organizational context through ownership assignments, business definitions of metrics, data quality indicators, and usage patterns. **** ## Accelerating discovery and development Organizations that manage metadata effectively experience tangible improvements in how teams work with data. Data discovery becomes dramatically faster when teams can search a centralized catalog rather than asking colleagues or hunting through scattered documentation. Instead of spending hours or days tracking down the right dataset for a new analysis, analysts can search using business terms, browse by domain or owner, and examine column-level details to determine fitness for purpose. This acceleration extends beyond individual productivity. When [dbt](https://www.getdbt.com/product/what-is-dbt) generates metadata during transformation runs (capturing lineage, test results, and execution statistics automatically), teams build a comprehensive view of their data landscape without additional manual effort. The metadata becomes a living resource that stays current with the codebase, rather than documentation that quickly becomes outdated and unreliable. Understanding improves when metadata provides clear definitions, ownership information, and usage examples. New team members can onboard more quickly when they can explore the data catalog to understand what datasets exist and how they're used. Cross-functional collaboration becomes more effective when business stakeholders and technical teams share a common reference point for discussing data assets. ## Enabling impact analysis and reducing risk Perhaps the most critical operational benefit of metadata management is the ability to perform impact analysis before making changes. When teams can see which downstream reports, dashboards, and models depend on a particular dataset, they can assess the consequences of modifications and communicate proactively with affected stakeholders. This visibility fundamentally changes how data teams approach development work. [Column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage?version=1.11) shows precisely which source fields contribute to each downstream column, enabling root cause analysis when issues occur. If a report shows unexpected values, teams can trace lineage upstream to identify where problems originated, examine execution metadata to see whether recent runs failed or took longer than usual, and review quality metadata to determine which tests failed and when problems first appeared. This comprehensive view reduces mean time to resolution and prevents incidents from cascading through dependent systems. The alternative (making changes without understanding dependencies) leads to broken dashboards, incorrect reports, and eroded trust in data systems. Teams become hesitant to make necessary improvements because they can't predict the impact. Development cycles slow as engineers manually trace dependencies or wait for issues to surface in production. Metadata management transforms this reactive approach into a proactive one where teams can move quickly with confidence. ## Supporting governance and compliance Governance requirements drive many metadata management initiatives, particularly in regulated industries. [Regulations like GDPR](https://www.getdbt.com/industry/financial-services) require organizations to track where sensitive data resides and how it flows through systems. Audit requirements demand documentation of data transformations and access patterns. Metadata management provides the foundation for meeting these obligations systematically rather than through manual, error-prone processes. Access control metadata defines who can view, modify, or use different data assets. Role-based permissions, row-level security policies, and data classification tags ensure that sensitive information remains protected while enabling appropriate self-service access. When auditors ask who has accessed sensitive data, the metadata catalog provides the answer. When regulations require documentation of data processing activities, the catalog serves as the source of truth. Beyond compliance, metadata management enables effective data governance programs. Organizations can identify which tables contain personally identifiable information, establish ownership and stewardship models, and enforce data quality standards. The metadata layer provides the visibility necessary to implement governance policies at scale, rather than relying on tribal knowledge or incomplete spreadsheets. ## Optimizing performance and cost Performance optimization relies on metadata about execution patterns and resource consumption. Understanding which models consume the most resources, which queries run most frequently, and where bottlenecks occur enables targeted improvements. Teams can identify inefficiencies in orchestration configurations, reduce infrastructure costs, and improve data freshness based on concrete execution data rather than guesswork. Operational metadata reveals patterns that aren't visible from examining code alone. A transformation might appear efficient in isolation, but execution metadata shows it's consuming disproportionate resources or creating downstream bottlenecks. Query performance metadata identifies which models would benefit most from materialization strategy changes or indexing improvements. This data-driven approach to optimization ensures engineering effort focuses on changes that deliver meaningful impact. For organizations running data platforms at scale, these [optimizations translate directly to cost savings](https://www.getdbt.com/product/cost-optimization). Cloud data warehouse costs correlate strongly with compute consumption, and metadata about execution patterns enables teams to reduce waste without sacrificing functionality. Understanding usage patterns also informs capacity planning and helps teams make informed decisions about infrastructure investments. ## Building trust through transparency [Trust in data systems](https://www.getdbt.com/product/build-trust-in-data-and-data-teams) depends on transparency about data quality, freshness, and provenance. Metadata management makes this transparency possible by surfacing quality metrics, test results, and lineage information alongside the data itself. When users can see that a dataset passed its quality checks, was updated recently, and derives from trusted sources, they can use it with confidence. Data quality monitoring generates metadata about the health of datasets through test results, freshness checks, and validation metrics. This operational metadata helps teams detect issues quickly and build confidence in analytical outputs. Rather than discovering data quality problems when reports show unexpected values, teams can monitor quality metrics proactively and address issues before they impact downstream consumers. Lineage tracking shows data provenance, enabling users to understand where data originated and how it was transformed. This visibility proves invaluable for explaining how metrics are calculated, understanding why values changed, and assessing whether data is appropriate for a particular use case. When business stakeholders can trace a dashboard metric back through transformation logic to source systems, they develop confidence in the numbers they're using to make decisions. ## Integrating metadata into development workflows The most successful metadata management implementations integrate metadata generation into development workflows rather than treating it as a separate process. When metadata management is separate from development, it quickly becomes outdated as code evolves. Embedding metadata generation into continuous integration pipelines means documentation updates alongside code changes. [dbt](https://www.getdbt.com/product/dbt) exemplifies this integrated approach by defining metadata in version-controlled YAML files alongside transformation code. Teams document models, columns, and metrics in the same pull requests where they modify transformation logic. This tight coupling ensures metadata stays current and makes documentation a natural part of the development process rather than an afterthought. For more information on dbt's approach to metadata, see the [dbt documentation](https://docs.getdbt.com). The [Discovery API in dbt](https://www.getdbt.com/blog/introducing-the-discovery-api) enables querying comprehensive metadata about projects, making it accessible to downstream tools and applications. This programmatic access transforms metadata from static documentation into dynamic, actionable information that can power custom applications, automated alerting, and integration with data catalogs. Teams can build workflows that leverage metadata to automate governance checks, generate custom reports, or trigger notifications based on lineage relationships. ## Addressing scale and complexity As data environments grow, metadata management faces challenges of scale and complexity. Large organizations may have hundreds of thousands of tables and millions of columns. Metadata systems must handle this volume while remaining responsive for search and browsing. [Lineage graphs](https://docs.getdbt.com/docs/explore/explore-projects?version=1.11#project-lineage) can become overwhelming when they include every possible dependency, requiring strategies for filtering, aggregating, and presenting information at appropriate levels of detail. Automation becomes essential at scale. Rather than manually documenting table structures, organizations should automatically ingest this information from data platforms. External metadata ingestion capabilities extend catalog coverage beyond transformation-managed assets by connecting directly to data warehouses and including tables, views, and other resources that exist outside transformation pipelines. This creates comprehensive catalogs that represent the full data landscape. Federated responsibility distributes metadata maintenance across teams rather than centralizing all documentation work. Organizations should empower data producers to document their own assets, with the metadata layer supporting team-specific properties and flexible schemas. The meta configuration in dbt allows teams to add custom metadata properties to models, columns, and other resources, enabling this federated approach without rigid centralized schemas. Learn more about [dbt's metadata capabilities](https://www.getdbt.com/blog/leveraging-dbt-metadata-in-data-management). ## Looking forward Metadata management continues to evolve as data environments grow more complex and AI applications introduce new requirements around model training data, feature definitions, and prediction explanations. The distinction between technical and business metadata is blurring as modern systems increasingly combine structural information with business context, quality metrics, and usage patterns in integrated views. For data engineering leaders, the question isn't whether to invest in metadata management, but how to build capabilities that scale with organizational needs. This means choosing approaches that integrate metadata management into existing workflows, establishing clear ownership and governance processes, and building habits that keep metadata current and useful. The organizations that succeed treat metadata as a first-class concern, not an afterthought: embedding it into development practices, governance programs, and operational workflows. When implemented effectively, metadata management transforms data from a scattered collection of tables into a well-organized, discoverable, and trustworthy asset that the entire organization can leverage. The key lies in automation, integration, and cultural commitment to maintaining metadata quality. Teams that master these elements position themselves to extract maximum value from their data assets while managing associated risks effectively. For teams using dbt, exploring [data catalog integration](https://www.getdbt.com/discover/understanding-data-catalogs) and [data product management best practices](https://www.getdbt.com/blog/data-product-management) provides additional context for building comprehensive metadata management capabilities. ## Metadata management FAQs **What is metadata management?** Metadata management refers to the processes, tools, and practices organizations use to organize, control, and leverage information about their data assets. It encompasses managing multiple dimensions of information including structural metadata (table names, column names, data types), operational metadata (update times, transformation durations, job statuses), lineage metadata (data flows through systems), and business metadata (ownership, definitions, quality indicators, usage patterns). This management layer helps organizations understand the characteristics, structure, lineage, and context of their data rather than just the data values themselves. **Why is metadata management important?** **What are the key components of metadata management?** The key components of metadata management include structural metadata describing technical characteristics like table and column names, operational metadata capturing process information such as update times and execution statistics, lineage metadata tracking how data flows through systems from upstream sources to downstream models, and business metadata adding organizational context through ownership, definitions, and quality indicators. Effective metadata management also requires integration into development workflows rather than treating it as a separate process, automation to handle scale and complexity, and federated responsibility that distributes maintenance across teams while maintaining comprehensive catalog coverage of the full data landscape. --- --- title: "Why ETL is still essential for modern data pipelines" description: "ETL consolidates fragmented data, enforces quality, and satisfies compliance requirements modern organizations depend on." url: "https://www.getdbt.com/blog/why-etl-matters-data-pipelines" date: "2026-03-12" authors: ["Joey Gault"] categories: ["Pulse"] --- # Why ETL is still essential for modern data pipelines ## The core problem: data fragmentation Modern organizations generate data across countless systems. Marketing teams work in platforms like HubSpot and Google Ads. Sales teams track opportunities in Salesforce or other CRMs. Product teams instrument application databases. Finance teams manage transactions in ERP systems. Each system serves its purpose well in isolation, but business questions rarely respect these boundaries. When an executive asks about customer lifetime value, the answer requires combining data from sales systems, product usage databases, and support platforms. When a marketing leader wants to understand campaign ROI, the analysis demands integrating ad spend data with conversion tracking and revenue attribution. These questions are impossible to answer when data remains scattered across disconnected systems. ETL emerged as a solution to this fragmentation. By systematically extracting data from source systems, transforming it into a consistent format, and loading it into a centralized warehouse, ETL creates what data teams call a "single source of truth" — one place where all organizational data comes together in a queryable, reliable format. ## Data quality and consistency Raw data from source systems is messy. Date fields use different formats. Customer identifiers vary across platforms. Required fields contain null values. Product names are spelled inconsistently. One system tracks revenue in cents while another uses dollars. Without addressing these inconsistencies, any analysis built on this data will be unreliable at best and dangerously misleading at worst. The transformation phase of ETL applies the business logic needed to clean and standardize data before it reaches the warehouse. This ensures that downstream users — whether analysts building dashboards or data scientists training models — work with consistent, reliable datasets regardless of the original source format. When every team starts from the same clean foundation, organizations avoid the common problem of different departments reporting conflicting numbers for the same metric. This focus on data quality becomes even more critical as organizations scale. A startup with three data sources might manage inconsistencies manually, but an enterprise with hundreds of data sources needs systematic processes to maintain quality. ETL provides that systematic approach, encoding [data quality rules](https://docs.getdbt.com/docs/build/data-tests) that run automatically with every pipeline execution. ## Governance and compliance requirements For organizations in regulated industries, ETL serves a critical governance function. Financial institutions must comply with regulations around transaction reporting and audit trails. Healthcare providers must protect patient information under HIPAA. Retailers handling customer data must meet privacy requirements like [GDPR and CCPA](https://www.nist.gov/privacy-framework). Traditional ETL workflows allow organizations to transform or mask sensitive data before it enters the warehouse. A healthcare provider might hash patient identifiers during the transformation phase, ensuring personally identifiable information never lands in the warehouse in raw form. A financial institution might apply fraud detection rules and data validation checks before loading transaction data, creating an auditable trail of how data was processed. This pre-load transformation provides stronger control over how sensitive or regulated data is handled. By encoding compliance requirements directly into ETL pipelines, organizations reduce the risk of accidental exposure and make it easier to demonstrate regulatory compliance during audits. ## Performance optimization ETL also addresses performance concerns that become critical at scale. When data is transformed before loading into the warehouse, queries run faster because the heavy lifting has already been done. Business intelligence tools retrieve pre-aggregated metrics without performing expensive calculations at query time. Dashboards load quickly because the underlying data is already in the right format. This performance benefit matters most when supporting large numbers of concurrent users. If hundreds of analysts query the same warehouse simultaneously, pre-transformed data reduces compute load and keeps costs manageable. While modern cloud warehouses have impressive computational power, there's still value in doing work once during the ETL process rather than repeatedly at query time. ## The evolution to ELT Despite these benefits, traditional ETL has real limitations that have driven many organizations toward [ELT (Extract, Load, Transform)](https://docs.getdbt.com/docs/introduction). The rise of cloud-native data warehouses with massive computational power has fundamentally changed the economics of data transformation. In an ELT workflow, raw data loads into the warehouse first, then transforms using the warehouse's own compute resources. This reversal unlocks several advantages: raw data becomes available immediately, even before transformations complete; teams can iterate on transformation logic without reprocessing data from source systems; and different teams can transform the same raw data in different ways to serve different use cases. [dbt](https://www.getdbt.com/product/what-is-dbt) has made ELT workflows practical by providing [version control, testing, documentation](https://docs.getdbt.com/best-practices), and deployment capabilities for transformations that happen inside the warehouse. This software engineering-inspired approach treats data pipelines as code — with all the benefits of [continuous integration, automated testing](https://docs.getdbt.com/docs/deploy/continuous-integration), and collaborative development. ## When ETL still matters The shift toward ELT doesn't make ETL obsolete. Certain scenarios still call for transforming data before it enters the warehouse. Organizations handling highly sensitive personally identifiable information often need to hash or mask that data before loading to meet compliance requirements. Some data governance frameworks require transformation logic to be applied and audited before data reaches the warehouse. Many organizations adopt hybrid approaches — using traditional ETL for regulated or high-risk data while leveraging ELT for more flexible analytics workflows. A healthcare provider might use ETL to mask patient identifiers before loading, then use [dbt](https://www.getdbt.com/product/dbt) to build analytics models on top of that de-identified data. This hybrid model provides the control compliance requires alongside the agility analytics development demands. ## Building for the future Whether implementing traditional ETL, modern ELT, or a hybrid approach, the fundamental need is the same: systematically consolidate data from disparate sources, ensure its quality and consistency, and make it available for analysis. The specific technical implementation matters less than treating data transformation as a critical engineering function that requires proper tooling, testing, and governance. For data engineering leaders, the question isn't whether to implement data integration processes — it's how to implement them in a way that balances governance requirements with the need for speed and flexibility. The organizations that succeed with data recognize transformation as a core competency. They [implement version control for transformation logic](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview), build automated testing into their pipelines, separate development and production environments to enable safe experimentation, and monitor pipeline health proactively. These practices apply whether transformations happen before or after loading. The goal is always the same: turning raw data into reliable, actionable insights. ETL — in its traditional form or its modern ELT evolution — is how organizations achieve that at scale. [Get started with dbt for free](https://www.getdbt.com/signup) to bring engineering best practices to your transformation workflows, or [talk to our team](https://www.getdbt.com/contact) about building the right architecture for your organization. ## FAQs **What is ETL (extract, transform, load)?** ETL is a data integration process that systematically extracts data from source systems, transforms it into a consistent format, and loads it into a centralized warehouse. This creates a "single source of truth" where all organizational data comes together in a queryable, reliable format — solving the problem of data fragmentation across disconnected systems like CRMs, marketing platforms, product databases, and ERP systems. **How does ELT differ from ETL, and when is ELT advantageous?** [ELT](https://docs.getdbt.com/docs/introduction) reverses the traditional ETL process by loading raw data into the warehouse first, then transforming it using the warehouse's own compute resources. Raw data becomes available immediately before transformations complete, teams can iterate on transformation logic without reprocessing data from source systems, and different teams can transform the same raw data in different ways. ELT has become practical with cloud-native data warehouses and tools like [dbt](https://www.getdbt.com/product/what-is-dbt) that provide version control, testing, and documentation for in-warehouse transformations. **When should an organization choose ETL over ELT?** Choose traditional ETL when handling highly sensitive personally identifiable information that needs to be hashed or masked before loading to meet compliance requirements, or when governance frameworks require transformation logic to be applied and audited before data reaches the warehouse. Many organizations adopt hybrid approaches — ETL for regulated or high-risk data, ELT for flexible analytics workflows. A healthcare provider might use ETL to mask patient identifiers before loading, then use ELT to build analytics models on top of that de-identified data. --- --- title: "How a global investment firm reduced runtimes by 30–40% with the dbt Fusion engine" description: "NBIM cut runtimes 30–40% in 3 months with the dbt Fusion engine and State-Aware Orchestration—without heavy optimization." url: "https://www.getdbt.com/blog/how-a-global-investment-firm-reduced-runtimes-by-30-40-with-the-dbt-fusion-engine" date: "2026-03-11" authors: ["Elaine Green"] categories: ["Product"] --- # How a global investment firm reduced runtimes by 30–40% with the dbt Fusion engine **** Norges Bank Investment Management (NBIM) manages Norway’s sovereign wealth fund, one of the largest in the world. Overseeing more than two trillion dollars in assets across global markets, NBIM holds about 1.5% of all listed companies worldwide. To support these investments, NBIM runs a sophisticated data platform that serves investment teams across Oslo, New York, and Singapore. But as NBIM’s data practice scaled to more than 150 dbt developers across 30+ decentralized projects, the team began to feel the strain. In this post, you’ll learn how NBIM addressed these challenges by adopting the dbt Fusion engine and State-Aware Orchestration (SAO). NBIM’s journey reveals that the best path to efficiency can start with making life easier for your developers. ## A smooth rollout of the dbt Fusion engine in 3 months When NBIM began exploring Fusion, its primary concern was improving the developer experience. SAO was viewed as a bonus. “We’re doing a lot of work with AI agents and coding tools like Cursor,” explains _​​_Øyvind Barsnes Eraker, Senior Data Engineer_ _at NBIM. “We wanted to bridge a gap between technical data engineers and business-oriented teams who write code but may struggle with code structure.” Rather than migrating its most complex legacy projects immediately, NBIM’s data team started with smaller projects built around fewer, well-structured models. In just a few months, NBIM set up five projects in production using Fusion and SAO. The projects focus on: - **Platform metadata:** These projects track usage of dbt, Snowflake, Fivetran, and Tableau, helping the team make decisions around data deprecation and cleanup. - **Communications insights:** The team can easily access PR, webpage, and social media data that informs NBIM’s communications strategy. - **Investment data:** NBIM can use data from external vendors to create a single source of truth for investment teams. “We were intentional about expanding Fusion into projects at a sustainable pace while onboarding new users,” reflects Øyvind. “It’s been a great experience with no significant hiccups at all.” ## Faster runtimes with a better developer experience By focusing on projects that can translate into quick wins, the data team has avoided migrating technical debt. Thanks to the thoughtful Fusion rollout, they’ve already experienced improvements in their workflows. “We’re seeing better feedback loops, faster parsing times, and improved linting,” says Øyvind. “Our less technical stakeholders, like portfolio managers, are creating higher-quality projects, too.” Importantly, since moving projects to Fusion, the team has observed faster end-to-end runs. They typically complete 30-40% faster, even without any dedicated optimization efforts. “SAO enables us to lighten up on the orchestration so that we can run models and pipelines more often without seeing runtime increases,” says Øyvind. “We’re confident we’ll see cost optimizations downstream too.” As the team looks ahead, they view Fusion as mission-critical for reducing SLA risk on NBIM’s most complex projects. “Our larger pipelines still run in a very batch-oriented way,” says Øyvind. “We have teams in Singapore that run and validate the pipelines so the data is ready as early as possible for portfolio managers across different timezones. That creates very tight SLAs around when data has to be delivered.” For these projects, there’s little room for delays. Fusion could allow the team to move away from nightly batches toward more continuous data delivery—helping the business access insights faster. ## Rethinking data delivery on a global level For NBIM’s data team, transitioning to Fusion initially was about accessing better tools. Their primary goal was to help business users write better code. But implementing Fusion quickly evolved into an overall performance upgrade. By improving developer workflows, Fusion has enabled faster data delivery with improvements in efficiency. "For us, the developer experience is still first and foremost,” concludes Øyvind. “But everything else on top of that—in terms of data timeliness, better testing, and better orchestrations—is a huge bonus. We expect dbt Fusion will continue to bring a lot of value to the business.” _Øyvind joined the keynote at Coalesce 2025 to discuss NBIM's experience with dbt:_ [Watch video](https://youtu.be/KhBsI2LQQ90?si=czrfetwOhQfpmr3R&t=5327) --- --- title: "Effective strategies to enhance data quality management" description: "Improve data quality with testing, metrics, automation, and a scalable governance framework." url: "https://www.getdbt.com/blog/enhance-data-quality-management-strategies" date: "2026-03-11" authors: ["Joey Gault"] categories: ["Pulse"] --- # Effective strategies to enhance data quality management ## Understanding data quality dimensions Before implementing quality management strategies, teams need a framework for defining what "quality" means in their specific context. Data quality can be broken down into several key dimensions that provide a comprehensive view of data health. **Accuracy** measures how closely data reflects reality. If your business sold 198 new subscriptions today, that exact number should appear in your data systems. Accuracy issues often stem from conflicting upstream sources, outdated information, buggy transformation logic, or technical failures in data pipelines. **Completeness** ensures you have all required records and fields needed to answer business questions. This dimension should be defined during the planning stages of any analytics project, as completeness requirements vary based on use case. A voluntary customer survey might be valuable with 50% missing values, while critical financial data could be considered incomplete with even 2% missing values. **Consistency** verifies that data remains uniform across systems and throughout its lifecycle. This dimension is particularly critical in scenarios like healthcare, where inconsistent patient identifiers across systems could lead to missing allergy information or medication records (potentially life-threatening situations). **Validity** confirms that data values are correct for their column types and fall within acceptable ranges. An integer representing a month should only contain values between 1 and 12. String fields should match expected formats, whether that's properly-formatted ZIP codes, valid JSON, or correctly-structured GUIDs. **Freshness** measures whether data has been updated within target timeframes. Different use cases demand different service level agreements. Weekly sales reports might tolerate a one-day SLA, while real-time order tracking systems may require updates within an hour. **Uniqueness** prevents the chaos caused by duplicate records with slightly different information. Maintaining uniqueness requires defining robust primary keys and enforcing unique values, particularly challenging when data flows across multiple systems. **Usefulness** is often overlooked but critically important. Data that generates no business value (so-called "dark data") represents wasted resources. Some estimates suggest as much as 55% of corporate data sits unused, consuming storage and compute resources while delivering zero return. ## Implementing testing throughout the data lifecycle Testing represents the cornerstone of effective data quality management. Rather than treating testing as an afterthought, leading data teams integrate quality checks at every stage of the data pipeline. ### Testing during development When creating new [data transformations](https://www.getdbt.com/discover/understanding-data-transformation), teams should test both raw source data and newly transformed datasets. With raw data, you're assessing the baseline quality and determining how much cleanup work lies ahead. Key tests include verifying primary key uniqueness and non-nullness, checking that column values meet basic assumptions, and identifying duplicate rows. As you layer on transformations (cleaning, aggregating, joining, and implementing business logic), the potential for errors multiplies. Testing transformed data should verify that primary keys remain unique and non-null, row counts are correct, joins haven't introduced duplicates, and relationships between upstream and downstream dependencies align with expectations. Using [dbt](https://www.getdbt.com/product/what-is-dbt), teams can create [generic data tests](https://docs.getdbt.com/docs/build/data-tests) that can be reused across multiple projects, significantly reducing the effort required to maintain comprehensive test coverage. ### Testing during code review Before merging transformation changes into production code, running tests provides an essential guardrail. When using git-based workflows, this process invites peer review, helps debug errors, and ensures that new code meets quality standards before entering the codebase. No data model or transformation should reach production without being tested against established standards and reviewed by other team members. ### Testing in production Once transformations are deployed, automated testing on a regular schedule becomes critical. Production environments are dynamic: engineers push new features that change source data, business users add fields in CRM systems that break transformation logic, and ETL pipelines experience issues that push duplicate or missing data into warehouses. Automated tests ensure that data teams discover these issues before end users do. ## Establishing data quality metrics Defining and tracking quantitative metrics enables organizations to measure quality improvements over time and identify gaps requiring attention. [Effective metrics span multiple dimensions of data quality](https://www.getdbt.com/blog/data-quality-metrics). Incident-related metrics provide visibility into data reliability. Track the total number of data incidents, time to detection, time to resolution, and table health scores based on incidents per table. These metrics help identify both systemic issues and specific problem areas requiring focused attention. Accuracy and completeness metrics include the number of empty or incomplete values, data transformation error rates, and the ratio of tests passed to tests failed over time. These measurements provide concrete evidence of whether quality is improving or degrading. Freshness metrics track hours since last data refresh, data ingestion delays, tables with the most recent or oldest data, and minimum, maximum, and average data delays. These metrics ensure that data meets the timeliness requirements of its intended use cases. Usability metrics measure data importance scores, the number of users accessing specific tables or queries, and the percentage of dark or unused data. These metrics help teams prioritize their efforts on data that actually drives business value. Operational metrics such as dashboard uptime, table uptime, data storage costs, and time to value help teams understand the broader impact of data quality on business operations. ## Building a data quality culture Technology and processes alone cannot ensure data quality. Organizations need to cultivate a culture where everyone who works with data feels responsible for its accuracy and usefulness. A data quality culture starts with organizational alignment on the connection between value and quality. This typically takes the form of documented principles (for example, "we prioritize accuracy to maintain client trust"). Common KPIs and shared definitions ensure everyone assesses quality consistently. Making quality an integral part of daily workflows reinforces its importance. This includes requiring comprehensive tests for every new dataset, setting standards for data classification, providing publicly available quality dashboards, and streamlining processes for reporting and resolving errors. When quality checks are seamlessly integrated into existing workflows rather than treated as additional overhead, adoption increases dramatically. ## Leveraging the Analytics Development Lifecycle (ADLC) The [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) provides a framework that unites processes and tools to enable continuous quality improvements through rapid development cycles. In the planning phase, technical and business stakeholders identify existing quality issues and prioritize them. For example, a team might discover that duplicate sales records prevent accurate purchasing trend analysis. They would then establish procedures for reconciling records, preventing duplicates, and creating metrics to monitor correctness. During development, data engineering teams create transformation pipelines that produce clean datasets meeting consumer requirements. They also build tests to verify quality in both pre-production and production environments. The test and deploy phase leverages source control and pull requests for internal code review before deployment. Continuous integration and continuous deployment (CI/CD) processes test quality management code in staging environments before production release. In the operate and observe phase, data consumers use the new datasets while reporting any issues back to engineering teams. Simultaneously, data teams track metrics and alerts to identify potential problems before they cause downstream failures. This cycle repeats continuously, with each iteration delivering new quality improvements. ## Automating quality management with modern tools Implementing comprehensive quality management requires significant effort, particularly when building everything from scratch. Modern data transformation tools can dramatically reduce this burden. [dbt](https://www.getdbt.com/product/dbt) provides capabilities that streamline quality management across the entire lifecycle. Teams can create [data models](https://docs.getdbt.com/docs/build/models) that import data from multiple sources, cleaning and transforming them into analysis-ready datasets. Documentation can be added directly to models and published automatically with each production deployment, providing detailed information on data origins and meaning. dbt supports both built-in tests like not-null checks and custom tests implementing domain-specific quality requirements. Version control integration ensures all changes are tracked and reviewed, with development work isolated in branches to prevent production impacts. Job scheduling and orchestration capabilities enable teams to regularly run models and tests, bringing data changes into production while continuously performing quality checks. Unlike tools that separate transformation and testing, dbt enables automating both in unified pipelines. [CI/CD support](https://docs.getdbt.com/docs/deploy/continuous-integration) automatically runs jobs based on check-ins or completed pull requests, testing changes in pre-production before user exposure. [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) serves as a data catalog where producers and consumers can find existing datasets and documentation, as well as trace data lineage to verify origins and troubleshoot upstream issues. Monitoring and alerting capabilities track metrics and fire alerts in response to test failures, ensuring teams learn about problems immediately rather than after stakeholders discover them. ## Managing trade-offs and priorities Data engineering leaders must recognize that perfect quality across all dimensions simultaneously is often neither achievable nor necessary. Trade-offs are inevitable. Defining aggressive timeliness targets might conflict with accuracy or completeness goals. Making data more accessible across the organization may conflict with security requirements for sensitive datasets. The nature of specific data and its business use should drive not just which metrics teams track, but how much importance they assign to each. When organizations need quote-to-cash systems ensuring financial data accuracy, they must place additional weight on accuracy metrics and invest more time in testing. The obligation extends beyond ensuring numbers are correct to proving they are correct. Prioritization should focus on the data that drives the most business value. Tracking usage statistics helps identify which datasets warrant the most rigorous quality management and which might be candidates for archival or deletion. ## Conclusion Enhancing data quality management requires a comprehensive approach combining clear frameworks, rigorous testing, quantitative metrics, cultural commitment, and appropriate tooling. By understanding quality dimensions, implementing testing throughout the data lifecycle, establishing meaningful metrics, building a quality-focused culture, and leveraging modern automation tools, data engineering leaders can transform data quality from a persistent challenge into a sustainable competitive advantage. The journey toward high-quality data doesn't happen overnight. It requires sustained effort, organizational commitment, and continuous iteration. However, the payoff (trusted data that enables confident decision-making and drives business value) makes the investment worthwhile. When data quality is high and trust is strong, data teams achieve a flow state where they can focus on delivering new insights rather than constantly firefighting quality issues. ## Data quality management FAQs **What is data quality?** Data quality refers to how well data serves its intended purpose across multiple dimensions. It encompasses accuracy (how closely data reflects reality), completeness (having all required records and fields), consistency (uniformity across systems), validity (correct formats and acceptable ranges), freshness (timely updates), uniqueness (no duplicates), and usefulness (generating actual business value). Quality data enables confident decision-making and drives business outcomes, while poor quality data can cost organizations millions annually in wasted resources and flawed decisions. **Why is data quality important?** **How can organizations ensure data accuracy and reduce errors? ** Organizations can ensure data accuracy through comprehensive testing at every stage of the data lifecycle. This includes testing raw source data and transformed datasets during development, implementing code review processes before production deployment, and running automated tests in production environments on a regular schedule. Teams should verify primary key uniqueness, check that column values meet expected assumptions, ensure joins haven't introduced duplicates, and validate that relationships between data dependencies align with expectations. Modern tools can automate much of this testing, integrating quality checks seamlessly into daily workflows rather than treating them as additional overhead. --- --- title: "Data movement patterns explained (ETL, ELT, CDC & more)" description: "ETL, ELT, batch, CDC, reverse ETL—learn the key data movement patterns and when to use each." url: "https://www.getdbt.com/blog/data-movement-patterns" date: "2026-03-10" authors: ["Joey Gault"] categories: ["Pulse"] --- # Data movement patterns explained (ETL, ELT, CDC & more) ## The evolution from ETL to ELT The most fundamental shift in data movement patterns has been the transition from [ETL (Extract, Transform, Load) to ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/etl-vs-elt). This change reflects broader architectural trends driven by cloud data warehouses and their separation of compute from storage. ### ETL: the legacy approach [ETL emerged when data storage and compute were expensive and tightly coupled](https://www.getdbt.com/blog/extract-transform-load). In this pattern, data is extracted from source systems, transformed in a separate processing layer, and then loaded into the target warehouse. The transformation happens before loading because warehouses were too slow and constrained to handle heavyweight processing themselves. This approach minimizes storage costs by ensuring only transformed data lands in the warehouse. However, it introduces significant limitations. Transformations are scattered across different systems, leading to repeated work as teams implement ad hoc queries for each new project. The pattern is inflexible and difficult to scale, particularly as data volumes grow and use cases multiply. ### ELT: the modern standard [ELT reverses this order](https://www.getdbt.com/blog/extract-load-transform). Raw data is extracted from source systems and loaded directly into the warehouse, where transformations happen using the warehouse's native compute power. This approach leverages the scalability and flexibility of cloud storage and computing, making it far easier to handle large volumes of data and growing numbers of use cases. The advantages are substantial. ELT enables a more organized data architecture with transformations performed directly in the warehouse, creating a streamlined and efficient process. Tools like dbt have emerged specifically to support this pattern, bringing software engineering best practices (version control, testing, documentation, and modularity) to the transformation layer. However, ELT introduces its own challenges. As organizations scale, the number of dashboards, tables, sources, and data products increases, leading to complex warehouse environments. Without proper tooling to manage this complexity, issues like data inconsistency, lack of documentation, and version control problems can undermine trust and usability. ## Batch processing: the workhorse pattern Batch processing remains the dominant pattern for data movement in most organizations. Data is extracted and loaded on a schedule (hourly, daily, or at some other interval) and transformations run in discrete jobs. The batch hub-and-spoke architecture uses a central database as a hub that receives scheduled extracts from various source systems. The hub processes the data and pushes curated outputs downstream to BI tools or reporting marts. This pattern remains common in environments where source systems are on-premises or where compliance requires tight control over scheduling. Batch workflows have clear strengths. They're well-understood, relatively simple to implement, and sufficient for many analytical use cases. When you're building dashboards to analyze revenue trends or cohort behavior, getting new data once a day may be entirely adequate. The limitations become apparent as organizations mature. Rigid batch workflows result in lengthy refresh windows, making real-time analytics challenging or impossible. As jobs grow, monolithic scripts become harder to debug or extend. Without modularity, parallel development is constrained, and even small changes can trigger full pipeline reruns. ## Change data capture and streaming [Change Data Capture (CDC)](https://docs.getdbt.com/best-practices/how-we-handle-real-time-data/2-incremental-patterns#when-to-use-cdc) represents a fundamentally different approach to data movement. Rather than periodically extracting full or incremental snapshots, CDC detects and synchronizes changes as they occur in source systems, supporting near-real-time pipelines. Two primary CDC approaches exist. Log-based CDC reads directly from database transaction logs, offering low-latency and low-overhead synchronization. Trigger-based CDC emits change events from within the application when direct log access isn't available. The added complexity of CDC is justified when data freshness is essential. Use cases like personalization, fraud detection, and operational reporting benefit significantly from reduced latency. A delayed view in these scenarios can lead to poor decisions or missed opportunities. To make CDC pipelines reliable, teams must design for exactly-once delivery and handle schema changes proactively. Transforming events into incremental models within the warehouse maintains clean, testable downstream logic. Tools like dbt support incremental materialization, processing only new or changed records after the initial model build, which reduces runtime and compute costs while maintaining data quality. The major cloud data warehouses have begun supporting constructs that enable more real-time flows. Snowflake emphasizes streams functionality, while BigQuery and Redshift focus on materialized views. Both approaches move in the right direction, though neither fully solves the real-time challenge today. ## Data virtualization and federation Data virtualization allows teams to query data across systems without physically moving it. This pattern is useful for quick proofs of concept or when working with data that cannot be moved due to compliance or privacy requirements. However, virtualization poses considerable challenges. Live queries can introduce latency and performance risk. Joins across systems are often fragile and slow. Permissions don't always flow downstream, complicating governance. For most use cases, if data needs to be queried regularly, joined with other sources, or analyzed deeply, it's better to load it into a centralized platform like a data warehouse or lakehouse. Virtualization works best as an exception rather than a core pattern. ## Reverse ETL: completing the feedback loop An emerging pattern that's gaining significant traction is [reverse ETL (moving data from the warehouse back into operational systems)](https://www.getdbt.com/blog/reverse-etl-vs-etl). This represents a fundamental shift in how organizations think about data movement. Traditionally, data flows one way: from operational systems into the warehouse where it's analyzed. If that analysis is going to drive action, a human must pick it up and act on it manually. [Reverse ETL](https://www.getdbt.com/blog/reverse-etl-playbook) changes this by feeding transformed data directly back into operational tools. The use cases are compelling. Customer support staff can see key user behavior data directly in their help desk product. Sales professionals can access product usage data inside their CRM. Marketing teams can trigger automated messaging flows based on product clickstream data, all without maintaining separate event tracking implementations. This pattern makes the data warehouse not just an analytical endpoint but a central nervous system that powers operational business systems. For teams writing transformation code in dbt today, this means their work will increasingly power production systems, not just internal analytics. This makes the job both more challenging and more impactful. ## The modern data lake pattern A newer pattern gaining adoption is loading data directly into open table formats like Iceberg or Delta in object storage, rather than into proprietary warehouse storage. This approach combines the organizational structure of a data warehouse with the flexibility and cost efficiency of a data lake. In this pattern, data is loaded into formats like Iceberg in S3, creating an organized, warehouse-like structure in the data lake. Different query engines can then operate on top of this data (Databricks, Snowflake, Athena, or others) without duplicating storage. This pattern relies on open catalogs to manage metadata and enable multiple engines to query the same data. It offers significant advantages: customers avoid paying for storage multiple times, can use different compute engines for different workloads, and maintain flexibility in their technology choices. The pattern is still maturing. Early adopters are using it successfully, but it's not yet turnkey. Making this widely adopted will require continued ecosystem development to simplify implementation and improve interoperability. ## Choosing the right pattern No single pattern fits all use cases. The right choice depends on latency requirements, data volumes, organizational maturity, and specific business needs. Batch processing remains appropriate for most analytical workloads where daily or hourly updates suffice. CDC and streaming make sense when real-time data drives operational decisions. Reverse ETL becomes valuable when data needs to flow back into operational systems to automate business processes. The modern data lake pattern suits organizations with diverse compute needs and large data volumes. The key is understanding that these patterns aren't mutually exclusive. Mature data organizations often use multiple patterns simultaneously, choosing the right approach for each data source and use case. ## Building for the future As the data ecosystem continues to evolve, several trends are clear. Latency will continue to decrease as streaming capabilities mature. Data will increasingly flow bidirectionally, powering both analytics and operational systems. Open formats and standards will provide more flexibility in choosing compute engines. For data engineering leaders, this means building with modularity and flexibility in mind. Invest in tools that support multiple patterns. Establish clear ownership and data contracts. Implement testing and observability from the start. Treat analytics code like application code, with version control and CI/CD. The patterns for data movement have matured significantly, but they continue to evolve. Understanding the major patterns, their trade-offs, and their appropriate use cases positions data teams to build infrastructure that scales with their organization's needs. *** **Related resources:** - [Data transformation best practices](https://www.getdbt.com/blog/data-transformation-best-practices) - [Understanding ETL vs ELT](https://www.getdbt.com/blog/etl-vs-elt) - [Data integration guide](https://www.getdbt.com/blog/data-integration-guide) - [Getting started with dbt](https://docs.getdbt.com/docs/introduction) ## ## Data movement FAQs **What is the purpose of data movement?** Data movement enables organizations to extract data from source systems and make it available for analysis, reporting, and operational use. The purpose has evolved beyond just feeding analytical endpoints; modern data movement patterns create a central nervous system that powers both analytics and operational business systems. This allows organizations to transform raw data into actionable insights, automate business processes, and drive decisions across customer support, sales, marketing, and other operational functions. **When should you use the broadcast pattern instead of migration for moving data between systems?** Batch processing with a hub-and-spoke architecture is appropriate when you need to push data from central systems to multiple downstream destinations on a scheduled basis. This pattern works well for analytical workloads where daily or hourly updates are sufficient, such as building dashboards to analyze revenue trends or cohort behavior. However, when data freshness is essential for operational decisions (like personalization, fraud detection, or real-time reporting), Change Data Capture (CDC) and streaming patterns become more appropriate as they synchronize changes as they occur rather than on fixed schedules. **What are the most common data integration patterns and when should you use them?** The most common patterns include ELT (Extract, Load, Transform), batch processing, Change Data Capture, and reverse ETL. ELT is the modern standard for most analytical workloads, loading raw data directly into cloud warehouses where transformations happen using native compute power. Batch processing remains appropriate when daily or hourly updates suffice. CDC and streaming make sense when real-time data drives operational decisions and reduced latency is critical. Reverse ETL becomes valuable when transformed data needs to flow back into operational systems like CRMs or help desk tools to automate business processes. Mature organizations often use multiple patterns simultaneously, choosing the right approach for each specific data source and use case. --- --- title: "How data transformation improves data quality and analysis" description: "Learn how transformation methods improve data quality, consistency, and analysis at scale with dbt." url: "https://www.getdbt.com/blog/data-transformation-improves-data-quality" date: "2026-03-09" authors: ["Joey Gault"] categories: ["Pulse"] --- # How data transformation improves data quality and analysis ## Understanding data transformation [Data transformation](https://www.getdbt.com/product/develop) is the process of converting one materialized data asset—such as a table or view—into another purpose-built for analytics through SQL or Python. Transformation creates structure from unorganized data, applies business logic consistently, and ensures analysts work from reliable foundations rather than wrestling with raw source systems. The transformation process follows four key stages. Discovery and profiling assess data structure, quality, and characteristics to identify anomalies and inconsistencies. Cleansing corrects inaccuracies, fills missing values, and removes duplicates. Mapping aligns data structures to target system requirements, converting data types and reorganizing fields as needed. Storage loads transformed data into centralized repositories like data warehouses, where it's available for analysis and reporting. Modern data transformation typically occurs within [ELT (Extract, Load, Transform)](https://docs.getdbt.com/docs/introduction) architectures, where data transforms after loading into the destination warehouse. This approach has largely replaced traditional ETL because cloud computing makes it more cost-efficient to load data before transformation. Raw data becomes immediately available to everyone with warehouse access, and teams with different needs can transform it to their specific requirements. ## How transformation methods improve data quality Transformation's most direct impact on analysis is improving data quality. [Low-quality data costs organizations an estimated 20–30% of their revenue](https://www.gartner.com/smarterwithgartner/how-to-improve-your-data-quality), creating downstream problems that compound throughout the analytics process. Transformation addresses these quality issues systematically through three core methods. **Data cleaning** finds and fixes errors and inconsistencies — correcting malformatted values, filling missing entries, and eliminating duplicate records. Without this foundational work, analysts spend time investigating anomalies that stem from data quality issues rather than genuine business insights. Clean data means faster analysis cycles and more reliable conclusions. **Normalization** transforms data into standard ranges or formats to ensure consistency and comparability across different sources. A global retail company might normalize transaction data by converting all currency values to USD, enabling accurate financial reporting across regions. This standardization eliminates distortions that arise when comparing data measured in different units or scales. **Validation** verifies that data meets specified criteria before it's eligible for analytics use. Common validation checks include format verification (ensuring phone numbers follow consistent patterns), uniqueness constraints (preventing duplicate customer IDs), completeness checks (confirming no critical fields are empty), and range validation (flagging values outside expected parameters). These checks catch problems early, before they propagate into reports and dashboards that inform business decisions. ## Creating consistency at scale As organizations grow, maintaining consistency across datasets gets harder. Different teams may use divergent naming conventions, apply SQL standards inconsistently, or implement varying testing approaches. That inconsistency creates duplicative work, misaligned metrics, and unclear data relationships that undermine analytical accuracy. Transformation addresses these challenges through standardization and centralization. Rather than letting each analyst implement their own version of key business metrics, transformation codifies definitions in centralized, version-controlled locations. When revenue calculations, customer segmentation logic, or operational KPIs exist in a single authoritative source, everyone works from the same definitions — and conflicting reports about supposedly identical metrics stop happening. The value of consistency extends beyond avoiding confusion. Standardized transformation logic becomes reusable across different analytical projects. Instead of repeatedly solving the same data preparation problems, teams reference foundational work completed by others. This modularity reduces duplication, improves maintainability, and makes dependencies explicit through clear lineage tracking. ## Enabling advanced analytics and integration Transformation makes data suitable for advanced analytical techniques. [Machine learning and AI models](https://www.getdbt.com/product/ai) are only as good as the data they're trained on. These approaches require large volumes of high-quality, consistently formatted data — exactly what a well-structured transformation layer provides. Data enrichment enhances internal data with external sources to create deeper insights. A retailer might enrich shipment data with real-time weather information to predict delivery delays and improve customer communication. This augmentation transforms basic operational data into predictive intelligence that drives proactive decision-making. Integration merges data from different sources into unified datasets that enable comprehensive analysis. Combining CRM systems, online store accounts, and loyalty program databases creates 360-degree customer views that would be impossible from any single source. This integrated perspective reveals patterns and relationships that stay hidden when data remains siloed in separate systems. Transformation also supports reverse ETL workflows by joining multiple datasets into enriched data models. This enables seamless integration into operational systems, putting timely insights in the tools stakeholders use daily — not just in analytical dashboards. ## ELT and the transformation layer The shift from ETL to ELT has fundamentally changed how transformation supports analysis. In legacy ETL, transformation occurred before loading, often in separate systems with limited scalability. Transformations scattered across different platforms, leading to repeated, inconsistent work as teams implemented ad hoc queries for each new project. ELT reverses this order — loading raw data into warehouses first, then transforming it there. This leverages cloud infrastructure scalability and flexibility, making it easier to handle growing data volumes and expanding use cases. More importantly, ELT enables a more organized data architecture where transformations occur in a centralized location using consistent tooling. A well-implemented transformation layer — a network of transformations and routines that process data automatically — provides several concrete advantages. It prevents conflicts between analyses and data silos. It provides a single, authoritative base of [dbt models](https://docs.getdbt.com/docs/build/models) ensuring everyone works from the same definitions and standards. It eliminates redundant data preparation work, reducing costs and improving speed. This transformation layer also creates reusable, complex datasets that speed up reporting. Rather than repeatedly cleaning data and calculating metrics manually, automated transformation generates accurate datasets without duplicated effort. Analysts spend less time on data preparation and more time on actual analysis. ## Implementing transformation best practices with dbt [dbt](https://www.getdbt.com/product/what-is-dbt) is a SQL-first transformation workflow that lets teams deploy analytics code using software engineering best practices, giving data teams the control and visibility needed to deliver reliable data products. dbt enhances transformation workflows through several capabilities. Modular transformation logic enables reusable SQL that ensures consistency and reduces redundancy across different data models. [Automatic documentation](https://docs.getdbt.com/docs/build/documentation) generates transparency for all transformations, making collaboration across teams easier. [Integrated testing](https://docs.getdbt.com/docs/build/data-tests) and [version control](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview) ensure transformations are reliable and changes are tracked. dbt's integrated development environment simplifies development and eliminates infrastructure management burdens, letting teams focus on transformation logic rather than maintaining technical infrastructure. As organizational needs grow, dbt scales accordingly. For data engineering leaders, dbt addresses governance and compliance requirements through [centralized access control](https://www.getdbt.com/product/dbt) and detailed documentation. Automated testing, version control, and documentation reduce error risks and ensure data accuracy — freeing teams to focus on high-value work rather than infrastructure maintenance. [Sign up for dbt for free](https://www.getdbt.com/signup) and start building reliable transformation workflows, or [talk to our team](https://www.getdbt.com/contact) about what the right setup looks like at your scale. ## Real-world impact The practical benefits of effective transformation are evident in how organizations apply these capabilities. [Nasdaq](https://www.getdbt.com/case-studies/nasdaq) leveraged dbt to overcome data engineering bottlenecks, significantly reducing the time required to produce business-critical reports. [Siemens](https://www.getdbt.com/case-studies/siemens) implemented dbt to manage complex transformations across global operations, maintaining consistency in data definitions across different regions and departments. These examples show that transformation methods don't just support analysis — they make it possible. Without proper transformation, these organizations would face fragmented data, inconsistent metrics, and analytical bottlenecks that prevent timely decision-making. ## Conclusion Data transformation creates the foundational conditions that make reliable analysis possible: improving data quality through systematic cleaning, validation, and normalization; enabling consistency through standardized and reusable transformation logic; preparing data for machine learning and AI; and integrating disparate sources into comprehensive analytical datasets. For data engineering leaders, robust transformation isn't optional — it's how you deliver value from data investments. Organizations that implement systematic transformation approaches and apply engineering best practices to their workflows are better positioned to generate competitive advantages from their data. Explore [dbt documentation](https://docs.getdbt.com/docs/introduction) and [best practices for reliable data workflows](https://docs.getdbt.com/best-practices) to go deeper. ## FAQs **What is data transformation?** Data transformation is the process of converting one materialized data asset — such as a table or view — into another purpose-built for analytics through SQL or Python. The process follows four key stages: discovery and profiling to assess data structure and quality, cleansing to correct inaccuracies and remove duplicates, mapping to align data structures to target system requirements, and storage to load transformed data into centralized repositories like data warehouses. **How does data transformation improve data quality and consistency?** Transformation improves quality through three systematic methods. Data cleaning fixes errors, corrects malformatted values, fills missing entries, and eliminates duplicates so analysts focus on genuine business insights rather than data quality issues. Normalization transforms data into standard ranges or formats to ensure consistency across sources — such as converting all currency values to a single standard. Validation verifies data meets specified criteria through format verification, uniqueness constraints, completeness checks, and range validation. For consistency at scale, transformation codifies key business metrics in centralized, version-controlled locations, ensuring everyone works from the same definitions and eliminating conflicting reports. **What are the key differences between ETL and ELT, and when is each approach better?** ETL (Extract, Transform, Load) transforms data before loading it into the destination system, often in separate systems with limited scalability — leading to scattered, inconsistent transformations across platforms. ELT (Extract, Load, Transform) loads raw data into warehouses first, then transforms it there. ELT has largely replaced ETL because cloud computing makes it more cost-efficient to load before transforming. ELT leverages cloud scalability, handles growing data volumes more effectively, and enables a centralized transformation architecture where consistent tooling applies across all datasets — with raw data immediately accessible to everyone with warehouse access. --- --- title: "Effective strategies to improve data quality across your organization" description: "Proven strategies to improve data quality with testing, governance, and scalable analytics workflows." url: "https://www.getdbt.com/blog/strategies-improve-data-quality" date: "2026-03-06" authors: ["Joey Gault"] categories: ["Pulse"] --- # Effective strategies to improve data quality across your organization ## Understanding the scope of the problem The financial impact of poor data quality is substantial. Research indicates that companies lose millions annually due to data quality issues, while a significant portion of corporate data fails to meet basic quality standards. Beyond the direct costs, poor data quality creates a ripple effect: business users question the data, trust in the data team erodes, and stakeholders begin conducting their own ad hoc analyses with inconsistent methodologies. The root cause often lies in reactive approaches to data quality. Too many organizations wait for issues to surface in BI tools or be caught by end users. By that point, decisions may have already been made on faulty data, and the damage to credibility is done. Effective data quality enhancement requires a fundamental shift from reactive firefighting to proactive prevention. ## Establishing a comprehensive data quality framework A data quality framework provides the foundation for systematic improvement across your organization. Rather than treating data quality as a series of one-off fixes, a framework establishes consistent principles, standards, and processes that teams can apply throughout the data lifecycle. Your framework should address multiple dimensions of data quality. Accuracy ensures that data values reflect reality. If your business sold 198 subscriptions today, that's what should appear in your data. Completeness verifies that all required records and fields exist to answer business questions. Consistency maintains alignment across upstream and downstream systems as data flows through pipelines. Validity checks that values conform to expected formats and ranges. Freshness ensures data updates within acceptable timeframes for each use case. Uniqueness prevents duplicate records from corrupting analyses. The key is recognizing that not all dimensions matter equally for every dataset or use case. Your framework should help teams identify which quality dimensions are critical for their specific business context and prioritize accordingly. This targeted approach prevents teams from getting overwhelmed trying to achieve perfect quality across every possible dimension. ## Integrating testing throughout the data lifecycle Testing represents the practical implementation of your data quality framework. Rather than treating testing as a final checkpoint before deployment, effective organizations integrate testing at multiple stages of the data lifecycle. Testing should begin with raw source data as soon as it lands in your warehouse. At this stage, you're assessing the baseline quality of your data and understanding the cleanup work required. Common tests include checking primary key uniqueness and non-nullness, verifying that column values meet basic assumptions, and identifying duplicate rows. These early tests help you catch upstream issues before investing effort in transformation. As you [transform data (cleaning, aggregating, joining, and implementing business logic)](https://www.getdbt.com/blog/analytics-engineering-transformation), the complexity increases and so does the potential for error. Testing transformed data should verify that primary keys remain unique and non-null, row counts align with expectations, joins don't introduce duplicates, and relationships between upstream and downstream dependencies match your assumptions. The specific tests you implement should reflect your business requirements rather than following a generic checklist. Testing during pull requests adds a crucial peer review layer. Before merging changes into your production codebase, other team members can review your code, help debug errors, and ensure new models meet your quality standards. This collaborative approach builds transparency and catches issues that individual developers might miss. Production testing provides ongoing validation after deployment. Even well-tested code can encounter issues when source systems change, business users add new fields, or ETL pipelines experience problems. Automated tests running on regular schedules ensure that you (not your end users) discover these issues first. Tools like dbt enable you to [implement automated testing](https://www.getdbt.com/blog/data-quality-testing) that runs continuously in production environments. ## Implementing the Analytics Development Lifecycle The [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) provides a structured approach for embedding data quality into every stage of analytics work. This iterative process ensures that quality considerations inform decision-making from initial planning through ongoing operations. The planning phase brings together technical and business stakeholders to identify data quality issues and establish priorities. For example, if duplicate records in sales data are preventing accurate purchasing trend analysis, the team defines procedures for reconciling records, preventing future duplicates, and creating metrics to monitor correctness. This upfront alignment ensures everyone understands both the problem and the success criteria. During development, data engineers create transformation pipelines that produce clean datasets meeting consumer requirements. They also build tests to validate quality in both pre-production and production environments. This parallel development of transformation logic and quality tests ensures that validation is built in rather than bolted on. The test and deploy phase leverages version control and CI/CD processes to review changes internally before release. By testing data quality management code in pre-production environments, teams catch issues before they impact production users. This staged approach significantly reduces the risk of deployments while maintaining development velocity. The operate, observe, discover, and analyze phase encompasses both consumption and monitoring. Data consumers use the clean datasets to create reports and applications, reporting any issues back to engineering teams. Simultaneously, data teams track metrics and alerts to identify potential problems proactively. This continuous feedback loop drives ongoing improvement. ## Leveraging automation and tooling Manual data quality processes don't scale. As data volumes grow and pipelines multiply, organizations need robust tooling to automate testing, monitoring, and documentation. Modern data transformation tools like dbt provide capabilities specifically designed to support enterprise data quality programs. [Automated testing in dbt ](https://docs.getdbt.com/docs/build/data-tests)allows you to define tests once and run them consistently across environments. Built-in tests handle common scenarios like null checks and unique constraints, while custom tests enable domain-specific validation. Generic tests can be reused across multiple projects, reducing duplication and ensuring consistency. For example, you might create a test verifying that refund amounts are never negative, then apply that test to every table containing financial transactions. Version control integration ensures that all changes to data models, transformations, and tests are tracked and reviewed. Isolating in-development changes in branches allows engineers to work on new features without affecting production. Pull request workflows enable peer review before merging changes, catching issues early when they're easiest to fix. Job scheduling and orchestration capabilities enable regular execution of data models and tests, bringing changes into production and continuously performing quality checks. Unlike tools that separate transformation and testing, dbt allows you to automate both in unified pipelines, simplifying operations and reducing the chance of gaps. [Data cataloging through dbt Catalog](https://www.getdbt.com/product/dbt-catalog) helps both producers and consumers find existing datasets and related documentation. Column-level lineage enables tracing data back to its source, which is invaluable for troubleshooting upstream issues. When quality problems arise, lineage helps you quickly identify root causes and fix them at their source rather than applying downstream patches. ## Building organizational capabilities Technology alone doesn't create high-quality data. Effective data quality enhancement requires building organizational capabilities that span technical and business functions. Data domain owners must be involved in validating that data conforms to business requirements. Without their input, technical teams may optimize for the wrong quality dimensions or miss critical business logic. Establishing clear ownership and accountability prevents data quality from becoming everyone's responsibility and therefore no one's responsibility. Each dataset should have a designated owner responsible for its quality, with clear escalation paths when issues arise. This ownership model ensures someone is always monitoring quality metrics and responding to alerts. Documentation plays a crucial role in data quality by ensuring that consumers understand what data represents and how it was calculated. Auto-generated documentation that updates with every deployment keeps information current without requiring manual maintenance. When documentation lives alongside the data models themselves, it's more likely to stay accurate as the code evolves. Metrics and monitoring provide visibility into data quality trends over time. Rather than relying on anecdotal reports of issues, quantitative metrics enable data leaders to track improvements, identify problem areas, and justify investments in quality initiatives. Common metrics include the total number of data incidents, time to incident detection, time since last data refresh, and test pass rates for critical tables. ## Creating sustainable improvement cycles Data quality enhancement isn't a one-time project but an ongoing practice. The most successful organizations treat data quality as a continuous improvement discipline, regularly identifying new use cases and addressing them through iterative development cycles. Each cycle should be short enough to deliver tangible value quickly while building toward longer-term quality goals. Start by identifying high-impact use cases where quality improvements will deliver clear business value. Perhaps a critical executive dashboard relies on data with known accuracy issues, or a customer-facing application occasionally displays incorrect information. Prioritizing these visible, consequential use cases builds momentum and demonstrates the value of quality investments. As you address initial use cases, capture learnings and incorporate them into your framework. Which tests proved most valuable? What quality dimensions mattered most to stakeholders? Where did you encounter unexpected challenges? These insights should inform your approach to subsequent use cases, creating a flywheel effect where each cycle makes the next one more efficient. Celebrate wins and share them broadly across the organization. When data quality improvements enable better decisions, prevent costly mistakes, or unlock new capabilities, make sure stakeholders know about it. These success stories build support for continued investment and help shift organizational culture toward viewing data quality as a shared priority rather than just a technical concern. ## Moving forward Enhancing data quality across an organization requires a multifaceted approach that combines frameworks, testing, automation, and organizational capabilities. There's no silver bullet, but there is a clear path forward: establish consistent standards, integrate testing throughout the data lifecycle, leverage modern tooling to automate quality checks, and build a culture of continuous improvement. For data engineering leaders, the imperative is clear. The cost of poor data quality (both financial and reputational) is too high to ignore. By implementing systematic approaches to data quality enhancement, you can build trust with business stakeholders, enable better decision-making, and create a foundation for advanced analytics and AI initiatives. The organizations that treat data quality as a strategic priority rather than a tactical afterthought will be the ones that successfully leverage data as a competitive advantage. Learn more about [building a data quality framework](https://www.getdbt.com/blog/how-to-choose-a-data-quality-framework) and [getting started with data quality management](https://www.getdbt.com/blog/getting-started-with-data-quality-management) to begin your organization's data quality journey. ## Data quality FAQs **Why is data quality important?** **What are the primary dimensions of data quality?** **How do data governance, validation rules, and continuous monitoring work together to maintain high data quality over time?** These elements work together through an integrated approach across the data lifecycle. Validation rules are implemented through automated testing at multiple stages (from raw source data through transformation and into production), checking for issues like null values, duplicates, and business logic errors. Continuous monitoring tracks quality metrics and alerts in production environments, ensuring teams discover issues before end users do. Data governance provides the organizational framework by establishing clear ownership and accountability for each dataset, involving business domain owners in validation, and creating escalation paths when issues arise. This combination of technical automation, ongoing observation, and organizational structure creates sustainable improvement cycles where learnings from each iteration inform future quality enhancements. --- --- title: "How AI improves data lineage at scale" description: "Discover how AI accelerates data lineage with automated docs, testing, and scalable governance." url: "https://www.getdbt.com/blog/ai-data-lineage" date: "2026-03-04" authors: ["Joey Gault"] categories: ["Pulse"] --- # How AI improves data lineage at scale ## The lineage challenge: scale and complexity Data lineage provides a holistic view of how data moves through an organization, where it's transformed, and how it's consumed. Most commonly, this takes the form of a [directed acyclic graph (DAG)](https://www.getdbt.com/blog/dag-use-cases-and-best-practices) showing sequential workflows at the dataset or column level, along with a catalog of data asset origins, owners, definitions, and policies. For teams using [dbt](https://www.getdbt.com/product/what-is-dbt), lineage is automatically inferred from the relationships between data sources and models. The DAG visualizes upstream dependencies (nodes that must come before a current model) and downstream relationships showing what's impacted by changes. This automatic generation is crucial: manual lineage tracking is time-consuming, error-prone, and quickly becomes outdated. However, as dbt projects scale with organizational growth, challenges emerge. The number of sources, models, macros, seeds, and exposures inevitably grows. With thousands of nodes in your DAG, auditing for inefficiencies or WET (write everything twice) code becomes difficult. Complex workflows compound these difficulties, particularly when granularity shifts from table-level to column-level lineage. Describing a data source's movement as it's filtered, pivoted, and joined with other tables requires specificity that traditional lineage systems don't always provide. These scaling challenges create bottlenecks for data teams. Engineers spend hours tracing data flows manually when issues arise. Impact analysis for upstream changes becomes guesswork. Documentation falls behind as teams prioritize shipping features over maintaining metadata. The very tool meant to make your work easier (the lineage graph) can become overwhelming without the right support. ## AI-assisted development: building better lineage faster [AI can address lineage challenges](https://www.getdbt.com/blog/how-ai-changes-data-pipelines) by accelerating the development work that creates and maintains your lineage graph. Large language models have proven adept at generating transformation code that experienced engineers can refine, test, and deploy faster than writing from scratch. When integrated into your analytics workflow, AI copilots help teams produce higher-quality transformations with less manual effort. This directly improves lineage quality because better transformation code (with clearer logic, consistent naming conventions, and proper documentation) creates more understandable and maintainable lineage graphs. Consider how AI-assisted development works in practice with tools like dbt Copilot. Rather than writing SQL from scratch, data producers can generate inline SQL using natural language descriptions. This shortens development time while ensuring generated code follows organizational naming conventions and best practices. New contributors can use natural language to generate working SQL, while experienced engineers rely on AI to refine complex logic or apply bulk edits across projects. This capability expands who can confidently contribute to data pipelines, supporting broader data democratization. When more team members can build transformation code, your lineage graph becomes more comprehensive and reflects a wider range of data flows across the organization. AI assistance also helps with code refinement and optimization. Engineers can use AI to apply project-wide edits, generate complex regex patterns, or enforce custom SQL style guides. This consistency reduces review time and maintains high-quality code across your analytics project, which translates directly to cleaner, more reliable lineage. ## Automated documentation: making lineage understandable Even the most comprehensive lineage graph falls short if data consumers can't understand what they're looking at. Documentation bridges this gap, providing shared context that improves discoverability and increases confidence in how data is used. Well-documented models help analysts, stakeholders, and new team members quickly understand where data comes from and how key fields are calculated. But for large or legacy projects, creating documentation manually represents a daunting lift. This is where AI can provide substantial value. AI can analyze SQL logic, historical query patterns, and model metadata to generate documentation automatically. It surfaces plain-language explanations for complex logic or obscure field names, giving teams a head start on documentation that would otherwise require hours of manual work. For data lineage specifically, AI-generated documentation enriches each node in your DAG with context. Instead of just seeing that model B depends on model A, users can understand why that dependency exists and what transformations occur between them. This contextual layer makes lineage graphs genuinely useful for troubleshooting and impact analysis rather than just visual representations of connections. The documentation process becomes more sustainable when AI handles the initial generation. Teams can then refine and expand documentation organically over time as they work with the data, rather than facing the overwhelming task of documenting everything from scratch. ## AI-powered testing: validating lineage accuracy [Data lineage](https://www.getdbt.com/blog/what-is-data-lineage) is only valuable if it's accurate. Testing ensures that the transformation code underlying your lineage graph works as expected and that the relationships shown in your DAG reflect reality. AI can automatically generate test suites based on your [dbt models](https://docs.getdbt.com/docs/build/models) and their schema relationships. Rather than manually writing tests for each transformation, teams can use AI to create comprehensive test coverage that validates data quality at every stage of the analytics development lifecycle. This automated test generation serves multiple purposes for lineage. First, it catches issues earlier in the development process, before incorrect transformations make it into your lineage graph. Second, it provides confidence that the dependencies shown in your DAG are accurate and that data flows as expected. Third, it reduces the manual burden of test creation, making it more likely that teams will maintain robust test coverage as projects scale. When tests run throughout your workflow (during development, on pull request checks, and before production deployment) they create continuous validation of your lineage. This ongoing verification ensures that your lineage graph remains a reliable source of truth rather than drifting out of sync with reality. ## Semantic layers and AI: consistent metrics definitions Data lineage extends beyond transformation code to include how data is ultimately consumed. A [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) defines metrics in a consistent, centrally governed way, ensuring everyone in the organization uses the same definitions for business-critical measures. AI can accelerate semantic layer development by generating metric definitions automatically and recommending common metrics based on your data models. This scaffolding helps teams build comprehensive semantic layers faster, which in turn creates richer lineage graphs that extend all the way to business metrics and dashboards. The combination of semantic layers and AI also improves lineage comprehension. When your lineage graph shows not just technical transformations but also how those transformations connect to business metrics, stakeholders can better understand the impact of data changes. An engineer modifying a source table can see exactly which business metrics will be affected, enabling more informed decision-making. [Context protocols like Model Context Protocol (MCP)](https://www.getdbt.com/blog/mcp) further enhance this capability by allowing AI systems to access enterprise metadata and data models. This contextual awareness means AI tools can provide accurate, governed answers about data lineage without compromising security or governance. ## Implementing AI for lineage: practical considerations While AI offers significant benefits for data lineage, implementation requires thoughtful integration into existing workflows. AI copilots work best when embedded within a mature, collaborative analytics process like the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle), a framework that helps teams build and manage analytics code at scale with speed and quality. The ADLC includes structured processes and checkpoints ensuring all transformations shipped to production are trustworthy and aligned with business needs. On the technical side, it ensures every transformation is defined as code, version-controlled, peer-reviewed, tested before deployment, and monitored in production. Adding AI outside this framework won't improve code quality and might increase risk. Every piece of AI-generated code must be reviewed and tested thoroughly before reaching production. This is why integration with tools like [dbt](https://www.getdbt.com/product/dbt) is crucial: it provides built-in guardrails including version control, automated testing, deployment pipelines, and automatically generated documentation. Data engineering leaders should also consider granularity requirements when implementing AI-assisted lineage. Determine whether you need table-level lineage or column-level lineage, and ensure your AI tools can support the required detail level. Look for automation that captures lineage as part of your data pipeline rather than requiring manual updates. Integration with existing data management tools is equally important. Lineage should be easy to access and use for all stakeholders, from data engineers to business users. Scalability matters too: your lineage solution needs to handle large, complex data flows and adapt to changing requirements as your organization grows. ## The future of AI and data lineage As AI capabilities advance, the relationship between AI and data lineage will deepen. AI will become embedded throughout the data lifecycle rather than applied as a separate layer. Natural language interfaces will make lineage exploration more accessible to non-technical users who can ask questions about data flows conversationally. AI systems will increasingly optimize themselves, automatically adjusting to changing data patterns and business requirements. This autonomous operation will reduce maintenance needs and help lineage systems adapt to evolving data landscapes without constant manual intervention. The boundaries between different types of data products will blur as reporting, analysis, prediction, and automation merge into unified experiences. Lineage will extend beyond technical data flows to encompass the full journey from raw data to business decisions, with AI helping teams understand and optimize every step. For data engineering leaders, the opportunity is clear: AI can help you build more comprehensive, accurate, and maintainable data lineage systems. The key is integrating AI thoughtfully into structured workflows that maintain quality and governance while accelerating development. ## Conclusion Data lineage remains fundamental to modern data work, enabling root cause analysis, impact assessment, and data transparency. As data landscapes grow in complexity, maintaining comprehensive lineage becomes both more critical and more challenging. AI offers practical solutions to these challenges by accelerating transformation development, automating documentation generation, enabling robust testing, and supporting semantic layer creation. When integrated into structured analytics workflows with appropriate guardrails, AI helps teams build lineage systems that scale with organizational growth. The organizations that succeed will treat AI as a core component of their data strategy, using it to enhance rather than replace human expertise. By combining AI capabilities with strong governance frameworks and collaborative processes, data engineering leaders can deliver lineage systems that truly serve their organizations, making data more discoverable, trustworthy, and valuable for everyone. *** **Related resources:** - [Getting started with data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) - [What is data lineage, and why do you need it?](https://www.getdbt.com/blog/what-is-data-lineage) - [dbt Catalog for lineage visualization](https://www.getdbt.com/product/dbt-catalog) - [Column-level lineage in dbt Explorer](https://docs.getdbt.com/docs/collaborate/explore-projects) - [How we structure our dbt projects](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview) ## AI and Data lineage FAQs **How can AI automate end-to-end data lineage mapping across hybrid and multi-cloud environments?** **What key features should enterprises look for when evaluating AI-powered data lineage solutions?** **How does AI-powered data lineage improve data governance and streamline manual, error-prone processes?** AI-powered data lineage eliminates time-consuming manual tracking by automatically generating transformation code, documentation, and tests. It produces higher-quality transformations with clearer logic, consistent naming conventions, and proper documentation, creating more understandable and maintainable lineage graphs. AI analyzes SQL logic and metadata to generate plain-language explanations for complex transformations, making lineage graphs genuinely useful for troubleshooting and impact analysis. By automating test generation and running continuous validation throughout the development workflow, AI ensures lineage accuracy and prevents incorrect transformations from entering production. This reduces the manual burden on data teams while maintaining robust governance and quality standards. --- --- title: "How AI is transforming modern data pipelines" description: "AI is reshaping data pipelines—driving real-time data, automation, and AI-assisted modeling at scale." url: "https://www.getdbt.com/blog/how-ai-changes-data-pipelines" date: "2026-03-02" authors: ["Joey Gault"] categories: ["Pulse"] --- # How AI is transforming modern data pipelines ## The new requirements AI places on data pipelines Traditional data pipelines were designed for a different era. They focused on batch processing, static dashboards, and structured reporting with predictable workloads. AI applications demand something fundamentally different. They require real-time or near-real-time data ingestion to stay accurate. They need continuous data flow rather than scheduled batch updates. They depend on automated model retraining as new data arrives. Without infrastructure designed for these requirements, AI systems learn from outdated or low-quality data, leading to poor predictions and costly mistakes. The [shift from ETL to ELT architectures](https://www.getdbt.com/blog/etl-vs-elt) laid important groundwork for this transition. By loading raw data into cloud warehouses first and transforming it using the warehouse's processing power, ELT enables the parallel processing and iterative reprocessing that AI workloads require. But AI pushes these requirements further. GenAI applications and autonomous agents don't just need clean data; they need it continuously updated, rigorously tested, and delivered with clear lineage and governance. Consider the range of AI use cases now entering production: customer service chatbots that provide human-like interactions, fraud detection systems that identify suspicious transactions in real time, code generation tools that accelerate software development, and autonomous systems that operate without human intervention. Each of these applications fails without pipelines that deliver trustworthy, timely data at scale. ## Core components of AI-ready pipelines Building pipelines that support AI requires rethinking several fundamental components. Data ingestion must handle diverse sources (databases, APIs, event streams, and unstructured formats) with minimal latency. Unlike batch-oriented pipelines that can wait for scheduled runs, AI systems require data to flow continuously or in near-real-time to maintain accuracy. Transformation becomes more critical and more complex in AI contexts. Raw data is never AI-ready out of the box. It requires cleaning, structuring, and testing before it can power models reliably. [dbt](https://www.getdbt.com/product/what-is-dbt) has become the industry standard for data transformation in modern enterprises by defining reproducible transformations as modular SQL models rather than lengthy scripts. This modular approach ensures consistency across AI workflows and allows pipelines to continuously deliver high-quality, structured datasets. By using version-controlled, testable transformations, dbt reduces engineering overhead and speeds up iteration, both essential for AI development cycles. Feature engineering represents a distinct challenge for AI pipelines. Unlike standard pipelines that focus on predefined metrics, AI pipelines must uncover meaningful patterns within data for predictive accuracy. Feature engineering transforms raw data into the specific inputs that AI models use to make predictions. With dbt, teams can build reusable, version-controlled feature sets using SQL and keep them updated incrementally as new data arrives. This approach integrates naturally with platforms like Snowflake's feature store, ensuring consistent, governed datasets for machine learning workflows. Model training and fine-tuning depend entirely on the quality and timeliness of transformed data. Retrieval-Augmented Generation enhances Large Language Model output by injecting real-time context from internal datasets. dbt supports these processes by automating data transformation and providing incremental updates, ensuring models can be refreshed efficiently without full reprocessing. Monitoring and feedback loops become non-negotiable in AI pipelines. Model performance must be monitored continuously to prevent degradation. Data drift detection catches shifts in data quality before they corrupt model outputs. dbt's column-level lineage improves the auditability and reliability of AI pipelines by making it possible to trace any data point back to its source and understand every transformation applied along the way. Security and governance take on heightened importance when data feeds AI systems that make consequential decisions. dbt embeds governance into the transformation layer with version control, role-based access, and automated testing. It integrates with platforms like Alation and Atlan to expose lineage, ownership, and compliance metadata, helping organizations meet AI audit and regulatory requirements. ## How AI transforms pipeline development The most striking change AI brings to data pipelines isn't just in what pipelines must deliver; it's in how they're built. AI is disrupting the work of data engineering itself by automating many routine and complex tasks. This shift doesn't replace data engineers; it augments them, allowing them to focus on higher-value work while AI handles repetitive tasks. Code generation represents the most visible change. Data transformations traditionally required writing SQL or Python to select data from sources and reshape it for specific business use cases. AI can now generate these statements from natural language descriptions, handling simple queries and complex joins with equal facility. This assistance benefits junior engineers learning the craft and senior engineers tackling complex queries without wasting time on syntax peculiarities. Testing has always been essential but often gets shortchanged when deadlines loom. Everyone knows they should write comprehensive tests, but the overhead involved means testing sometimes gets left out. AI can generate basic tests for new or revised data models, eliminating much of this upfront work. That reduces psychological barriers to creating adequate test coverage and frees engineers to focus on refining tests that bring true value to data quality. Documentation suffers from the same problem as testing: everyone knows it's important, but it's time-consuming to create and maintain. AI can generate descriptions for tables and fields based on their names, context, and similar assets in the project. When you have hundreds of fields to document, this automation provides an initial draft that engineers can check into source control and gradually improve over time. Good documentation makes data more discoverable and usable while building confidence in its validity and accuracy. The [dbt Copilot ](https://www.getdbt.com/product/dbt-copilot)integrates with every step of the data engineering workflow, using your own data (its relationships, metadata, and lineage) to automate routine tasks and implement essential practices like testing and documentation. Besides generating artifacts for data pipelines, dbt Copilot can enforce code consistency using custom style guides, ensuring that AI-generated code follows your organization's standards. ## Architectural implications AI's impact on data pipelines extends beyond individual tasks to fundamental architectural decisions. The importance of frameworks and standards increases dramatically in an AI-centric world. While AI can theoretically write code in any language or style, heterogeneous codebases become intractable for both humans and AI systems. Code bases that are concise, homogeneous, and use well-documented standards are far more comprehensible to AI systems. This reality makes [frameworks like dbt](https://www.getdbt.com/product/dbt) even more valuable. AI systems know how to build reliable dbt pipelines because the framework is well-documented and widely used, with extensive examples in the training data. Standardized frameworks also emit well-understood error messages, which improves AI's ability to diagnose and fix issues. The promise of a truly consistent codebase becomes achievable because AI adapts infinitely to whatever standards you establish, with no learning curve. Observability requirements also intensify. Modern data stacks are complex and fragmented, making it difficult to gain visibility across the entire data landscape. Without observability, teams cannot detect anomalies, trace root causes, assess schema changes, or ensure reliable data for AI applications. The consequences of poor observability compound in AI contexts because errors propagate silently through models and into business decisions. Scalability challenges that were manageable with traditional pipelines become critical with AI workloads. Large monolithic scripts are difficult to debug and maintain. Pipeline bottlenecks delay insights. Manual processes don't scale with business growth. AI pipelines require modular, version-controlled transformations that make it easier to isolate and fix errors. Incremental processing (transforming only new or updated data) reduces costs and improves efficiency. State-aware orchestration optimizes workflows by running models only when upstream data changes, eliminating redundant executions. ## Best practices for AI-ready pipelines Building pipelines that reliably support AI requires implementing several foundational practices. Eliminating "garbage in, garbage out" starts with automated validation tools that detect anomalies, missing values, and inconsistencies before data reaches AI models. [dbt's built-in testing](https://docs.getdbt.com/docs/build/data-tests) and observability ensure data integrity throughout the pipeline. Data transformation must clean, structure, and optimize data for AI workflows, with dbt automating these transformations to ensure consistency and reproducibility. Automation should extend to every possible aspect of pipeline operation. Manual processes slow down AI pipelines and introduce errors. dbt enables pipeline automation from development to production, with the [dbt Fusion engine](https://www.getdbt.com/product/fusion ) dramatically accelerating deployment and feature engineering through 30X faster SQL parsing speeds. Automated data quality checks catch errors before they impact AI models. Built-in observability helps maintain data quality and reliability at scale. Cloud services provide the elastic scale, cost efficiency, and reliability that AI pipelines require. dbt offers cloud-native scalability with platforms like Snowflake, BigQuery, and Databricks, enabling dynamic scaling that ensures AI models receive optimized, high-performance data. The cloud-hosted dbt platform manages deployments, schedules jobs, and integrates with CI/CD tools without manual maintenance. Leveraging AI to build AI pipelines creates a virtuous cycle. AI tools in dbt streamline workflows by automating repetitive tasks and accelerating development. dbt Copilot acts as an AI assistant that generates SQL queries, documentation, tests, metrics, and semantic models using natural language prompts. dbt Canvas provides AI-powered visual editing that accelerates model development in a governed environment. dbt Insights enables analysts to explore and analyze data efficiently using AI-powered queries that align with dbt's metadata and governance framework. Security cannot be an afterthought. AI pipelines must be secure, compliant, and auditable to protect sensitive data. dbt supports role-based access permissions, ensuring only authorized users can modify transformations. Lineage tracking and built-in data governance with support for tagging PII/PHI and enforcement of data policies help businesses maintain transparency and regulatory compliance. ## The changing role of data engineers These technological shifts are reshaping the data engineering profession itself. Many tasks that data engineers spend time on today (authoring transformation code, writing tests and documentation, defining metrics, monitoring production jobs, and resolving incidents) are becoming heavily AI-enabled. The efficiency gains will be substantial, potentially reducing time spent on these activities by 20%, 50%, or more. This doesn't make data engineers obsolete. It pushes them in three directions: toward the business domain, toward automation, or toward the underlying data platform. Data platform engineers focus on the infrastructure that pipelines are built on (performance, quality, governance, and uptime). Automation engineers sit alongside data teams and build business automations around data insights, turning insights into action. Domain-focused data engineers act as enablement and support for the insight-generation process, owning datasets and liaising with stakeholders. The value data engineers provide to businesses won't diminish. But the way the job is done will change fundamentally. Engineers will spend less time on repetitive coding and more time on strategic work that drives business outcomes. They'll have more work to do than ever, but it will be higher-leverage work that commands greater recognition and compensation. ## Looking forward The transformation of data pipelines through AI is already underway. The foundational technologies (reasoning models, chain of thought, inference-time compute, agentic workflows) are here. Open frameworks like dbt have become widely deployed, making it possible to create framework-specific AI tooling that delivers immediate value. The commercial incentive to innovate in this space is high, and attention from companies of all sizes is intense. For data engineering leaders, the imperative is clear: build pipelines with AI requirements in mind from the start. Adopt frameworks and standards that enable AI assistance. Implement observability and governance that scale with AI workloads. Invest in automation that compounds over time. And prepare your teams for a future where AI augments their capabilities and elevates their impact. The data pipelines of 2028 will look fundamentally different from those of 2024. Organizations that embrace this transformation will build more reliable, more scalable, and more valuable data infrastructure. Those that resist will find themselves struggling to meet the demands that AI applications place on data systems. The choice isn't whether AI will change data pipelines; it's whether you'll lead that change or be forced to catch up. *** **Learn more about building AI-ready data pipelines:** - [Understanding data pipelines](https://www.getdbt.com/blog/understanding-data-pipelines) - [AI data pipelines: Critical components and best practices](https://www.getdbt.com/blog/ai-data-pipelines-critical-components) - [Understanding AI data engineering](https://www.getdbt.com/blog/understanding-ai-data-engineering) - [dbt Cloud Semantic Layer](https://www.getdbt.com/product/semantic-layer) - [Data transformation with dbt](https://www.getdbt.com/analytics-engineering/transformation) ## AI data pipelines FAQs **What is an AI data pipeline?** An AI data pipeline is a modern data infrastructure designed to meet the specific requirements of AI applications. Unlike traditional pipelines that focused on batch processing and static reporting, AI pipelines handle real-time or near-real-time data ingestion, provide continuous data flow rather than scheduled updates, and support automated model retraining as new data arrives. They must deliver trustworthy, timely data at scale to power applications like customer service chatbots, fraud detection systems, code generation tools, and autonomous systems. **How do AI pipelines differ from traditional data pipelines?** **How does generative AI enable self-updating ETL pipelines that automatically adapt to schema changes?** Generative AI transforms pipeline development by automating code generation, testing, and documentation. AI can generate SQL transformations from natural language descriptions, automatically create tests for data models, and produce documentation for tables and fields based on context. AI tools can enforce code consistency using custom style guides and diagnose issues through well-understood error messages. This automation allows pipelines to adapt more quickly to changes, as AI can regenerate code and tests when schemas evolve, reducing the manual overhead traditionally required to maintain pipelines through structural changes. --- --- title: "Write once, analyze anywhere: Omni + the dbt Semantic Layer" description: "Unify metrics across BI and AI with Omni + the dbt Semantic Layer. Define once, query everywhere, eliminate semantic drift." url: "https://www.getdbt.com/blog/omni-dbt-semantic-layer" date: "2026-02-27" authors: ["Roxi Pourzand"] categories: ["Learn"] --- # Write once, analyze anywhere: Omni + the dbt Semantic Layer If you've ever spent an afternoon debugging a dashboard only to realize the "Active Users" definition in your BI tool doesn't match the one in your data warehouse, you know the pain of semantic drift. And now, as more teams rely on AI-powered tools to explore and summarize data, that inconsistency only compounds—AI can only be as trustworthy as the metrics it’s built on. It’s the exact problem the dbt Semantic Layer was built to solve: define your metrics centrally, in code, and query them consistently across your entire stack. Today, we're thrilled to celebrate a massive milestone in that mission: building on top of its existing dbt integration,** Omni has officially launched its first-class integration with the dbt Semantic Layer.** This is a big deal for analytics engineering teams because it allows you to reuse the logic you already have - wherever it is, and saves you a lot of repetitive work and debugging. [Omni](https://omni.co) has always shared our philosophy of bringing software engineering best practices to BI, and this native integration takes that alignment to the next level. For example, our team might define a core metric like ARR once in dbt, then explore it instantly in Omni, whether that’s in a dashboard, an ad hoc query, or an AI-driven workflow that summarizes performance trends for leadership. We use Omni internally at dbt Labs to explore our own business metrics, including AI-powered analysis, so we’ve seen firsthand how valuable it is when governed definitions are available everywhere teams ask questions. Here is a look at how it works, why this integration will save you time and headaches for all your traditional workflows, and how it will help your team as it accelerates AI adoption. ## How it works: Your dbt metrics, native to Omni On top of Omni's existing dbt [integration](https://docs.omni.co/integrations/dbt), this new [release](https://docs.omni.co/integrations/dbt/semantic-layer) deeply understands your semantic definitions. By pointing Omni at your dbt semantic layer project, your central definitions are automatically mapped directly into the Omni data model. Here’s what that looks like in practice: - **Metrics:** Your Simple, Ratio, and Derived metrics in dbt automatically map to corresponding views and measures in Omni. No recreating the math. - **Dimensions and Entities:** dbt dimensions map directly to Omni dimensions, while your dbt entities map to Omni relationships, automatically building the correct join logic behind the scenes. - **Zero context switching:** The descriptions, labels, and metadata you painstakingly curated in your dbt .yml files appear right in the Omni UI. When business users explore data, they have full context without ever leaving their workflow. To get started, you just need to enable the dbt semantic layer integration in your Omni connection. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/6e8d737e4d01fda4003def81a089875c6424bfba-1200x773.png) ## A massive step for open semantics (and the OSI spec) At Coalesce 2025, we talked a lot about the future of interoperable semantics. Recently, we published the first version of the **Open Semantic Interchange (OSI) [specification](https://github.com/open-semantic-interchange/OSI)** alongside partners like Snowflake, Databricks, and Salesforce. Our shared goal with [OSI](https://open-semantic-interchange.org/) is to create a vendor-neutral, extensible model for representing metrics, dimensions, and relationships so they can be interpreted consistently across any tool or AI application. No more vendor lock-in; no more recreating business logic in five different platforms. Omni's native integration is exactly the kind of workflow the OSI vision is built to enable. It proves that when you give teams more flexibility, they win. You author your business logic once in dbt, and your downstream tools—like Omni—simply _know_ how to read it, safely and consistently. It’s a beautifully DRY approach to analytics. **** ## Try it out We’re excited to see the Omni team bring this integration to life. We’ve already experienced the benefits of Omni internally, and we know analytics engineers are going to love the seamless experience of governing metrics in dbt and exploring them in a best-in-class BI platform. Ready to stop rewriting your metric definitions? Check out the [Omni documentation](https://docs.omni.co/integrations/dbt/semantic-layer) to get started with the dbt Semantic Layer integration today. --- --- title: "How Zscaler cut PR review time by 90% using dbt context and multi-agent AI (OpenAI)" description: "Zscaler built an AI-powered, multi-agent PR review system that uses dbt’s structured context." url: "https://www.getdbt.com/blog/how-zscaler-cut-pr-review-time-dbt-context-multi-agent-ai" date: "2026-02-25" authors: ["Hrishi Kulkarni", "Chakshu Mehta"] categories: ["Product"] --- # How Zscaler cut PR review time by 90% using dbt context and multi-agent AI (OpenAI) Zscaler built an **AI-powered, multi-agent PR review system** that uses **dbt’s structured context (metadata + lineage + CI signals)** to automate governance at scale. The result: **90% less reviewer time** and a projected **2,100 engineering hours saved annually**. Zscaler is a leading cloud-based cybersecurity company. A pioneer of [“Zero Trust”](https://www.zscaler.com/resources/security-terms-glossary/what-is-zero-trust) security architecture, protecting thousands of organizations from cyber attacks and data loss. As the company expanded, so did its enterprise data platform (built on Snowflake, dbt, and Matillion). One part of the scale hit hardest: pull request reviews. Governance was essential, but manual review became the bottleneck. We'll explore how Zscaler's data team built a multi-agent PR review system they internally called PRISM (PR Review Intelligence System Mentor), powered by OpenAI and MCP tools (dbt, GitHub, and Snowflake). Their multi-agent PR system was built to turn governance into an automated agent workflow,** **transforming their pull request (PR) process and reducing reviewer time by 90%. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a504f49d7509ebf62201407f84905d6831286e41-512x288.png) ## The double-edged sword of self-service Like many companies, Zscaler started with a centralized model. Every data request flowed through a single data team to ensure consistency and quality control. But as Zscaler grew, the data team struggled to keep up with data requests. Simple data pulls took weeks — a frustrating experience that undermined trust for stakeholders. To help the team scale, Zscaler transitioned to a self-service analytics model with dbt. The central data team became a center of enablement focused on building foundational data layers. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/e977787d97be473e93a50972f4a4b20ab55af14b-512x286.png) Initially, the shift significantly increased data velocity. Business teams were empowered to create their own transformations and took greater ownership of their analytics. But self-service created a new problem: **governance didn’t scale.** Two forces collided: - **Review (pull requests) volume exploded.** As the number of contributors and dbt models grew, the data team became inundated with 900-1,000 PR reviews every quarter, each requiring careful evaluation, which became overwhelming. - **Review complexity grew.** Every manual review was time-consuming and required a deep understanding of pipelines that the reviewers hadn’t built: were freshness and data quality tests defined? Did models include proper documentation? Would changes break downstream dashboards? The data team found itself constantly context-switching and devoting hours to educating contributors on best practices. “We thought we had built a Self-Service Paradise. But enabling self-service can be a double-edged sword,” reflects Rahan Raman, Head of Enterprise Data Platform at Zscaler. “It turned out we had turned the data team into a help desk for peer reviews.” Self-service had solved the velocity problem. But without a sustainable way to enforce standards, governance had become the new bottleneck. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/05bc563978a200cb89dd3bc176cc96a71b6e5e7e-512x287.png) ## Building a multi-agent PR review system that run on dbt’s structured context To built their PR agent, Zscaler used a LangGraph-based multi-agent orchestrator to automate code reviews for dbt models. Unlike generic coding agents, their PR agent works because it’s a domain-aware reviewer that understands Zscaler’s warehouse, dbt project structure, dependencies, and standards, because dbt exposes that context. "AI can help automate governance, reduce review burden, and educate contributors, but without real context, it’s just noise,” says Rishi Varahagiri, Senior Data Engineer at Zscaler. “When AI has the right dbt context (lineage, CI performance metrics, and validated compilation), it can give targeted, meaningful PR review feedback instead of generic suggestions.” ### The context layer: what dbt gives the agents to reason with The system gathers dbt structured context by pulling from four key sources: 1. **dbt Discovery APIs**, which expose downstream lineage for every model in a PR and establishes dependency context (what’s upstream, what’s downstream, what could break). 2. **dbt CI jobs**, which automatically validate changes on every PR and provide performance signals like execution time, bytes scanned, and partitions scanned—turning CI into a baseline for optimization decisions. 3. **dbt Cloud APIs**, which compile and validate AI-generated suggestions before they’re surfaced to developers—reducing the risk of “hallucinated” changes. 4. **Snowflake Query Insights**, (query plans and operator statistics) show how queries execute and where time and resources are spent—ground truth for performance tuning. ## The end-to-end workflow: specialized agents that review, enforce, and improve PRs With context in hand, now, let’s look at how it comes together inside their multi-agent system. Everything starts with the pull request. When a developer opens a PR for a dbt model change, the **context collection process** begins immediately. Their agent pulls: - The file diffs (what changed) - The dbt CI job execution results (did it run end-to-end? What were the metrics?) - Query execution insights from Snowflake (what does the plan show? where are costs concentrated?) If the dbt CI job fails, the process **stops there**. No “AI review” layered on top of a failing build—CI is the gate. Once that context is assembled, the **LangGraph-based multi-agent orchestrator** takes over. Each agent is specialized and performs a specific task: - **Linter agent:** Automatically checks structural best practices: naming conventions, SQL & Python formatting, folder structure. - **Governor agent:** Enforces governance requirements: documentation, tags, metadata policy, owner, groups, and checks for critical patterns like missing freshness configuration or incomplete documentation. - **Impact analyzer:** Maps downstream lineage and shares exactly what’s impacted—including dependent models and dashboards. - **Optimizer-tester-healer trio:** refactors long-running queries _only when necessary_ and offers performance improvements based on CI and warehouse execution signals. - **Test Reviewer Agent** ensures before-and-after results match when optimized code is proposed. - **Self-Healing Agent** fixes code errors/bugs in generated or refactored SQL and hands it back to validation. Finally, their agent logs every action into an audit table, capturing recommendations, outcomes, and workflow behavior for observability and adoption metrics. That log also acts as “memory,” so when a PR gets updated multiple times, their agent can understand what it already recommended and when, instead of repeating itself. Then everything shows up where developers already work: as **GitHub PR comments**. **Zscaler’s multi-agent workflow doesn’t just “review.” It _acts_.** It can propose fixes, validate them, and present changes that developers can accept quickly. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/544a8dce94133be042154f22981da4634ca24806-512x288.png) #### What developers experience PR reviews move from slow, manual, and inconsistent to fast, contextual, and mostly automated—without leaving GitHub. When a developer opens a PR for a dbt model change, dbt CI runs and the multi-agent review workflow kicks off. If CI fails, it stops. If it passes, the system posts review feedback directly as GitHub comments—so developers get guidance immediately, in the same place they already work. **What changes day-to-day:** - **Human-in-the-loop adoption: **When the workflow proposes a refactor, developers can merge the optimized code into their branch by simply commenting “Accept.” Less back-and-forth, less waiting, fewer long review cycles. - **Automated reviews with guardrails: **Their multi agent system auto-approves PRs that follow best practices, pass CI checks, and don't need optimization. It falls back to targeted comments when logic is too complex—so it doesn’t produce risky code - **Impact visibility: **It surfaces downstream lineage and exposure context so teams can see what’s affected—models, dashboards, and key metrics—before changes ship. In the “happy path,” the workflow can deliver meaningful performance gains. “You can see a 30% improvement in runtime—and this is fully vetted code with before-and-after results of the optimization,” says Rishi Varahagiri. “When a developer raises a pull request, two checks kick off immediately: the dbt CI check and the agent check. Once they complete, it automatically posts review comments on the PR, and developers can merge the optimized code by simply commenting ‘Accept’—so there’s less back-and-forth and no long review cycles.” The balance of automation where safe, human in the loop where uncertain, is what makes agentic automation scalable. ### Quantifiable time savings, faster reviews, and higher data quality Zscalers multi-agent workflow has proven itself as a strategic enabler and have seen measurable efficiency gains: - **90% reduction in reviewer time:** Decreased reviewer time by 90%, freeing data engineers to focus on higher-impact work. - **High volume PR handling.** In a single quarter, the system reviewed 956 PRs. - **2,100 engineering hours saved annually.** Zscaler projects a savings of 2,100 hours of annual time savings, which is the equivalent of one full-time engineer. “With our multi-agent PR reviewer and dbt context, we reduced review time by up to 90%. We handle about 900–1,000 PRs per quarter. We actually had 956 PRs reviewed by the agent last quarter,” says Ra, “And even if you assume 30 minutes per PR, that projects to about 2,100 hours saved per year, basically one full-time engineer. AI isn’t just a tool anymore, it’s a collaborator that turns bottlenecks into opportunities.” Perhaps most importantly, the central data team is no longer a bottleneck. Contributors get faster feedback, governance is consistently enforced, and overall data quality improves as the organization ships faster. For data teams wanting to build agentic automation that can truly operate inside your real system but facing similar governance challenges, feed them dbt’s structured context so they can make trustworthy decisions. When agents can see lineage, tests, metadata, CI results, and warehouse behavior, governance becomes automatable—and self-service becomes sustainable. Watch Zscaler’s full session to see how a context-driven, multi-agent PR reviewer reduces reviewer back-and-forth and automates governance in the PR itself. [Watch video](https://www.youtube.com/watch?v=VUxSaD7k6d4) Explore how dbt turns structured context into AI-powered workflows, whether you’re accelerating development with dbt Copilot and Agents or building your own agents with the dbt MCP Server, learn more [here](https://www.getdbt.com/product/ai) or [book a demo](https://www.getdbt.com/contact) today. --- --- title: "Data ins and outs for 2026: what data teams are keeping, cutting, and reconsidering" description: "Data ins and outs for 2026: simplify workflows, avoid over-engineering, and rethink how data teams use AI." url: "https://www.getdbt.com/blog/data-ins-and-outs-for-2026" date: "2026-02-20" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Data ins and outs for 2026: what data teams are keeping, cutting, and reconsidering In the latest episode of _The View on Data_, hosts Jerrie Kenney, Erica “Ric” Louie, and Faith McKenna do a New Year reset. But instead of setting perfect goals they’ll forget by February, they share their “ins and outs” for 2026: what they want more of in their work, what they’re done tolerating, and what they’re trying to do differently as data teams keep scaling (and as AI keeps… being AI). 🎧 Listen & subscribe: [Spotify](https://open.spotify.com/show/0Itxg3BoF6zaQdORfVuY6I?si=6133f88257624765) | [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-view-on-data/id1835755886) | [Amazon Music](https://music.amazon.com/podcasts/84c47dbe-adf7-4d33-a371-a1fe55255495/the-view-on-data) | [YouTube](https://www.youtube.com/playlist?list=PL0QYlrC86xQk2WktL3vbdB1FcWlpay7VQ) [Watch video](https://youtu.be/AVB08YC1Nh4) ## In: simplifying your work without dumbing it down Jerrie opened with a theme that’s easy to say and harder to do: simplify. Not because simplicity is trendy, but because complexity has a way of multiplying when you’re busy, growing, or adopting new tools. Over time, “just ship it” turns into a maze: more models, more layers, more documentation, more handoffs, more ways for someone new to get lost. The group’s point wasn’t “do less.” It was closer to “make it easier to understand what’s happening, and why.” In practice, that looked like: - Trimming unnecessary bloat in projects (especially long chains of intermediate models that nobody can explain) - Writing documentation that answers the questions people actually ask, instead of documenting everything just to prove you did it - Cleaning up metadata so future-you (and future teammates) can navigate without reverse engineering your entire stack - Choosing one shared way of working across teams so collaboration doesn’t feel like translating between two systems This ties back to the reality that most data teams are judged on: trust, speed, and cost. When your work is understandable, it’s easier to trust. When your processes are consistent, it’s easier to ship. When your stack is less chaotic, you spend less time rebuilding the same logic in five places. ## In: frameworks that stop you from solving the same problem 12 times Ric brought up something that comes up in almost every scaling data org: you can fix individual problems forever, or you can build a framework that prevents the problems from repeating. A broken dashboard is a problem. A revenue model no one trusts is a problem. A messy intake process that sends every request into a black hole is also a problem. But if those issues keep showing up, the real problem is usually the workflow around them. Ric talked about the friction that happens when teams work closely together but operate in completely different modes (like Kanban in one corner, sprints in another). Sometimes that split works. Sometimes it just guarantees a constant series of “wait, who owns this?” conversations. The “in” here was alignment: fewer parallel processes, clearer paths from request to production, and shared habits that help teams move faster without creating governance nightmares later. It’s very ADLC in spirit: treat data like software, and build a workflow that supports quality over time, not just speed today. ## In: making work fun again (seriously) Faith shared an idea she’s been thinking about a lot as someone who teaches technical topics: learning is only engaging when it feels alive. She jokingly framed it as “learning is boring unless it’s gossip,” but the underlying point was real. People remember what feels interesting, human, and relevant. They forget what sounds like a monotone tutorial they’re forcing themselves through. They also talked about how “fun” can show up in normal work, not just in training. PR reviews that feel like humans wrote them. A little personality in collaboration. Team culture that doesn’t treat every interaction like a compliance audit. If you’re building enterprise software and the whole experience feels joyless, it’s harder to stay curious. And if curiosity dies, you get stagnation, burnout, and a team that stops experimenting. ## In: using AI to get unstuck, not to outsource your thinking AI came up in a way that felt grounded: it’s helpful, and it also makes it easier to create nonsense at scale. Jerrie shared a workflow that a lot of people will recognize. When you can’t turn your thoughts into a clear plan, it’s often easier to talk it out than it is to write it. So she’ll “narrate” the problem to ChatGPT and ask it to organize the mess into something usable. That’s the good version of AI support: getting momentum when you’re stuck, or turning scattered thoughts into a structured first draft. The group was also clear that this only works if you keep ownership of what you’re saying. If you let AI generate documentation (or solutions) that you don’t actually understand, you’re not saving time. You’re pushing complexity downstream to whoever has to debug it later. ## Out: over-engineering (especially when AI makes it easier) When they moved to “outs,” the biggest one was the obvious opposite of simplify: over-engineering. They called out the specific flavor of over-engineering that’s getting worse right now: AI-generated complexity. You ask for a solution to a simple problem, and suddenly you have a query that looks like it was designed to win an argument on the internet instead of run reliably in production. The underlying warning was pretty simple: - If you can’t explain what it does, you shouldn’t ship it - If it takes five layers to get to something that should be straightforward, there’s probably a clearer approach - If you’re building clever abstractions mostly because they’re clever, you’re creating future maintenance work for someone (possibly you) It’s also a cost issue. More complexity often means more compute, more rebuilds, and more time spent on rework instead of delivering useful data products. ## Out: “That’s too technical for me” Faith’s “out” was personal and honestly pretty relatable: she’s done saying “that’s too technical for me.” Her point wasn’t that everyone needs to become a platform engineer overnight. It was that this phrase can turn into a reflex that stops you from even trying. And in data, almost everything feels technical until you’ve had a few reps with it. It also turned into a broader conversation about culture: some corners of the data internet make it easy to feel behind, especially when the loudest voices are early adopters who live for the newest tooling. The counterpoint the group offered was healthier: - You can learn it, even if it takes longer - You’re allowed to ask questions without knowing the perfect term - “I don’t know yet” is a normal state, not a personal failure ## Episode takeaways A few things the episode kept circling back to: - Simplification is work, but it pays you back every time someone new touches your project. - If your team keeps tripping over the same problems, you probably need a shared framework more than you need another patch. - AI can help you think and draft faster, but it can also help you create bad complexity faster. - You don’t have to be the most technical person in the room to be effective. You do have to stay curious and keep learning. - Saying “I don’t know” is often the most competent thing you can say, as long as you follow it with “and here’s how I’ll find out.” As they wrapped, the hosts challenged listeners to write their own work ins and outs for 2026 and share them. What are you keeping because it makes you better, calmer, and faster? What are you dropping because it creates noise, rework, or insecurity? --- --- title: "How AI accelerates and improves data modeling" description: "Learn how AI accelerates SQL modeling, testing, documentation, and performance optimization in modern data teams." url: "https://www.getdbt.com/blog/ai-data-modeling" date: "2026-02-12" authors: ["Joey Gault"] categories: ["Pulse"] --- # How AI accelerates and improves data modeling ## The data modeling bottleneck Raw data is never ready for analytics or AI workloads out of the box. It requires [transformation](https://www.getdbt.com/blog/data-transformation), cleaning, testing, and documentation to shape it into formats suitable for driving business decisions. This work, the heart of data engineering, frequently becomes a chokepoint for creating new production-ready datasets. The pressure on data teams has only intensified with the emergence of generative AI use cases, which demand large volumes of high-quality data to produce useful results. Data engineers struggle to meet this demand using traditional methods alone. The good news is that generative AI itself can help address this challenge. ## What AI data engineering means for modeling AI data engineering leverages large language models (LLMs) trained on massive amounts of data, combining them with information from existing pipelines (database schemas, data models, tests, documentation, and metrics) to produce first drafts of artifacts based on natural language descriptions. Engineers can then refine, test, and deploy these outputs. This approach cuts workload by automating the creation of assets that constitute a data pipeline. Rather than replacing data engineers, AI augments them. The results mirror what's happening in software engineering more broadly, where developers using AI assistance report significantly faster task completion and reduced time-to-deployment. ## Accelerating the modeling workflow ### Generating transformation code Data transformations require selecting data from multiple sources and reshaping it into formats suitable for specific business use cases. This involves writing SQL or Python code that integrates data from multiple tables while correcting underlying issues like malformed fields or missing values. AI can generate the base SQL statements data engineers need for their transformations, from simple queries to complex regular expressions and bulk edits on existing code. This assistance proves particularly valuable for junior engineers, though even senior engineers benefit when performing complex queries. By issuing instructions in natural language and allowing AI to write the code, engineers avoid wasting time looking up SQL syntax peculiarities. When working with [dbt](https://www.getdbt.com/product/what-is-dbt), engineers can [use AI to create](https://www.getdbt.com/product/ai) transformation files that implement broader data modeling designs. These individual files serve as building blocks for the complete architecture, and AI can accelerate their creation while maintaining consistency with established patterns. ### Automating documentation [Documentation](https://docs.getdbt.com/docs/build/documentation) is essential for data discoverability and usability, yet it often gets short shrift when deadlines loom. Good documentation tells downstream consumers where data comes from, how to use it, and how calculations were derived. This increases data confidence and makes datasets more accessible. If you use dbt for transformations, you already get automated [data lineage](https://www.getdbt.com/product/dbt-catalog) generation. AI data engineering goes further, creating descriptions for tables and fields based on their names, context, and similar assets in your projects. When you have hundreds of fields to document, this capability becomes extremely valuable. AI-generated documentation provides an initial cut of descriptions for all tables and fields. Engineers can check these into source control, where they and team members can gradually improve the documentation over time. This approach eliminates the psychological barrier of starting from a blank page while ensuring documentation exists from day one. ### Building comprehensive tests Code doesn't always work as intended, and it may encounter issues when dealing with edge cases like values outside expected ranges or malformed data. [Building tests](https://docs.getdbt.com/docs/build/data-tests) for data transformation code provides confidence that transformations work as expected under various circumstances. These tests can take several forms: unit tests that validate small portions of data model logic, data tests that ensure generated data is sound, and integration tests that verify the entire project end-to-end. Testing is one area that often gets deprioritized during crunch time, despite everyone knowing its importance. [AI data engineering](https://www.getdbt.com/blog/ai-data-engineering) can generate basic tests for new or revised data models, eliminating much of the upfront coding overhead. This reduces psychological barriers to creating adequate test coverage and frees engineers to focus on refining tests in ways that bring true value to dataset quality. Rather than spending time writing boilerplate test code, engineers can concentrate on edge cases and business logic validation that requires domain expertise. ### Defining metrics and semantic models AI can assist in defining consistent metrics that provide global availability across your organization for key values. A [semantic layer](https://www.getdbt.com/product/semantic-layer) framework defines common representations of data using standard business terminology, translating SQL or Python into common business language. This ensures consistency and democratizes data access by making key values available to all stakeholders. Defining a semantic layer requires tools for creating and exposing new metrics globally. With [AI data engineering](https://www.getdbt.com/blog/traditional-to-ai-data-engineering), you can generate these models automatically and even ask the AI engine to recommend useful metrics based on your data transformation definitions. This capability helps teams identify valuable metrics they might not have considered and ensures metrics align with the underlying data structures. ## Optimizing model design and performance ### Converting between platforms In data migration projects, adapting functions from one data warehouse to another (such as moving from [Redshift](https://www.getdbt.com/data-platforms/redshift) to [Snowflake](https://www.getdbt.com/data-platforms/snowflake) or [BigQuery](https://www.getdbt.com/data-platforms/bigquery) to [Databricks](https://www.getdbt.com/data-platforms/databricks)) can be time-consuming. Each platform has its own syntax, operators, and native functions requiring manual adjustments for compatibility. AI can automate much of this process by converting code between platforms while preserving the original logic. This speeds up migrations and helps maintain consistency across environments. For data modeling work, this means existing models can be adapted to new platforms without complete rewrites, reducing the risk of introducing errors during migration. ### Improving performance and readability Beyond syntax conversion, AI can suggest improvements for performance and readability in data models. It can refactor subqueries into c[ommon table expressions (CTEs)](https://docs.getdbt.com/best-practices/how-we-style/2-how-we-style-our-sql#functional-ctes), reduce duplicated logic, and simplify joins. These optimizations reduce execution time, improve maintainability, and make models easier to understand without changing underlying logic. For larger projects with complex transformations, small optimizations scale into meaningful gains in both performance and team collaboration. When models are easier to read and understand, new team members can contribute more quickly, and debugging becomes faster. Performance improvements translate directly to lower compute costs and faster query results for end users. ### Designing dimensional models Designing scalable data models is challenging, especially when working with multiple tables. AI can suggest optimized structures for fact and dimension tables, helping build more efficient and scalable models. Based on source system diagrams or descriptions, AI can propose dimensional model designs that align with best practices. Once the structure is defined, AI can create mapping tables between transactional and dimensional models, generate corresponding dbt models using transactional tables as sources, and suggest naming conventions and folder structures to organize projects. This assistance is particularly valuable when establishing new subject areas or onboarding team members who are less familiar with dimensional modeling techniques. ## Navigating the challenges AI-generated code isn't always accurate. Sometimes it's incorrect or produces outputs that don't align with business requirements. Studies have shown that AI assistance can help enforce best practices in some areas while potentially encouraging shortcuts in others, such as security considerations. The key is treating AI as an assistant within a mature analytics workflow process, such as the [Analytics Development Lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle). Processes like this ensure alignment with business objectives and establish checkpoints to verify code quality and conformance to best practices. Data teams should always validate AI output carefully, reviewing generated code and documentation for accuracy. Testing before deployment is essential. Running tests on any generated or modified code prevents issues from reaching production. Teams should never input confidential data, credentials, or proprietary business logic into AI tools, maintaining appropriate security boundaries. Using AI output as a learning opportunity helps teams understand new techniques and patterns rather than simply copying code without comprehension. The quality of AI output often depends on prompt clarity and specificity, so teams should iterate and refine their prompts to improve results. ## Building with AI-assisted modeling dbt provides a consistent approach to developing, testing, deploying, and documenting data through SQL and YAML code, with rigorous testing and verification processes. When integrated with AI capabilities through tools like [dbt Copilot](https://www.getdbt.com/product/dbt-copilot), teams can automate routine tasks while maintaining the governance and quality standards that production data products require. AI assistance can enforce code consistency using custom style guides, ensuring that generated models align with team conventions. This consistency becomes increasingly important as data warehouses grow and more team members contribute to the codebase. Clear patterns make projects immediately comprehensible to anyone navigating the code. For data engineering leaders, the strategic question isn't whether to adopt AI assistance for data modeling, but how to implement it effectively. The teams that succeed will be those that establish clear conventions early, maintain discipline as systems grow, and treat AI as a tool that amplifies human expertise rather than replaces it. The goal remains unchanged: creating data warehouses that serve as reliable foundations for analytics, where business users find data intuitive to work with, where data teams can efficiently build and maintain transformations, and where architecture can evolve alongside changing business needs. AI assistance makes achieving this goal faster and more sustainable, but it still requires thoughtful data modeling as a first-class concern. *** _Learn more about [data modeling fundamentals](https://www.getdbt.com/blog/understanding-data-modeling) and explore how [AI is transforming data engineering workflows](https://www.getdbt.com/blog/understanding-ai-data-engineering)._ ## Data modeling FAQs **What is a data model?** A data model is a structured representation of how data is organized, stored, and transformed to support business use cases. It involves selecting data from multiple sources and reshaping it into formats suitable for analytics and AI workloads through SQL or Python code that integrates data from multiple tables while correcting issues like malformed fields or missing values. **Why is data modeling important for AI?** Data modeling is critical for AI because raw data is never ready for AI workloads out of the box. It requires transformation, cleaning, testing, and documentation to be useful. Generative AI use cases in particular demand large volumes of high-quality data to produce useful results. Proper data modeling ensures datasets are reliable, well-documented, and structured in ways that make them accessible and trustworthy for driving AI-powered business decisions. **How can AI-driven tools reduce or eliminate the custom code traditionally required for customer data modeling and cleaning?** AI-driven tools can generate base SQL statements and transformation code from natural language descriptions, automating much of the manual coding work. They can create transformation files, generate documentation for tables and fields, build comprehensive tests, and even suggest optimized structures for fact and dimension tables. This reduces the time engineers spend writing boilerplate code and looking up syntax, allowing them to focus on refining outputs and addressing complex business logic that requires domain expertise. --- --- title: "How dbt Labs reduced dbt-related compute costs by 64% with Fusion and state-aware orchestration" description: "With intelligent orchestration and optimization, we achieved a 64% reduction in compute costs and simplified our job architecture." url: "https://www.getdbt.com/blog/dbt-compute-cost-reduction-fusion-state-aware-orchestration" date: "2026-02-12" authors: ["Brandon Thomson", "Pat Kearns", "Ken Ostner"] categories: ["Product"] --- # How dbt Labs reduced dbt-related compute costs by 64% with Fusion and state-aware orchestration **** When you're running analytics for a company that builds analytics infrastructure, expectations are high. Our internal analytics project processes roughly 1,500 models across multiple jobs on various schedules, serving everything from executive dashboards to operational metrics. As our data estate grew, so did our compute costs—and we needed a solution that would let us meet our data freshness requirements without breaking the budget. The answer came from our own product roadmap: migrating to [dbt’s next-generation Fusion engine](https://www.getdbt.com/product/fusion) and implementing [State-Aware Orchestration (SAO)](https://www.getdbt.com/product/dbt-state). What we didn't anticipate was just how transformative this journey would be. Through a combination of platform migration, intelligent orchestration, and thoughtful optimization, we achieved a 64% reduction in compute costs while actually simplifying our job architecture. This is the story of how we got there—including the false starts, lessons learned, and recommendations for teams considering a similar path. ## The challenge: Balancing freshness and cost Before our migration, we were running on legacy dbt infrastructure with complex job configurations that had accumulated over years. We had multiple projects, intricate package dependencies creating technical debt, and an ever-growing compute bill. The fundamental tension was clear: our stakeholders needed fresh data to make decisions, but every model run cost money. We were rebuilding models whether they needed it or not, reprocessing unchanged views, and running transformations even when the source data hadn't been updated. We needed intelligence in our orchestration—a way for dbt to understand what _actually_ needed to run. At a high level, State-Aware Orchestration changes the fundamental question dbt asks: Traditional orchestration asks: _“Is it time to run?”_ SAO asks: _“What actually changed?”_ If nothing upstream has changed—or if a model already satisfies its freshness requirement—dbt simply reuses the existing result. No rebuild. No wasted compute. That shift sounds subtle, but at scale, it’s transformative. ## Phase 1: The migration to Fusion (May–August 2025) We kicked off our Fusion migration on May 21st, 2025. The timeline was aggressive but achievable: we completed two smaller projects by May 28th and had our entire data estate running on Fusion by July 10th. This phase wasn’t about savings yet. It was about unlocking the foundation that makes intelligent orchestration possible. The upgrade process wasn't without challenges that you’d expect when migrating to an Alpha version. This is why we at dbt Labs make it a practice to use our software first before even rolling it out to customers in a Beta stage. We encountered parse and compile issues that required "whack-a-mole" fixes, dealt with package updates cascading through dependencies, and coordinated with package maintainers to ensure compatibility. The key technical changes included moving custom configs under `config.meta`, adding `arguments` to tests due to spec changes, and relocating several config properties from top-level to config blocks. Fortunately, dbt's [Autofix tool](https://www.notion.so/24fbb38ebda78096acd8c0fd1afae8e0?pvs=21) handled most of these transformations effectively. For our largest project (Internal Analytics), we created a manual fix branch that fed improvements back into the Autofix script, then used the automated tool for the final migration. One critical decision: we maintained the ability to revert between Fusion and hosted dbt using the same configuration. This flexibility proved essential during July when we hit production stability issues. As a Fusion Canary project—essentially an advanced staging environment—we experienced bouts of instability as new versions rolled out. When issues arose, we'd roll back to the prior stable version or switch back to Mantle while the Fusion team marked the problematic version as bad and fixed it in the next release. By August, we had stabilized on Fusion and completed the migration phase. ## Phase 2: Activation and immediate wins (August–October 2025) With Fusion stable, we turned our attention to State-Aware Orchestration. We treated this as a control experiment: same jobs, same schedules, same models—just with state awareness turned on. The results were immediate: **9% cost savings** against our data warehouse with zero configuration effort. SAO delivered these savings by making intelligent decisions about what actually needed to run. It stopped rebuilding unchanged views, skipped models when sources had no new data, and reused approximately 35% of models daily. We switched our jobs to state-aware run methods, conducted smoke tests to verify behavior, and watched the compute costs drop. This was our "activation" phase—proof that the technology worked and delivered value out of the box. But we knew there was more optimization potential if we were willing to tune our approach. ## Phase 3: The optimization journey (October 2025–January 2026) Emboldened by our initial wins, we set out to optimize further. Our first approach seemed logical: map source data frequency to dashboard SLAs, building a detailed matrix of requirements and implementing directory-level configs in `dbt_project.yml`. This "right-to-left" approach failed spectacularly. We discovered we lacked documented SLAs, had unclear source freshness thresholds, and created significant mental overhead managing mixed configuration levels. Worse, our configuration was still tied to our old, complex job architecture. We pivoted to simplification. Instead of engineering an elaborate system, we aligned on business needs: The majority (90%) of our models just needed to be fresh on a daily (24-hour) basis. The remaining models had varying freshness requirements, from 1 hour to 7 days. We found that we could cover all of these refresh intervals across only 2 jobs, compared to the 10 jobs we had before. Our configuration became elegantly simple. The project-wide default specified 24-hour freshness: ```sql # dbt_project.yml ... models: +freshness: build_after: count: "{{ var('build_after_count', 1) }}" period: "{{ var('build_after_period', 'day') }}" updates_on: any ``` This single default replaced dozens of schedule-based assumptions we had previously encoded in job logic. But why variables? Originally, we thought we could get away with only one job - an hourly job. When we had the default project config set to one day, we found that the time of day at which that model runs could drift throughout the week. We quickly realized that we had lost our ability to guarantee the last complete day of data when our business users logged on in the morning. As a result, we decided on three jobs: a weekly job (for capturing long-range late-arriving events), a daily job, and an hourly job. The intention was that our daily job always runs after midnight UTC, and that it would build all of our models, with any level of freshness, **only if there was new data.** These variables allow us to overwrite the project-default freshness to ensure this is possible. Here's what our job config looks like for our daily job: ```sql dbt build --vars '{build_after_count: 0, build_after_period: minute}' ``` In the above command, we are telling dbt to “build all models that have ANY new data since they last ran.” For some of our sources, we just don’t see new data every day! Not many people register for Coalesce in January before registration has opened up, for example. Compared to running this daily job without SAO, we are still managing to build 20% fewer models, and we can guarantee last complete day data. For operational models requiring hourly updates, we added model-specific overrides directly in the model configuration. This ensures that throughout the day, models will update on the agreed upon cadence (intra day updates). ```sql # example_model.yml ... - name: example_model description: A model that updates hourly. config: freshness: build_after: count: 1 period: hour updates_on: any ``` We hit a snag with `updates_on: all`—applying it too broadly caused our hourly job to exceed one-hour runtime, creating job queuing and missed runs. We rolled back to more selective use, reserving this setting for specific incremental models where all upstream dependencies updated on a regular cadence. ## The results: 64% cost reduction and simplified operations The numbers speak for themselves. Our initial Fusion enablement delivered 9% savings. Tuned SAO configurations added another 55%. Total reduction: **64% in compute costs**. But the benefits extended beyond the bottom line. We simplified our job architecture dramatically—consolidating many disparate jobs into three main workflows (hourly, daily, weekly). Our mental model became cleaner. Our configurations aligned with actual business needs rather than technical constraints. A significant portion of our models now get reused across runs, eliminating unnecessary rebuilds. Our data remains fresh for stakeholders, but we're only paying for the transformations that truly need to happen. The biggest gains didn’t come from more rules—they came from fewer jobs and clearer intent. **** ## What's left to do Some challenges remain, particularly around seeding data for new models. Since SAO might not detect a need to use new data initially (if sources haven't hit their SLA yet), we occasionally encounter logic issues on first runs that resolve themselves subsequently. We need a better workflow for new model builds after PR merges. We're also working to make our jobs more targeted, minimizing SAO's overhead by explicitly excluding sources we know should be skipped rather than having SAO monitor everything. Even intelligent orchestration has computational costs, albeit much reduced compared to running an unchanged model. ## Lessons learned: What we'd tell other teams If you're considering a similar journey, here's what we learned: - **Start simple.** Our "activation" phase delivered 9% savings with no tuning. Don't over-engineer your initial SAO configurations—get the quick wins first, then iterate. - **Let business needs drive technical decisions.** Our elaborate right-to-left mapping failed because we led with technical constraints instead of business requirements. Two freshness tiers (daily and hourly) work better than elaborate matrices. - **Be cautious with `updates_on: all`.** This setting has power but requires thoughtful application. Monitor job runtimes to prevent queuing and missed runs. It can have unintended (or intended!) effects when using views or when multiple tables are referenced, but update at vastly different frequencies. - **Maintain rollback capability during early phases.** The ability to revert between Fusion and Mantle saved us during production instability. - **Migration is manageable.** Autofix handles most technical changes. The path isn't always linear, but it's navigable. - **Simplicity beats complexity.** Your job architecture can actually get simpler, not more complex. We went from many jobs to three main workflows and have better outcomes. The journey from 1,500 models on legacy infrastructure to a 64% cost reduction wasn't just about adopting new technology—it was about rethinking how we orchestrate data transformations. Fusion and State-Aware Orchestration gave us the tools, but the real wins came from aligning our technical implementation with business needs and having the courage to simplify. Intelligent orchestration doesn’t just save money and drive efficiency, it also provides clarity. For teams running large dbt projects, the promise is real. The path requires iteration, but the destination is worth it. --- --- title: "Maximize the business value of your data platform with dbt" description: "Learn how tools like Fusion’s state-aware orchestration deliver real business value." url: "https://www.getdbt.com/blog/dbt-business-value" date: "2026-02-11" authors: ["David Tishgart"] categories: ["Learn"] --- # Maximize the business value of your data platform with dbt Every data leader today faces a familiar paradox: as technology becomes more efficient, costs don't typically fall - they often rise. This phenomenon, known as [Jevons Paradox](https://www.npr.org/sections/planet-money/2025/02/04/g-s1-46018/ai-deepseek-economics-jevons-paradox), explains why fuel-efficient cars don't reduce driving costs (we just drive more), and why lower BI costs don’t shrink analytics costs (we just create more dashboards). We observe this pattern at dbt Labs, too. Every day across our platform, teams run nearly a million jobs that power analytics products and AI workflows, with the majority running on an hourly cadence. But most of the time, the models supporting these runs haven't changed. That means users are rebuilding thousands of models that look exactly the same the second time they run. At scale, this is inefficient, costly, and not optimized to drive the best possible outcome. We need to think about data work, not in terms of model runs or data products produced, but in terms of business value delivered and return on investment generated. **** ## Understanding the ROI equation ROI is a fairly simple equation: value divided by cost. Every decision we make about data - tools, people, pipelines - ultimately rolls up to this equation. ![ROI equation](https://cdn.sanity.io/images/wl0ndo6t/main/09d04cb2f4b46d8ed3affddf1067f4d8c6f3d399-1252x608.png) There are three levers we can pull to improve ROI. The first is obvious: **increase the value our work delivers**. That could mean faster insights, higher adoption, or more trusted data - anything that grows the value numerator. The second lever is **cost**. You might think the goal is to reduce or cut costs, but in many cases, you just can't. You're already paying for tools. So, on the cost measure, what you want to do is make your use more efficient. Increase efficiency, reduce rework, and drive more adoption of the tools you already paid for. That's how you optimize the denominator. The third lever is **improving how we measure ROI itself**. If we can't accurately define value, cost, and ROI, we really can't effectively measure ROI. Everything we do should fall into one of these three buckets. ## Building trust through quality When we talk about value, we often start with what our teams make: models, dashboards, data products. But none of that matters unless it creates something the business can actually feel. Value goes up when a decision gets made faster, when a forecast is more accurate, or when a customer experience improves. All of this is based on **trust **- trust in decision-making tools, in forecasts, and trust from users. The challenge is not to build more - it's to make what we build matter more by building trust. Before you can establish trust, you need to consistently deliver quality. The way to prevent data quality issues is the same way software engineers prevent bugs before they reach production: through a well-defined lifecycle, such as the [Analytics Development Lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle). With dbt, [it starts with testing](https://docs.getdbt.com/docs/build/data-tests). Every [dbt model](https://docs.getdbt.com/docs/build/models), every transformation is tested automatically, so you catch broken logic before it hits the business. [Continuous integration](https://docs.getdbt.com/docs/deploy/continuous-integration) takes it a step further - every change is validated in isolation before it hits production. That gives teams confidence that what they deploy is clean and stable. [Documentation](https://docs.getdbt.com/docs/build/documentation) and [data lineage](https://www.getdbt.com/blog/what-is-data-lineage) make that trust visible. People can see where the data came from, what changed, and who changed it. They stop guessing and can trace every number back to its source. [Peer review via pull requests](https://www.getdbt.com/blog/how-to-review-an-analytics-pull-request) creates shared accountability. When someone has looked at your work, asked questions, and approved it before it goes live, you can feel good about not just the models, but the downstream data products those models create. Even when we build great systems, things go wrong. The difference between a trusted data team and an untrusted one is how we communicate around those issues. In software, users expect visibility. When an app has an outage or service is degraded, you get a status page, not silence. That transparency builds credibility, even when the system is down. We should take the same approach with data. If something breaks, we need to make it visible to the people who consume data. Don't bury it in a ticket - embed it as a trust signal that says this data is healthy, or this dashboard has problems. This is where health tiles built into tools like [Tableau](https://tableau.com) become essential. When users know we're watching and we're honest about issues, they trust us more. Silence erodes confidence, but visibility builds it. ### Speed as a value driver Another side of the value equation is speed. With the new [dbt Fusion](https://www.getdbt.com/product/fusion) engine, we've completely changed the game for developer experience. Native SQL comprehension provides real-time feedback whether you're writing in [VS Code](https://docs.getdbt.com/docs/install-dbt-extension), [dbt Studio](https://docs.getdbt.com/docs/cloud/studio-ide/develop-in-studio), or [dbt Canvas](https://docs.getdbt.com/docs/cloud/canvas). You stay in flow with Intellisense, real-time error detection, and hover previews. Beyond the developer experience, there's pure speed. Fusion speeds parse time by 30x or more compared to dbt Core. As you're building trusted data, you're also working much, much faster. ## Optimizing costs through intelligent orchestration Think about how a warehouse processes data - it understands a single SQL statement at a time. It's efficient within that boundary, but it has no concept of the larger picture. Fusion broadens that understanding. It knows how queries connect across your data models and teams. That means it can avoid running redundant transformations, skip unchanged work, and orchestrate intelligently. [State-aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-about) is a concrete example. Today, when you run dbt build, everything selected will be built, even if the models haven't changed. With state-aware orchestration powered by Fusion, we know exactly what code or data has changed, and Fusion only rebuilds what's needed. Rather than rebuilding an entire DAG, we only build the models that would actually change or need to be rerun. The efficiencies extend beyond orchestration to testing. Today in dbt, we test everything at the end of every model build. With efficient testing, we only run tests that are needed. If you have a unique test that already passes upstream, and there isn't any logic or code introduced that invalidates it, we'll just reuse that same test. This speeds up run times and optimizes dbt build costs by avoiding unnecessary computation. The results are significant. We're seeing roughly 10%+ of compute saved with state-aware orchestration, another 15%+ with tuned configurations, and another 4% more with advanced configurations - upwards of 29% saved on compute with state-aware orchestration. ## Stop looking at vanity metrics! Are we measuring the right things? In many cases, the answer is probably no. It's easy to measure new dashboards built, number of tables created, or how much data we're working with. These are fun numbers to play around with, but they add no real value. These numbers don't tell you whether your business is improving, getting faster, getting smarter, or making better decisions. They don't influence the value equation. They're just easy to measure. They're called **vanity metrics**. Let's resolve to stop looking at vanity metrics and try something else. How about: - **Time to onboard** - how many hours between a new hire and their first commit? - **Time to insight** - how long from first commit to production-ready? - **Optimizing compute** - how are we using models more efficiently with our runs and tests? These are examples of how we can optimize our measurements and help our customers optimize their ROI in the process. ## A framework for quantifying business value At dbt, we're focused on helping customers use data more effectively to support critical initiatives. This means partnering with customer leadership and practitioners to align on business initiatives and validate the financial benefit expected from dbt. The work involves identifying business priorities, then documenting how dbt saves time and money and helps achieve those priorities. The final output is a measurable business case that quantifies return on investment based on cost and time savings, while highlighting how dbt supports the business. Every customer leader talks about embedding reliable data-driven decision-making across their business. At a high level, these efforts boil down to three common themes. - **Supporting revenue**. - **Enabling efficiency and cost savings**. - **Building and modernizing technology.** ## Building the business case When building a value case, there's a consistent framework that follows how customers evaluate dbt. First is tooling and maintenance. What's the cost of the dbt platform versus the current tooling? For self-hosted customers, we quantify the total cost of ownership for maintenance and scale. For native warehouses, it’s time saved on stored procs and notebooks: time spent using your tool rather than maintaining it. Next is warehouse optimization: maximizing warehouse value through improved governance and eliminating redundant model builds with features like state-aware orchestration. Finally is operational time savings - improving collaboration, accelerating pipeline builds, and improving data quality. This frees up time to be more productive and focus more hours on business impact. All these cost savings and freed-up hours are allocated toward broader business initiatives: driving revenue, supporting cost savings, and modernizing tech stacks to reduce risk. ![How practitioners support business initiatives using dbt](https://cdn.sanity.io/images/wl0ndo6t/main/b1e071e7b8bcd80e04edcdb17b495de4ae76b54a-1510x844.png) **** ## Three habits to start today Considering your own business initiatives, here are habits to get the most ROI from your dbt investment. 1. **Bring engineering discipline to analytics.** Version control, CI/CD, testing, and ownership - this is the foundation of value. If you're already a dbt user, you're doing these things and probably crushing this discipline. 2. **Optimize your costs. **Lineage, orchestration, and remediation make your spend predictable. Limiting or eliminating model rebuilds, unnecessary rework, and questions from confused teammates will save costs and headaches. 3. **Make ROI visible to leadership**. Think about how to connect data analytics and AI investments to outcomes. Speak in the language of the business by focusing on the measurables that matter. When you do this, you're not just managing a data platform - you're delivering quantifiable business value that transforms how your organization operates and builds trust. --- --- title: "Bring structured context to Snowflake Intelligence with dbt" description: "Power reliable Cortex and Snowflake Intelligence experiences with dbt and Snowflake Semantic Views." url: "https://www.getdbt.com/blog/bring-structured-context-to-snowflake-intelligence-with-dbt" date: "2026-02-11" authors: ["Chakshu Mehta", "Dave Connors", "Luis Leon"] categories: ["Product"] --- # Bring structured context to Snowflake Intelligence with dbt Snowflake Intelligence brings agentic and conversational experiences directly into Snowflake. It is powered by Cortex capabilities such as Cortex Analyst and Cortex Search, and its AI agents connect to governed assets, including semantic views and models, search services, and other tools. That architecture reinforces a familiar point for data teams: these AI experiences are only as reliable as the structure and semantics they are grounded in. In [Parts 1](https://www.getdbt.com/blog/bring-structured-context-to-conversational-analytics-with-dbt) and [2](https://www.getdbt.com/blog/bring-structured-context-to-agentic-data-development-with-dbt) of this series, we introduced the idea of a "structured context layer" and why AI systems need it to produce reliable, governed, cost-efficient outcomes. In Part 3, we will look at how dbt helps you build that structure in Snowflake, and how that structure directly improves the quality of [Snowflake Intelligence](https://www.getdbt.com/blog/what-is-snowflake-intelligence-anyway) experiences. ## The problem: AI can't infer meaning from raw schemas Large language models (LLMs) don't know how your models relate, what your metrics actually mean, which joins are valid, or what’s production-ready versus experimental. If all you give an AI system is a database schema, it’s left to infer meaning from names and patterns. When semantics aren’t explicit, the model has to guess. That leads to incorrect SQL, inconsistent metrics, ungoverned access paths, and answers you can’t trace back to trusted definitions. ## What Snowflake Intelligence needs to be accurate To answer business questions reliably, [Cortex Analyst](https://www.snowflake.com/en/developers/guides/getting-started-with-cortex-analyst/) needs explicit rules for what fields mean, which joins are valid, and how metrics should be computed. Snowflake [Semantic Views](https://docs.snowflake.com/en/user-guide/views-semantic/overview) are designed to provide that structure. They are schema-level objects that package up the definitions that matter most—metrics, dimensions, relationships, and the surrounding context—so Cortex Analyst has something concrete to follow when translating a question into SQL against your physical tables. One operational detail that matters: semantic views aren’t just Cortex configuration. They’re Snowflake objects with normal lifecycle and access controls. That means teams can validate them, promote them through environments, and manage who can use them, rather than configuring semantics once and hoping they don’t drift. ## dbt builds the governed foundation Snowflake-level governance solves only part of the problem. Your warehouse changes constantly, and so does the AI stack around it: new agents, new interfaces, new retrieval patterns, new places where definitions get copied or reinterpreted. Semantics are only trustworthy if they stay aligned with the models underneath and remain consistent as new consumers show up. Snowflake Semantic Views are great for making semantics enforceable at runtime **inside Snowflake**. But most teams want more than Snowflake-native enforcement. They want semantics to live in a code-first system of record: centrally owned, reviewed like software, **and reusable everywhere they need it**—**from BI and embedded analytics to AI**—not trapped inside a single product experience. That is where dbt fits. dbt is where meaning gets defined and kept correct over time: in code, with version control, review, tests, lineage, and repeatable deployments, so those definitions stay interoperable across your stack, even as your underlying warehouse or AI stack evolves. In dbt, the "structured context layer" is the governed, machine-readable context your team already builds as part of analytics engineering, made reliable and portable to AI. Practically, dbt's structured context layer includes: - **Semantic definitions**, including metrics and the models they depend on - **Lineage**, so systems can understand how models connect and where data comes from - **Contracts, tests, and CI signals**, so downstream consumers know what is valid and what changed - **Documentation and ownership**, so definitions have meaning and accountability - **Freshness and operational metadata**, so you can reason about trust and timeliness - **Policies and business rules**, so governance and logic are centrally maintained dbt then makes that context available in the right shape for the job. For interactive or agentic workflows, the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) can give AI tools governed access to dbt-managed assets (including the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) via the [MCP query tool](https://docs.getdbt.com/docs/dbt-ai/about-mcp#semantic-layer)), so agents can discover metrics and models—and pull lineage, documentation, and trust signals—before generating queries. And on Snowflake, when you want Cortex Analyst to use the same governed definitions natively, you can publish that context into Snowflake Semantic Views directly from your dbt project (via supported tooling/packages), giving Cortex a Snowflake-native semantic object to follow. ### **Where the dbt Semantic Layer fits** The [dbt Semantic Layer](https://www.getdbt.com/blog/how-the-dbt-semantic-layer-works) is part of that structured context layer. It is where teams define governed metrics in code, once, and reuse them across many consuming surfaces without re-implementing business logic each time. Under the hood, it is powered by [MetricFlow](https://www.getdbt.com/blog/open-source-metricflow-governed-metrics), the open-source SQL generation engine that compiles metric definitions into efficient SQL to [guarantee accuracy, consistency, and performance](https://www.getdbt.com/blog/why-your-ai-will-fail-without-a-semantic-layer) while enforcing the rules that usually get lost in translation (ie grain, aggregation behavior, join paths, and semantic constraints). And [dbt Fusion](https://docs.getdbt.com/docs/fusion/about-fusion) ensures that SQL executes consistently and efficiently on your warehouse. This means AI retrieves both the correct meaning and the correct computation of metrics like “revenue” without the model wasting tokens trying to infer missing logic. That matters because it keeps metric results consistent even when the query shape changes across tools or AI-generated prompts. This becomes especially valuable when the same metric needs to show up reliably across BI, embedded analytics, and AI experiences—without three teams quietly maintaining three “almost identical” versions. **Why the dbt Semantic Layer matters for AI-enabled analytics** 1. **A single, auditable definition of truth**: Metrics live in code, get reviewed, and ship through the same workflow as your models. When someone asks, “Why is revenue down?” you can point to one definition and its lineage, not reconcile competing interpretations. 2. **Lower-cost, more stable query behavior**: MetricFlow compiles metrics into SQL that’s optimized and semantically correct. When agents generate lots of queries (and retries), optimization and consistency matter: fewer full scans, fewer “wrong shape” queries, fewer corrective loops. 3. **An interface you can reuse everywhere**: Instead of teaching every tool (or every agent) how to compute metrics, you expose metrics through a consistent central interface. When you add a new AI surface, you can reuse governed definitions rather than rebuilding logic. Together, these give AI systems something they usually lack: semantics that are trustworthy, portable, and operationally maintained. ## How dbt and Snowflake Semantic Views work together Teams often want two things at the same time: 1. A Snowflake-native semantic object that Cortex Analyst/Snowflake Intelligence can consume directly 2. A code-first system of record for semantics that stays governed and reusable across **more than one** consuming surface Those map cleanly to two complementary layers: ### dbt is the system of record The dbt Semantic Layer is where semantics get defined, reviewed, tested, and maintained over time, alongside the models they depend on. That’s what keeps meaning stable as the warehouse changes and new consumers get introduced. It also means you can surface the same governed context beyond Snowflake Intelligence. For example, via the dbt MCP server, any AI agent outside of Snowflake can discover which metrics exist, read documentation, understand lineage, and check trust signals before they attempt query generation. This means agent interactions are more reliable and reduce token overhead. ### Snowflake Semantic Views are the Snowflake-native consumption layer Semantic views express business concepts in a format Snowflake Intelligence expects. They describe metrics, dimensions, relationships, and context so Cortex Analyst can interpret questions and generate SQL reliably. In practice, many teams will author and govern the source of truth in dbt, _then_ publish the right shape of semantics into Snowflake where Cortex can use it directly. Here's a practical workflow: 1. Define metrics and semantic models in dbt 2. Test and validate them with tests/contracts/CI and review 3. Deploy curated models to Snowflake 4. Create Snowflake Semantic Views directly within dbt as part of your DAG 5. Cortex Analyst consumes those views to generate SQL dbt provides the governance, testing, and version control that makes Snowflake Intelligence trustworthy. Semantic Views provide the Snowflake-native interface that Cortex Analyst reads. ## Where to start If you are deciding where to start, here is a lightweight way to think about it: When Cortex operates inside the guardrails defined by dbt's structured context layer, you unlock: - **Correctness**: Cortex queries follow official metric and schema definitions - **Governance and security**: Access, contracts, and trust signals are enforced end-to-end - **Cost efficiency**: Smaller prompt payloads, fewer warehouse scans, fewer retries, and more optimized SQL paths reduce AI operational costs - **Explainability**: AI answers can be traced back to models, tests, and metadata - **Consistency across tools**: the same semantics power BI, AI, and embedded use cases ### Future-proofing with OSI: Building beyond Snowflake Intelligence As teams adopt more BI and AI tools, semantic definitions tend to get duplicated across systems. The [Open Semantic Interchange (OSI)](https://open-semantic-interchange.org/?_fsi=JBu1651l) aims to reduce that duplication by providing [a vendor-neutral spec](https://github.com/open-semantic-interchange/OSI) for sharing semantic definitions between systems. OSI doesn't replace platform-native semantics like Snowflake Semantic Views. It provides a portable layer that can help you avoid re-authoring semantics from scratch each time you add a new consuming surface. This portability matters because as teams add BI tools, embedded analytics, and other AI-enabled applications, OSI lets them define semantics once in dbt, then generate the appropriate native representations for each consumer—whether that's Snowflake Semantic Views, semantic models for Tableau, etc. dbt is in the process of building a dbt Semantic Layer to OSI converter so teams can export governed definitions into an OSI-compliant artifact. Once mature, this will enable smoother interoperability between dbt semantic definitions and Snowflake Semantic Views, allowing teams to define metrics in dbt and generate Snowflake-native semantic views with less manual work. ## Bringing structured data to Snowflake Intelligence without a rewrite Most teams don’t start from a blank slate. You can converge on a combined approach incrementally. ### If you started with Snowflake Semantic Views 1. **Bring the underlying tables under dbt ownership:** Use dbt to build and maintain the curated tables/view your Semantic Views depend on, with documentation, testing, and controlled releases. 2. **Manage Semantic View changes like code:** Define Snowflake Semantic Views as part of your dbt DAG using the [dbt_semantic_view](https://hub.getdbt.com/Snowflake-Labs/dbt_semantic_view/latest/) package. This ensures Semantic View changes are version controlled, reviewable, and promoted across environments alongside your dbt deployments. 3. **Add the dbt Semantic Layer where you need to reuse beyond Snowflake Intelligence:** For metrics that need to stay consistent across BI + AI + embedded surfaces, define them once in dbt Semantic Layer and reuse them rather than re-implementing logic across tools. _**Note for teams using Fusion:** Semantic Views are Snowflake-specific SQL. To keep them in your dbt project (and benefit from dbt governance), set `static_analysis: off` for those files so Fusion doesn’t attempt to analyze them._ ### If you started with the dbt Semantic Layer 1. **Keep dbt as the system of record for governed metric definitions:** Define metrics in code, manage change through review, and keep consistent over time across tools and teams. 2. **Publish to Snowflake Semantic Views for Snowflake Intelligence (choose one path):** If you want change control and release discipline, manage Semantic Views in dbt; if you want the fastest setup with Snowflake-managed generation, use Autopilot. 1. **Recommended path: Manage Semantic Views in dbt:** Use the [dbt_semantic_view](https://hub.getdbt.com/Snowflake-Labs/dbt_semantic_view/latest/) package to create and deploy the Semantic Views you need, so the Snowflake-facing contract ships with the same CI/CD, review, and environment promotions as the rest of your dbt project. 2. **Alternative path: Snowflake [Semantic View Autopilot](https://docs.snowflake.com/en/user-guide/views-semantic/autopilot):** Use the new Semantic View Autopilot tool to generate Snowflake Semantic Views directly from your dbt Semantic Layer definitions. This Snowflake-managed tool will read your dbt Semantic definitions and write the corresponding Snowflake Semantic View. This can be a quick way to get Snowflake-native objects without manually authoring them, but the resulting view lifecycle is driven by the Snowflake tool rather than your dbt deployments. 3. **Start narrow, represent the subset needed for Snowflake Intelligence in Semantic Views, then expand:** Use the [dbt_semantic_view](https://hub.getdbt.com/Snowflake-Labs/dbt_semantic_view/latest/) package to publish the metrics, dimensions, relationships, and context required for Cortex Analyst and Snowflake-native consumption first, and broaden coverage as usage grows and patterns stabilize. 4. **Connect Custom Applications to dbt & Cortex via MCP for richer agent behavior:** Use the [dbt MCP Server](https://github.com/dbt-labs/dbt-mcp) to expose dbt semantics—metrics, metadata, lineage, and trust signals—to any AI client application. This allows AI systems to ground their responses in governed definitions and provides transparency around data quality. For [example](https://github.com/dbt-labs/streamlit_mcp_cortex/tree/main), you can build a custom Streamlit application within Snowflake that calls both Cortex agents and the dbt MCP Server simultaneously. This enables users to ask data questions and immediately see whether the underlying data is tested, fresh, and trusted, making AI outputs more explainable and actionable. This video walks through an end-to-end workflow for powering reliable Cortex and Snowflake Intelligence experiences with dbt and Snowflake Semantic Views: [Watch video](https://youtu.be/PYaue8UY7J4) ## **Building reliable AI with dbt and Snowflake** Snowflake Intelligence brings AI experiences directly into your warehouse. dbt provides the structured context that makes those experiences traceable and reliable. Together, they make AI outputs more accurate, auditable, and cost-efficient. And as your AI strategy expands beyond Snowflake Intelligence (new agents, new interfaces, new applications), dbt keeps semantics centralized and reusable so you’re not rebuilding meaning in every new system. The result: accurate, governed, and cost-efficient AI that works for everyone. If you're interested in learning more about bringing structured context to Snowflake Intelligence with dbt, [get a demo](https://www.getdbt.com/contact) today or speak with your account representative. --- --- title: "Do you need a data observability platform?" description: "Not sure if you need a data observability platform? Learn the signs, tradeoffs, and business impact of investing in one." url: "https://www.getdbt.com/blog/do-you-need-a-data-observability-platform" date: "2026-02-10" authors: ["Joey Gault"] categories: ["Pulse"] --- # Do you need a data observability platform? ## Understanding the problem space Modern data stacks have fundamentally changed the landscape of data operations. While cloud-native architectures and tools like [dbt](https://www.getdbt.com/product/what-is-dbt) provide unprecedented flexibility and power, they've also expanded the surface area for potential failures. Data now flows from hundreds of sources through various ingestion tools, gets transformed through multiple stages, and ultimately feeds dozens of downstream applications and dashboards. Each step represents a potential failure point, and the interconnected nature of these systems means issues can propagate quickly and unpredictably. The fragmentation inherent in modern data architectures compounds these challenges. Organizations typically use different systems for ingestion, transformation, orchestration, and consumption. This creates visibility gaps that make it difficult to understand what's happening across the entire data estate, particularly when problems span multiple systems or when root causes lie upstream from where symptoms appear. Data engineering leaders consistently face questions they can't answer within reasonable timeframes: Why isn't my model up to date? Is my data accurate? Why is my model taking so long to run? How should I materialize and provision my model? These questions highlight the gap between having data infrastructure and truly understanding how that infrastructure performs. When you can't answer these questions quickly, you're operating with significant blind spots that expose your organization to risk. ## Assessing your need for observability The decision to invest in a data observability platform depends on several factors specific to your organization's situation. Consider the scale and complexity of your data operations. If you're managing hundreds of data sources, running thousands of transformations, and supporting dozens of downstream applications, the manual effort required to maintain visibility becomes unsustainable. The larger and more complex your data estate, the more essential comprehensive observability becomes. The business criticality of your data systems matters significantly. When data directly drives operational processes (automated marketing campaigns, fraud detection systems, recommendation engines, or real-time pricing), the tolerance for data quality issues approaches zero. These use cases are far less forgiving than traditional batch reporting. A recommendation engine fed with stale or incorrect data produces poor recommendations immediately, creating direct business impact. Your current pain points provide clear signals about observability needs. If your team spends significant time firefighting data issues rather than building new capabilities, if business users frequently report discrepancies in dashboards before your team detects them, or if you struggle to identify the root cause of data quality problems, these are indicators that your current visibility is insufficient. The maturity of your data organization also influences this decision. Teams adopting distributed ownership models or [data mesh architectures](https://www.getdbt.com/blog/what-is-data-mesh) need observability to maintain quality and reliability across decentralized teams. Clear visibility into data lineage, quality metrics, and performance characteristics enables domain teams to take ownership of their data products while maintaining organizational standards. ## What observability platforms provide A comprehensive data observability platform typically consists of several interconnected components. Data collection mechanisms capture metadata about data pipelines, transformations, and quality metrics, serving as the raw material for all observability insights. Monitoring capabilities form the reactive component, detecting anomalies and alerting teams when issues arise. These systems track data freshness, volume changes, schema evolution, and quality degradation across the entire pipeline. Testing frameworks provide the proactive element, allowing teams to define expectations for data behavior and catch issues before they reach production. When integrated properly with development workflows, testing creates a safety net that prevents many issues from ever affecting end users. Performance monitoring adds another crucial dimension, tracking execution times, resource utilization, and system bottlenecks to help teams optimize pipelines and make informed decisions about infrastructure scaling. The relationship between observability platforms and transformation tools like dbt represents a particularly powerful combination. dbt provides the foundation for data transformation, offering models that clean raw data and create high-quality datasets. The [dbt testing framework](https://docs.getdbt.com/docs/build/data-tests) verifies data quality through automated checks, while [dbt documentation](https://docs.getdbt.com/docs/collaborate/documentation) creates consistent, well-documented models that serve as single sources of truth. When observability platforms monitor [dbt models](https://docs.getdbt.com/docs/build/models) and pipelines in production, they can detect anomalies and ensure ongoing data accuracy. The real power emerges from how these tools work together. When monitoring systems detect anomalies indicating serious data quality issues, teams can create corresponding dbt test cases that prevent pipelines from proceeding if the same conditions occur again. This integration shifts responsibility for data quality upstream, enabling business users to address issues at their source rather than waiting for data engineering intervention. ## Alternatives to dedicated platforms Before committing to a dedicated observability platform, consider that teams can build significant observability capabilities using native artifacts from their transformation tools. dbt generates detailed artifacts after every run, test, or build command, containing granular information about model execution, test results, and pipeline performance. These artifacts serve as a rich data source for custom observability solutions. The project manifest provides complete configuration information for dbt projects, while run results artifacts contain detailed execution data for models, tests, and other resources. When combined with [data warehouse](https://www.getdbt.com/blog/future-of-the-modern-data-stack) query history, these artifacts enable deep insights into model-level performance that can inform optimization decisions. Teams have successfully built lightweight ELT systems that ingest artifact data into their data warehouses, then use dbt itself to transform this metadata into structured models that power dashboards and alerting systems. This approach leverages existing infrastructure and skills while providing customizable observability tailored to specific organizational needs. The key components of such a system include orchestration that reliably captures artifacts regardless of pipeline success or failure, storage that preserves historical artifact data for trend analysis, modeling that transforms raw artifacts into actionable insights, and alerting that notifies relevant stakeholders when issues arise. This DIY approach requires more upfront investment but offers complete control and customization. ## Measuring the business impact The business value of comprehensive data observability can be substantial and measurable. Organizations implementing robust observability practices often see dramatic improvements in both cost efficiency and data reliability. Performance monitoring frequently reveals opportunities for significant cost reduction through identification of inefficient queries, unused models, and optimization opportunities. Teams have reported reductions in cloud data warehouse credit usage of 70% or more by systematically addressing performance issues identified through observability platforms. **** Beyond cost savings, observability frameworks enable organizations to scale while keeping costs stable. Despite adding many more models and bringing additional data sources online, teams can maintain stable job execution times while decreasing cost per unit of computation through systematic optimization. This demonstrates how effective observability can support growth without proportional increases in operational overhead. Data quality improvements are equally significant. When organizations integrate new data sources, observability platforms initially detect spikes in anomalies as teams learn to work with new data. However, the systematic conversion of anomalies into test cases leads to corresponding decreases in anomalies over time, creating self-improving systems that become more robust with experience. The strategic value extends beyond immediate cost and quality metrics. Teams with comprehensive observability can confidently make changes to their pipelines, knowing that issues will be detected and addressed quickly. This confidence enables more rapid iteration and innovation in data products and analytics. Rather than constantly fighting fires and rebuilding fragile systems, teams with solid observability foundations can focus on delivering business value through innovative data products and insights. ## Implementation considerations Successful data observability requires more than just technology; it demands organizational commitment and cultural change. The most effective implementations treat observability as a core competency rather than an afterthought, investing in the tools, processes, and skills necessary to maintain visibility into data systems at scale. Effective alerting represents one of the most critical aspects yet is often implemented poorly. The goal is to provide timely, actionable notifications to the right people without creating alert fatigue. Best practices include implementing domain-specific tagging that allows alerts to be routed to appropriate team members based on model ownership. Alerts should include sufficient context for debugging, including error messages, model names, and timestamps, enabling recipients to quickly understand and address issues. Training team members to leverage observability tools for incident routing and data discovery creates a self-service culture that reduces the burden on data engineering teams. When business users understand how to interpret observability data and respond to alerts, they can often resolve issues without escalating to technical teams. The combination of proactive testing and reactive monitoring creates more resilient systems than either approach alone. Transformation frameworks catch many issues before they reach production, while observability tools detect the problems that slip through. When important anomalies are detected, creating test cases that prevent pipeline execution until upstream issues are resolved shifts responsibility appropriately and prevents the propagation of known data quality problems. ## Making the decision The decision to invest in a data observability platform ultimately depends on your specific circumstances. If you're managing complex data operations at scale, supporting business-critical use cases, and struggling with visibility gaps, a dedicated platform likely makes sense. The investment pays dividends through cost savings, improved reliability, and increased trust in data systems. However, if your data operations are relatively straightforward, you have strong existing monitoring practices, or you have the technical capacity to build custom solutions, you may be able to achieve sufficient observability without a dedicated platform. Starting with dbt's native capabilities and gradually expanding as needs grow represents a pragmatic approach. The key is to honestly assess your current state, understand the risks you're exposed to, and evaluate whether your existing capabilities are sufficient to maintain the reliability and trust your organization requires. As data becomes increasingly central to business operations, organizations that master data observability (whether through dedicated platforms or custom solutions) will have significant advantages in their ability to make reliable, data-driven decisions at scale. For more information on building reliable data systems, explore the [dbt documentation](https://docs.getdbt.com/) and [dbt Learn](https://learn.getdbt.com/) resources to understand how transformation best practices integrate with observability strategies. ## Data observability FAQs **What is a data observability platform?** **What are the primary features of a data observability platform?** **How can I integrate data observability into my existing pipeline?** Integration can be achieved through several approaches. The most powerful combination involves connecting observability platforms with transformation tools like dbt, where the platform monitors models and pipelines in production to detect anomalies and ensure ongoing data accuracy. When monitoring systems detect issues, teams can create corresponding test cases that prevent pipelines from proceeding if the same conditions occur again, shifting responsibility for data quality upstream. Alternatively, teams can build custom observability solutions using native artifacts from their transformation tools, creating lightweight systems that ingest artifact data into data warehouses and use existing infrastructure to transform metadata into structured models that power dashboards and alerting systems. --- --- title: "dbt Core v1.11 is GA" description: "dbt Core v1.11 is GA with first-class UDFs, clearer project validation, and performance improvements across adapters." url: "https://www.getdbt.com/blog/dbt-core-v1-11-is-ga" date: "2026-02-10" authors: ["Grace Goheen", "Sara Gawlinski"] categories: ["Product"] --- # dbt Core v1.11 is GA The dbt Core team wrapped 2025 with a final release of v1.11. This release continues to drive forward the language of dbt, adding new features and fixes for teams who want to standardize logic across their data stack, understand when they're adhering to the dbt standard and when they're doing something custom, and benefit from better debugging and performance optimizations. Check out our [upgrade guide](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.11) for more information, and read on for a deeper dive into the community efforts and conversations that helped bring these features to life. If you’d rather hear it straight from the team, we’ll also be hosting a live virtual release event to walk through v1.11, share roadmap context, and answer questions. [Save your seat here](https://www.getdbt.com/resources/webinars/dbt-core-1-11-live-release-updates-roadmap). Let’s get into what’s new. ## ⭐️ The star of the show: User-defined functions (UDFs) With v1.11, dbt introduces support for defining and managing user-defined functions (UDFs) directly within your dbt project. For years, the community has used macros to encapsulate reusable logic. Macros are powerful, but they live only inside dbt. UDFs extend the same principle of DRY ("don't repeat yourself") code by allowing you to register custom functions as **objects inside your warehouse**, so the same transformation logic can be used _everywhere_ in your data ecosystem. **Check out our FAQ page [here](https://docs.getdbt.com/faqs/Project/udfs-vs-macros)** for a deeper breakdown of when to use a UDF instead of a macro. Out-of-the-box support for UDFs in dbt has been a long time coming. Our CEO Tristan Handy opened up the original feature request back in 2016. ![The original feature request for UDFs in 2016](https://cdn.sanity.io/images/wl0ndo6t/main/337e137db0b3721fd5fc407e5578355adcb38196-1164x370.png) Since then, community members have come up with a variety of workarounds to half-bake UDFs into dbt - `pre-hooks`, `on-run-start` hooks, even SQL headers. A couple of community members took it one step further and developed a strategy of defining UDFs as models via a custom materialization (looking at you [@Brabster](https://github.com/brabster) and [@anaghshineh](https://github.com/anaghshineh)). Stories like this really highlight the magic of dbt. Its extensibility allows the community to experiment, [share their findings](https://github.com/brabster/dbt_materialized_udf), and have the feature organically grow in adoption and maturity. The most successful experiments eventually make it into the dbt standard, just like UDFs now have in v1.11. After almost a decade of letting this discussion bake, as different community members figured out what worked best and where the thorny edges were, it became clear these workarounds weren’t quite right. It was time to make UDFs an official part of the dbt standard. In the months leading up to the `v1.11` release, we had great community [discussions](https://github.com/dbt-labs/dbt-core/discussions/11851) on GitHub and live on Zoom, all of which helped to shape the feature. Thank you to everyone who participated. We’re excited to finally deliver dbt-managed UDFs to all of you. ### What UDFs enable When you define a function in dbt, it becomes a node in your DAG. That means dbt manages building and updating your UDF in the warehouse before the model(s) that references it. Let's say you want to create a function that checks if a string represents a positive integer. (The regular expression here is fairly simple, but we know there are more-complex regexes hanging out in many dbt projects.) First, create a file in the `functions/` directory that contains your logic: ```sql # functions/is_positive_int.sql REGEXP_INSTR(a_string, '^[0-9]+$') ``` Second, define your function's name, configs, and properties in a corresponding YAML file: ```yaml # functions/schema.yml functions: - name: is_positive_int description: > My UDF that returns 1 if a string represents a naked positive integer (like "10", "+8" is not allowed). config: schema: udf_schema database: udf_db arguments: - name: a_string data_type: string description: The string that I want to check if it's representing a positive integer (like "10") returns: data_type: integer ``` Now, you can reference your function in a dbt model: ```sql # models/my_model.sql select maybe_positive_int_column, {{ function('is_positive_int') }}(maybe_positive_int_column) as is_positive_int from {{ ref('a_model_i_like') }} ``` When compiled, the `{{ function('is_positive_int') }}` is replaced by the fully-qualified UDF name `udf_db.udf_schema.is_positive_int`. Finally, run `dbt build` to materialize your UDF in the warehouse before the model(s) that references it. ![UDF as part of a DAG](https://cdn.sanity.io/images/wl0ndo6t/main/666ef8e8d3158e52c7597ef5c1743554c9928826-1192x472.png) User-defined functions bring warehouse-level extensibility into the dbt workflow, reduce duplication, and enable cross-tool logic reuse. ### What’s new for UDFs since beta If you tuned in to Coalesce, you already saw [a demo (on rollerskates) of UDFs in action](https://www.youtube.com/watch?v=aMUAQjqTKtc&t=1116s). ![Jeremy and Grace on skates on stage demoing UDFs](https://cdn.sanity.io/images/wl0ndo6t/main/489dfdb43f75fb7a67a43a7a5ece5999bd09a343-2048x1365.png) So what's new since October? In the final release of v1.11, you can now define UDFs that run Python logic, when supported by your data warehouse and adapter. Simply define your logic in a python file in your `functions/` directory, and provide additional python-specific configurations such as [`runtime_version`](https://docs.getdbt.com/reference/resource-configs/runtime-version) and [`entry_point`](https://docs.getdbt.com/reference/resource-configs/entry-point). This makes it possible to reuse complex transformations, calculations, or logic that would be difficult or verbose to express in SQL. Additionally, you can now configure: - [`default_value`](https://docs.getdbt.com/reference/resource-properties/function-arguments#default_value)s for arguments - [`volatility`](https://docs.getdbt.com/reference/resource-configs/volatility) to describe how predictable the function output is - [aggregate](https://docs.getdbt.com/reference/resource-configs/type#aggregate) functions that operate on multiple rows and return a single value Do you have additional extensions of UDFs you want dbt to support? [Open up an issue](https://github.com/dbt-labs/dbt-core/issues) in the dbt Core repo! [Javascript UDFs](https://github.com/dbt-labs/dbt-core/issues/12332) and [pypi packages](https://github.com/dbt-labs/dbt-core/issues/12041) are already on our radar :) ### UDF support expectations UDF support varies by adapter and execution environment. UDFs are currently supported across BigQuery, Snowflake, Redshift, Postgres, and Databricks, with some limitations depending on the platform. Python UDFs are currently only supported in Snowflake and BigQuery when using dbt Core. Read more about UDFs, including prerequisites and limitations in the [UDF documentation](https://docs.getdbt.com/docs/build/udfs). ## The authoring experience is getting stricter (in a good way) A focus of v1.11 is to make dbt authoring more explicit and predictable. The goal isn’t to be pedantic, it’s to make sure projects fail _early and clearly_, instead of succeeding quietly and behaving strangely later. If you’ve ever found yourself asking, _“Why did dbt accept this config but not actually do what I expected?”_, this work is for you. To get technical about it for a moment, the dbt language spec is now codified by a set of strongly-typed [jsonschemas](https://github.com/dbt-labs/dbt-jsonschema/tree/main/schemas/latest_fusion). These are automatically generated and define what is acceptable “dbt code”. We introduced many deprecations warnings for invalid code in v1.10 (see [GitHub discussion](https://github.com/dbt-labs/dbt-core/discussions/11493)), with a subset of those deprecation warnings (the ones that _use_ the new jsonschemas) gated behind a behavior change flag to give users a migration window to update their project code. But starting in v1.11, **dbt now warns you by default when your project code doesn't align to the standard dbt spec.** These warnings help you proactively catch and clean up things like: - [misspelled](https://github.com/dbt-labs/dbt-core/issues/8942) or deprecated config keys - invalid top-level properties - missing `+` prefixes If you’ve ever had a project where a subtle config typo didn’t show up until late in development (or only appeared as wonky downstream behavior), you’ll feel this improvement immediately. Importantly, this **does not mean dbt is taking away customization**. That kind of flexibility is what enabled our community to experiment with UDFs as a custom materialization for the past decade ;) You can still add custom configuration in dbt, those values just need to live under the `meta` key moving forward. ```yaml models: - name: my_model config: meta: my_custom_thing: is_awesome ``` If you’re not yet ready to resolve these issues, you can use the [warn_error_options](https://docs.getdbt.com/reference/global-configs/warnings) configuration to silence specific categories temporarily while you update your code on your own timeline. ```yaml flags: warn_error_options: silence: - CustomTopLevelKeyDeprecation - CustomKeyInConfigDeprecation - CustomKeyInObjectDeprecation - MissingPlusPrefixDeprecation - SourceOverrideDeprecation ``` _Note: dbt Core v1.11 surfaces invalid code as warnings. In dbt Fusion, they’re treated as **errors**. This gives teams a safe migration window to clean up code before they upgrade to Fusion. Check out the _[dbt-autofix](https://github.com/dbt-labs/dbt-autofix)[ tool](https://github.com/dbt-labs/dbt-autofix)_ to autofix many of these!_ ## Adapter-specific improvements Our adapter releases also bring the ecosystem forward. Highlights include: ### BigQuery: batched metadata-based source freshness dbt can now issue a single batch query when calculating metadata-based source freshness (instead of one query per source) for BigQuery. This reduces overhead and can dramatically speed up freshness checks for projects with lots of sources. [Enable it with](https://docs.getdbt.com/reference/global-configs/bigquery-changes#the-bigquery_use_batch_source_freshness-flag): ```yaml flags: bigquery_use_batch_source_freshness: true ``` Thanks [@adamcunnington-mlg](https://github.com/adamcunnington-mlg) for the help stress-testing this new approach. ### Snowflake: Iceberg + dynamic table clustering Snowflake users get improved support for: - basic Iceberg table materialization via a Glue catalog–linked database - `cluster_by` supported on dynamic tables—closing an annoying gap for teams adopting dynamic tables and wanting more control over how data is organized and optimized ### Spark: better retry handling for PyHive New profile configs improve reliability for async query polling and connection issues: - `poll_interval` - `query_timeout` - `query_retries` ### Databricks: snapshots and real-world deletes Snapshots now support `hard_deletes='new_record'` , which enables users to track deletions from snapshots as auditable events. Shoutout to [@randypitcherii](https://github.com/randypitcherii) for the initial implementation. ## Bug bashing during ‘De_bug_-cember’ A healthy open source project isn’t just about shipping new features, it’s also about doing the unglamorous work of making things more stable and predictable. In December, the dbt community and maintainers leaned into “De_bug_-cember,” working through long-standing issues across parsing, execution, logging, and error messages. Thanks for the contributions from community members [@edgarrmondragon](https://github.com/edgarrmondragon), [@asiunov](https://github.com/asiunov), [@rjspotter](https://github.com/rjspotter), [@avasireddi3](https://github.com/avasireddi3), [@mattogburke](https://github.com/mattogburke), [@D3nn3](https://github.com/D3nn3), and [@mjsqu](https://github.com/mjsqu)! Core v1.11 rolls up 35 De_bug_-cember fixes into one release. If you want to see the running commentary (and celebrate the wins), check out `#dbt-core-development` in the [Community Slack](https://www.getdbt.com/community/join-the-community). ![Debug-cember post in the dbt community slack](https://cdn.sanity.io/images/wl0ndo6t/main/993d36f51aaf8d8126b26305142b82332036dd4d-1316x708.png) ## Upgrading to v1.11 A few reminders as you plan your upgrade: - dbt Labs is committed to backward compatibility for all versions 1.x. Any behavior changes are accompanied by behavior flags to provide a migration window. - If you’re on dbt platform release tracks, **Latest** and **Compatible** already have access to the newest Core features. - Read through our [upgrade guide](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.11) for more information ## What’s next If you want to read more about what’s next for dbt Core, check out [our latest roadmap post over in github](https://github.com/dbt-labs/dbt-core/blob/main/docs/roadmap/2025-12-magic-to-do.md) (musical-theater-references guaranteed). To hear directly from the people doing this work, you’re invited to join the dbt Core product, engineering, and developer experience teams for a live virtual event on February 19. We’ll cover what shipped in v1.11, why it matters in practice, share roadmap context, and answer your questions live. [Save your seat here](https://www.getdbt.com/resources/webinars/dbt-core-1-11-live-release-updates-roadmap). And as always, we want to hear from you: What custom solutions are you maintaining today that you wish dbt handled out of the box? What workflows still feel painful, even though you have to deal with them all of the time? Because the pattern is the same as it’s always been: you experiment, you share, we learn, and the standard evolves. To everyone who contributed to a discussion, filed an issue, proposed a PR, tested a beta: it’s time to celebrate. dbt isn’t dbt without you. --- --- title: "Deliver reliable AI with the dbt Semantic Layer and dbt MCP Server" description: "Learn how to provide structured context to AI systems with a dbt-powered MCP server." url: "https://www.getdbt.com/blog/dbt-mcp-server-reliable-ai" date: "2026-02-05" authors: ["Stephen Robb"] categories: ["Learn"] --- # Deliver reliable AI with the dbt Semantic Layer and dbt MCP Server AI is changing how we use data at a pace we've never seen before. There's a new tool every day that requires us to redo entire workflows. Data engineering teams that once focused solely on pipelines and transformations are now on the front lines of AI strategy. They're expected to deliver intelligent, reliable data products faster than ever before. And somehow, at the same time, they need to make sure those systems are trustworthy, scalable, and secure. Just as dbt revolutionized data transformation in the cloud era by adding better processes around ETL - modular, version-controlled, testable workflows - we're now doing the same thing in the AI space. In this article, I’ll delve into how you can use [the dbt Model Context Protocol (MCP) Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server), combined with the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl), to deliver data to AI and ensure unparalleled accuracy. **** ## The AI context gap Our customers are moving incredibly fast, building everything from agentic workflows and internal copilots to customer-facing chatbots. However, the underlying AI stack is changing even faster. Each week, there's a new breakthrough, a new model, a smarter framework, or even a new vector store. That velocity means AI and data teams are constantly adapting, moving from [Claude](https://claude.ai) to [Cursor](https://cursor.com/), from [Pinecone](https://www.pinecone.io/) to [semantic layers](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl). Many are retraining systems just to keep up. Each of these shifts introduces friction: new integrations, rewrites, tuning, and duplication. There's a lot of time, risk, and slowing of innovation caused by this wide variety of tools. At the end of the day, a lot of the fundamental questions are the same. What does this metric mean? Is this definition still up to date? Where did this number come from? Can we trust this data? For enterprise AI, the bottleneck isn't compute power or data volume. It's context. Specifically, a structured, governed context. Large language models can't reason over raw or fragmented data. They need a strong foundation. Without structured definitions and shared business logic, AI systems get stuck guessing. And when they guess, they start hallucinating. That's where you see all the data problems of asking a question and getting the wrong answer. Take an internal chatbot, for example. A user asks, "What's our revenue for enterprise customers in Q2?" To answer that question, the model has to understand a few things: - What defines an enterprise customer? - Which up-to-date table or model includes revenue logic? - Does that revenue table account for cancellations, ticket sales, or returns? - Are there any other considerations for how you answer that question? Without structured, governed context, the large language model might generate SQL that runs but returns the completely wrong answer with complete confidence. You would think the machine is correct without being able to validate it. Or consider an agentic workflow generating an entire weekly order summary without a version-controlled definition of "order." It could pull that definition from the wrong table, double-count returns, or blatantly miss critical logic. That could result in incorrect results for a question as simple as "What was my revenue last quarter?" ## dbt as the control plane for AI What's needed isn't more tools. It's a unifying [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) for your AI systems: one place where your business logic, transformations, tests, and documentation live. Structure it once, then apply it everywhere. That's where we see dbt fitting in. With dbt as the control plane for your AI, you structure your data once and use AI everywhere. AI systems need access to structured context. That’s primarily what dbt offers: the ability to create data about your data that can then be used for querying, for AI, for all these other use cases. Those AI systems are going to pull that context so we can get to a more centralized, governed platform and workflow. The idea is that we can remove a lot of that context that exists in fragments or tribal knowledge and lock it down in code so it's reliable across your entire organization. dbt is already the standard for creating high-quality, governed datasets from your warehouse. It captures rich metadata: model lineage, test coverage, and centralized metric definitions. It's also bringing performance and cost efficiency. Rather than connecting each AI workflow to its own source, you can let dbt centralize those transformations, metrics, and documentation in one layer. This reduces queries, decreases compute spend through unified models and state-of-the-art orchestration, and provides faster response times for your end users. With dbt, you get cross-platform flexibility. dbt works across all major data warehouse systems, including [Snowflake](https://snowflake.com), [Databricks](https://databricks.com), and [BigQuery](https://cloud.google.com/bigquery). You're able to build fast across your entire stack without sacrificing consistency. Another important consideration is security and governance, which is one of the big things holding back production grade AI projects. dbt helps close that gap. With every model, test, and transformation, we can validate that it's logged, [versioned](https://docs.getdbt.com/docs/cloud/git/version-control-basics), and auditable. This means you can meet enterprise-grade requirements for compliance and data protection with AI. ## Introducing the dbt MCP Server Everything I've covered so far has been about providing the context that's required. So, how does dbt make it really easy to integrate your tools? That's where the [Model Context Protocol](https://modelcontextprotocol.io/docs/getting-started/intro) (MCP) fits in. Tools like LangChain and Semantic Kernel can directly query the [dbt Semantic Layer](https://www.getdbt.com/blog/semantic-layer-introduction), [data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage), and tested models via API so that your AI systems don't just access the data, they also understand it. While the AI tools may change - GPT-4 to Claude, chatbots to agents - your foundation shouldn't change. Or, at least, it doesn't have to if you're using dbt. The MCP unifies all dbt assets and AI applications. We make that accessible for you in two ways: both a local connection and a remote connection. The local option is pretty self-explanatory. It runs on your laptop alongside your dbt project. It's a fantastic option for local development with tools like Cursor or Claude, and it empowers your agent to create your dbt code right on your machine. The second option is to run it remotely. We have an incredibly straightforward setup where you can plug and play your dbt MCP Server with any of your AI tools and access them through any web application. It means you can connect multi-agent, multi-user systems more easily than ever before. You'll see a wide variety of MCP integrations with many of our partners in the future. ## Real-world impact: the M1 Finance case study A lot of what I've shared has been theoretical conversation, but I wanted to support that with a real case study. By introducing dbt MCP, [M1 Finance](https://m1.com/) was able to reduce its engineering bottlenecks and dramatically improve its efficiency. Their teams didn't have to wait for specialized resources to move projects forward. It gave them a clear path to reduce some of the biggest blockers to adopting AI basically, hallucinations. With structured, validated access to those systems, the AI could act on real authoritative data. With that additional context, they were finally confident that the answers and outputs they received were accurate, reliable, and safe. This lets them unlock true scale, not just a proof of concept or an experiment. ## A customer journey: Galaxy's Edge Travel Company To illustrate how this works in practice, let me walk you through a customer example we created. Imagine Galaxy's Edge Travel Company, the largest tourism operator in the outer rim of space. They offer a wide range of services, including star cruiser vacation packages, droid-assisted lodging experiences, lightspeed-enabled transportation, and holotable concierge services. Galaxy's Edge wanted to build an AI concierge—a holo guide—a much cooler version of Alexa or Siri. This holo guide can answer any questions instantly: pricing, availability, loyalty points, packaging rules, and bundling recommendations. You ask it what you want to know, and it spits out the answers immediately and always consistently. The challenge? They have the data for all of these systems, but it's scattered across multiple platforms. They have star cruiser manifests in JSON events captured from their hyperspace travel system, droid service logs in semi-structured information, resort booking information stored in structured tables, and loyalty point balances that come through microservice API dumps. While you likely don’t work for an intergalactic travel tourism operator, these systems probably resonate with systems in your environment. ## Understanding the context gap The holo guide doesn't understand the business. It's a brand new AI system. It can show up, but it doesn't know anything about their systems. This isn't because the AI is bad, it's because their enterprise context is a mess. If you take all this data in disparate systems and put an AI on top of it, it's not going to produce very good results. For instance, calculations such as total trip cost vary across systems, producing different results. Room names could be "Deluxe Pod," "D-Pod," or "Pod Deluxe," and they could all refer to the same room, but AI doesn't know that unless we provide context. Loyalty tiers aren't joined correctly, causing the AI to hallucinate discounts and produce inconsistent pricing. Packaging availability depends on complex business logic. All of these things have to be taken into account to ensure the holo guide consistently produces the right results. The organization recognizes that AI is only as good as the structured context we provide. What they need is a data control plane and a structured context layer built on something like dbt. ## Building the solution: three key steps To power their AI concierge, Galaxy's Edge needed to implement three critical components: **Data modeling.** First, they created consistent facts and dimensions for trips, packages, customers, and droid services. They curated this data, breaking it up into facts and dimensions in an easy-to-understand manner. **The Semantic Layer.** They defined metrics like total trip cost, occupancy rate, and loyalty-eligible balance that are always correct. The Semantic Layer also defines the model's joins to help it understand how joins are performed, adds business-friendly naming, and stores calculations and metrics. **The MCP Server.** This is how they expose governed dbt context directly to the AI agent—in this case, the Holo Guide. When someone asks, "How many loyalty points will I earn if I add the Holocron Discovery Tour to my three-night star cruiser stay?" the system produces accurate results. For AI to be reliable, it needs to know the definition of a Holocron Discovery Tour. It needs to pull semantic metrics on cost, loyalty accrual, and discount rules. It needs to join the customer profile or customer dimension, and it needs to apply real business logic. To see a walkthrough of how to implement this in practice, [sign up for our recent webinar and watch the full demo](https://www.getdbt.com/resources/webinars/delivering-reliable-ai-with-the-dbt-semantic-layer-and-dbt-mcp-server). ## What's coming next: dbt agents The dbt MCP Server is just one part of our AI strategy. We’re also working to build several [agents directly inside](https://www.getdbt.com/product/dbt-agents) the platform. We're developing tools such as an analyst agent, discovery agent, observability agent, developer agent, and even more over the upcoming months that will solve specific parts of the [Analytics Development Lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle) directly within dbt. The first one I've tried is [dbt Insights](https://docs.getdbt.com/docs/explore/dbt-insights), which lets you use natural language querying to generate SQL and get results. These agents will be among the easiest and most powerful ways to use dbt AI directly within the dbt platform. ## Conclusion: structured context is the foundation AI systems are only as good as the structured context you provide them. Without that foundation - without clear definitions, governed metrics, documented transformations, and tested models - your AI tools will hallucinate, produce inconsistent results, and ultimately fail to deliver business value. With dbt as your data control plane and the dbt MCP Server connecting your AI tools to that governed context, you can: - Build reliable AI experiences that scale - Reduce hallucinations by providing clear business logic - Accelerate development by automating tedious tasks - Deliver trusted data products that your business can depend on Whether you're building internal chatbots, customer-facing AI concierges, or agentic workflows, the foundation is the same: structured, governed, reliable data. And that's exactly what dbt delivers. --- --- title: "How data mesh solved centralized data management challenges" description: "Data mesh decentralizes data ownership and reduces bottlenecks with self-service platforms and federated governance." url: "https://www.getdbt.com/blog/how-data-mesh-solves-centralized-data-challenges" date: "2026-02-04" authors: ["Joey Gault"] categories: ["Pulse"] --- # How data mesh solved centralized data management challenges ## The core problems with centralized data management Traditional centralized architectures create several persistent challenges that compound as organizations grow. Understanding these problems is essential to appreciating how data mesh addresses them. ### Bottlenecks slow innovation In centralized models, data teams become organizational bottlenecks. Every dashboard request, data model, or schema update funnels through the same team, creating long queues and delays. As business units wait for their data needs to be met, the organization's ability to gain insights slows, and innovation stalls. This bottleneck effect intensifies as data volume and complexity increase, making it nearly impossible for a single team to keep pace with demand across multiple departments. ### Distance from domain expertise Centralized data engineering teams often lack deep knowledge of the business domains they serve. They may not fully understand what constitutes "good" data for finance, marketing, or sales. This distance between data processors and data owners results in long load-test-fix cycles that delay project delivery. The complexity of these systems means few people understand how all the parts work together, making troubleshooting and optimization difficult. ### Data silos persist Despite centralization efforts, data silos continue to proliferate. Frustrated by bottlenecks, some teams build their own solutions, resulting in shadow IT and fragmented data landscapes. Even when data exists in a central repository, teams struggle to discover what's available, assess its quality, or understand whether it meets their specific needs. This fragmentation makes cross-departmental analysis difficult and prevents organizations from developing a holistic view of their operations. ### Scalability limitations As organizations grow, centralized systems struggle to scale. The volume and variety of data can quickly overwhelm central infrastructure. A single data engineering team must manage data and business logic for every team across the organization, creating an unsustainable burden. Without a way to distribute this load, performance and efficiency suffer. ## How data mesh decentralizes ownership Data mesh fundamentally reimagines data architecture by shifting from technology-driven to business-driven ownership. Rather than a single team managing all data, each business domain (such as finance, sales, or marketing) takes responsibility for its own data and its complete lifecycle. ### Domain-driven data ownership The foundational principle of data mesh is that individual business domain teams should own their own data. This concept, built on domain-driven design principles, aligns responsibility with business function rather than technology. Finance owns financial data, sales owns sales data, and each team manages its own pipelines, transformations, and quality controls. This shift brings clear ownership and demarcation. In centralized architectures, determining who owns a given dataset is often unclear. With domain-driven ownership, teams register their ownership in a data catalog, eliminating ambiguity. This clarity accelerates decision-making and reduces coordination overhead. Domain ownership also distributes the burden that previously fell on a single team. Instead of one data engineering group managing everything, individual teams handle their own data products and pipelines. This distribution directly addresses the bottleneck problem, enabling teams to deliver new solutions without waiting in queue. ### Treating data as a product Data mesh introduces the concept of data products: well-defined, self-contained units of data that solve specific business problems. A data product might be as simple as a table or report, or as complex as a machine learning model. What distinguishes a data product is how it's managed and exposed to others. Data products include interfaces that define how teams expose data to others, specifying columns, data types, and constraints. They incorporate contracts that serve as written specifications for these interfaces, which all teams can use to validate conformance. Versioning enables teams to introduce new revisions while supporting previous versions for backward compatibility. Access rules define who can see what data, ensuring sensitive information like personally identifiable information remains protected. This product-oriented thinking addresses the brittleness of traditional systems. Without contracts or versions, downstream systems have no elegant way to handle change. An alteration to a data type or text field format can break any application or report that depends on that data. Data products with versioned contracts give partner teams time to adapt, preventing unexpected breakages. ## Enabling self-service while maintaining governance Data mesh's success depends on balancing autonomy with oversight. Domain teams need independence to move quickly, but organizations require consistent standards for quality, security, and compliance. Data mesh achieves this balance through two key components: self-serve data platforms and federated computational governance. ### Self-serve data platforms Domain teams manage their own data products from end to end, including ingestion, transformation, quality testing, and analytics. However, it doesn't make sense for every team to build this infrastructure independently. That would result in redundant contracts, incompatible tooling, and wasted resources. A self-serve data platform provides the tools domain teams need to work independently. This platform includes data storage solutions, ingestion tools, transformation capabilities like dbt, orchestration systems, security controls, and data catalogs for registering and discovering data products. By standardizing these tools centrally while making them available to all teams, organizations enable scalability without sacrificing consistency. The platform approach also reduces the learning curve. Using tools like dbt for transformation requires less ramp-up time because it leverages SQL, a language most engineers and analysts already know. This familiarity accelerates adoption and reduces friction. ### Federated computational governance Without proper governance, data mesh could devolve into data anarchy. Federated computational governance prevents this by tracking and managing compliance centrally while allowing domain teams to operate independently. While teams own their data products, the data platform and corporate governance team enforce standards through automated policies. Data governance automation ensures that new data products meet requirements for security, quality, privacy, and compliance. Many of these checks run automatically, reducing manual labor and enabling governance at scale. When issues arise, the domain team (as the data owner) is responsible for responding and fixing compliance problems, such as classifying unclassified values or removing sensitive information from logs. This federated approach maintains the speed benefits of decentralization while ensuring organizations can meet regulatory requirements like GDPR and maintain consistent data quality standards across all domains. ## Practical benefits for data engineering leaders For data engineering leaders evaluating architectural approaches, data mesh offers concrete operational improvements that address long-standing pain points. ### Faster time to market Domain teams possess deeper knowledge of their own data and own their pipelines and business logic. This ownership means they can deliver new solutions faster than if they had to hand implementation to a centralized team. The reduction in handoffs and the elimination of queue time translate directly to accelerated project delivery. ### Improved data quality Teams closest to the data understand it best and are best positioned to make quality decisions. Centralized teams often lack the context needed to determine what constitutes "good" data for each domain. Returning these decisions to domain owners results in better decision-making and higher data quality across the organization. ### Reduced costs While building a data mesh architecture requires upfront investment, most organizations find the effort pays for itself. Savings come from multiple areas: reduced friction between business units and IT, less time spent searching for quality data, automated governance reducing manual effort, and elimination of redundant data and processes. Organizations using data mesh have reported significant cost savings (in some cases, millions of dollars driven back into the business). ### Better resource utilization By distributing data responsibilities, organizations reduce strain on central teams and enable more efficient resource allocation. Data engineering teams can focus on platform capabilities and enablement rather than becoming overwhelmed with implementation requests. Domain teams gain the autonomy to solve their own problems without waiting for scarce central resources. ## Implementation considerations Data mesh represents a significant shift in how organizations approach data management. Success requires more than technical changes; it demands cultural transformation and careful planning. Organizations benefit most from data mesh when they've hit limits with traditional architectures. Common indicators include centralized data engineering teams that have become bottlenecks preventing quick project launches, or spikes in downstream errors due to lack of product-oriented thinking. Companies should have a general understanding of which data domains belong with which teams, whether organized by usage or business function. Buy-in from all participants is critical. This includes C-suite sponsorship and, importantly, buy-in from data engineers, analytics engineers, domain teams, business analysts, and product managers. Without organizational alignment, the transition can create frustration and encourage further shadow IT. A solid training program is essential. All stakeholders should understand what the shift entails and receive proper training on new tools and processes. Domain teams particularly need guidance on what data ownership means and how to manage pipelines with new toolsets. Starting small often works best. Data platform teams should gather requirements broadly but begin implementation with a single domain team. After onboarding that team and incorporating feedback, they can progressively onboard additional teams, iterating on toolsets and processes along the way. ### The role of modern tools Data mesh leverages many existing data technologies (object storage, data warehouses, data lakes) but uses them differently. The difference lies in who has access and how access is federated across domains. Tools like dbt play a central role by enabling domain teams to build, validate, test, and run data pipelines independently while maintaining consistency through features like model contracts. Data catalogs become increasingly important in data mesh environments, serving as the single source of truth for all data sources and data products. They enable discovery across a distributed network of heterogeneous domain owners and support governance through data classification, quality tracking, and lineage visualization. Orchestration tools define when datasets should be used and under what conditions. Data governance software implements automated enforcement of policies. Self-service reporting tools enable teams to run their own analytics after discovering data through the catalog. Together, these components create an ecosystem that supports both autonomy and alignment. ## Conclusion Data mesh addresses the challenges of centralized data management by fundamentally rethinking how organizations structure data ownership and responsibility. Rather than funneling all work through a single team, it distributes ownership to domain experts while maintaining governance through federated automation and self-serve platforms. For data engineering leaders, data mesh offers a path to scale data operations without proportionally scaling central teams. It reduces bottlenecks, improves data quality, accelerates time to value, and creates more sustainable operating models. While implementation requires careful planning and cultural change, the architectural approach represents a natural evolution in managing data at scale. Organizations considering data mesh should evaluate whether they've encountered the limits of centralized approaches and whether they have the organizational readiness to support distributed ownership. For those facing bottlenecks, quality issues, or scalability challenges, data mesh provides a proven framework for moving forward. ### Related resources - [The 4 principles of data mesh](https://www.getdbt.com/blog/four-principles-of-data-mesh) - [Getting started with data mesh](https://www.getdbt.com/blog/getting-started-with-data-mesh) - [Understanding data mesh architecture](https://www.getdbt.com/blog/data-mesh-architecture) - [Data quality testing with dbt](https://www.getdbt.com/product/data-quality) - [Learn more about dbt Mesh](https://www.getdbt.com/product/dbt-mesh) ## Data mesh FAQs **What is a data mesh?** **Why has data mesh become so popular? ** Data mesh has gained popularity because it addresses critical problems that organizations face with traditional centralized data architectures. As companies grow, centralized data teams become bottlenecks that slow innovation, creating long queues for dashboard requests and data models. Additionally, central teams often lack deep domain expertise, leading to quality issues and long troubleshooting cycles. Data mesh solves these challenges by distributing ownership to domain experts who understand their data best, enabling faster time to market, improved data quality, reduced costs through automation and elimination of redundancies, and better resource utilization across the organization. **When to adopt data mesh? ** --- --- title: "Why analytics engineering isn't just data modeling" description: "Analytics engineering extends beyond modeling to collaboration, testing, documentation, and scaling analytics workflows." url: "https://www.getdbt.com/blog/why-analytics-engineering-isnt-just-data-modeling" date: "2026-02-02" authors: ["Joey Gault"] categories: ["Pulse"] --- # Why analytics engineering isn't just data modeling [Analytics engineering](https://www.getdbt.com/resources/the-analytics-development-lifecycle) is often misunderstood as simply “data modeling” — a task focused on building tables and defining relationships. In reality, analytics engineering is a broader discipline that blends technical craft with software engineering practices and organizational enablement. Analytics engineers transform and test data, document lineage and logic, implement collaboration workflows, and help scale analytics work across teams. ## The full scope of analytics engineering Analytics engineers transform, test, deploy, and document data. They apply software engineering best practices like version control and continuous integration to analytics code. This work extends well beyond simply designing tables and defining relationships between entities. **** The role emerged from a practical need. Traditional data teams operated in a cycle where analysts gathered requirements from stakeholders, submitted requests to data engineers, and waited (sometimes for weeks) for new data pipelines or fixes. Data engineering teams accumulated months-long backlogs while business users grew frustrated with delays. Analytics engineers broke this bottleneck by taking ownership of the transformation layer, using tools like dbt to build data pipelines that were both less complicated and less fragile than previous approaches. ### Core responsibilities beyond modeling Analytics engineers maintain clean analytics code at scale. This means checking code into source control, versioning transformation models, creating and running data transformation tests, ensuring code follows [DRY principles](https://www.getdbt.com/blog/dry-principles) by bundling reusable components into packages, and using CI/CD to automatically and safely ship data changes to production. These practices come directly from software engineering, adapted for the specific challenges of data work. Documentation and data definitions form another critical area of responsibility. Analytics engineers ensure that data users and other engineers understand what purpose data serves, where it comes from, and how it was derived. This documentation work directly enables data self-service across the organization. Training business users on tools represents yet another dimension of the role. Analytics engineers bridge the gap between data engineers and business users through brown bags, one-on-one training, and other teaching modalities. They show data analysts and others how to find data and leverage it in reports, reducing dependency on centralized teams. ## The organizational impact The value analytics engineering delivers extends far beyond technical outputs. When organizations introduce an analytics engineering practice, they typically see improvements across multiple dimensions. ### Reducing engineering backlogs When all requests for new data projects or fixes must flow through the data engineering team, that team becomes a bottleneck. Adding analytics engineering means adding a practice focused solely on creating new data pipelines and transformations. Instead of relying on data engineering for every data change, teams can hire analytics engineers whose sole responsibility is managing their specific data needs. ### Improving data velocity and quality Data engineering team members juggle multiple responsibilities, which often creates large lags between requests and fulfillment. If a business user has an issue with something the central data team delivered, it means making yet another request and returning to the back of the queue. By contrast, an analytics engineer's sole focus is developing new data sets for their teams. They can ship changes to users in days instead of weeks or months. This faster iteration cycle also improves data quality. Analytics engineers embedded with their teams develop deep domain expertise in their specific business area. They understand the business model and data needs in detail, making them more efficient at implementing requirements. They can get answers to hard questions more easily than members of a centralized data engineering team who support multiple domains. ### Enabling infrastructure evolution Data engineers can absolutely build data pipelines, but that doesn't mean it's the best use of their time. Most data engineering teams would rather focus on providing new capabilities for analytics engineers and data analysts. When analytics engineers take on more data pipeline work, it frees data engineers to focus on infrastructure improvements like creating self-serve tools for initializing new data products, improving data governance capabilities, optimizing query performance, and reducing pipeline costs. Since infrastructure improvements benefit the entire company, this work often has high return on investment. Consider a change that reduces the time to run CI/CD jobs by 10 minutes. In a company with hundreds of data pipelines, this represents significant time and cost savings. ## The technical craft: modularity and maintainability While analytics engineering encompasses organizational responsibilities, the technical craft remains central to the role. The shift from monolithic SQL scripts to modular, version-controlled transformations fundamentally changed what's possible in data work. ### From monoliths to modules Before modern data transformation tools, the most reliable approach to data modeling involved writing massive SQL files (sometimes 10,000 lines or more) or splitting logic across separate SQL files or stored procedures run in sequence with Python scripts. Few people in the organization would be aware of these scripts, so even when someone needed similar transformations, they'd start from source data rather than leveraging existing work. With modular approaches enabled by tools like dbt, every producer or consumer of data models can start from foundational work that others have done before them. When foundational data models are referenced in multiple places rather than rebuilt from scratch each time, the dependency graph becomes much easier to follow. Teams can clearly see how layers of modeling logic stack upon each other and where dependencies lie. ### Naming conventions and project structure A solid naming convention makes data modeling projects easy to navigate. Without one, teams may rebuild models that had already been reviewed and published, or rejoin data in duplicative, low-performance ways. Naming conventions define both the types of models used (source, staging, marts) and what types of transformation each model type handles. Staging models clean up and standardize raw data coming from the warehouse, providing consistency for downstream models. They're typically a one-to-one reflection of raw sources, with light transformations like field type casting, renaming columns for readability, and filtering out deleted records. When anything changes in source data, fixing the staging layer ensures changes flow into downstream models without manual intervention. Data mart models apply business logic to build core data assets used directly in downstream analysis. These typically include dimension and fact tables with heavier transformations: joins of multiple staging models, `CASE WHEN` logic, and window functions. Organizations may also use intermediate models to break up complex logic or base models to join related staging models before moving downstream. ### Readability and reusability Individual data models should remain readable (ideally around 100 lines of code). Models shorter than 100 lines typically avoid overly complex joining, either by limiting the raw number of joins or by joining in simple ways. When someone opens the model, they can quickly understand what it does and how they might modify or extend it. SQL files can become long and tedious to read, especially when handling tasks like unioning multiple source tables or generating date spines. This is where macros become essential: they allow teams to invoke modular blocks of SQL from within individual files. Leveraging macros and packages built by others enables teams to sprint out data models quickly while keeping them easily readable. ## Where transformation ends and analysis begins Defining where data modeling effort ends and where analysis begins helps teams avoid over-engineering transformations. If end users write SQL or use analysis tools that join tables for them, the final data modeling output can be generalized fact and dimension tables. Users can freely mix and match these to analyze various aspects of the business. If end users don't write SQL or analysis tooling is limited in terms of self-serve joining, analytics engineers may need to curate datasets that answer specific business questions by joining together multiple fact and dimension tables into wide tables. The goal in transformation work is to bring data together and standardize prep work, not to pre-build every analysis or complex aggregation that may come up in the future. ## Building for the team Analytics engineering enables broader participation in data work. Once a project has clearly defined structure, new contributors can produce high-quality SQL modeling regardless of their "technical" background. A solid folder structure and clear model review process provide enough guardrails to keep transformation code quality high. This democratization of data work represents one of analytics engineering's most significant impacts. By creating accessible, well-documented data products and establishing clear patterns for how to build on them, analytics engineers enable more people across the organization to work directly with data. This reduces bottlenecks, improves data literacy, and ultimately helps organizations become more data-driven. ## The path forward Analytics engineering is not just data modeling. It's a practice that combines technical craft with organizational enablement, software engineering principles with business domain expertise, and infrastructure thinking with end-user empowerment. For data engineering leaders evaluating whether to build an analytics engineering practice, the question isn't whether your team needs people who can model data; it's whether you need people who can transform how your entire organization builds, maintains, and uses data. To learn more about building an analytics engineering practice, explore resources on [what analytics engineers do](https://www.getdbt.com/what-is-analytics-engineering), review [analytics engineering best practices](https://www.getdbt.com/resources/the-analytics-development-lifecycle), or check out the [dbt Learn training catalog](https://www.getdbt.com/dbt-learn) to develop your team's skills. ## Analytics engineering FAQs **Data engineer vs analytics engineer vs data analyst: which role fits you?** Analytics engineers sit between data engineers and data analysts, focusing specifically on the transformation layer of data work. They transform, test, deploy, and document data using software engineering best practices. Data engineers typically focus on infrastructure and providing new capabilities like self-serve tools, data governance, and query optimization. Data analysts work with the data products that analytics engineers create to perform analysis and generate insights. Analytics engineers enable both groups by building modular, well-documented data pipelines that reduce bottlenecks and improve data velocity. **But who ensures everything is connected?** Analytics engineers ensure everything is connected by maintaining the transformation layer that sits between raw data and analysis-ready datasets. They build modular data models that reference foundational work, creating clear dependency graphs that show how layers of modeling logic stack upon each other. Through solid naming conventions, project structure, and documentation, they make it easy for anyone in the organization to understand what data exists, where it comes from, and how it was derived. This enables data self-service across the organization and ensures that teams can leverage existing transformations rather than rebuilding from scratch. **Data analyst vs data engineer vs data scientist: what's the difference?** Data analysts use analysis-ready data to generate insights and answer business questions, often working with curated datasets and reporting tools. Data engineers build and maintain the infrastructure that ingests, stores, and processes raw data, focusing on capabilities like pipeline optimization, cost reduction, and governance. Analytics engineers bridge these roles by owning the transformation layer: they take raw data that data engineers provide and transform it into clean, documented, modular datasets that analysts can use. They apply software engineering practices like version control and CI/CD to analytics code, enabling faster iteration and higher quality data products. --- --- title: "What defines a control plan for data infrastructure" description: "Learn how a data control plane centralizes metadata, governance, & collaboration to make distributed data architecture manageable." url: "https://www.getdbt.com/blog/what-defines-a-control-plan-for-data-infrastructure" date: "2026-02-02" authors: ["Joey Gault"] categories: ["Pulse"] --- # What defines a control plan for data infrastructure Modern data environments are sprawling with data scattered across platforms, tools, and domains, and teams using different workflows to build and consume insights. That complexity creates governance gaps, fragmented metadata, and brittle pipelines. A [data control plane](https://www.getdbt.com/discover/understanding-the-data-control-plane) offers a unifying architecture: centralizing the coordination of governance, orchestration, observability, and metadata while enabling collaboration across technical and business users. In this article, we explore what defines a control plane, the capabilities that distinguish it from traditional architectures, and why it’s foundational for scalable, trusted analytics. ## The architectural foundation A [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) is an abstraction layer that sits across your data stack, unifying capabilities that traditionally existed in separate tools. Rather than managing orchestration in one system, observability in another, and cataloging in a third, a control plane centralizes these functions while connecting the metadata that flows between them. This architectural approach differs fundamentally from both data planes and traditional control planes. A [data plane](https://www.ibm.com/think/topics/control-plane-vs-data-plane) handles the actual movement and processing of data (executing queries, transforming datasets, and storing results). A control plane in the traditional sense manages configurations and policies without touching the data itself. A data control plane bridges these concepts by creating a centralized hub that manages data workflows, governance, and metadata while maintaining awareness of what's happening across your entire data estate. The distinction matters because modern data infrastructure is inherently distributed. Data lives in multiple platforms, teams work across different domains, and consumption happens through varied interfaces from BI tools to AI systems. A control plane provides the connective tissue that makes this distributed architecture coherent and manageable. ## Core capabilities that define a control plane Three fundamental capabilities define an effective control plane for data infrastructure: flexibility, collaboration, and trustworthiness. These characteristics directly address the challenges that data engineering leaders identify as their biggest obstacles. ### Cross-platform flexibility A control plane must operate across diverse data platforms and cloud environments without creating vendor lock-in. This flexibility enables distributed teams to work with the tools and platforms that best serve their needs while maintaining centralized governance and consistent workflows. When business logic is abstracted into a flexible control plane, organizations can optimize spend across platforms and adapt as the market evolves. [dbt](https://www.getdbt.com/product/what-is-dbt) exemplifies this approach through its support for multiple data warehouses and cloud providers. Teams can build transformation logic once and deploy it across different platforms, or manage complex projects that span multiple data environments. This interoperability becomes increasingly important as organizations adopt multi-cloud strategies and [data mesh](https://www.getdbt.com/blog/data-mesh-getting-started) architectures where different domains may operate on different platforms. ### Governed collaboration Data development cannot remain the exclusive domain of data engineers. A control plane must make analytics workflows accessible to users with varying technical backgrounds while maintaining appropriate governance and quality standards. This democratization accelerates delivery by reducing bottlenecks and enables domain experts to participate directly in building data products. The [Analytics Development Lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle) provides a framework for this collaboration. Similar to how software engineering adopted the SDLC to break down silos between developers and operations teams, the ADLC promotes collaboration between data producers and data consumers. A control plane implements this framework through features like version control, code review workflows, and interfaces tailored to different user personas. dbt supports this through multiple development environments. Data engineers work in [dbt Studio](https://www.getdbt.com/product/dbt) or [VS Code](https://docs.getdbt.com/docs/install-dbt-extension), analysts use visual interfaces, and business users discover data through catalog tools. All of these interfaces operate on the same underlying codebase with consistent governance, testing, and documentation. ### Trustworthy outputs Perhaps most critically, a control plane must ensure that data products are accurate, well-tested, and observable. Trust in data doesn't happen by accident; it requires systematic approaches to quality, clear ownership, and transparency into how data is created and maintained. This means building testing and validation directly into development workflows rather than treating them as separate activities. It means providing [column-level lineage](https://www.getdbt.com/product/dbt-catalog) so teams can trace data from source to consumption and quickly identify root causes when issues arise. It means automating documentation so knowledge doesn't live solely in people's heads. And it means surfacing data quality signals to consumers so they can assess freshness and reliability before making decisions. ## Metadata as the foundation What truly distinguishes a control plane from a collection of tools is how it handles metadata. A control plane doesn't just store metadata; it makes metadata actionable by connecting information across the entire analytics workflow. Consider what happens when a data model changes. In a fragmented tool landscape, that change might break downstream dashboards without anyone knowing until a stakeholder reports incorrect numbers. With a control plane that maintains comprehensive metadata and lineage, the system can identify affected downstream assets, run tests to validate the change, and notify relevant stakeholders before the change reaches production. This metadata awareness extends beyond technical lineage to include business context. A control plane should understand not just how data flows through transformations, but what that data means, who owns it, and how it's being used. The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) exemplifies this by centralizing metric definitions alongside transformation logic, ensuring that business logic is defined once and consumed consistently across BI tools, embedded applications, and AI systems. ## Enabling the Analytics Development Lifecycle (ADLC) A control plane exists to support a mature analytics practice, not simply to manage infrastructure. The Analytics Development Lifecycle provides the process framework, while the control plane provides the technological foundation that makes that process practical at scale. **** The ADLC encompasses eight stages: planning analytics products, building them, testing, deploying to production, operating production systems, ensuring reliability, and making data products discoverable. Each stage requires specific capabilities that a control plane must provide. During development, teams need environments where they can safely build and test changes without affecting production. They need [CI/CD workflows](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud) that automatically validate changes before they're merged. They need the ability to reuse modular components rather than rebuilding logic from scratch. In production, teams need orchestration that runs jobs efficiently, observability that surfaces issues quickly, and the ability to roll back changes when problems occur. They need cost visibility to optimize warehouse spend and performance metrics to identify bottlenecks. For discovery and consumption, stakeholders need catalogs that help them find relevant data, documentation that explains what data means, and interfaces that enable self-service access to governed datasets. A control plane integrates these capabilities rather than forcing teams to stitch together disparate tools. ## From tool sprawl to unified workflows The emergence of control planes as an architectural pattern reflects the maturation of the modern data stack. Early cloud data platforms solved fundamental problems around storage and compute. As organizations scaled their analytics practices, specialized tools emerged for orchestration, observability, cataloging, and semantic modeling. This specialization drove innovation but created new problems. Metadata became fragmented across tools with no centralized way to connect it or take holistic action. Teams spent time integrating tools rather than delivering value. Costs multiplied as organizations paid for overlapping capabilities across multiple vendors. A control plane consolidates these fragmented capabilities into a cohesive platform. Rather than managing orchestration in one tool, observability in another, and cataloging in a third, teams work within a unified environment where these capabilities are deeply integrated. This consolidation reduces complexity, lowers costs, and enables more sophisticated workflows that leverage metadata across the entire analytics lifecycle. ## The role of AI in control planes As organizations adopt AI and machine learning, the requirements for data infrastructure intensify. [AI systems demand high-quality, well-governed data with clear lineage and explainability](https://www.getdbt.com/blog/ai-data-transformation). They require semantic understanding of what data represents, not just technical schemas. And they need interfaces that enable both humans and AI agents to work with data safely. A control plane designed for the AI era must provide structured context that AI systems can leverage. This includes comprehensive metadata about data models, tests that validate data quality, and semantic definitions that explain business logic. When this context is centralized and accessible, AI copilots can generate more accurate code, conversational analytics can provide trustworthy answers, and AI agents can take action with appropriate guardrails. dbt addresses this through features like [dbt Copilot](https://www.getdbt.com/product/dbt-copilot), which uses AI to accelerate development while maintaining governance, and integrations with AI systems that consume data through the semantic layer. The control plane ensures that AI systems work with the same trusted, governed data that powers traditional analytics. ## Practical implementation considerations For data engineering leaders evaluating control plane solutions, several practical factors warrant consideration. The platform should integrate with your existing data infrastructure rather than requiring a wholesale replacement. It should support your team's preferred development workflows while providing pathways for less technical users to participate. And it should provide clear ROI through reduced tooling costs, improved efficiency, and faster time to value. Migration from legacy systems to a modern control plane requires careful planning but need not be disruptive. Approaches like replatforming and refactoring enable organizations to modernize iteratively, delivering value quickly while managing risk. The key is starting with a comprehensive assessment of existing pipelines, understanding complexity and dependencies, and developing a phased migration plan that maintains business continuity. Organizations that successfully implement control planes report substantial benefits: 50-80% reductions in transformation costs, dramatically faster development cycles, fewer data quality incidents, and the ability to reallocate resources toward strategic initiatives like AI. These outcomes stem not from any single feature but from the holistic approach that control planes enable. ## Conclusion A control plane for data infrastructure is defined by its ability to unify fragmented capabilities, centralize and activate metadata, and support collaboration across diverse teams and platforms. It provides the architectural foundation for mature analytics practices that deliver trusted data at the speed modern businesses require. For data engineering leaders, adopting a control plane represents a strategic choice about how their organizations will work with data. It's a shift from managing disparate tools to orchestrating unified workflows, from reactive firefighting to proactive quality management, and from siloed development to collaborative data product delivery. The organizations that thrive in the AI era will be those that can ship trusted insights quickly, govern data effectively across distributed teams, and adapt as technologies and business needs evolve. A well-designed control plane makes this possible by providing the coordination layer that modern data infrastructure demands. To learn more about implementing a data control plane, explore [dbt](https://www.getdbt.com/product/dbt) and the [Analytics Development Lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle). ## Control plan FAQs **What is a data control plane? ** A data control plane is an abstraction layer that sits across your data stack, unifying capabilities that traditionally existed in separate tools like orchestration, observability, and cataloging. It creates a centralized hub that manages data workflows, governance, and metadata while maintaining awareness of what's happening across your entire data estate. Unlike a data plane that handles actual data movement and processing, or a traditional control plane that only manages configurations, a data control plane bridges these concepts by providing connective tissue that makes distributed data architecture coherent and manageable. **What are the core capabilities that define an effective data control plane?** **How does metadata function as the foundation of a data control plane?** A data control plane makes metadata actionable by connecting information across the entire analytics workflow, rather than just storing it. It maintains comprehensive metadata and lineage that can identify affected downstream assets when changes occur, run tests to validate changes, and notify relevant stakeholders before changes reach production. This metadata awareness extends beyond technical lineage to include business context, understanding not just how data flows through transformations, but what that data means, who owns it, and how it's being used, ensuring business logic is defined once and consumed consistently across all systems. --- --- title: "Data marts vs. Data products: What's the difference?" description: "Understand how data marts and data products serve different roles in modern data platforms—and why you need both." url: "https://www.getdbt.com/blog/data-marts-vs-data-products" date: "2026-02-02" authors: ["Joey Gault"] categories: ["Pulse"] --- # Data marts vs. Data products: What's the difference? As data organizations mature, they often start using similar terms to describe very different ideas. Two of the most commonly conflated concepts are data marts and data products. Both are critical to delivering reliable, scalable analytics, but they serve different purposes and solve different problems. Understanding how they differ helps teams design better data architectures, clarify ownership, and align technical work with business outcomes. In this article, we’ll break down what data marts are, what data products represent, and how the two work together to support trusted, self‑service data at scale. ## Understanding data marts Data marts represent a specific layer in the data transformation pipeline. In [dbt projects](https://docs.getdbt.com/docs/build/projects), marts form the final layer of core transformations, where staging models and intermediate models come together to create business-defined entities. Each mart represents a specific entity or concept at its unique grain: an order, a customer, a payment, or a click event. The defining characteristic of marts is their entity-grained structure. A customers mart contains all useful data about customers at the customer level. An orders mart maintains individual orders as its core grain, even while incorporating data from other entities like users or products. This granular approach provides flexibility for downstream analysis without pre-aggregating data or answering specific business questions within the transformation layer itself. In modern data warehousing, where storage costs less than compute, marts embrace denormalization. Rather than maintaining strict normalization like traditional Kimball star schemas, marts borrow and add relevant data from other concepts. Building the same data in multiple places (such as including order information within a customers mart) proves more efficient than repeatedly rejoining these concepts during analysis. This represents a fundamental shift in how data teams balance storage against computational efficiency. ### Structural characteristics of marts Marts follow consistent organizational patterns within dbt projects. They typically group by department or business area (finance, marketing, operations) rather than by source system. File naming uses plain English based on the entity: `customers.sql`, `orders.sql`, `payments.sql`. This clarity helps anyone navigating the project understand what each mart contains. The models themselves are materialized as tables or [incremental models](https://docs.getdbt.com/docs/build/incremental-models-overview) rather than views. This approach gives end users faster query performance and reduces costs by avoiding repeated computation of entire model chains. Teams generally start with views, move to tables when query performance degrades, and finally implement incremental materialization when full table rebuilds become too slow. Marts avoid building the same concept differently for different teams. Creating `finance_orders` and m`arketing_orders` typically signals an anti-pattern, though legitimate exceptions exist (such as when finance requires government reporting that diverges from standard revenue measurement). The key is ensuring these represent genuinely separate concepts rather than departmental perspectives on identical data. ## Understanding data products Data products represent a fundamentally different concept. Rather than describing a technical layer in the transformation pipeline, data products describe how teams manage and reason about their most important business processes in data. A data product groups related data assets (often spanning multiple marts, staging models, and source tables) into a cohesive unit with clear ownership, quality standards, and business purpose. The data product lens provides a management framework for business-critical data. Teams organize data products into functional areas like business intelligence, finance, or operations. Each product receives a priority level indicating its importance, which determines how urgently teams address issues. This organizational structure brings transparency: stakeholders can instantly see whether errors exist on or upstream of critical data products. ### Ownership and accountability Data products establish clear ownership patterns that extend beyond the data team. While analytics engineers might own core transformation models, ownership distributes across the organization based on business function. Finance teams receive notifications about issues with finance data products. Operations managers get alerted when problems affect operational datasets they're responsible for maintaining. This distributed ownership model changes how organizations think about data quality. Rather than treating data quality as solely the data team's responsibility, data products create shared accountability. Business stakeholders become active participants in maintaining the reliability of data assets they depend on. ### Reliability and monitoring Data products incorporate comprehensive monitoring and testing strategies. Teams combine dbt tests with automated anomaly detection to catch both known issues (through explicit tests) and unexpected problems (through pattern-based monitoring). This dual approach significantly reduces detection time; organizations report resolving most issues within minutes rather than hours. The reliability workflow for data products spans detection, resolution, and learning. When issues occur, teams declare incidents that become part of a knowledge base. This historical record helps teams respond faster when similar problems recur, turning incident management into organizational learning. ## Key differences in practice The distinction between marts and data products manifests in several practical ways. Marts represent technical artifacts: SQL files that transform data according to specific patterns and conventions. Data products represent business concepts: groupings of related data assets managed as coherent units with defined quality standards and ownership. Marts exist within a single layer of the transformation pipeline. A customers mart or orders mart lives in the models/marts directory of a dbt project. Data products span multiple layers, encompassing source tables, staging models, intermediate transformations, and multiple marts that together support a business function. Marts follow technical naming conventions like `dim_customers` or `fct_orders` that indicate their role in dimensional modeling. Data products use business-oriented names that reflect their purpose: "Customer Analytics," "Financial Reporting," or "Operational KPIs." The scope differs fundamentally. A single data product might include dozens of marts alongside their upstream dependencies. Conversely, a single mart might contribute to multiple data products serving different business functions. ## Architectural implications These differences shape how teams architect their data platforms. Marts require careful attention to grain, denormalization strategy, and materialization approach. Teams must decide when to build separate marts versus when to reference existing ones, balancing performance against maintainability. Data products require governance frameworks that marts alone don't address. Teams need processes for defining product boundaries, assigning ownership, setting quality standards, and managing incidents. This governance layer sits above the technical transformation work, providing business context and accountability structures. The relationship between marts and data products isn't hierarchical; it's complementary. Well-designed marts provide the building blocks for reliable data products. Clear data product definitions guide decisions about which marts to build and how to structure them. Teams need both concepts to deliver trusted data at scale. ## The role of the Semantic Layer The distinction between marts and data products becomes more nuanced when considering [dbt's Semantic Layer](https://www.getdbt.com/product/semantic-layer). Without the Semantic Layer, teams typically build heavily denormalized marts optimized for direct consumption by business users. With the Semantic Layer, teams maintain more normalized marts, allowing MetricFlow flexibility in how it combines and aggregates data. This architectural choice affects how data products are constructed. In projects without the Semantic Layer, data products might consist primarily of wide, denormalized marts ready for immediate analysis. With the Semantic Layer, data products include normalized marts plus the semantic definitions that specify how to calculate metrics and combine entities. ## Practical guidance for data leaders [Data engineering](https://www.getdbt.com/blog/data-engineering) leaders should think about marts and data products as addressing different concerns. Invest in marts when focusing on transformation architecture: how to structure models, what grain to maintain, how to optimize performance. Invest in data products when focusing on business value: what data assets matter most, who owns them, how to ensure reliability. Both concepts deserve attention, but at different times and for different purposes. Early in a data platform's lifecycle, establishing solid mart patterns matters most. As the platform matures and complexity grows, layering on data product thinking helps manage that complexity and maintain business alignment. The most successful data organizations use marts as their technical foundation and data products as their management framework. Marts provide the modular, well-tested building blocks. Data products provide the business context, ownership, and quality standards that turn those building blocks into trusted assets stakeholders rely on. Understanding this distinction helps data leaders communicate more effectively with both technical teams and business stakeholders. When discussing transformation architecture with analytics engineers, talk about marts. When discussing data strategy with business leaders, talk about data products. Both conversations matter, but they serve different purposes in building a mature data platform. #### Related resources: - [Marts: Business-defined entities](https://docs.getdbt.com/best-practices/how-we-structure/4-marts) - Learn best practices for structuring marts in dbt projects - [Understanding data modeling](https://www.getdbt.com/blog/understanding-data-modeling) - Explore the fundamentals of data modeling and transformation - [Modular data modeling techniques](https://www.getdbt.com/blog/modular-data-modeling-technique) - Discover how to build maintainable, modular data models - [Building reliable data products](https://www.getdbt.com/blog/building-reliable-data-products-dbt-cloud-synq) - See how organizations implement data products in practice ## Data mart vs Data product FAQs **How do data marts differ from data products in purpose, ownership, and lifecycle?** Data marts are technical artifacts that represent a specific layer in the data transformation pipeline, containing business-defined entities at a unique grain (like customers or orders). They follow technical naming conventions and exist within a single transformation layer. Data products, in contrast, are management frameworks that group related data assets (often spanning multiple marts, staging models, and source tables) into cohesive units with clear business purpose. While marts typically have technical ownership within analytics engineering teams, data products establish distributed ownership across the organization, with business stakeholders like finance teams or operations managers receiving accountability for their respective data domains. Data products also incorporate comprehensive monitoring, incident management, and quality standards that extend beyond the technical scope of individual marts. **When should an organization choose a departmental data mart over a domain-owned data product?** Organizations should focus on building marts when addressing transformation architecture concerns: determining how to structure models, what grain to maintain, and how to optimize performance. This is particularly important early in a data platform's lifecycle when establishing solid technical foundations matters most. As platforms mature and complexity grows, layering on data product thinking becomes valuable for managing that complexity and maintaining business alignment. The choice isn't either/or: marts serve as the technical foundation providing modular, well-tested building blocks, while data products provide the business context, ownership structures, and quality standards that turn those building blocks into trusted assets. **What quality, discoverability, and self-service guarantees do data products provide that traditional data marts typically lack?** Data products incorporate comprehensive monitoring combining explicit dbt tests with automated anomaly detection to catch both known and unexpected issues, significantly reducing detection time from hours to minutes. They establish reliability workflows spanning detection, resolution, and learning, with incidents becoming part of a knowledge base for faster future responses. Data products also provide clear priority levels indicating business importance, which determines how urgently teams address issues, bringing transparency so stakeholders can instantly see whether errors exist on critical data assets. This governance layer (including defined product boundaries, assigned ownership, quality standards, and incident management processes) sits above the technical transformation work that marts handle, providing business context and accountability structures that marts alone don't address. --- --- title: "What the Open Semantic Interchange (OSI) spec means for metrics, semantics, and AI" description: "An open OSI spec for semantic portability is now available, plus dbt's commitment to operationalize semantics." url: "https://www.getdbt.com/blog/the-osi-spec-updates" date: "2026-01-29" authors: ["Dave Connors"] categories: ["Product"] --- # What the Open Semantic Interchange (OSI) spec means for metrics, semantics, and AI The pace of AI adoption across the data ecosystem is accelerating, and semantic context is quickly becoming critical infrastructure for trustworthy AI-enabled analytics. At Coalesce 2025, we shared our [commitment to open, interoperable semantics](https://www.getdbt.com/blog/dbt-labs-affirms-commitment-to-open-semantic-interchange-by-open-sourcing-metricflow): we introduced [MetricFlow as an open-source project](https://www.getdbt.com/blog/open-source-metricflow-governed-metrics) and announced dbt Labs’ participation in [the Open Semantic Interchange (OSI)](https://open-semantic-interchange.org/?_fsi=JBu1651l), alongside partners like Snowflake, Databricks, and Salesforce, working toward a shared specification. Since then, dbt Labs has worked closely with the OSI partners to publish an initial OSI specification, so businesses can standardize semantic context before semantics get defined differently across tools. Today, the [first version of the OSI specification](https://github.com/open-semantic-interchange/OSI) is available in an open-source, Apache 2.0–licensed repository. OSI defines a vendor-neutral, extensible model for representing semantic layer constructs, including datasets, metrics, dimensions, relationships, and context, so they can be interpreted consistently across tools, platforms, and AI-enabled applications. ### Why OSI matters to analytics engineering teams Organizations often define business metrics in multiple places. If you've ever rebuilt the same metrics across tools, debated which "active users" definition is correct, or watched dashboards drift after silent changes, you understand the problem OSI is designed to address. OSI provides a vendor-neutral interchange format for semantic definitions—metrics, dimensions, datasets, relationships, and the context needed to interpret them—so those definitions can be transferred between tools without being re-authored or re-interpreted. ### dbt Labs' focus: Operationalizing semantics dbt complements the OSI spec by making semantic definitions operational: teams can define and govern metrics in the dbt Semantic Layer, and execute them consistently with MetricFlow. OSI provides the interchange format to move those definitions across tools (like Snowflake and Tableau). Together, they let teams author once and adopt interoperable semantics incrementally. Looking ahead, dbt is committed to using the new OSI spec to power semantic interoperability from the dbt Semantic Layer across the broader analytics and AI ecosystem. Keep an eye out for updates! ### A shared goal is semantics you can trust anywhere We're proud to be collaborating on OSI and committed to building in the open with the OSI community and partners. The specification will only get better with real-world input from teams defining metrics today. If you want to participate in OSI, the best ways to engage include: - Visit the [OSI project site](https://open-semantic-interchange.org/?_fsi=JBu1651l) to learn more. - [Review the OSI repository](https://github.com/open-semantic-interchange/OSI?_fsi=JBu1651l) and the current specification. - Start a [discussion](https://github.com/open-semantic-interchange/OSI/discussions) in the OSI community forums if you have a use case or proposal. - Contribute [directly](https://github.com/open-semantic-interchange/OSI?_fsi=JBu1651l) via pull requests, including tooling and converters. --- --- title: "What makes a great dbt Summit presentation? (And why yours belongs on stage)" description: "Share your data story at dbt Summit. CFP now open. Real problems, real solutions. Submit by March 31." url: "https://www.getdbt.com/blog/what-makes-a-great-dbt-summit-presentation" date: "2026-01-27" authors: ["Daniel Poppy"] categories: ["Community"] --- # What makes a great dbt Summit presentation? (And why yours belongs on stage) The call for papers (CFP) for dbt Summit is now open, and we hope to hear from you. You might notice a new name for this event. Coalesce is now dbt Summit. Same event, same vibe, new name. dbt Summit is the world's largest gathering of dbt users to shape the future of data & AI. We’re looking for real stories from real practitioners. If that’s you, the dbt community wants to hear your story. ## Who we are looking for You don’t need to be a data superstar to speak at dbt Summit. But spoiler: you might become a data superstar after you speak at dbt Summit. We welcome, and encourage, first-time speakers. And we value the diverse perspectives that reflect the dbt community. dbt is the trusted standard for analytics engineering in the age of AI, and is used by 80,000+ companies and one million engineers. From multinational companies to data teams of one, we want practitioners and data leaders with real problems and real solutions that improve data quality, accelerate workflows, and optimize data costs. ## What makes a great proposal The dbt Summit call or papers is open through March 31. We want you to share your data story, and we want to make sure you are able to put a spotlight on the work you do. We have two great options for sessions: ### Breakout sessions - **Format**: Traditional knowledge sharing sessions with slide presentations - **Duration**: 30 minutes - **Delivery style**: Presented content (can be single presenter, co-presenter, or 3-person max panel) - **Speakers**: Can be delivered by dbt community members, dbt Labs employees, or sponsors - **Purpose**: Share best practices, technical deep dives, use cases, and how-tos ### **Peer exchange sessions** - **Format**: Facilitated, discussion-based sessions focused on peer-to-peer learning - **Duration**: 60 minutes - **Delivery style**: Interactive conversations with **no formal presentation slides** - **Purpose**: Shared problem-solving and exchange of real-world experiences to foster community building and knowledge sharing ### At its core, a great dbt Summit session proposal: 1. Starts with a specific problem 2. Shows how you solved (or tried to solve) the problem 3. Identifies what you learned along the way and what you’d do differently in hindsight, and 4. Teaches the audience something new that they can take back to their teams. It’s important to **keep it real**. Be authentic, use your own words and lived experience. **Avoid marketing speak**; we’ve all heard enough of that for a lifetime (and I say that as someone in marketing). **Think about the “before and after” of your story**, and even better if you can provide a measurable impact: data platform optimization metrics, hours saved resolving data errors, whatever you’ve been able to demonstrate to your organization that your data team is making an impact. Steer clear of product pitches; dbt Summit attendees want to learn. Speaking of attendees, think about your audience in your proposal. See if you can answer: - Who’s this presentation for? - Why will they stick around for your talk? - What’s the ONE THING you want attendees to walk away knowing or being able to do? ## **Session topics we’re excited about** There are so many great data stories we’ve seen and heard at dbt events. [Take a look at some of last year’s sessions](https://www.youtube.com/playlist?list=PL0QYlrC86xQmBW4S7JRCBeVQP7q2dFH4I). This year there are a few areas we want to focus on, based on the requests of the dbt community: - Analytics development best practices - Data modernization (those "before and after" stories) - Scaling dbt at the enterprise - AI and the structured context layer - Building agentic workflows - Self-service analytics with governance - Cost optimization and efficiency There’s so much room to work in those topics. Regardless of the specific challenges and wins of your data team, there will be a lot relatable to other data teams. Just remember to keep it specific and show concrete steps. ## **Never done this before? Perfect.** Every year, we get data practitioners and leaders speaking about their work publicly for the first time. We love this. It’s an absolute joy to put a spotlight on the great work being done, and we’re here to help you share your data story. dbt Summit track leaders work with speakers to develop their sessions, provide feedback during session prep, and help ensure your talk will have the highest impact possible with the dbt community. Peer exchange sessions will be paired with a dbt Labs employee to help moderate the discussion. The community wants YOUR unique perspective. Not sure if your idea fits? Join #dbt-summit-pitch-party in dbt Community Slack. ## **What you get as a dbt Summit speaker** - First, a comped dbt Summit conference pass (value: $1,695). - Access to a dbt Summit track leader to help your story shine. - Exclusive promotion kit for your network to highlight your work - And most importantly, a platform to share your data story with the largest dbt community gathering in the world. ## Submit your dbt Summit talk now The dbt Summit call for papers (CFP) is now open. Don’t wait, submit your talk now. the CFP closes March 31. Submit here: [https://sessionize.com/dbt-summit2026](https://sessionize.com/dbt-summit2026) Remember, the dbt community wants to hear from YOU. We can’t wait to hear your story. Questions? Reach out to [dbtsummit@dbtlabs.com](mailto:dbtsummit@dbtlabs.com). --- --- title: "Centrally defined metrics: The key to AI success" description: "Spaghetti metrics won’t work in the age of AI. Here’s how to create and manage a centralized, consistent metrics framework." url: "https://www.getdbt.com/blog/centrally-defined-metrics" date: "2026-01-26" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Centrally defined metrics: The key to AI success Everyone wants to leverage artificial intelligence to accelerate data-driven decision-making. The blocker, as always, is the data itself. While generative AI is new tech, it doesn't change the age-old computing principle of GIGO (Garbage In, Garbage Out). AI outputs are only as good as the data inputs. That's why, according to a [Salesforce study on the state of data and analytics](https://www.salesforce.com/content/dam/web/en_us/www/documents/research/state-of-data-analytics.pdf), 86% of business leaders believe high-quality data is essential for AI success. Even before the advent of AI, however, your business has likely struggled with generating accurate metrics. Every team has its own approach to calculating fundamental metrics. That's confusing even for humans. When conflicting metrics are fed to AI systems, it can lead to poor decision making and ultimately stall AI adoption. Before launching AI initiatives, it's imperative to have a single source of truth for your company's key metrics. In this article, we'll discuss how to build a centralized, governed semantic layer that feeds into your AI solutions and sets you up for success. ## Why AI needs high-quality data and context Artificial intelligence and machine learning use cases require high-quality, structured data, and governed context to function properly. Unlike traditional BI, where humans review dashboards and can spot obvious errors before making decisions, AI models (particularly AI agents) often act autonomously through automation. This makes subtle errors far more risky. Our [2024 State of Analytics Engineering report](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024) found that the majority of analytics engineers are now managing data for AI model training. This shift reflects how quickly AI adoption has moved from experimental to production-critical across industries. It's critical to get this data right. AI decisions now affect customer-facing products, automated systems, and business strategy. When an AI chatbot pulls from inconsistent revenue definitions across departments, it might tell Sales that Q3 was a banner quarter while telling Finance that performance was flat. These contradictions erode trust not just in the AI system, but in your data infrastructure as a whole. Consider what happens when an AI-powered pricing algorithm uses outdated cost metrics, or when a customer service chatbot references product availability data that hasn't been refreshed in hours. The speed and scale at which AI operates means these errors can affect thousands of decisions before anyone notices. These errors can also be hard (if not impossible) to track down once injected. That's because AI systems such as [Large Language Models (LLMs)](https://www.cloudflare.com/learning/ai/what-is-large-language-model/) are [complex neural networks that operate stochastically](https://dev.to/kuldeep_paul/how-to-debug-llm-failures-a-practical-guide-for-ai-engineers-5c0b). It's not always clear how they arrive at their decisions—and developers can't step through their decision-making processes using normal application debugging techniques. What makes this especially challenging is that AI doesn’t just need “clean data.” It needs **context it can trust**: what a metric means, how it’s calculated, what it can be sliced by, what it can be joined to, and whether it’s approved for use. Without governed context, AI systems will fill in the gaps, often incorrectly. ## The root problem: metrics chaos Modern data infrastructure has become increasingly complex. According to [Forrester research](https://www.forrester.com/blogs/the-bi-fabric-baby-is-slowly-but-surely-growing-up/), 61% of organizations use four or more business intelligence tools, and 25% use ten or more. Each tool becomes a silo with its own metric definitions, creating what we call "metrics chaos." Your revenue metric is calculated one way in Tableau, another way in Looker, and yet another way in your custom applications. Each definition made sense at the time it was created, but as they've diverged, no one can definitively say which is correct. Finance uses one number, Sales uses another, and your executive dashboard shows a third. This chaos has several root causes that compound over time. ### Version drift Metrics evolve as your business changes. You might update how you calculate customer lifetime value to exclude certain transaction types, or modify your churn definition to better reflect your new subscription model. The problem is that old definitions persist in some tools while new ones appear in others. There's no single source tracking which definition is current, which is deprecated, and which teams are using which version. Your marketing team might be optimizing campaigns based on metrics that your data team abandoned months ago. ### Duplication and inconsistency Teams repeatedly rebuild the same metrics across different tools and platforms. Each time they rebuild, subtle differences in calculation logic creep in. Maybe one version rounds differently, or includes a slightly different set of transaction types, or uses a different lookback window. The names might differ, too. What Finance calls "gross revenue" might be what Sales calls "total bookings" and what the executive team knows as "sales volume." Or the names might be obscure and technical—like metric_rev_001—rather than defined in clear business language. When stakeholders see conflicting KPIs in different dashboards, trust in data erodes quickly. ### The maintenance nightmare When metrics definitions are scattered across multiple systems, updating them becomes a monumental task. Changing a single metric might require touching five different tools, each with its own syntax, testing environment, and deployment process. Testing changes becomes impractical at scale. You'd need to verify that the updated metric works correctly in every system that references it, and that downstream reports and dashboards still function as expected. Teams start avoiding making necessary updates simply because the complexity isn't worth the effort. Your metrics calcify, becoming increasingly disconnected from business reality. ## Why traditional BI doesn't work for AI Traditional business intelligence tools were designed with a [human-in-the-loop](https://hai.stanford.edu/news/humans-loop-design-interactive-ai-systems) approach. Someone reviews a dashboard, considers the numbers in context, spots anomalies, and decides whether to take action. Humans serve as the final filter for data quality. [AI agents](https://www.ibm.com/think/topics/ai-agents) eliminate that filter. They make decisions and take actions without human review, at a speed that makes manual oversight only possible after the fact, via auditing and logging. An AI-powered pricing algorithm might adjust thousands of products before anyone realizes it's using outdated cost data. A marketing AI might send millions of personalized messages based on incorrect segmentation rules. This speed of decision-making means errors propagate instantly across your business. Traditional BI tools also weren't designed to provide the explicit, machine-readable metric definitions that AI tools need. A dashboard might be sufficient for a human analyst. By contrast, an AI agent needs structured metadata about what each metric means, how it's calculated**, **what it can be sliced by, what it can be joined to, and when it's appropriate to use. The [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) emphasizes treating data as software—with versioning, testing, and deployment processes. AI implementation exposes technical debt in your data infrastructure that was previously invisible. When humans were the consumers, they could work around data quality issues. AI systems can't. ## Solution: Centralized, governed metrics via a semantic layer A [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) provides a unified, business-friendly representation of your data. Think of it as a translation layer that sits between your raw data and your analytics tools, defining metrics once in clear business terms. But for AI, the goal isn’t “just a semantic layer.” What AI actually needs is governed semantic models: machine-readable business definitions that are version-controlled, tested, documented, reviewed through CI/CD, access-controlled, and auditable, so AI systems can rely on approved logic instead of guessing. The architecture follows a hub-and-spoke model: you define metrics once in the semantic layer, and they're consumed everywhere—from traditional BI tools to embedded applications to AI agents. Metrics are defined in code alongside your data transformations, making them version-controlled and testable using the same principles you apply to your data pipelines. The semantic layer exposes these metrics via APIs and the Model Context Protocol (MCP) server to all downstream consumers. For example, you might define your revenue metric once. That same definition then powers your Tableau dashboards, your embedded analytics application, and your AI chatbot. When you need to update the definition, you change it in one place, and the change propagates everywhere. This means these governed definitions stay interoperable as your BI tools, applications, and AI systems evolve, so you don’t have to re-implement metric logic every time your stack changes. This centralization doesn't just reduce duplication—when it’s combined with governance, it makes metrics governance practical. With governed semantic models, definitions live in one place where they can be version-controlled, tested, documented, reviewed and deployed through CI/CD, access-controlled, and auditable. You can track which definition is current, who approved it, when it was last updated, and which systems are consuming it. ## The benefits of centrally defined, governed metrics for AI Centralized metrics unlock several critical capabilities for AI systems that directly impact business value and operational efficiency. **Consistency** means AI agents always access current, approved metric definitions. Your AI doesn't need to guess which revenue calculation is correct or reconcile conflicting definitions. It uses the canonical version that your data team has validated and governance team has approved. **Governance** becomes granular and practical. You can implement role-based access controls at the metric level, ensuring that sensitive financial metrics are only available to authorized AI systems. You can require approval workflows for metric changes, and audit who accessed which metrics and when. **Auditability** provides the transparency you need for regulated industries and high-stakes decisions. You can track exactly which performance metrics your AI systems are using and when they access them. If an AI makes a questionable decision, you can trace it back to the specific metric definitions and data it relied on. **Velocity** improves dramatically. You can ship new metrics to AI applications without rebuilding integrations. Your AI chatbot automatically gains access to new product metrics the moment you add them to your semantic layer. There's no separate deployment process for each consuming system. **Interoperability** keeps you flexible. You can adopt new tools—new BI platforms, new agent runtimes, new applications—without re-implementing metric definitions each time. Your governed semantic models remain the stable source of truth. ## Driving centrally defined, governed metrics with dbt dbt acts as a [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction), driving data quality and trust regardless of where your data lives. With dbt, you define semantic models and metrics on top of those governed dbt assets in the _modeling layer_ (your dbt project). Crucially, in dbt these semantic definitions are **in code.** Semantic models and metrics are configured in **YAML files within your dbt project repo.** The workflow follows analytics engineering best practices your team already uses for transformations. You define transformations in SQL, test them quickly on dev machines with [dbt Fusion](https://www.getdbt.com/product/fusion)—which delivers up to 30x faster parsing—then add testing and documentation before deploying changes. You manage deployment through [CI/CD pipelines](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud), ensuring that every metric change is reviewed, tested, and versioned. That’s what turns “a semantic layer” into governed context that can safely power both BI and AI The [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) provides purpose-built tooling for defining a single source of truth for organizational metrics, grounded in semantic models. You specify how to calculate each metric, which dimensions it can be sliced by, and what aggregations are valid. The semantic layer then translates metric requests into the appropriate, optimized SQL for your warehouse, including the joins and aggregations needed to answer a question correctly. Key features include: - **Metric definitions and semantic models as code (YAML)**: Define calculations and manage business logic once using familiar SQL syntax with reviews and change history. - **Multi-dimensional modeling**: Specify which dimensions, entities, relationships, and filters apply to each metric so AI and BI isn’t guessing. - **Query optimization**: The semantic layer generates efficient, correct SQL regardless of how metrics are consumed - **Access controls**: Govern which teams can access which metrics - **Version management**: Track changes to metric definitions over time through Git and CI/CD. You can expose dbt semantic layer governed metrics and context to AI systems using the [dbt MCP Server](https://www.getdbt.com/blog/mcp). This allows language models to query your metrics directly, with the dbt semantic layer ensuring AI systems always use the correct, approved definitions. ## Best practices for getting started with a semantic layer You don't need to centralize every metric on day one. Start with the critical few that drive key decisions and business objectives. **Identify 5-10 metrics** that matter most to your business. These might include revenue, customer acquisition cost, churn rate, and other KPIs that executives review regularly. Focus on metrics where inconsistency currently causes real problems—perhaps, for example, where Finance and Sales regularly disagree on numbers. **Apply ADLC principles** to metric development. Plan which metrics you'll centralize first and why. Develop the metric definitions in code, with clear documentation about calculation logic, and test them against known data to verify accuracy. Then, [use a CI/CD deployment process](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud) to safely deploy changes to production and monitor their behavior. **Integrate with AI workflows** from the start. Expose metrics through APIs that your LLMs can consume. Use connectors such as the [dbt Model Context Protocol (MCP) server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) to connect MCP-aware clients to your semantic layer, ensuring that your AI applications always have access to your single source of truth and governed context. **Define a governance framework** that grows with your semantic layer. Establish clear ownership for each metric and create an approval process for metric changes—perhaps requiring sign-off from both the data team and the relevant business stakeholder. Set up [production testing](https://www.getdbt.com/blog/build-trust-through-data-testing) and [data health signals](https://docs.getdbt.com/docs/explore/data-health-signals) so you can issue alerts on metric quality issues, like sudden unexpected changes in values or failed data tests. ## From firefighting to proactive AI enablement Organizations that want to succeed with AI must first get their metrics under control. A governed semantic layer isn't a nice-to-have. It's foundational infrastructure. Without it, you'll spend your time firefighting data quality issues and explaining why different systems show different numbers, rather than delivering AI-powered insights that drive business value. The good news is that you don't need perfect data to start. Begin with your most critical metrics, apply consistent governance, and expand coverage over time. AI adoption won't wait for you to clean up every data silo and reconcile every definition. But it will expose your data quality issues faster and more publicly than ever before. Start small, but start now. The organizations that thrive in the AI era will be those that build on a foundation of trustworthy, centrally defined, governed metrics. To learn more about how to transform your metrics chaos into AI-ready infrastructure, [contact us for more info](https://www.getdbt.com/contact) on how dbt can help you build the governed semantic layer your AI systems need. --- --- title: "Apache Iceberg and the catalog layer" description: "Everything you ever wanted to know about open table formats with a member of Apache Iceberg and Apache Polaris." url: "https://www.getdbt.com/blog/apache-iceberg-and-the-catalog-layer" date: "2026-01-25" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Apache Iceberg and the catalog layer In this episode of The Analytics Engineering Podcast, Tristan talks with Russell Spitzer, a PMC member of Apache Iceberg and Apache Polaris and principal engineer at Snowflake. They discuss the evolution of open table formats and the catalog layer. They dig into how the Apache Software Foundation operates. And they explore where Iceberg and Polaris are headed. If you want to go deep on the tech behind open table formats, this is the conversation for you. A lot has changed in how data teams work over the past year. We’re collecting input for the [2026 State of Analytics Engineering Report](https://forms.gle/KBU9smukSfiK1g4W7) to better understand what’s working, what’s hard, and what’s changing. If you’re in the middle of this work, your perspective would be valuable. [Take the survey](https://forms.gle/Jc54NuP96qekHU9j7) _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [Youtube](https://www.youtube.com/playlist?list=PL0QYlrC86xQm83Q9deiy4euEnbw8ceu3I) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) [Watch video](https://youtu.be/wLH-vADSwaw) ## Key takeaways ### Tristan Handy: You spend a lot of your time thinking about Iceberg and Polaris. Give the audience background on how you found yourself in this niche of high‑volume analytic data file formats. **Russell Spitzer:** It’s a bit random. I started at DataStax on Apache Cassandra as a test engineer and quickly got drawn into analytics. I saw big compute clusters and wanted to be involved. A coworker, Piotr, noticed Spark 0.9 and began a Spark–Cassandra connector. That got me into Spark. Over six to seven years I focused on moving data between Cassandra and Spark and into other systems. The interoperability problem across distributed compute frameworks was compelling. This was pre‑Apache Arrow and pre‑table formats. We were just putting Parquet files everywhere and no one quite knew what they were doing. Pre‑Spark, people explored DSLs like Apache Pig. Eventually the industry converged on SQL for end‑user interfaces. I later applied to Apple for the Spark team. ### Helping build Apple’s Spark infra, or working directly on Spark? Apple has an open-source Spark team and a Spark‑as‑infra team. I was trying to join the open source team, pushing Apple’s priorities into the project and supporting Spark as a service. During interviews, Anton—another Iceberg PMC—convinced the hiring manager I should join the data tables team, essentially Apple’s Apache Iceberg team. They ambitiously planned to replace lots of internal systems with Iceberg. Iceberg existed but was early (Netflix started it around 2018/2019; I joined Apple in 2020). At Apple it was Iceberg all the time; convincing teams to move off older stacks, adopting open‑source‑as‑a‑service to save money, and getting onto ACID‑capable foundations. We were successful. ### Migrations are hard. How did you make it accessible? We replaced complicated bespoke reliability fixes with Iceberg. In Hive/HDFS, small‑file problems lead teams to write custom compaction and locking. Removing that toil is a big win. For big orgs, migration is a long‑term investment with ongoing engineering cost. For smaller companies, the key is offloading runtime responsibilities—ideally to SaaS—so engineers aren’t in the loop. Open source limits lock‑in so you can move between systems. Most companies are paid to deliver business value, not to build data infra. dbt is a great example of avoiding hand‑rolled pipeline code. Same logic applies to table/file formats. ### Let’s talk Apache governance. What’s a PMC? How do projects run? Apache projects aren’t owned by one company. Influence is earned by contributing to the community. The PMC governs merges, releases, membership. People move companies; the project stays with them. The goal is to make the project broadly useful. There’s no CEO dictating roadmap and no company can change the license. Most big projects—Spark, Kafka, Iceberg, Flink—are maintained by employees of companies with vested interests, but governance is consensus‑driven. Vetoes are for technical issues (security, future‑limiting design), not ideology. ### Is Iceberg for the top 20 tech companies or for everyone? Not everyone needs Iceberg. OLTP belongs elsewhere. But for analytics, we should move past raw Parquet partition trees with folder‑name partitioning. In the Hadoop era, lakes were dumping grounds; schema evolution was painful. Many are still moving from CSV to Parquet. Over time, better encodings and table formats become default. Decoupling compute and storage changes everything versus co‑located HDFS. Defaults tuned for HDFS (like 128MB Parquet files) don’t always hold for S3. We want elastic storage and compute; no one wants to pay for compute because storage grew. ### Walk us through Iceberg versions. v1: transactional analytics—ACID commits instead of fragile Hive/HDFS patterns. v2: row‑level operations—logical deletes via delete files so you don’t rewrite 10M‑row data files to remove one row; later compaction physically purges (key for GDPR). v3: expanded types—geospatial and variant for semi‑structured data; Variant was standardized across vendors and Parquet so everyone can write/read consistently. v4: two thrusts—streaming and AI. Reduce commit latency, make retries faster under contention. Historically writes took 10–20 minutes, so commit latency didn’t matter. For streaming (writes every minute/five), it does. We’re evolving commit and REST catalog protocols so clients can specify intent (add these files, ensure these exist, then delete those) and let the catalog resolve conflicts server‑side. On AI: Iceberg doesn’t yet serve some vector/image‑heavy patterns well. We’re exploring changes in Iceberg, Parquet, or both, without breaking existing tables. ### Talk about Polaris and the catalog layer. Polaris is an Apache incubator project (PPMC). Incubation proves we operate like an Apache project (community‑driven, trademarks donated). Iceberg defines the REST catalog spec/client; Polaris implements a catalog that speaks that spec. Many of us work across projects (Parquet, Iceberg, Polaris), which helps align boundaries. ### Horizon, Polaris, external catalogs—what’s the story? We’re simplifying: Snowflake can act as an Iceberg REST catalog, or you can use an external REST catalog. External can be Polaris (managed by Snowflake or self‑hosted) or another REST implementation. Interoperability means everything talks the same REST. ### What is Polaris trying to be best at? A broad, interoperable lakehouse catalog. It can act as a generic Spark catalog (HMS replacement) and aims to support multiple table/file formats. Architectural choices differ (KV vs. relational storage, where transactions live, policy enforcement vs. recording, identity integration). Polaris aims for base implementations that are pluggable—e.g., AWS/GCP/Microsoft identity. ### Identity and scope—where does the catalog stop? There’s a “business catalog” for discovery/listing versus a “system catalog” that must know table layout to govern access. Polaris can vend short‑lived credentials for the exact directory of a table’s files for a load operation; that requires understanding layout. Purely relational metadata often needs to delegate that decision. ### Will identity/grants slow broad adoption? Possibly. But many once‑complex things become default—compressed files, columnar formats, soon encryption. With collaboration (like Variant), we’ll land broadly accepted patterns. ## Chapters 00:01:30 — Guest welcome and interview start 00:02:00 — Russell’s path: DataStax Cassandra, Spark connector, interoperability 00:05:20 — Joining Apple’s Iceberg team and early Iceberg momentum 00:06:20 — Why migrations resonated: replacing bespoke Hive/HDFS compaction/locking 00:09:10 — Apache governance 101: PMCs, consensus, and corporate influence 00:15:40 — How decisions land without votes; when vetoes apply 00:18:30 — Who needs Iceberg and where it fits 00:22:20 — Lake → lakehouse and warehouse → lakehouse in the cloud era 00:25:20 — Iceberg versions: v1 transactions, v2 row‑level ops (GDPR), v3 types 00:28:10 — Standardizing Variant across vendors and Parquet 00:31:10 — Iceberg v4 goals: streaming commit/retry improvements and AI use cases 00:33:40 — Commit latency and server‑side conflict resolution 00:37:20 — Polaris as an Apache incubating project (PPMC) 00:39:30 — Iceberg REST catalog spec and Polaris implementation 00:42:30 — Clarifying Snowflake Horizon, Polaris, and external REST catalogs 00:45:10 — What Polaris aims to be best at; pluggable identity providers 00:48:00 — Identity scope: business vs. system catalogs and credential vending 00:51:00 — Will identity/grants slow mass adoption? 00:52:50 — Wrap‑up --- --- title: "How Sweetgreen transformed unstructured data into conversational analytics with dbt and AI" description: "Sweetgreen built trustworthy conversational analytics by standardizing data and metrics with dbt and AI." url: "https://www.getdbt.com/blog/sweetgreen-conversational-analytics-dbt-ai" date: "2026-01-21" authors: ["Hrishi Kulkarni", "Chakshu Mehta"] categories: ["Product"] --- # How Sweetgreen transformed unstructured data into conversational analytics with dbt and AI Data is supposed to settle debates. But at Sweetgreen, a popular restaurant chain, it would sometimes start them. Stakeholders either didn’t have access to the right insights, or they ran into inconsistencies. As Sweetgreen looked toward self-serve AI-powered analysis, the need for a consistent foundation became a high priority. They underwent a massive data transformation using dbt to standardize the business logic behind its reporting. By redesigning its data models around the core actions of its business, Sweetgreen created a single source of truth with consistent metric definitions, automated tests, and documentation that made data easier to trust and use. Today, Sweetgreen’s business teams can leverage conversational AI to self-serve data insights by asking questions in plain English and getting consistent, reliable answers. In just one year, trust in the data has grown, and teams are making faster and more confident decisions. Let’s take a look at how Sweetgreen did it. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/5c47073daf5e790dd1afa2ccdd9139716e833b3d-512x277.png) ### A persistent data bottleneck Before implementing dbt, Sweetgreen struggled to access consistent insights. Sankalp Vatsh, Analytics Lead at Sweetgreen, explains that the problem was threefold: - **Multiple sources of truth.** Different ingestion paths and databases often produced different answers for the same metric. “The exact same analytics dashboard might show gross sales as $13 million in one widget and $14 million in another,” says Sankalp. “Instead of understanding what actually drove $14 million in sales that week, the business and data teams would lose days to reconciling numbers.” - **Bespoke data flows.** Every new business question required building a new script, pipeline, or one-off data flow — turning even simple analysis into a bottleneck. “We never had an enterprise data model before,” says Sankalp. “Whenever someone asked the data team a new question, we would just build a new data workstream.” - **Manual processes.** The data team relied on Google Sheets that were manually ingested into data pipelines. If a user made a mistake upstream, the error eventually appeared downstream in the dashboards. With no easy way for stakeholders to trace what went wrong, data analysis became a time-consuming search for errors. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f88b5751a10d15ca44c4252bf4fd3b94d6f042ae-512x319.png) The result was a fragmented data ecosystem that generated inconsistent metrics. And without a clear view into how the business was performing, Sweetgreen couldn’t trust the data — let alone act on it. “In the restaurant industry, every day matters,” says Sankalp. “If it takes even a few days to gather insights, the week is already over and the opportunity to improve performance is gone.” ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/e291e9f5e6676adf9b4829327c217be8b6d1e578-512x311.png) ### Redesigning data models around the business To fix their data issues, the team made the bold choice to overhaul Sweetgreen’s data model from the ground up. “We thought about the core actions of a restaurant: a customer walking in to place an order for a certain product on a certain date,” reflects Sankalp. “That’s what we needed to build the new data model around.” To create its new enterprise data model, Sweetgreen turned to dbt Labs to turn business logic into reusable, tested, and documented models. With dbt, Sweetgreen established a governed foundation that made everything downstream—from KPIs to semantic models—easier to build and maintain: - **Fact tables **captured ‌core business events and served as the source of truth for measures (e.g., dollars sold, units sold). - **Dimension tables **added standardized context and hierarchies (e.g., store → city → region). - **dbt’s Semantic Layer** built on top of these tables to standardize KPIs and aggregation logic so metrics stayed consistent across dashboards and conversational analytics. “dbt helped standardize our KPIs so that we can all work from a single source of truth,” says Sankalp. “For the first time, our business users had direct visibility into the data.” The redesign gave Sweetgreen a true single source of truth for metrics—and, for the first time, business users could see how data flowed from ingestion all the way to the reporting layer. Using dbt’s lineage and documentation, stakeholders could understand where the numbers came from instead of guessing. The data team also added custom, model-specific quality checks on top of dbt’s built-in tests to catch issues early and prevent bad data from reaching the business. The result was consistent metrics, higher confidence, and more self-serve exploration without needing to loop in the data team. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/61dd22ddc6e375362e6bb94c904b839112debb3f-512x262.png) ### Enabling self-service with Claude MCP and the dbt Semantic Layer Next, Sweetgreen turned to the matter of creating a true AI-powered self-service experience for their business users. The data team had tried multiple business intelligence (BI) tools in the past (like Tableau, PowerBI, or ThoughtSpot) but kept running into a deeper cultural problem. “People don't want to learn a new tool,” says Sankalp. “People absolutely are very used to talking to people to get answers. Natural-language tools are the way to move forward.” Sweetgreen chose to use Claude as their AI tool to deliver conversational analytics to their business users. To ensure reliable AI outputs, they leveraged Claude, MCP, and the dbt Semantic Layer so the right domain business users get the right answers without bypassing governance. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/b64d35bb42c426d4ac094f8adde32cd0534ff2cd-512x421.png) In the diagram, the yellow nodes represent the governed semantic metrics, the canonical definitions for KPIs like gross sales, net sales, refunds, and loyalty. These are defined once in the dbt Semantic Layer and composed onto the underlying sales summary model. Because Claude was reading directly from the dbt Semantic Layer (which enforced consistent metric definitions), users received the same accurate, reliable answers grounded in version-controlled, tested metrics, not hallucinations or unverified numbers. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a7246c1de6aab1594581fea7ffa9ac0dccbba7f0-512x284.png) Say that a user wants to know more about Sweetgreen’s recently launched loyalty program. If they ask, “What loyalty analysis can I do?” Claude reviews the available semantic models and provides multiple analyses (e.g., by channel, venue, time, or customer behavior) that stakeholders may not have considered. “Instead of reaching out to us with questions, people can play around with the data themselves to find answers,” says Sankalp. “It has freed us up from spending so much time just answering questions.” ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/7f6bb5603891ef209f43bf71de86b101895b9db9-512x310.png) **** ### Faster insights, greater trust Sweetgreen rolled this out in phases by domain. The data team started with high-impact areas like sales, built the fact/dimension foundation, and then layered semantic metrics on top. Each semantic model stayed scoped (e.g., COGS separate from financial metrics) while still using shared definitions and quality checks. As each domain moved over, it eliminated conflicting numbers from different sources because the dbt Semantic Layer became the single, canonical place to get the metric everywhere, including in Claude. “Data transformation is also about culture change,” says Sankalp. “We started with two or three team champions who would try it out and could show the rest of their team how much faster they were working.” And the impact showed up immediately in time saved. _“By leveraging AI and the dbt Semantic Layer, self-service analysis has become a 30-minute job, compared to the old process of reaching out to the data team and waiting two weeks for the bandwidth to get an answer,”_ says Sankalp. What once required a data-team work order can now be handled directly by business users through AI-powered conversational analytics with Claude and governed semantic metrics. “Compared with a year ago, we’re delivering faster insights and more consistent metrics for the business,” reflects Sankalp. “In that time, the business has gained a lot of confidence in the data we provide.” Since launching, the impact has been clear: teams get answers faster, and the data team is no longer the bottleneck. With governed metrics and consistent definitions, stakeholders see the same numbers across tools, with far less inconsistency and confusion. That shift has turned the data team from gatekeepers into enablers, letting business users self-serve insights with confidence. Next, Sweetgreen plans to migrate the remaining dashboards off the legacy Airflow-based layer onto these dbt models so reporting is consistent end to end—across dashboards, repeatable analytics, and automation. Today, the conversational experience is limited to Claude Desktop and a small pilot group of ~15–20 users. The team is intentionally starting with just 3–5 semantic models as they scale, staying vigilant about governance so AI responses remain grounded in verified metrics and don’t drift into hallucinations. “We are no longer the bottlenecks or the sole bearers of all data questions,” concludes Sankalp. “Instead, we have become enablers.” [Watch video](https://youtu.be/w0BJtge0NpU) --- --- title: "The role of the semantic layer in data governance and security" description: "See how a semantic layer strengthens governance, security, and trusted analytics at scale." url: "https://www.getdbt.com/blog/semantic-layer-data-governance-security" date: "2026-01-21" authors: ["Joey Gault"] categories: ["Pulse"] --- # The role of the semantic layer in data governance and security ## The governance challenge in modern data environments The proliferation of data tools has created significant governance challenges for data teams. Organizations today commonly use four or more business intelligence tools, with a quarter using ten or more. When metric definitions and business logic live within each of these disparate tools, maintaining consistent governance becomes nearly impossible. Different teams develop their own definitions for critical metrics like revenue or customer lifetime value, leading to conflicting reports and eroding trust in data. This fragmentation creates several governance problems. Data teams struggle to track where sensitive information flows across the organization. Updating access controls requires changes across multiple systems. When business logic changes, teams must manually propagate updates to every tool, creating opportunities for errors and inconsistencies. The result is governance that's reactive rather than proactive, with data teams constantly firefighting issues rather than preventing them. ### Centralized governance through the semantic layer A semantic layer fundamentally changes this dynamic by centralizing metric definitions and business logic in a single location. Rather than defining what "active customer" means separately in Tableau, Looker, and Power BI, data teams define it once in the semantic layer. This definition then flows consistently to every downstream tool that queries it. **** This centralization creates a natural governance checkpoint. When all data access flows through the [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction), data teams gain visibility into how metrics are being used across the organization. They can track which teams access which data, identify potential compliance risks, and ensure that changes to business logic propagate consistently everywhere. If a metric definition needs to change, updating it in one place automatically refreshes it across all applications. The semantic layer also enables [version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) for data definitions, treating them like software code. Data teams can track who made specific changes to metric definitions, when those changes occurred, and why they were necessary. This audit trail is crucial for compliance and helps teams understand how business definitions evolve over time. If a change introduces problems, teams can roll back to previous versions while investigating the issue. ### Implementing role-based access controls Security in modern data environments requires granular control over who can access what data. The [semantic layer acts as a centralized enforcement point](https://www.getdbt.com/blog/semantic-layer-introduction) for role-based access controls, implementing permissions that follow users regardless of which tool they use to access data. Consider a global organization with regional sales managers. Through the semantic layer, data teams can define access policies that automatically limit each manager to their own region's data. When the EMEA sales manager logs into any connected BI tool, they see full sales data for European countries but cannot access data from other regions. These same access controls apply whether they're viewing a dashboard, running an ad-hoc query, or accessing data through an embedded analytics application. This centralized approach to access control offers significant advantages over tool-specific permissions. Data teams define security policies once rather than recreating them in every downstream application. When an employee changes roles or leaves the organization, updating their permissions in the semantic layer immediately affects all connected tools. This reduces the risk of orphaned access and ensures consistent security enforcement. The semantic layer can also implement more sophisticated security patterns like data masking. Sensitive fields such as customer personal information can be automatically masked or redacted based on user roles. A marketing analyst might see aggregated customer behavior patterns while a customer service representative sees full customer details needed to resolve support issues. These policies are enforced at the semantic layer level, ensuring they apply consistently regardless of how users access the data. ### Protecting sensitive data at scale As organizations adopt AI and machine learning, protecting sensitive data becomes even more critical. The semantic layer provides guardrails that ensure AI systems query only approved, governed, and contextualized metrics. Rather than giving AI agents direct access to raw data tables, organizations can expose carefully curated metrics and dimensions through the semantic layer. This approach reduces the risk of AI systems inadvertently exposing sensitive information or generating insights based on incorrect data interpretations. The semantic layer enforces the same access controls and business logic for AI applications as it does for human users, creating consistent governance across all data consumption patterns. For organizations subject to regulatory requirements like [GDPR](https://www.getdbt.com/industry/financial-services) or [HIPAA](https://www.getdbt.com/industry/healthcare), the semantic layer provides a centralized point for implementing compliance controls. Data teams can define which fields contain personally identifiable information, implement retention policies, and ensure that data access aligns with regulatory requirements. When auditors ask who accessed specific data, the semantic layer provides comprehensive logs of all data access patterns. ### Balancing accessibility with control The true power of the semantic layer lies in its ability to democratize data access while maintaining strict governance. By translating complex database structures into business-friendly concepts, the semantic layer enables self-service analytics without requiring users to understand SQL or navigate complex data models. A sales manager can build reports using familiar terms like "customer lifetime value" without knowing the underlying calculation logic or which tables to join. This accessibility doesn't come at the expense of control. Behind the scenes, the semantic layer enforces all relevant security policies, applies appropriate data masking, and ensures users only access data they're authorized to see. Business users experience the freedom of self-service analytics while data teams maintain centralized governance. The semantic layer also reduces the risk of users accidentally misinterpreting data. By embedding business definitions and metadata directly into the data model, it provides context that helps users understand what metrics mean and how they should be used. When someone queries "active customer," they see not just the data but also the definition: "A customer who has made a purchase in the last 90 days." This shared understanding reduces errors and improves the quality of insights. ## Implementing governance with dbt [dbt](https://www.getdbt.com/product/what-is-dbt) provides a semantic layer that integrates [governance directly into the transformation workflow](https://www.getdbt.com/product/governance). Data teams define metrics, access controls, and business logic using YAML configuration files that live alongside their dbt models in version control. This approach brings software engineering best practices to data governance, enabling code review, testing, and collaborative development of governance policies. Because [dbt's semantic layer](https://www.getdbt.com/product/semantic-layer) is built on top of dbt's transformation framework, governance policies benefit from the same development workflow as data transformations. Changes to metric definitions or access controls go through pull requests where team members can review and discuss them. Automated tests can verify that governance policies work as intended before they reach production. This systematic approach reduces the risk of governance gaps and ensures that policies are well-documented and understood. The integration with dbt also means that governance extends throughout the data pipeline. Data teams can implement data quality tests, document data lineage, and define access controls all within the same framework. This comprehensive approach to governance ensures that data quality, security, and accessibility are considered at every stage of the analytics workflow. ### The strategic value of governed self-service For data engineering leaders, the semantic layer represents a strategic investment in scalable governance. As organizations grow and data use cases multiply, manual governance approaches become unsustainable. The semantic layer provides the infrastructure needed to govern data at scale while enabling the self-service analytics that business stakeholders demand. This balance between control and accessibility directly impacts business outcomes. When stakeholders trust that they're working with accurate, consistent, and properly governed data, they make better decisions faster. Data teams spend less time responding to ad-hoc requests and more time on strategic initiatives. Security and compliance risks decrease as governance becomes systematic rather than reactive. The semantic layer also future-proofs governance as new tools and use cases emerge. Whether the next wave of analytics involves embedded applications, AI agents, or technologies not yet invented, the semantic layer provides a consistent governance framework that adapts to new consumption patterns. Data teams can confidently enable new use cases knowing that existing governance policies will apply automatically. For organizations serious about data governance and security, the semantic layer has evolved from a nice-to-have to an essential component of modern data architecture. It provides the centralized control point needed to govern data at scale while enabling the accessibility that drives business value. As data continues to grow in volume and importance, this combination of governance and accessibility will only become more critical. ## Semantic layer FAQs **What is a semantic layer and how does it simplify data access and analysis for nontechnical users?** A semantic layer is a centralized abstraction that translates complex database structures into business-friendly concepts, enabling users to work with familiar terms without needing to understand SQL or navigate complex data models. It allows business users like sales managers to build reports using terms such as "customer lifetime value" or "active customer" without knowing the underlying calculation logic or which database tables to join. The semantic layer embeds business definitions and metadata directly into the data model, providing context that helps users understand what metrics mean and how they should be used, reducing errors and improving insight quality while maintaining strict governance and security controls behind the scenes. **How does the semantic layer differ from the data layer?** The semantic layer sits above the data layer and serves a fundamentally different purpose. While the data layer consists of raw database tables and structures, the semantic layer acts as an abstraction that centralizes metric definitions and business logic in a single location. Rather than requiring users to interact directly with complex database schemas, the semantic layer provides a business-friendly interface with consistent definitions that flow to all downstream tools. It serves as a governance checkpoint where all data access is controlled, monitored, and secured, while the data layer simply stores the underlying information without inherent business context or access controls. **What are the benefits of a semantic layer?** A semantic layer provides multiple critical benefits for organizations. It enables centralized governance by defining metrics and business logic in one place rather than across multiple tools, ensuring consistency and eliminating conflicting reports. It implements role-based access controls that follow users across all connected applications, reducing security risks and simplifying permission management. The semantic layer democratizes data access by allowing self-service analytics while maintaining strict governance, enabling business users to work independently without compromising data security. It also provides comprehensive audit trails for compliance, reduces manual work for data teams through automated policy enforcement, and future-proofs governance as new tools and use cases emerge by providing a consistent framework that adapts to new consumption patterns. --- --- title: "Introducing ADE-bench: measuring how AI agents perform data work" description: "Learn how ADE-bench leverages dbt to produce realistic data benchmarks against all popular LLMs and agents." url: "https://www.getdbt.com/blog/ade-bench-dbt-data-benchmarking" date: "2026-01-21" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Introducing ADE-bench: measuring how AI agents perform data work As AI coding agents reshape software engineering, a critical question remains unanswered: how well do these tools actually perform analytics and data engineering work? While benchmarks like [SWE-bench](https://www.swebench.com/) have become standard for measuring coding agent performance in general software development, the data community has lacked a rigorous way to evaluate how models handle the specific challenges of data work. [ADE-bench](https://github.com/dbt-labs/ade-bench), created by Benn Stancil (founder of Mode) in collaboration with dbt Labs, fills this gap. The benchmark evaluates AI agents on real-world analytics and data engineering tasks using actual dbt projects, databases, and the kinds of messy problems data practitioners face daily. Early results show significant performance differences across models and configurations. In particular, they show how [dbt Fusion](https://www.getdbt.com/product/fusion), in conjunction with the [Model Context Protocol (MCP)](https://docs.getdbt.com/docs/dbt-ai/about-mcp), provides substantial improvements in agent accuracy and efficiency. Jason Ganz, Director of Developer Experience and AI at dbt Labs, and Stancil [recently held a webinar to discuss](https://www.getdbt.com/resources/webinars/analytics-data-engineer-bench) why data work needs its own benchmark, how ADE-bench works, and what the initial results reveal about the current state of AI agents in analytics engineering. ## Why data work needs its own benchmark "At the beginning of the year, the question was: are LLMs and agents useful for software engineering?" Ganz explained. "Pretty quickly this year, we’ve passed that. These things are useful. They're very useful. But how exactly they're useful and what exactly they're good at is still something that we don't exactly know." One thing we _do_ know is that benchmarks have become key to measuring coding agent quality. Software engineering has rigorous benchmarks, such as SWE-bench, that every new model release references. These provide rigorous ways to measure how well a model performs specific tasks, allowing teams to evaluate model progress over time. One of the key philosophies of dbt is that it’s critical to apply [the lessons learned from software engineering to data work](https://www.getdbt.com/resources/the-analytics-development-lifecycle). However, there are also many ways that data is _not_ like software engineering - data-specific workflows, tooling, etc. If coding agents are coming to data work, we need rigorous ways to measure their quality. "The key thing is we don't know exactly how this is going to go," Ganz said. "We don't know how fast. We don't know what tasks agents are good at. Right now, data teams are kind of operating blind about how good these models are at actually performing the tasks that you all do every single day." ## The reality of data work versus the demo The gap between how analytics work is often presented and how it actually happens is stark. Every demo follows the same script for analytics projects: you look at a chart, see some numbers, point out a weird blip, write SQL queries, do fancy analysis, pull up tables, find insights, then publish a great report. The business turns around, and everyone's happy. These demos always feature tight, well-structured problems where everything is neat and clean, with amazing needle-in-the-haystack solutions that solve business problems. But for most people in data work, this isn't reality. "Typically, instead of looking at dashboards with this amazing little outlier, you look at very broken-looking dashboards," Stancil explained. "Or dashboards that don't really say anything. Or dashboards where there's just a thing that is broken, and the actual problem you have is your boss reaches out to you and says it's broken, and you have to fix it." Software engineers can now tackle similar debugging challenges with AI tools. That’s left data engineers and analysts asking, “Can we have one of those, too?” Several companies are trying to solve this problem. But the real answer to whether these tools can solve data problems is that we don't know. We're not actually sure whether these things work. Individual companies might have things that work quite well for them, but, broadly, we don't have a clear sense of how well these solutions can solve data problems. ## The gap in existing benchmarks [When Claude released Opus 4.5](https://www.anthropic.com/news/claude-opus-4-5), their benchmark results included nothing related to data tasks. The same was true for [Gemini](https://gemini.google.com/) 3 and [GPT-5](https://openai.com/index/introducing-gpt-5/). Until we have something that helps us understand whether models are good at solving data problems, everyone has to test individually. Some existing benchmarks attempt to help. [Spider](https://spider2-sql.github.io/) is one of the major text-to-SQL benchmarks, featuring problems with basic data sets. They ask simple questions like, "What is the most popular product in California?" "The work that they have done has been tremendously helpful to figure out how well these things can write SQL," Stancil acknowledged. "But again, this isn't that representative of the types of business problems that most people are trying to solve. In reality, we have projects with hundreds of tables and broken projects." ## How ADE-bench works [ADE-bench](https://github.com/dbt-labs/ade-bench) is designed to represent real-world data work with messy projects featuring hundreds of tables, often with multiple tables that could conceivably represent the same entity. This stands in contrast to the tight schemas designed as toy problems for AI. The benchmark consists of three major components. First, every task has a [dbt project](https://docs.getdbt.com/docs/build/projects) complete with [models](https://docs.getdbt.com/docs/build/models), [macros](https://docs.getdbt.com/docs/build/jinja-macros), and [configuration files](https://docs.getdbt.com/reference/configs-and-properties) representative of real-world setups. Second, there are databases with actual data. Many problems aren't just in the dbt project itself. Sometimes the data itself is broken, or data types don't match. ADE-bench includes [DuckDB](https://duckdb.org/) databases—essentially collections of [Parquet](https://parquet.apache.org/) files—that provide data to run projects on. Third, there are the tasks themselves designed to be solved by an agent. These tasks correspond to real-world data problems - some specific, some vague. For example, a task might simply state "it's broken," similar to what an executive might say to a Slack bot. When you run ADE-bench, the system spins up a [Docker container](https://www.docker.com/resources/what-container/) that creates a sandbox environment with the dbt project, database, and configuration. The task is then presented to the chosen agent, which attempts to solve the problem. The agent works in the sandbox to figure out what's wrong, using whatever agent/LLM you designate. When it completes, it brings back its solution. ADE-bench then tests whether the solution works using primarily [dbt tests](https://docs.getdbt.com/docs/build/data-tests). Tests can check if an orders total table has correct metrics, verify that a materialization changed from a table to a view, or compare agent-generated models to answer keys. You can also run scripts before or after the agent runs to inject problems, such as removing commas, deleting intermediate models, or introducing chaos representative of real-world scenarios. ## Performance results across models and configurations ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/967c9e63089bfd79614bd3a285edb9ed2285074f-2172x1212.png) We ran ADE-bench against a wide variety of LLMs, as well as against different combinations of [dbt Core](https://docs.getdbt.com/docs/core/installation-overview), Core with the [Model Context Protocol (MCP) server](https://docs.getdbt.com/docs/dbt-ai/about-mcp), and [the dbt Fusion engine](https://www.getdbt.com/product/fusion) with the MCP server. ADE-bench assessed both test pass rate, total cost of running the tests, and the total runtime. The results show that using [OpenAI's Codex](https://openai.com/codex/) coding agent with GPT-5.1, the tool passed 56% of the tests, at a cost of around $14.90. But the absolute pass rate isn't the key metric—it's about establishing a baseline for comparison. "I would not consider the 56 percent here to be meaningful," Stancil said. "The point is a baseline of how well things do relative to one another." [Claude Sonnet 4.5](https://www.anthropic.com/news/claude-sonnet-4-5) performed almost identically to Codex with the exact same pass rate and nearly identical cost. When testing across different database setups, [Snowflake](https://snowflake.com) proved slightly harder than DuckDB, which makes sense because agents don't have direct access to the underlying data and must wait for queries to return. Using dbt Core with the MCP server improved the pass rate from 50% to 54%. When using dbt Fusion instead of dbt Core, the pass rate went up significantly because Fusion allows agents to interact with data more effectively. "You can see that in the run times," Stancil noted. "It actually spends a little bit more work doing this, but it actually solves at a higher rate." Combining Fusion with MCP didn't increase the pass rate further but made the process faster. Fusion provides the biggest gains in pass rates, while MCP provides the biggest gains in efficiency. Testing different model versions showed clear progression. Claude Haiku 3, an older model, only solved 10 percent of tasks. [Haiku 4.5](https://www.anthropic.com/claude/haiku) had a much higher pass rate but still lower than more expensive models. Opus matched Sonnet 4.5's pass rate, though thinking models tend to spiral and over-solve some problems. ## What's next and how to contribute [ADE-bench is available as an open-source project on GitHub](https://github.com/dbt-labs/ade-bench), providing a flexible scaffold for adding new tasks, projects, and trying different approaches. Contributing tasks is one of the most valuable ways to help. More tasks that are representative of real-world problems make the benchmark better. Contributing new projects and databases is also tremendously helpful, though more complex due to permissions and privacy considerations. The community can also contribute agent configurations. While most people won't build foundational models, adding context within environments can test how well agents solve problems. This might include documentation that guides the agent in thinking about analytics problems, or additional research steps before attempting solutions. dbt Labs is focused on two key areas: improving the dbt language framework and engine, and building better MCP tooling. The dbt Fusion engine provides a substantial boost in agent accuracy due to its SQL comprehension capabilities. Future improvements include column-level lineage, better schema information, and enhanced type information. "This is not something that we expect to figure out on our own," Ganz said. "This is something that we expect to figure out in partnership with the teams actually solving this." For teams interested in deploying agents, ADE-bench lessons can help you understand what you can do to make your agents perform well. For the broader industry, having a shared sense of what models are good at and what they struggle with helps build systems that deploy models safely and securely. The ADE-bench community meets in [the dbt Slack channel #tools-ade-bench](https://www.getdbt.com/community/join-the-community). You can try out the benchmark yourself and contribute to shaping how AI transforms data work over the coming years. Meanwhile, to see ADE-bench in action and hear answers to pressing questions_ _asked by other customers, [watch a replay of the webinar](https://www.getdbt.com/resources/webinars/analytics-data-engineer-bench). --- --- title: "29%+ warehouse savings: How the dbt Fusion engine drives cost efficiency" description: "You’re probably overspending on data compute. We know we were. Here’s how we fixed it." url: "https://www.getdbt.com/blog/dbt-fusion-cost-efficiency" date: "2026-01-21" authors: ["Kathryn Chubb"] categories: ["Product"] --- # 29%+ warehouse savings: How the dbt Fusion engine drives cost efficiency Data pipelines have a cost efficiency problem. We should know. We deal with it first-hand ourselves at dbt Labs. Every day, we run over 895,000 scheduled jobs. Most run every hour. The problem is, most don't _need_ to be run. The underlying data hasn't changed. That leads us to rebuild thousands or even tens of thousands of [dbt models](https://docs.getdbt.com/docs/build/models) that just output the same results. That wastes compute costs—and it wastes the time of data engineers who spend cycles managing jobs instead of improving the models themselves. That's why we looked to state-aware orchestration as a solution. With state-aware orchestration, you rebuild a model only if it needs to be rebuilt. The potential cost savings with this approach are massive, running anywhere from 10% up to 64%. Implementing state-aware orchestration required, however, that we rewrite the dbt engine from the ground up. Let's take a look at how the old approach to building data models leads to waste, how state-aware orchestration helps, and how [the new dbt Fusion engine](https://www.getdbt.com/product/fusion) uses state-aware orchestration to cut development cycles and slash compute costs. **** ## The orchestration problem—or, why freshness doesn't have to mean waste dbt models map one or more source tables to new destination tables with relations between them, [creating a directed acyclic graph (DAG)](https://www.getdbt.com/blog/dag-use-cases-and-best-practices). Today, when you run dbt build even on a single DAG, it rebuilds every model in the DAG. In the model below, the stg_orders and stg_customers models in our staging layer, the int_orders model in our intermediate layer, and the dim_customers and cust_orders models in our mart layer are all rebuilt, regardless of the state of the underlying source_orders and src_customers tables. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/27901fabeff45263e72732834cd01cc215f4036f-1438x515.png) State-aware orchestration works differently. With state-aware orchestration, we keep a cache at the model level. This means we can recognize that, in the current model, we only need to rebuild three of the models, whereas we can reuse the other two (stg_orders and int_orders) whose source data hasn't changed. This leads to a whole lot less wasted compute. We've seen that our customers are able to reuse 30% of their models using this approach. ## 10% savings just by using Fusion The question is, how is the dbt Fusion engine able to do this? Fusion moves dbt from being a stateless piece of software to being a stateless engine with real-time model state. What this means is that, as jobs run, we keep a real-time cache of the entire environment with each model's hash data and code state. If nothing's changed, we simply don't build it, and we get lower costs across every build. It also means that we can make real-time decisions about what to build as we traverse across the DAG. It no longer matters which job runs which model because all of these are reading and writing from that same shared state. And if two jobs build the same model at the same time, we're actually going to wait for that first one to finish. This reduces the complexity because those models can no longer clash across different jobs. So, in a first build of a DAG with 13 models, you might see 13 models built, taking a total of 49 seconds to run. If you kick off another build just a few minutes later, you might see it takes only 27 seconds for that entire job. That's because we only run the models whose sources have changed. In the case shown below, this means rebuilding only two out of 13 models. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/87cf391307c3a08157333e94032c08893bd8a838-1600x682.png) This is really the power of what we can do with state-aware orchestration. What we've seen is that from our customers and our own data team, simply by turning on state-aware orchestration, they've been able to save 10% plus on the data warehouse compute. ## Advanced configurations: Turning 10% into 30% or 50% That's a great cost reduction. But we can go further by fine-tuning our runs using three basic approaches. ### Run on SLAs, not on timers You can use advanced configurations to customize the behavior of both how sources and models behave in state-aware orchestration. This helps you deal with situations like when a model might be needed less often, even if its sources are fresh. A data source might update every hour, for example, but the report it feeds is only reviewed once every week. You can fine-tune your runs further by asking for a given model: Do I want to wait until all of the dependencies upstream have built? Or simply do it whenever there's new data on any? ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a397352c1671f6a86257dcfc113b69f42ca00912-1255x712.png) As you fine-tune these configurations, you're able to dramatically increase how much is reused. That leads to fewer builds against the actual data warehouse as you align your builds with the SLAs of the business. ### Customize behavior based on business needs With these advanced configurations, you can define what it means for a source to be fresh. There are three ways to approach this. First, freshness might be tied to a specific column. You can do that with [the loaded_at_field](https://docs.getdbt.com/reference/resource-properties/freshness) in your dbt models, or create a custom SQL query where you pull that metadata from an Iceberg table. You could put in a custom combination of number of rows loaded and a datetime in order to refine the definition of "fresh" in a highly granular manner. Second, you can create model SLAs for situations where sources refresh more frequently than the business needs data. You can use a [build_after setting](https://docs.getdbt.com/docs/deploy/state-aware-setup#advanced-configurations) in your dbt models to wait until enough time has passed, even when there is new data to incorporate. build_after: {count: 2, period: hour} Third and finally, updates_on lets you specify whether we should wait on all upstreams to be updated or simply build whenever any of them are. The magic thing about dbt is that because this all exists in your dbt project YAML, you can incorporate this as sensible defaults at the project level and then override those across any folder, subfolder, or even individual model that really allows you to fine-tune your project to the exact configurations that make the most sense. ### Efficient testing: Only run the tests that are needed You can optimize further by optimizing how often you test. Traditionally, with dbt Core, we run a test set at the end of every model that builds. This is great because it helps ensure data quality by ensuring tests pass as data changes throughout our [data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage). But it also means that we're probably overusing tests. With efficient testing using the dbt Fusion engine, you only run the tests that are actually needed. Suppose you have a unique test that passes upstream. With dbt Fusion's column-aware, semantic understanding, if no downstream model logic invalidates that test, it reuses the passing test rather than rerunning it. This reduces runtime and lowers costs by avoiding unnecessary computation. Using Fusion's semantic understanding of the SQL, we don't just _skip_ the tests. We **reuse** them because we know the test is statistically guaranteed to pass. There isn't any SQL logic that will invalidate any unique joins or WHERE conditions that depend on the individual test run upstream. When we run those tests, instead of executing each one as a separate query against the data warehouse, we bundle them into a single query and execute it in a single pass. This further maximizes compute efficiency. ## Proving the value to your business Now, here's how you can prove these results to the business. With state-aware orchestration, we show you the models built and reused across your entire account. The dbt platform's reporting shows you the impact your changes have on your runtimes and on the number of models built. On the individual jobs, we show you that same view of models built versus reused. You can fine-tune those advanced configurations to determine how to meet the SLAs for when the business needs the data and how often. If you want to go above and beyond, you can also write this cost savings data in dollar or credit terms directly to your data warehouse. This lets you analyze and prove those cost savings wherever you work. More importantly, this helps you make friends with finance. When it comes time to ask for that additional headcount, you can show them and your management team exactly how much you've saved on your data warehouse bill. ## Real results: Our journey and customer success In summary, enabling state-aware orchestration delivers 10%+ in cost reduction. Fine-tuning those advanced configurations can give you an additional 15% or more. Efficient testing can add another 4%, for a total of 29%. dbt Labs was the first user of Fusion and the first user of state-aware orchestration in production. We're a 9-year-old dbt project with 108 contributors. To be honest, historically, cost optimization wasn't our focus. But as we've grown (and as our data warehouse bill has grown along with us), it's come under scrutiny. When we implemented state-aware orchestration, we first enabled it and then reconfigured our jobs into two freshness-based tiers. We achieved these results without rewriting any SQL code. (I say SQL here because obviously we did write some YAML in order to put in those build after configs across our project and defaults, and then override those for more frequent tiers where needed.) What we saw here was pretty incredible. We saw a 63% improvement in average job runtime. For example, our incremental job went from 3.5 hours on average to 25 minutes—that's 88% faster. We saw 75% of models reused daily. That's gone from 9,000 built to only 2,200. And most importantly, we saw a 64% annual data platform savings of all of our dbt-related spend. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a9aa1168f1643c50c8d5a2c7151b045c921b6740-1600x519.png) The hope is that, with state-aware orchestration, advanced configurations, and features such as efficient testing, everyone can achieve the same results in reducing their data warehouse spend and improving the efficiency of their data pipelines. ## What's next As we look ahead, we're really working on bringing that cost data from the warehouse directly into the dbt platform. We can show you how much you're saving without any additional work on your part. We want to build even slimmer CI using the power of Fusion and that column-level awareness we get through our semantic understanding to create an even slimmer set of models that are actually changed for your CI jobs. We're working on creating AI agents to automate configurations and save you both time and money as you fine-tune your dbt projects. And we're investing in even more automatic cost savings like dynamic indexing across your tables so that when you turn on cost efficiency features with Fusion, you get the maximum out of the box across your warehouse. ## Get started today Anyone using dbt with an eligible project who's on an enterprise-tier plan can turn this on today. You simply need to change your production settings from dbt latest to dbt Fusion. You can then check the boxes in your job settings and start benefiting from those immediate cost savings through state-aware orchestration and from fine-tuning your model builds as you need across your project. If you're not currently using dbt, [sign up today for a free account](https://www.getdbt.com/signup) and step through [the Fusion getting started guide](https://docs.getdbt.com/docs/fusion/get-started-fusion) to see how Fusion can optimize your data transformation workflows and workloads across your entire enterprise. --- --- title: "Closing the context gap: Cribl’s blueprint for trusted AI with dbt + Omni" description: "AI analytics fail without context. See how Cribl solved this to deliver trusted, AI-powered insights to 700+ users." url: "https://www.getdbt.com/blog/cribl-dbt-omni-ai-analytics" date: "2026-01-21" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Closing the context gap: Cribl’s blueprint for trusted AI with dbt + Omni Data leaders are under real pressure to show progress with AI. One of the top priorities we hear is to deliver trusted conversational analytics safely and cost-effectively. That's a real challenge for data teams. Why? Because the solution is more complicated than "point your [Large Language Model](https://www.ibm.com/think/topics/large-language-models) (LLM) at the data." Many AI-driven analytics projects fail not because the model can't write SQL (AI can write SQL very well). They fail because they lack the right context to understand the business. This includes things like metric definitions, business logic, and up-to-date signals about the data's quality and freshness. Not having that information creates a context gap for the model. You can't just point your LLM at raw data. Even if the AI model finds the right data, it still lacks the context to define your specific metrics. Lacking context, it's forced to guess. The consequences are real: stakeholders lose trust because outputs are incorrect or inconsistent. Risk goes up because you can't easily govern how data is being queried. Token and compute costs expand without adding business value. All of this stalls AI adoption across your data teams. This is where dbt and Omni come in. Here's how these two platforms work together to ensure your AI is grounded in a single source of truth that's accurate, up-to-date, and validated. **** ## How dbt and Omni close the context **gap** dbt Labs provides a structured context layer that centralizes governed analytics under a single [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) that's: - [Version-controlled](https://docs.getdbt.com/docs/cloud/git/version-control-basics); - Has [models](https://docs.getdbt.com/docs/build/models) and [tests](https://docs.getdbt.com/docs/build/data-tests); - Supports [defined metrics](https://www.getdbt.com/product/semantic-layer); - Contains end-to-end [data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage); and - Provides visibility into [data freshness](https://docs.getdbt.com/docs/deploy/source-freshness) [Omni Analytics](https://omni.co/) is the business intelligence platform that brings together flexible data exploration with governed, consistent metrics. When someone asks Omni's AI, "What was Q4 revenue?" it uses the approved revenue definition and trusted models, relying on tests and freshness rather than AI hallucinations. With dbt's [MCP server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server), you can expose that metadata and governance consistently across multiple AI systems through API integrations. Whether connecting to OpenAI, Omni's AI agents, or any other agentic workflow, you can always use the same metric definitions and guardrails. The result is reliable SQL, auditable AI-powered outputs, controlled token costs, and better AI adoption. ## Building self-service at scale: Cribl's story [Cribl](https://cribl.io/) is the data engine for IT and security. Using Cribl, companies can route, transform, and reduce logging, metric, and trace data using Cribl's platform as their universal translator. Self-service data is really important at Cribl. 77% of the company uses data - more than 700 people. These users expect this data to be accurate, so it must be high-quality and governed. One of Cribl's success metrics last year was data quality. That's why it focuses on making sure its 150 data producers and heaviest data users (like marketing) are set up for success through scalable workflows. Even with the best support, however, data will always have problems. Data quality isn't about "perfect" data. It's about catching problems, alerting the correct people, and fixing them before stakeholders notice. Because that's where you lose trust. ### The technical foundation: dbt as the transformation layer Cribl ingests data from about 48 different data sources across its ecosystem. In their raw data warehouse, they have three to four thousand tables. However, they don't expose all those to their BI tool. Business context would be lost, and it's an absurd number for users to navigate. This is why Cribl uses dbt as its data control plane and data transformation layer. No table is exposed to end users unless it is processed by dbt's open-source transformation engine. This enables Cribl to govern data access through automated validation, ensuring that all metrics and as much semantic knowledge as possible reside in dbt as a single source of truth. That way, when people ask how many sales opportunities are open, for example, they always get the same answer: creating scalable, data-driven workflows. ### Moving to Omni for AI-ready analytics Cribl recently migrated from Looker to Omni after four years as a Looker customer. Migration wasn't about pricing - it was about business value and AI-powered capabilities. Cribl was impressed with Omni's AI readiness and its ability to function as a trusted AI platform. They could leverage the enrichment available in dbt (e.g., YAML files and descriptions) and push it to Omni, where it can be used to train the AI models. Cribl uses hub and spoke teams - i.e., data producers and data consumers. Spoke teams don't have access to dbt. However, in the Omni developer layer, they can see all the dbt models. Cribl didn't have that integration before. This lets people use tools like Google Sheets to create their own measures without relying on the data team. Cribl's tech stack focuses on importing data into Snowflake, a cloud-native data platform, where it transforms it with dbt. From there, it loads it into Omni as a single source of truth and governance layer. There, employees can not only build dashboards - they can leverage Omni's AI support to query that data reliably through natural language interfaces. ![Tech stack](https://cdn.sanity.io/images/wl0ndo6t/main/224d9d11f393c0a6286a28276d142b9893635911-1762x1026.png) ### Automating documentation with AI to improve data quality AI is only as good as your data. That's why Cribl has numerous initiatives tied to fine-tuning Omni's responses and optimizing their AI-driven workflows through automation. Field descriptions are a key part of quality. For example, Cribl manages work in Jira, a SaaS tool used by software development teams. However, they don't use Jira's "closed date," relying instead on a number of other "Resolved" fields. Accurately describing these through data modeling is key so that when a user asks Omni to operate on "closed" tickets, Omni's AI can translate that into the correct fields. Issues like this led Cribl to realize it was imperative to document its dbt layer effectively using [dbt's built-in documentation support](https://docs.getdbt.com/docs/build/documentation). But when your data models are complex, that's easier said than done. When you have five dbt models, it's easy to document them. When you have 230, it's a lot harder—and automation becomes essential for maintaining scalable workflows across the lifecycle. This is when Cribl decided to use generative AI to help AI. They figured that, if they could pass sufficient context to an LLM, it could automatically generate field-level descriptions, descriptions for new tables and models, and even descriptions for schemas, and consistently update them as the models evolve through their lifecycle. To accomplish this, Cribl passes all its fields and business logic to [OpenAI](https://openai.com/) through APIs to generate useful descriptions. The first passes were also…well, a learning experience! The data engineering team struggled with hallucinations and all the familiar problems of using generative AI. They gradually incorporated techniques - such as limiting field documentation requests to 20 maximum per request, or passing exact column names instead of SELECT * to the LLM - that improved the output through validation. The ROI on the automation process has been incredible. The first attempt, for example, only cost around $20 to run on about 150 models. Compare that to the cost of spending hundreds of hours on manual documentation—a real-world example of AI tools delivering business value. Since then, Cribl has spun this off into a CI/CD pipeline that automates the entire workflow, running automatically with every commit to their dbt repository hosted on GitHub. Say someone adds a new dbt model. An AI agent runs and automatically creates a new YAML file with descriptions for that model through orchestration. The huge benefit here is that, when the data engineering team reviews the changes in their workflows, all of that metadata has already been created through automated processes. Once the change is approved and merged, it automatically gets updated into Cribl's branch for their repository, which automatically goes to Omni through API integrations. Omni supports this automation process with a Sync metadata button. You just click it, and Omni updates all the descriptions and changes automatically—streamlining the entire workflow across the lifecycle. Those updates are critical for making your trusted AI useful. Because at the end of the day, your AI is only as useful as the governance and validation you have underneath it. ## How the Omni and dbt integration works Omni's dbt integration is built on the premise that the two tools require a tightly integrated workflow with a seamless handoff. Omni supports many functions that enable this - everything from syncing metadata in both directions through APIs to ensuring that you can see your environments in dbt and Omni consistently across your data platform. ### Synchronized development This approach enables synchronized development workflows through streamlined processes. You can develop in dbt, use your dbt development schema to materialize models, then point Omni to that environment and make changes to dashboards without impacting production. This all happens within a source control branch in dbt and a corresponding branch in Omni. You no longer have to cross your fingers when you make a dbt change, waiting to see which dashboards broke. You can validate everything in a sandbox, then merge your dbt branch and Omni branch simultaneously through automated orchestration. This solves a common pain point for many teams while optimizing the development lifecycle. ### Model once, use everywhere Write a dbt model with foreign keys and descriptions. Omni automatically reads this information through API calls, along with your dbt documentation, and populates your Omni data model. Users leverage the same structures you built in dbt without redefining anything—creating scalable workflows that streamline collaboration across the data management ecosystem. You can also work in the other direction. Business users who don't spend time in dbt can define new fields or calculations in Omni through natural language queries and describe them, creating a seamless handoff between BI users and data engineers. This flexibility makes the ecosystem more accessible while maintaining governance and validation. ### Building trust at the point of consumption This metadata enables building trust at consumption - the dashboard or report. Omni can push [exposures](https://docs.getdbt.com/docs/build/exposures) to dbt automatically through API integrations. If you use exposures to index data lineage, you understand exactly what each dashboard depends on and what a change would impact. You can use dbt's embeddable tiles in dashboards to show data status and confirm that everything is up to date and accurate in real-time—providing validation at every step. Through exposure syncing, the dbt team knows what changes will impact and whether they'll break a dashboard. This benefits both sides: dbt exposes quality information into Omni through APIs, and teams using dbt understand the impact of their changes. This creates high-performance workflows for optimizing both development and operations across the lifecycle, streamlining the entire process. ## Making AI work for your users With this foundation in place, Omni's advanced AI truly shines as an AI-powered assistant for data teams seeking actionable insights. Omni lets you ask quick questions using natural language interfaces. One of my favorite new features has been their AI summary visualization, where you can literally just pass the data points. Whenever you open a dashboard, if there are any updates, Omni updates the summary in real-time. That's the kind of future I want to live in—where AI agents handle the heavy lifting. About 20% of Cribl users use Omni AI every month through the AI assistant interface. They're now aiming to reach 25%—demonstrating real-world adoption of AI agents in their data analytics ecosystem through scalable, self-service workflows. The next metric is: of those 25%, how many users are getting the answers they expect from ‌AI tools? Omni helps improve response accuracy by exposing the queries and prompts your users are asking through API integrations. The goal of the integration is to enable rapid iteration with a smooth, seamless workflow between Omni and dbt—supporting the full analytics development lifecycle through automated orchestration and streamlining processes. In an AI workflow, tuning your data model and context to answer questions accurately is the key to optimizing performance. A common example is your terminology. If someone asks about a specific term and you don't know what it means, you can look it up and update the data model through automated validation. Then, the next time somebody asks that question, the answer is ready - and correct. This turns tuning your AI models into a three-part workflow that streamlines the optimization process: - Create your data model with proper metadata and validation; - Add context through automated documentation; and - Review the questions that are being asked to optimize performance across use cases. You can have your users give thumbs up, thumbs down, and feedback through the AI assistant, and identify use cases where the AI lacks sufficient context. Then you can iterate on your data model, whether in Omni or in dbt or in both, to improve the quality of the questions and answers while optimizing the platform's performance. This continuous improvement cycle is essential for scalable AI adoption across the ecosystem, streamlining workflows and automating validation processes. ## Closing the loop With dbt and Omni, the cycle is closed. Build trusted data with dbt, then explore and access it in Omni through AI-powered analytics. Trust everywhere. This integration solves one of the most pressing challenges facing data teams today: how to deliver AI-powered analytics that stakeholders can actually trust. dbt Labs grounds AI in a structured, governed, version-controlled layer of context and semantics that streamlines workflows. That eliminates the guesswork and hallucinations that plague most conversational analytics implementations through automated validation and orchestration. The result? Faster answers. Lower costs through optimized workflows and automated processes. Higher trust through validation. And, ultimately, better AI adoption across your organization. --- --- title: "Incorporating version control into data analytics" description: "Version control brings reliability, collaboration, and accountability to analytics workflows as data teams scale." url: "https://www.getdbt.com/blog/version-control-data-analytics" date: "2026-01-21" authors: ["Joey Gault"] categories: ["Pulse"] --- # Incorporating version control into data analytics As analytics teams grow, so does the complexity of their work. What starts as a handful of SQL queries and dashboards quickly becomes a shared system of transformations, metrics, tests, and documentation that multiple people depend on every day. Without clear ways to track changes, review work, and recover from mistakes, analytics environments can become fragile and difficult to trust. Version control brings structure to this complexity by giving data teams a reliable way to collaborate, manage change, and treat analytics code with the same discipline as software. ## The case for version control in analytics Analytics code, whether SQL transformations, Python scripts, or data models, represents a significant organizational investment. This code transforms raw data into the insights that drive business decisions. Without [version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics), analytics teams face persistent challenges: duplicated work, inconsistent metric definitions, unclear data lineage, and difficulty collaborating across team members. Version control addresses these challenges by treating analytics code as what it truly is: a software asset that requires the same rigor applied to application development. When analytics code is version controlled, teams gain visibility into who changed what and when. Changes can be reviewed before deployment. Errors can be traced to their source and rolled back. Knowledge moves from individual analysts' heads into a shared, documented codebase. The benefits extend beyond individual productivity. Version controlled analytics enables teams to scale from one analyst to dozens without descending into chaos. It creates accountability through code review processes. It establishes a foundation for automated testing and continuous integration. Most importantly, it transforms analytics from a collection of isolated scripts into a coherent, maintainable system. ### Version control fundamentals for analytics teams At its core, version control for analytics follows the same principles as software development. Teams work with Git repositories that store all analytics code: data models, transformations, tests, and documentation. The repository serves as the single source of truth for how data is processed and analyzed. The basic workflow centers on branching. Analysts create separate branches to develop new features or fix issues. These changes remain isolated from production code until they're ready. When development is complete, changes go through review and testing before merging into the main branch. This branching model prevents untested code from affecting production systems while giving analysts freedom to experiment. Key Git concepts translate directly to analytics work. A commit represents a discrete change to analytics code, whether adding a new data model or fixing a calculation error. Branches allow parallel development; multiple analysts can work on different features simultaneously without conflicts. Pull requests formalize the review process, ensuring that changes are validated before deployment. Merges integrate approved changes back into the production codebase. For analytics teams using [dbt](https://www.getdbt.com/product/dbt), version control integration is seamless. [dbt projects are structured as Git repositories from the start](https://docs.getdbt.com/docs/cloud/git/git-configuration-in-dbt-cloud). All data models, tests, and documentation live in version-controlled files. When analysts develop in the [dbt IDE](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/develop-in-the-cloud), they're working directly with Git. The IDE provides git commands for creating branches, committing changes, and opening pull requests without requiring command-line expertise. ### Implementing protected branches and development workflows A mature version control strategy requires protecting production code from direct modification. In dbt, the main branch typically represents the production environment. This branch should be protected; analysts cannot commit changes directly to it. Instead, all changes flow through a structured development and review process. The protected branch model enforces discipline. When an analyst needs to make changes, they create a new branch from main. This development branch provides an isolated environment for building and testing. The analyst can iterate freely, committing changes as they progress. Other team members' work doesn't interfere, and production systems remain unaffected. Once development is complete, the analyst opens a pull request to merge their branch back into main. This triggers the review process. Other team members examine the code changes, checking for correctness, adherence to style guidelines, and potential downstream impacts. Automated tests run to validate that the changes don't break existing functionality. Only after review approval and passing tests can the changes merge into main and deploy to production. This workflow scales effectively. Small changes (fixing a typo in documentation) move through quickly. Larger changes (refactoring core data models) receive proportionally more scrutiny. The process remains consistent regardless of team size. Whether you have two analysts or twenty, the same branching and review workflow applies. ### Establishing code review practices Code review represents one of the most valuable aspects of version control for analytics. No analytics code should reach production without a second set of eyes reviewing it. This practice catches errors, shares knowledge across the team, and maintains code quality standards. Effective code review in analytics requires specific focus areas. Reviewers examine the business logic to ensure transformations correctly implement requirements. They verify that SQL follows established style conventions and best practices. They check that appropriate tests exist to validate the code's behavior. They consider downstream impacts: will this change break existing dashboards or reports? The review process also serves as knowledge transfer. Junior analysts learn from feedback on their code. Senior analysts share context about data sources and business logic. The entire team develops shared understanding of the analytics codebase. This shared knowledge reduces silos and makes the team more resilient. For code review to work, teams need clear expectations and accountability. Review turnaround time should be reasonable; pull requests shouldn't languish for days. Reviewers should provide constructive feedback focused on improving the code. Authors should be receptive to feedback and willing to iterate. The culture around code review matters as much as the technical process. ### Managing multiple environments through version control Version control enables analytics teams to maintain separate development, staging, and production environments. Each environment corresponds to a different state of the codebase. Development environments run code from feature branches. Staging environments run code from integration branches. Production environments run code from the protected main branch. This environment separation is critical for safe analytics development. Analysts need freedom to experiment and iterate without impacting production data or reports. Development environments provide that freedom. Analysts can test changes against production-like data, validate results, and refine their approach before promoting code to production. [dbt's integration with version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) makes environment management straightforward. When working in a development branch, analysts run dbt commands against their personal development schema. Changes build in isolation. Once changes merge to main, automated deployment processes run dbt against the production schema. The same code that was tested in development now populates production tables. The connection between Git branches and data environments creates clear boundaries. Production data comes from production code in the main branch. Development data comes from development code in feature branches. This mapping is explicit and auditable. Anyone can trace a production table back to the exact code version that created it. ### Handling merge conflicts in analytics code As teams grow and multiple analysts work simultaneously, merge conflicts become inevitable. Conflicts occur when two branches modify the same section of code in incompatible ways. Git cannot automatically determine which change should take precedence, requiring manual resolution. In analytics code, conflicts often arise in shared data models or configuration files. Two analysts might modify the same SQL model to add different columns. Or they might both update the same configuration file to add different sources. These conflicts must be resolved before code can merge. Resolving conflicts requires understanding both sets of changes and determining how to integrate them. Sometimes one change should take precedence. Sometimes both changes can coexist with minor adjustments. Sometimes the conflict reveals a deeper issue requiring discussion with stakeholders. The best approach to conflicts is prevention. Teams should communicate about planned changes to shared code. Larger refactoring efforts should be coordinated to minimize overlap. Keeping changes small and merging frequently reduces conflict likelihood. When conflicts do occur, resolving them quickly prevents them from compounding. ### Integrating testing with version control Version control and automated testing form a powerful combination. When code changes are committed to a branch, automated tests should run to validate correctness. This continuous integration approach catches errors early, before they reach production. For analytics code in [dbt](https://www.getdbt.com/product/dbt), testing happens at multiple levels. Unit tests validate individual model logic. Data tests check that actual data conforms to expectations: primary keys are unique, foreign keys have valid references, critical columns contain no nulls. Integration tests verify that changes don't break downstream dependencies. These tests run automatically when pull requests are opened. If tests fail, the pull request cannot merge. This enforcement ensures that only validated code reaches production. It shifts quality assurance left in the development process, catching issues when they're easiest to fix. The [testing capabilities built into dbt](https://docs.getdbt.com/docs/build/tests) make this integration seamless. Tests are defined alongside the models they validate, all stored in version control. When code changes, the relevant tests automatically run. Test results appear directly in pull requests, giving reviewers immediate feedback on code quality. ### Documentation as code Version control transforms documentation from an afterthought into an integral part of analytics development. When documentation lives in the same repository as analytics code, it stays synchronized with the code it describes. Changes to data models include corresponding documentation updates in the same commit. dbt treats documentation as code. Model descriptions, column definitions, and data dictionaries are written in [YAML files](https://docs.getdbt.com/reference/node-selection/yaml-selectors) stored in version control. These documentation files go through the same review process as SQL code. Documentation changes are visible in pull requests. Reviewers can verify that documentation accurately reflects code changes This approach solves the perennial problem of outdated documentation. Documentation doesn't lag behind code because they're updated together. The documentation visible to end users always reflects the current production codebase. When code is rolled back, documentation automatically rolls back with it. Version controlled documentation also enables collaboration. Multiple team members can contribute to documentation. Subject matter experts can add business context. Analysts can document technical implementation details. All contributions flow through the same review and approval process, ensuring documentation quality. ### Deployment automation and continuous delivery Version control enables automated deployment of analytics code. When changes merge to the main branch, automated processes can deploy those changes to production without manual intervention. This continuous delivery approach reduces deployment friction and accelerates the pace of analytics development. For dbt projects, deployment automation typically involves running dbt commands when code changes are detected in the main branch. A [CI/CD](https://docs.getdbt.com/docs/deploy/about-ci) system monitors the repository, detects merges to main, and triggers a dbt run. The run executes all modified models and their downstream dependencies, updating production tables with the latest code. This automation eliminates manual deployment steps that are error-prone and time-consuming. Analysts don't need to remember which models to run or in what order. The deployment process is consistent and repeatable. Deployments happen quickly after code merges, reducing the lag between development and production availability. Automated deployment also enables rollback capabilities. If a deployment introduces errors, the main branch can be reverted to its previous state. The deployment automation then runs again, restoring production to the last known good state. This safety net makes teams more confident in deploying changes frequently. ### Building a sustainable analytics codebase Version control is the foundation for building analytics systems that are maintainable over the long term. Analytics code often outlives the analysts who wrote it. Team members change roles, new analysts join, and the codebase continues to grow. Version control ensures that this evolution happens in a structured, traceable way. The commit history provides invaluable context for understanding why code exists in its current form. When an analyst encounters a confusing transformation, they can examine the commit that introduced it. The commit message explains the business requirement. The pull request discussion reveals the reasoning behind implementation choices. This historical context makes maintenance far easier. Version control also supports refactoring and technical debt management. As analytics requirements evolve, code needs to be restructured. Version control makes refactoring safer by providing rollback capabilities. It makes refactoring more collaborative by enabling code review. It makes refactoring more transparent by documenting what changed and why. For data engineering leaders, version control provides visibility into team productivity and code quality. Commit frequency, pull request cycle time, and code review participation are all measurable. These metrics help identify bottlenecks and opportunities for process improvement. They provide objective data for assessing team health. ## Conclusion Incorporating version control into data analytics represents a fundamental shift in how analytics teams operate. It moves analytics from ad-hoc scripting to disciplined engineering. It enables collaboration at scale. It creates accountability and auditability. It provides the foundation for automated testing and deployment. For teams using dbt, version control integration is built into the core workflow. The [dbt IDE](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/develop-in-the-cloud) provides accessible version control capabilities for analysts of all skill levels. The [Analytics Development Lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle) framework provides a comprehensive model for mature analytics workflows that incorporate version control at every stage. The transition to version controlled analytics requires investment in process, tooling, and culture. Teams need to establish branching strategies, code review practices, and testing standards. They need to train analysts on Git workflows and best practices. They need to build a culture that values code quality and collaborative development. The payoff is substantial. Version controlled analytics teams ship faster, with higher quality, and with greater confidence. They scale more effectively as they grow. They build analytics systems that are reliable, maintainable, and trustworthy. For data engineering leaders, incorporating version control into analytics is not just a technical decision; it's a strategic imperative for building world-class data organizations. ## Version control FAQs **How does a data version control system work?** A version control system for analytics works by storing all analytics code (including SQL transformations, Python scripts, data models, tests, and documentation) in Git repositories that serve as the single source of truth. The basic workflow centers on branching, where analysts create separate branches to develop new features or fix issues while keeping changes isolated from production code. When development is complete, changes go through review and testing before merging into the main branch. Each commit represents a discrete change to the code, branches allow parallel development by multiple team members, and pull requests formalize the review process to ensure changes are validated before deployment. **How do Git branches help data analysts collaborate without overwriting each other's work?** **Why use version control for data analysis?** Version control addresses persistent challenges that analytics teams face, including duplicated work, inconsistent metric definitions, unclear data lineage, and difficulty collaborating across team members. It provides visibility into who changed what and when, allows changes to be reviewed before deployment, enables errors to be traced and rolled back, and transforms knowledge from individual analysts' heads into a shared, documented codebase. Beyond individual productivity, version control enables teams to scale from one analyst to dozens without chaos, creates accountability through code review processes, establishes a foundation for automated testing and continuous integration, and transforms analytics from isolated scripts into a coherent, maintainable system that can be treated as a valuable software asset. --- --- title: "Data transformation vs. Data modeling: Key differences" description: "Explore the difference between data transformation and modeling—and how both shape scalable analytics." url: "https://www.getdbt.com/blog/data-transformation-vs-data-modeling" date: "2026-01-15" authors: ["Joey Gault"] categories: ["Pulse"] --- # Data transformation vs. Data modeling: Key differences Data transformation and data modeling are often mentioned in the same breath — but they play fundamentally different roles in modern analytics. One focuses on executing change; the other designs what that change should look like. Understanding the distinction helps data teams scale, standardize, and deliver trustworthy data products that support confident decision-making. ## Understanding data transformation Data transformation is the process of converting raw data from its original format into one readily usable by business decision-makers. This includes normalizing, cleaning, validating, and aggregating data to ensure it's ready for analysis. At its most practical level, transformation involves writing SQL or Python code that takes materialized data assets (tables or views) and converts them into purpose-built datasets for analytics. **** The transformation process typically unfolds across several stages. It begins with discovery and profiling, where teams assess data structure, quality, and characteristics to identify anomalies and inconsistencies. Cleansing follows, correcting inaccuracies, filling missing values, and removing duplicates. Data mapping then structures information according to target system requirements, converting data types and reorganizing fields as needed. Finally, transformed data loads into a central data store like a data warehouse, where it becomes available for analysis and reporting. Modern data transformation commonly occurs within [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) pipelines, where data is transformed after loading into its destination. This approach has largely replaced [traditional ETL methodologies](https://www.getdbt.com/blog/extract-transform-load) because cloud computing makes it more cost-efficient to load data prior to transformation. Raw data becomes immediately available to everyone with warehouse access, and teams with different needs can transform it however they see fit. The benefits of robust data transformation are substantial. It increases data quality by addressing malformatted values, redundancies, and inconsistencies that plague raw data. Transformation also produces organized, easy-to-use datasets that eliminate the need for analysts to reinvent the wheel with each new report. Perhaps most importantly, transformation paves the way for machine learning and AI workloads by providing the large volumes of high-quality data these probabilistic approaches require. ## Understanding data modeling Data modeling represents something fundamentally different: the architectural blueprint for how data should be organized, stored, and connected across an entire system. While transformation is the act of moving and changing data, modeling is the design discipline that determines what the end result should look like. It establishes repeatable patterns for how schemas and tables are structured, how models are named, and how relationships are constructed. **** The distinction between a data model and individual transformation files matters considerably. A data model is the complete blueprint (the architecture defining how all pieces fit together to tell the story of a business). Individual transformation files, such as those created in [dbt](https://www.getdbt.com/product/what-is-dbt), are the building blocks that implement this broader design. Think of the data model as the recipe, with transformation files as the ingredients that combine to create the final product. Modern data modeling typically organizes work into distinct layers, each serving a specific purpose. Staging models form the foundation, cleaning and standardizing raw source data through light transformations like casting field types and renaming columns. Intermediate models handle complex transformations that don't fit neatly into other layers, breaking complicated logic into manageable pieces. Mart models apply business logic to create core data assets for analysis, typically producing fact and dimension tables that represent measurable events and descriptive context. Several methodologies guide how teams structure their models. Dimensional modeling categorizes entities into facts and dimensions, optimizing for analytical workloads and aligning data structures with how businesses naturally think and operate. Data vault modeling abstracts entities into hubs, links, and satellites, excelling at tracking data changes in high-governance environments. Entity-relationship modeling focuses on business processes and how entities connect. Each approach offers different tradeoffs between complexity, flexibility, and performance. Well-designed data models determine whether business users trust and adopt data products. When end users lack confidence in data quality or find datasets difficult to navigate, they retreat to familiar tools like spreadsheets, creating silos and inconsistency. Proper data modeling addresses this by creating intuitive, navigable structures where relationships are clear, naming conventions are consistent, and business logic lives in centralized, version-controlled locations rather than scattered across individual reports. ## The relationship between data transformation and data modeling Data transformation and data modeling are not competing concepts but complementary disciplines that work in tandem. Data modeling provides the architectural vision (the blueprint for what your data warehouse should look like and how different entities should relate to each other). Data transformation provides the execution (the actual code and processes that implement that vision by moving and changing data). In practice, transformation is one technique among many that you'll use to realize your data model. Other transformation techniques include cleaning, aggregating, generalization, validation, normalization, and enrichment. Data integration (bringing data from multiple sources into a unified view) is itself a type of data transformation that often plays a central role in implementing dimensional or other modeling approaches. The modeling decisions you make directly influence how you approach transformation work. If you've designed a dimensional model with separate fact and dimension tables, your transformation code will focus on building those distinct entities and establishing the relationships between them. If you've opted for wide, denormalized tables to support less technical users, your transformations will emphasize pre-joining data and reducing the need for complex queries downstream. Similarly, the realities of your transformation process can inform modeling decisions. Performance considerations drive choices about normalization versus denormalization. When every analysis requires multiple joins that consume expensive compute resources, denormalized models may be more practical despite increased redundancy. When source systems change frequently, staging layers in your transformation pipeline provide defensive architecture that absorbs those changes and protects downstream models. ## Practical challenges and considerations Both data transformation and data modeling present distinct challenges that data engineering leaders must navigate. On the transformation side, creating consistency across multiple datasets proves difficult at scale. Teams must ensure datasets follow standardized naming conventions, SQL best practices, and consistent testing standards. Without this consistency, analysts risk duplicative work, misaligned timezones, and unclear data relationships that lead to inaccurate reporting. Standardization of core KPIs represents another transformation challenge. Key business metrics should be version-controlled, defined in code, and accessible within BI tools. When different teams generate conflicting reports due to inconsistent metric definitions, confusion and inefficiency follow. Modern transformation tools like dbt address this through features like the [Semantic Layer](https://www.getdbt.com/product/semantic-layer), which allows teams to create and apply the same metric calculation across different models, datasets, and BI tools. Data modeling introduces its own complexities. Determining whether an entity should be treated as a fact or dimension depends on analytical needs rather than rigid rules. The decision to create wide, denormalized tables versus maintaining separate fact and dimension tables similarly depends on context (specifically, the SQL skills of end users and the capabilities of their BI tools). These decisions require understanding both the data and stakeholder needs, two of the most difficult aspects of data work. Maintaining readability as models grow presents an ongoing challenge. Individual transformation files should remain concise enough for team members to quickly understand their purpose and logic. Modular SQL blocks help keep files readable by abstracting repetitive patterns into reusable components. Clear [naming conventions](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview) bring order to data warehouses, making project structure immediately comprehensible to anyone navigating the codebase. ## Building scalable analytics systems Managing both data transformation and data modeling at scale requires a unified approach. The nature of enterprise data (hundreds of sources including data warehouses, analytics tools, marketing platforms, relational databases, NoSQL stores, data lakes, and lakehouses) makes ad hoc management impossible. Without a single approach to manage transformation and integration across disparate systems, engineering teams end up implementing data pipelines in fragmented ways, using technologies and languages that lock them into vendor-specific solutions. This fragmentation creates numerous inefficiencies. Engineers reuse little code, solving the same problems redundantly in different languages. Organizations lack visibility into available data assets and running pipelines, making it impossible to ensure data quality and consistency across data stores or manage pipeline costs effectively. There's no consistency in how transformation and integration code is tested before reaching users, risking bad data in production datasets. Changes are rolled out in ad hoc fashion rather than through rigorous CI/CD processes that verify changes before release. A data control plane (a single toolset for managing all data transformations across the enterprise) addresses these challenges. Tools like [dbt](https://www.getdbt.com/product/dbt) enable data engineers, analytics engineers, and business users to model data transformations uniformly using SQL or Python code, avoiding vendor lock-in. This approach provides out-of-the-box features that support high-quality datasets: [version control](https://docs.getdbt.com/docs/cloud/git/connect-github) and peer reviews, support for creating DRY code that can be reused across projects, [testing](https://docs.getdbt.com/docs/build/data-tests) and [documentation](https://docs.getdbt.com/docs/collaborate/documentation) including automatically generated [data lineage](https://docs.getdbt.com/terms/data-lineage), and CI/CD deployment with automated testing in pre-production environments. The separation of concerns across transformation layers creates modularity that scales effectively. Rather than building monolithic transformations from raw data each time, practitioners reference foundational work completed by others. This reduces duplication, improves maintainability, and makes dependencies explicit through clear lineage. When combined with thoughtful data modeling that establishes clear architectural patterns, this approach creates data warehouses that serve as reliable foundations for analytics (where business users find data intuitive to work with, where data teams can efficiently build and maintain transformations, and where architecture can evolve alongside changing business needs). ## Conclusion Data transformation and data modeling are distinct but inseparable disciplines in modern analytics engineering. Transformation is the operational work of converting raw data into usable formats through cleaning, joining, aggregating, and applying business logic. Modeling is the architectural discipline of designing how data should be organized, structured, and related to tell the story of a business. Neither can succeed without the other. For data engineering leaders, the key is recognizing that both require deliberate investment and strategic thinking. Quick wins through ad hoc transformations may deliver initial value, but they don't scale as complexity grows. The cost of rebuilding fundamentally flawed architecture is substantial. Similarly, elegant data models that don't account for transformation realities (performance constraints, source system changes, user capabilities) will fail to deliver practical value. Success requires treating both data modeling and data transformation as first-class concerns, establishing conventions early, maintaining discipline as systems grow, and choosing tools that support engineering best practices at scale. When organizations get this right, they create analytics systems that are not just functional but truly transformative (enabling self-service analytics, supporting confident decision-making, and freeing data teams to focus on high-value work rather than repeatedly answering the same questions). The distinction between transformation and modeling matters because understanding it allows you to build better systems that serve both disciplines effectively. ## Data transformation vs. Data modeling FAQs **What is the difference between data modeling and data transformation in the modern data stack?** Data modeling is the architectural blueprint that determines how data should be organized, stored, and connected across an entire system. It establishes repeatable patterns for schema structure, naming conventions, and relationships. Data transformation, on the other hand, is the operational process of converting raw data from its original format into usable datasets through cleaning, normalizing, validating, and aggregating. Think of modeling as the recipe that defines what the end result should look like, while transformation is the actual cooking process that implements that vision through SQL or Python code. **At which stage of the data pipeline does data modelling occur compared to data transformation?** Data modeling occurs as an upfront design discipline that establishes the architectural vision before transformation work begins. It defines the overall structure and organization of your data warehouse. Data transformation then implements this vision across multiple stages within ELT pipelines: starting with staging models that clean and standardize raw data, moving through intermediate models that handle complex logic, and culminating in mart models that apply business logic to create final analytical assets like fact and dimension tables. The modeling decisions guide how transformation code is written at each stage. **Can robust data modeling reduce the need for heavy transformations, or do well-modeled schemas still require layered transformation pipelines to move data from raw to refined?** Well-modeled schemas still require layered transformation pipelines. Data modeling and transformation serve complementary but distinct purposes (modeling provides the architectural blueprint while transformation provides the execution). Even with excellent modeling, you still need transformation processes to clean raw data, handle source system changes, apply business logic, and move data through staging, intermediate, and mart layers. However, good modeling decisions can influence transformation complexity: for example, choosing denormalized tables may reduce downstream join requirements, and thoughtful staging layers can absorb source changes to protect downstream models. The separation of concerns across transformation layers actually creates modularity that scales more effectively than monolithic approaches. --- --- title: "Finding your people in data (and building a community that actually sticks)" description: "The View on Data explores how data leaders find community, rethink mentorship, and build connections." url: "https://www.getdbt.com/blog/finding-your-people-in-data" date: "2026-01-13" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Finding your people in data (and building a community that actually sticks) In the latest episode of **The View on Data**, hosts Jerrie Kenney, Erica “Ric” Louie, and Faith McKenna get real about one of the most underrated career skills in data: finding community, and redefining what mentorship can look like when you don’t have a magical “career sage” on speed dial. From local meetups and data Twitter chaos, to the dbt Community Slack, the group breaks down how relationships form, how they last (or don’t), and why finding community is important. 🎧 Listen & subscribe: [Spotify](https://open.spotify.com/show/0Itxg3BoF6zaQdORfVuY6I?si=6133f88257624765) | [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-view-on-data/id1835755886) | [Amazon Music](https://music.amazon.com/podcasts/84c47dbe-adf7-4d33-a371-a1fe55255495/the-view-on-data) | [YouTube](https://youtube.com/playlist?list=PL0QYlrC86xQk2WktL3vbdB1FcWlpay7VQ&si=XIDBd1XcoDIEhnm5) [Watch video](https://youtu.be/5oiJ2MVvDeM) ## What community‌ means when you work in data The conversation started with a deceptively simple question: **What does community mean in the data space?** For Faith, it clicked at dbt’s annual conference. It was the first time she met people who had the same day-to-day work struggles of debugging, stakeholder pressure, and the weird mental load of being “the data person.” Community stopped feeling like networking and started feeling like relief. Ric shared a different angle: when you’ve always worked inside dbt, it’s easy to take the community for granted. But once you compare notes with friends in other careers, you realize how rare it is to have a space where people will openly share patterns, lessons, and hard-earned shortcuts without gatekeeping. Jerrie’s path started locally: meetups, community college conferences, and the kind of sessions that only data people would attend voluntarily (like _“mistakes you’ve made in data”_). That local-first approach later expanded into online communities—Slack groups, Locally Optimistic, dbt Community, and everything in between. The punchline: community isn’t a platform. It’s a feeling. It’s finding people who understand your work without needing a 10-minute explanation first. ## Finding your people is part luck, part reps A recurring theme: you don’t always “strategically network” your way into community. Faith told the story of attending Coalesce solo, feeling awkward, and sitting next to people who looked friendly. Those strangers turned into real friends. Jerrie described the “try everything” phase: joining multiple Slacks, using Donut intros, and taking dozens of conversations, only to realize that not everyone will relate to your specific path. That’s normal. You’re not doing it wrong. You’re just sorting signal from noise. And Ric dropped one of the most useful truths in the episode: Sometimes you only know you’ve found the right people **after** you experience the wrong ones. ## Mentorship isn’t one person. It’s a toolkit Jerrie shared how she couldn’t find the traditional “advisor” style mentor during a career transition, so she reframed the goal. Instead of waiting for someone to map her future, she started prototyping her path, sharing what she learned, and letting the community respond. That shift turned mentorship into something more realistic: - skill-specific guidance - cheerleaders - feedback loops - peers who are learning the same thing at the same time Faith added a tactical insight: asking “Can I pick your brain?” rarely works. Asking a _specific_ question often does. Examples that land: - “You transitioned from teaching into data. How did you do that?” - “You always give me advanced SQL feedback. Where did you learn those patterns?” - “I’m applying to your team. What should I know about the role and the culture?” Mentorship becomes dramatically easier when you give people something concrete to respond to. ## Maintaining community without burning yourself out Faith shared a low-lift way she stays connected: a weekly “wins of the week” thread in Data Angels. It keeps relationships warm without requiring constant 1:1 meetings. (Also: wins range from “CFO loved my dashboard” to “I learned to ice skate,” which is exactly the right energy.) Ric was honest about capacity: for the last couple years, she’s had more bandwidth to invest in **personal** community than professional. And that’s fine. Community doesn’t have to look like being active everywhere all the time. The bigger idea the group circled: **kindness scales**. You don’t need a master plan. You need consistent, human interactions that leave doors open—so reconnecting later feels natural. ## Why community is bigger than career opportunities Yes, community can help you find your next job. But the group urged listeners to think bigger than that. Community is also: - a place to **nerd out** with people who actually care - a way to get **thought partnership** when you’re solving hard problems - a sanity saver when you’re a **team of one** - a buffer against isolation when work gets weird - a reminder that your struggles are _normal_ and not uniquely yours If your friends outside of data don’t understand why “we refactored the models and lineage finally makes sense” is a huge emotional win, your community will. ## Episode takeaways - Find people who make the work feel lighter. - Ask specific questions. Respect time. Get better answers. - Stop looking for one perfect mentor. Build a mentorship toolkit. - Stay connected in ways that fit your season. - Be kind. Keep doors open. Small interactions compound. As the episode wrapped, the hosts left a simple challenge: **find one person you admire and ask one specific question about how they got there.** See what happens. --- --- title: "AI agents and the data lake" description: "Lauren Anderson, the head of Okta's enterprise data platform, on why central governance and the semantic layer are so essential." url: "https://www.getdbt.com/blog/ai-agents-and-the-data-lake" date: "2026-01-11" authors: ["Daniel Poppy"] categories: ["Insights"] --- # AI agents and the data lake One of the interesting commonalities of AI and the data lake is that they both require new thinking around how we manage identity. For AI, the big question is how do agents interact with underlying data? For the data lake, the big question is how do we make open data stored outside the purview of any given data platform act like you’d expect? In this episode of The Analytics Engineering Podcast, Tristan talks with Lauren Anderson, who leads the enterprise data platform at identity company Okta. Lauren discusses how identity sits at the center of two seismic shifts in data—AI agents and the open data lake—and why central governance and a shared semantic layer are critical. She lays out how analytics engineers and data engineers should divide responsibilities as agents begin to write a growing share of analytical queries. A lot has changed in how data teams work over the past year. We’re collecting input for the [2026 State of Analytics Engineering Report](https://forms.gle/KBU9smukSfiK1g4W7) to better understand what’s working, what’s hard, and what’s changing. If you’re in the middle of this work, your perspective would be valuable. [Take the survey](https://forms.gle/Jc54NuP96qekHU9j7) _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [Youtube](https://www.youtube.com/playlist?list=PL0QYlrC86xQm83Q9deiy4euEnbw8ceu3I) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) [Watch video](https://youtu.be/sa-BJkM75TQ) ## Key takeaways ### Tristan Handy: Before we dive into the current day, can you share a little bit about your background and how you came to the role that you’re in today. **Lauren Anderson:** I’ve had a 20‑something year career at this point. I have basically spent my entire career in analytics some way, but my first data job was at a big bank. I won’t name it. There’s only a few big banks you could probably guess. I worked for the finance org and I did compensation planning and administration, with a side of sales tracking and analytics. I was part database analyst, part customer support for people that made a lot more money than I did. I was there for seven, seven and a half, eight years. Towards the end of it, I became the owner and creator and almost business architect for our brand‑new sales tracking data warehouse. At a very young age, I got to think about how relational databases should come together for the outcome of both analytics and reporting—dashboards and whatnot—but also operations, which was paying compensation every month. It got me super excited about this world of data and being able to architect pipelines and the end‑to‑end flow for real‑world outcomes. ### What do you think allowed you to be successful in that era? I often think the things that enabled success then aren’t the same as what make data folks successful today. When I took it over, we ran compensation out of an Access database. I was new, the person who designed it left, and there wasn’t much documentation. It worked the first month, then broke the second—right before a payroll deadline. I rebuilt it as a long series of SQL queries with inline comments and step‑by‑step checks that produced a clean file. That willingness to throw away the brittle thing and rebuild with clarity and documentation gave me early success. The meta‑skills:ability to learn, take chances, and figure out the best path—still apply, but the technology is completely different now. ### You’ve split time at Okta into two stints. How would you characterize the work? Okta was my first truly B2B company. I realized quickly B2B data is my sweet spot. I love thinking about customers as businesses and how business users interact with our products and features. Okta data is complex—many products, features, and highly configurable use cases—especially with large customers. That variety is exciting. In simpler retail flows you see a lot of the same patterns; in B2B, the variety is the appeal. ### What’s your current role? I lead our enterprise data platform, engineering, and architecture function. For enterprise data used to make business decisions, we own ingestion into the warehouse, transformations, and delivery—dashboards, reverse ETL to third‑party applications, other data stores, and internal apps. ### How big is the central function and how do you engage with the business? We’re about 50 people across data engineering and analytics/data science in a company south of 7,000 employees. We support every business unit. Engagement spans a maturity curve. One end is platform self‑service: teams land data via approved connectors, build transformations in dbt on our implementation, and build dashboards in Tableau we administer. Governance and roles are defined centrally, and teams assign people to those roles. The other end is a white‑glove model where we partner through the full lifecycle—question, discover existing assets, requirements, data work, build, interpretation, validation, and end‑of‑life of the data product. Our sweet spot is the middle: we own enterprise “gold” pipelines for company‑level metrics—monitored and governed—while domains build and later graduate via a path‑to‑production under stronger governance. ### Okta is known for identity and security. How does security‑first actually work in practice? Reinventing controls every time slows you down. We invest in repeatable frameworks. Any new source goes through third‑party risk review, classification, and decisions on masking or exclusions. We help teams through that; after a couple times, they can engage directly with risk while we stay in the loop and monitor. As our classifications and expectations got clearer, review cycles shrank from weeks to days. It’s not all roses—it takes time—but we all operate as security practitioners. That shared mindset builds trust and reduces corner‑cutting. ### How much do users need to know? We don’t expect everyone to know everything. We provide dbt frameworks and minimum testing standards, plus SMEs to guide teams. The culture is to ask when unsure. ### Will agents write more analytical queries than humans in the next 12–24 months? Macro, yes. For us, more like 24–36 months because we’re careful. The key is safe, ethical AI consistent with being a security company. ### How are you thinking about agent access? Central governance. Ideally, agents query centralized, agent‑ready stores. Run governance once: policies, roles for users and for data, tracking and logging on a central plane. The semantic layer is essential. Creating semantic views must get easier and more automated, and semantics should inform policy application. ### Why are agents different from humans in access patterns? Row‑level security to the extreme. Conversational intelligence data should be limited to what the requesting user can access. Aggregations could be broadly accessible with anonymization, but detailed content should remain constrained. You might also limit allowed functions on large unstructured objects. Identity for agents matters—Okta Secures AI looks at distinct identity patterns to secure agents across applications. ### Where are you with MCP and agent building? Early, building support and insight use cases. Progress is fast, but nothing broad in production yet. ### How should analytics engineers and data engineers participate? Analytics engineers should own semantics—tooling, vendor choices, onboarding use cases, and the shared business language. Data engineers should optimize for consistency and scale, notice overlap across agents, and provide a platform others can build on with confidence in governance and security. ### Will you standardize an agent development platform? Yes, in partnership with engineering and shared services. Our current pull skews to the business, so we’re leaning toward accessible, governed platforms that serve both business and engineering with central governance. ### Any assumptions you’re rethinking? Treating everything like a relational model. Many initial agent questions are intentionally simple, where speed and reasonable accuracy trump perfect sophistication. The important thing is to start, observe, and mature. ## Chapters 00:02:28 — From bank analytics to owning a sales DW 00:05:00 — Rebuilding brittle Access → SQL with documented checks 00:08:30 — Ops accountability then vs. optimization today 00:11:00 — TripIt, marketing analytics, and moving into tech 00:13:14 — Why B2B data became Lauren’s sweet spot 00:16:00 — Current role: ingestion → transform → delivery at Okta 00:18:10 — Operating models across business units and the path to production 00:22:20 — Security-first in practice: repeatable frameworks over friction 00:24:23 — Third‑party risk, classification, and shrinking review cycles 00:28:00 — Policies, masking, and the need for a central governance plane 00:30:20 — Frameworks for dbt, testing, and SME guidance 00:32:11 — Will agents outwrite humans? Macro yes; Okta timeline nuance 00:33:48 — Central governance and agent access patterns 00:37:19 — Semantic layer as bridge and policy carrier 00:41:00 — Function limits on unstructured data and Okta Secures AI 00:42:35 — Early MCP experimentation and support use cases 00:43:03 — Roles: analytics engineers (semantics) and data engineers (scale) 00:46:10 — Enabling an org-wide agent platform with shared governance 00:47:43 — Solve governance once, serve business and engineering 00:49:30 — Simpler questions first; rethinking relational assumptions --- --- title: "Understanding LSP: What it is, and what you can use it for" description: "Learn how the Language Server Protocol (LSP) works and how it enables smarter code editing." url: "https://www.getdbt.com/blog/language-server-protocol" date: "2026-01-09" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Understanding LSP: What it is, and what you can use it for A good code editor can make or break a developer’s productivity. Writing code in a blank text file without hints or warnings slows even experienced developers. Code editors provide features like error checking and suggestions that plain text editors don’t. These features depend on language tooling that parses and compiles code behind the scenes. Previously, language providers built separate extensions for each editor because editors exposed different APIs for language features. The Language Server Protocol (LSP) enables separating language servers to function independently of code editors. This enables language servers like [dbt LSP](https://docs.getdbt.com/docs/about-dbt-lsp) to provide the same language features for different development tools. In this article, we'll explore how LSP works, the problems it solves, and how dbt is leveraging LSP capabilities to make writing data transformation pipelines easier than ever before. ## What is Language Server Protocol (LSP)? LSP is an open protocol that standardizes communication between code editors and language tooling. It uses [JSON RPC](https://www.jsonrpc.org/specification) messages between the development tool acting as the client and a language server. The protocol defines a shared set of messages for language features like code completion and syntax highlighting. Microsoft developed LSP in 2016 for Visual Studio Code to eliminate the need for building separate language services. The goal was to enable a single language implementation to work across multiple development tools. This approach saves both language service providers and development tools from additional effort. The client sends notifications to the server about file changes or cursor positions. The language server processes these requests and responds with precise diagnostics and information. ## How does LSP help software developers? Developers faced [fragmented support and heavy performance loads](https://code.visualstudio.com/api/language-extensions/language-server-extension-guide) before LSP adoption. With LSP, they benefit from the following solutions: - Language providers no longer need to develop separate plugins for each editor. - Heavy tasks like code parsing and syntax analysis run in the language server instead of the editor. This keeps the editor lightweight and responsive. - Developers can more easily integrate linters, formatters, and refactoring tools into their editors. - A language server backend can be written in any programming language. It can then be integrated easily into a variety of tools. - Developers can experiment with various plugins and editors while keeping the same language features and plugin support. In short, LSP makes it easier to support special-purpose development languages without creating a new Integrated Development Environment (IDE) from scratch. ## How does LSP work? LSP defines client-server message exchange using JSON-RPC over stdio or socket transports. Each message contains headers specifying Content-Length and Content-Type before the payload. The payload encodes requests, responses, or notifications following [JSON RPC 2.0 rules](https://www.jsonrpc.org/specification). ![Communication between an LSP client and server](https://cdn.sanity.io/images/wl0ndo6t/main/99669ee86d5d3c30d1ca3237f508c8c0883b3313-1600x595.jpg) ### Communication start Communication starts when the editor launches the language server as a separate process. 1. The client sends an initialization request describing supported features and workspace settings. 2. The server replies with an initialization response declaring supported language features and capabilities. 3. After receiving the response, the client confirms readiness through an initialized notification. From this point on, both sides rely on their declared capabilities to exchange feature messages. ### Document management LSP keeps the server synchronized with editor changes in real time as developers edit files. 1. When a file opens, the client sends a didOpen notification with full document text. 2. As edits occur, the client streams incremental changes using didChange notifications. Closing a file triggers a didClose notification, which the server uses to release the document state. ### Feature interactions & results The client and server communicate through feature interactions triggered by specific actions within the editor. Features such as hover-over tooltips or go-to-definition clicks trigger this exchange. During the exchange, the client requests language intelligence, and the server returns the relevant results. 1. Feature interactions use request-response pairs for tasks such as completion, definition, or references. 2. The request includes the document [uniform resource identifier (URI)](https://www.techtarget.com/whatis/definition/URI-Uniform-Resource-Identifier), cursor position, and method-specific params object. Each request has a unique ID, allowing the client to match responses reliably. 3. Servers analyze documents independently and return diagnostics without executing queries remotely. 4. Editors display results immediately, keeping typing latency low and feedback continuously visible. The protocol supports multiple transports, including standard input output and TCP sockets. This workflow allows a single language server to provide consistent tooling across multiple editors. ## The limitations developers face without LSP diagnostics Real-time feedback and comprehension from a language server help developers address common challenges earlier. Data developers especially face these challenges in general SQL development. They become even more visible when managing complex data transformations or writing intricate SQL logic. Common challenges include: 1. **Runtime error discovery.** Errors in SQL queries or models only become apparent after running them against the data warehouse. This often results in additional compute costs. 2. **Previewing common table expression (CTE) outputs.** Developers comment out parts of queries to preview CTE outputs. This creates a slow and manual debugging process. 3. **Subtle changes impact downstream models.** Renaming columns or modifying joins may cause dependent models to break. There is no early visibility into these effects. 4. **Opaque dynamic SQL.** Some SQL queries are generated automatically, hiding the statements that actually run. This makes it difficult to trace changes across different platforms. 5. **SQL dialect differences.** Platform-specific syntax, like Snowflake versus Databricks, only becomes apparent at runtime. As a result, incompatible functions may go unnoticed during editing. Without these immediate diagnostics, debugging often becomes slower, more costly, and more manual. ## How dbt uses LSP and Fusion to power SQL development There are many SQL linters that can detect simple syntactical issues with SQL code. Until now, however, there hasn’t been a solid LSP implementation that could dynamically detect errors in your SQL code relative to your environment. dbt uses LSP to deliver language tooling features to the [dbt VS Code extension](https://docs.getdbt.com/docs/about-dbt-extension) via the high-performance [dbt Fusion engine](https://docs.getdbt.com/docs/fusion/about-fusion). Fusion is the next generation of our dbt Core engine for data transformation that natively understands SQL across multiple data warehouse engine dialects. Fusion supports most of (and, soon, all) of the core features of[ dbt Core](https://docs.getdbt.com/docs/introduction#dbt-core). This enables using [dbt models](https://docs.getdbt.com/docs/build/models) to transform data with a common syntax, no matter where it lives in your enterprise. Unlike dbt Core, Fusion can locally emulate your remote data warehouse environment and run queries against it without calling out to your remote stores. This means Fusion can detect issues with your code before you even check it in to source control. Here are the features dbt Fusion provides to developers via its use of LSP in VS Code: - **[Live error detection](https://docs.getdbt.com/docs/dbt-extension-features#live-error-detection).** Language server automatically validates your SQL code instantly as you type. LSP detects errors instantly in VS Code, giving developers full project awareness without querying your data warehouse. - **[Lightning-fast parse times](https://docs.getdbt.com/docs/dbt-extension-features#lightning-fast-parse-times).** The dbt Fusion engine processes even the largest projects quickly. It parses up to 30 times faster than [dbt Core](https://github.com/dbt-labs/dbt-core), accelerating the developer workflow even at scale. - **[Powerful IntelliSense](https://docs.getdbt.com/docs/dbt-extension-features#powerful-intellisense).** This feature provides intelligent autocompletion suggestions for SQL functions, model names, and macros. It autocompletes references and source calls, facilitating faster development workflows. - **[Instant refactoring](https://docs.getdbt.com/docs/dbt-extension-features#instant-refactoring).** You can easily rename models or columns using the LSP. Downstream references are automatically updated throughout the entire project, ensuring consistency and validity instantly. - **[Rich lineage context](https://docs.getdbt.com/docs/dbt-extension-features#rich-lineage-in-context).** Project metadata like column-level and table-level lineage is displayed directly within the editor. There’s no context switching or breaking the flow during development. - **[Live CTE previews](https://docs.getdbt.com/docs/dbt-extension-features#live-preview-for-models-and-ctes).** Developers can preview a specific CTE’s output directly. This happens right from inside the current dbt model file. This feature enables faster query validation and debugging cycles. - **[Hover insights](https://docs.getdbt.com/docs/dbt-extension-features#hover-insights).** You can see critical context without ever leaving your code. Contextual details about tables, columns, and functions are visible within the code editor. - **[View compiled code](https://docs.getdbt.com/docs/dbt-extension-features#view-compiled-code).** Developers can get a real-time, side-by-side preview of the SQL code. The compiled SQL reflects what dbt models generate and updates automatically as source files change. ## Leveraging LSP for easier data pipeline development LSP has fundamentally transformed the development and delivery of language support across the industry. Now that LSP is widely adopted, the focus is on enhancing and expanding its capabilities. The limitations of code editors no longer constrain language providers. They can develop sophisticated tooling to help users write more efficient, higher-quality code. dbt uses this standard to deliver a step-function improvement in building data pipelines with dbt Fusion. Experience faster development, fewer mistakes, and deeper project understanding. [Sign up for a free dbt account](https://www.getdbt.com/signup). Then, [get started with dbt Fusion](https://docs.getdbt.com/guides/fusion) and the [dbt VS Code extension](https://docs.getdbt.com/docs/install-dbt-extension) today. --- --- title: "The dbt community chronicle: December 2025" description: "2025 reflections, Community Awards, tools, top discussions, plus the 2026 State of Analytics Engineering survey." url: "https://www.getdbt.com/blog/the-dbt-community-chronicle-december-2025" date: "2025-12-29" authors: ["Bolaji Oyejide"] categories: ["Community"] --- # The dbt community chronicle: December 2025 What's changing in analytics engineering? The 2026 State of Analytics Engineering survey is now live, and your input helps define industry standards for titles, skills, compensation, and tooling. Hiring managers and industry leaders rely on this report. Take 10 minutes to ensure your reality is represented. 👉 [Add your perspective to the 2026 survey](https://forms.gle/Jc54NuP96qekHU9j7?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) **Pro tip:** Share with teammates. A stronger sample size means sharper benchmarks for everyone. ### Analytics data engineer benchmark is live On December 10th, over 700 data practitioners joined the quarterly community webinar featuring Benn Stancil and Jason Ganz for an introduction to the Analytics Data Engineer (ADE) benchmark. **Resources:** - [Watch the webinar replay](https://www.getdbt.com/confirmation/analytics-data-engineer-bench-webinar-recording?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) - [Explore the GitHub repository](https://github.com/dbt-labs/ade-bench?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) - Join the discussion in #tools-ade-bench ## Community contributions ### 2025 reflections: What data practitioners learned In [the dbt Community Slack](https://www.getdbt.com/community?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU), members shared skills, habits, and mindsets gained through the dbt community in 2025: **Eddy Zulkifly** 🇨🇦 _"Using a dev container to run dbt Fusion allows you to have both dbt Core and dbt Fusion running in the same machine..."_ **Marcelo Bour** 🇦🇷 _"Jumping into different threads with real-life data problems is a great way to learn new things in an unstructured but powerful way."_ **Jairus Martinez** 🇺🇸 _"The importance of lowering the access barriers to tools like LLMs and MCPs for agentic dbt dev workflows..."_ **Abraham Setiawan** 🇸🇪 _"I learned to trust incremental changes, adapt to the situation, and share the small wins."_ **Sebastian Freiman** 🇦🇷 _"Desde la buena experiencia que tuve en Coalesce, adquirí el hábito de ir si puedo a eventos in-person para conectar o re-conectar."_ **Angela Alves** 🇧🇷 _"O meetup foi bastante importante para mim esse ano, conectar com mais pessoas do meu país e conversar sobre os desafios que compartilhamos e aprendermos uns com os outros."_ ### Coalesce 2025 community awards The 2025 dbt Community Awards celebrated practitioners who make dbt more than code. These builders, mentors, and community leaders show up daily to help others become self-sufficient data practitioners. **Watch the highlights:** - [Community Keynote with Grace Goheen and Jeremy Cohen](https://www.youtube.com/watch?v=aMUAQjqTKtc&utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) (60 minutes) - [Community Awards Ceremony, with Bolaji Oyejide](https://youtu.be/I-DgySJ0Syg?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) (10 minutes) ### 2025 dbt community award winners 1. **Community Champion:** [Marcelo Bour](https://www.linkedin.com/in/marcelo-bour/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU), Analytics Engineer at Dynamic Data 🇦🇷 2. **dbt Ambassador:** [Bruno Lima](https://www.linkedin.com/in/brunoszdl/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU), Analytics Engineer at phData 🇧🇷 3. **Analytics Engineer of the Year:** [Silja Mardla](https://www.linkedin.com/in/siljamardla/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU), Analytics Engineering Manager at Bolt 🇪🇪 4. **Data Engineer of the Year:** [William Whelan](https://www.linkedin.com/in/william-whelan/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU), Data Engineer at Storypark 🇳🇿 5. **Data Analyst of the Year:** [Millie Symns](https://www.linkedin.com/in/millie-symns/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU), Senior BI Analyst at Justworks 🇺🇸 6. **Design Partner of the Year:** [Peter Empey](https://www.linkedin.com/in/peterempey/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU), Staff Analytics Engineer at GitHub 🇺🇸 7. **Code Contributor of the Year:** [Yu Ishikawa](https://www.linkedin.com/in/yuishikawa0301/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU), Principal Data Architect 🇯🇵 8. **Meetup Organizer of the Year:** [Karen Hsieh](https://www.linkedin.com/in/karenhsieh/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU), Principal Product Manager 🇹🇼 9. **Igniter of the Year:** [Shinya Takimoto](https://www.linkedin.com/in/shinya-takimoto-2793483a/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU), Analytics Engineer at 10X 🇯🇵 10. **Commit Comedian of the Year:** [Raz Widrich](https://www.linkedin.com/in/razwidrich/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU), Head of Brand at Euno 🇮🇱 Connect with these practitioners on LinkedIn to learn from their expertise. ### Community-built tools and packages The #i-made-this Slack channel showcases work-in-progress projects and learning from community builders. **Recent projects:** [**Data News Monitoring Platform**](https://getdbt.slack.com/archives/C01NH3F2E05/p1764942077222179?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) Alphonse Polycea built an automated data news monitoring site and welcomes feedback on features and bug reports. [**superset-chat Plugin**](https://getdbt.slack.com/archives/C01NH3F2E05/p1764685892552389?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) Egor Tarasenko released an open-source plugin unifying dbt lineage with Superset metadata, creating an AI-powered interface inside Superset. [**dbt-ml-eval Package**](https://getdbt.slack.com/archives/C01NH3F2E05/p1764529234586719?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) Matthew Senick published a package simplifying ML evaluation with one macro call generating 27 classification metrics and 9 regression metrics. [**dbt-meta CLI Tool**](https://getdbt.slack.com/archives/C01NH3F2E05/p1764417503476079?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) Pavel Filyanin built a command-line interface for instant metadata extraction from dbt artifacts. [**AI Agent for Data Connectors**](https://getdbt.slack.com/archives/C01NH3F2E05/p1764248607475619?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) Amaan Nawab created an AI agent team that builds data connectors automatically. [**upstream_prod Package**](https://getdbt.slack.com/archives/C01NH3F2E05/p1763475631233459?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) Lewis Davies shared a package eliminating the need for complete, up-to-date production data copies in development schemas, reducing time and warehouse costs. ## Community connections ### Finding your people in a 70,000+ member community **Most active channels:** - #advice-dbt-help - #dbt-fusion-engine - #tools-dbt-mcp - #db-snowflake - #jobs **Popular database channels:** - #db-snowflake - #db-bigquery - #db-databricks-and-spark - #db-redshift - #db-athena **Newest channels:** - #tools-ade-bench - #topic-agentic-analytics - #data-for-social-good ### December's most active community members These practitioners drive engagement, answer questions, and strengthen the dbt community: 1. [Marcelo Bour](https://www.linkedin.com/in/marcelo-bour/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) 2. [Jeremy Chia](https://www.linkedin.com/in/jjchia/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) 3. [Morgan Kerle](https://www.linkedin.com/in/morgan-kerle/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) 4. [Karthik Rajashekaran](https://www.linkedin.com/in/karthikrajashekaran/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) 5. [Matěj Novák](https://www.linkedin.com/in/novakmates/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) 6. [Tobie Tusing](https://www.linkedin.com/in/tobie-tusing-b4817b4a/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) 7. [Noel Gomez](https://www.linkedin.com/in/noelgomez/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) 8. [Amey Khedekar](https://www.linkedin.com/in/ameyakhedekar/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) 9. Gabriel Grunberg 10. [Vince Faller](https://www.linkedin.com/in/vincefaller/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) ### Top new members this month Welcome to practitioners making their presence felt in their first 30 days: 1. Sandeep Uikey 2. Daniel Gebreneset 3. Mitchell Johnson 4. Yi Hahn Pang 5. Mpho Kubeka 6. Michael Mabinuola 7. Viktor Osetrov 8. Elvin Shahsuvarli 9. Sarah Martens 10. Jonathan Markland 11. Ryan Donohue 12. Bruno Freitas **How to become a featured community member:** - Engage actively in Slack channels - Contribute to [dbt's GitHub repositories](https://github.com/dbt-labs?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) - Attend [local dbt meetups](https://www.getdbt.com/events?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) - Share dbt-related work [on LinkedIn](https://www.linkedin.com/company/dbtlabs/posts/?feedView=all&utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) - Participate in the [Data Career Transformations podcast](https://www.youtube.com/playlist?list=PL0QYlrC86xQkz_bHgNz3zFOmC3F1MuKZ8&utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) ## Career development ### dbt certification before year-end December is an ideal time to invest learning and development budgets before they expire. Consider advancing your dbt skills with certification. **Available courses:** - [dbt Canvas Fundamentals](https://learn.getdbt.com/courses/canvas-fundamentals?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) - [dbt Insights Fundamentals](https://learn.getdbt.com/courses/dbt-insights-fundamentals?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) Explore courses from beginner to advanced at [learn.getdbt.com](http://learn.getdbt.com) ## Join the dbt community The dbt community grows through contributions from practitioners worldwide. Whether you're building tools, sharing knowledge, organizing meetups, or advancing your skills, your participation strengthens the entire ecosystem. **Get involved:** - [Join the dbt Community Slack](https://www.getdbt.com/community?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) - Share projects in #i-made-this - [Attend or organize a local meetup](https://www.getdbt.com/events?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) - [Complete the State of Analytics Engineering survey](https://forms.gle/Jc54NuP96qekHU9j7?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) - [Earn your dbt certification](https://learn.getdbt.com/?utm_campaign=Community%20Newsletter&utm_source=hs_email&utm_medium=email&utm_content=2&_hsenc=p2ANqtz-8E71_bQZUFYba39-6NgBQqA3wFJg2JwKVPWYrLrYx70I3ScRFmrjE33T6-D1PWwooXmL-xluUTyRI3EngmjQaLxT3AndBL_w_J-EVo6VVUspByJXU) --- --- title: "Cloud vs on-premise data transformation" description: "Compare cloud and on-premise data transformation to choose the model that fits your team’s performance, security, and cost needs." url: "https://www.getdbt.com/blog/cloud-data-transformation" date: "2025-12-22" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Cloud vs on-premise data transformation Successful businesses anticipate customer needs and deliver them. Some businesses, however, go even further, anticipating these needs before customers even express them. The key to this lies in using data strategically. In fact, [65% of highly data-driven businesses report financially outperforming their peers](https://aws.amazon.com/blogs/smb/data-driven-smbs-financially-outperform-their-competitors-and-the-gap-is-widening/), nearly double the share of less data-driven companies. This shows that having the right data and knowing how to use it is one of a company's most substantial competitive advantages. At the core of this advantage lies data transformation: the essential process of converting raw datasets into meaningful insights that support decision-making. Teams can run these transformations in the cloud or on-premises. Each approach is beneficial depending on business needs. This article explores both options to help you select one that best supports your team's goals. ## What is cloud data transformation? Cloud data transformation refers to the process of transforming data entirely within cloud platforms, which offer on-demand compute on a pay-as-you-go pricing model. It uses cloud-based services to process, clean, and reshape data for downstream use. Cloud solutions don't require significant upfront costs or capital expenditure to create a massive data center. They also avoid the problem of overplanning capacity, where compute runs idle during less heavily-trafficked periods. As such, cloud computing enables automated data transformations for the large data volumes required by modern enterprises due to its scalable architecture and cost-effective, high-performance compute power. These cloud-based solutions provide highly scalable tools that require less setup time and manual work, with most cloud providers providing out-of-the-box support for common tools and technologies such as data warehousing, streaming, data processing, and more. Cloud data transformations run in cloud data warehouses like [Snowflake](https://snowflake.com) and [Amazon Redshift](https://aws.amazon.com/redshift/), where compute resources scale independently of data storage. Cloud service providers like [AWS](https://aws.amazon.com/) and [Microsoft Azure](https://azure.microsoft.com/) offer these cloud platforms with robust infrastructure. Teams can modernize legacy applications, migrate non-critical data, and streamline IT operations. ## What is on-premises data transformation? On-premises data transformation is the process of transforming data within the organization's local infrastructure. Organizations gain complete ownership of their environment. They design and implement custom security protocols that meet specific compliance requirements. Hosting and managing IT infrastructure in-house enables more direct control over hardware security. This helps ensure compliance with regulations and standards such as [GDPR](https://gdpr.eu/what-is-gdpr/), [PCI-DSS](https://www.pcisecuritystandards.org/), and [HIPAA](https://www.hhs.gov/hipaa/index.html). Teams can manage audit trails more effectively by keeping the organization's data on the company's own servers within private data centers. ## How cloud and on-premises data transformation differ ![Differences between cloud and on-premises data transformation](https://cdn.sanity.io/images/wl0ndo6t/main/d18006d8e863a0d040d8729c73f3d7d0a9120b60-1600x728.jpg) Both on-premise and cloud data transformation models affect how businesses manage their daily operations, including resource allocation, availability, security, and maintenance. Understanding the key differences helps you evaluate which approach best serves your business needs. ### Scalability and cost model **Cloud.** Cloud data transformation operates on an [operational expense (OPEX)](https://www.investopedia.com/terms/o/operating_expense.asp) model. Companies pay only for the resources they consume with pay-as-you-go pricing. Compute power and data storage adjust on their own to keep up with changing data volumes and processing needs. This saves businesses from investing in additional hardware. The scalability of cloud infrastructure allows you to optimize cloud costs by scaling resources up during peak workloads and down during quieter periods. When evaluating the total cost of ownership (TCO), cloud environments often prove more cost-effective for variable workloads. **On-premises.** On-prem solutions require [capital expenditures (CAPEX)](https://www.investopedia.com/terms/c/capitalexpenditure.asp) for hardware and ongoing maintenance. Managing high data volumes requires physical hardware upgrades and infrastructure investments. The downside is that organizations pay for capacity even when it isn't being used. For example, a company must still pay for servers it purchased to accommodate Christmas shopping traffic even after the Christmas season ends. However, on-premise setups can offer better pricing predictability and control over the IT infrastructure lifecycle. ### Deployment speed **Cloud.** Cloud-based transformations can be provisioned and executed quickly. They use automated workflows and prebuilt services to manage large datasets efficiently. Fast deployment provides businesses with a significant competitive advantage. One trade-off with cloud-based services is that latency between your network and your cloud provider may slow data-intensive operations. However, cloud environments support real-time data processing for most use cases, and low latency can be achieved through strategic placement of cloud resources. **On-premises.** The setup time for on-premises deployments is lengthy due to hardware configuration and manual processes. However, the main reward of on-premise implementations is the control that IT teams gain. Access to on-site servers is also orders of magnitude faster than access to cloud services, making on-prem ideal for low-latency workloads. ### Maintenance and management **Cloud.** Cloud services transfer infrastructure management responsibilities to the service provider, including updates, patching, and security monitoring. There's no overhead of physical servers to purchase or set up. This makes cloud solutions cost less and easier to manage. Cloud service providers handle routine upgrades, backups, and disaster recovery, allowing your IT team to focus on higher-value work. The cloud ecosystem provides automation tools that further reduce maintenance overhead. **On-premises.** On-premise systems require skilled IT staff for continuous maintenance, monitoring, and troubleshooting. The team has physical access to both hardware and software to resolve issues. While this requires more in-house resources, it provides complete control over the maintenance lifecycle and upgrade schedule. ### Security and control **Cloud.** Cloud providers offer many of the same security controls used in on-premises environments as fully managed services. They secure the underlying infrastructure, including data centers, hardware, and monitoring systems. Cloud security follows a [shared responsibility](https://learn.microsoft.com/en-us/azure/security/fundamentals/shared-responsibility) model. Organizations are still responsible for securing their own data and applications. Teams must configure identity and access controls while ensuring compliance with both internal policies and industry regulations. One consideration with cloud systems is that, by default, most workloads run on shared public cloud infrastructure. While cloud providers maintain a logical barrier between customers, this might not be enough reassurance for companies with highly sensitive data workloads. However, private cloud and hybrid cloud options provide additional security and data security layers for sensitive data. **On-premises.** On-premises transformation provides full control over data. With physical control over all hardware and infrastructure components, teams can implement customized security protocols and maintain full auditability. This makes it easier for businesses to comply with regulations that require strict handling of financial or personal data. On-premises infrastructure keeps sensitive data on-site with redundancy measures fully under your control, which can be critical for healthcare and other regulated industries. ### Flexibility and innovation **Cloud.** Cloud environments support rapid experimentation and integration of new tools in AI and analytics. Teams have the luxury of experimenting and testing at a much faster pace. This makes cloud platforms a perfect option to drive innovation and speed up deployments. Cloud computing offers a vast ecosystem of services and integrations. The multi-cloud and hybrid approach options give you flexibility to choose best-of-breed solutions without vendor lock-in. **On-premises.** On-premises environments are constrained by existing infrastructure limitations and slow adaptation cycles. Advanced analytics access is limited due to hardware restrictions and the workaround of software upgrades. On the flip side, having full control over the environment allows teams to customize processes when upstream workflows or data sources change. They can develop stable workloads, integrate seamlessly with legacy systems, and avoid service disruption risks. On-prem data stays within your controlled environment, which some organizations prefer for critical systems. ## Selecting the right data transformation model for your team Today's businesses operate increasingly in digital environments. Making the right decision between on-premises and cloud data transformation is becoming increasingly important. This decision can influence everything from operational efficiency to financial outlay. These considerations will help you select the right model for your team: - **Data sensitivity and location.** Businesses managing highly sensitive and confidential data, such as in healthcare or finance, often favor on-premises storage. This ensures no third-party vendor has access to sensitive data, like patient records in healthcare, which is particularly important for compliance. Keeping the organization's data on private servers ensures teams maintain full control and protection. - **Regulatory compliance.** Security and compliance are critically important in regulated sectors such as healthcare and finance. Standards like HIPAA, PCI DSS, and GDPR govern how data must be stored, processed, and accessed. Although cloud providers offer configurable deployment options to support compliance requirements, organizations with strict data sovereignty or local residency requirements often prefer on-premises infrastructure. Storing data in specific locations and maintaining direct control ensures compliance with local regulations and simplifies audits, particularly when sensitive data must remain within defined jurisdictions. - **Skills and management approach.** Teams should assess whether they have the technical expertise to manage infrastructure themselves. If not, they may rely on managed services. Cloud solutions help businesses with limited IT resources to maintain their infrastructure more easily. Consider what your IT team can realistically support given their current skills and capacity. - **Workflow complexity and automation.** The nature of your data transformation workflow matters. If your top goal is to automate and streamline processes, then cloud platforms provide integrated orchestration and scheduling. On-premises systems instead require extra configuration and customization to support complex data pipelines. Cloud-based automation can significantly reduce manual work. - **Scalability and flexibility.** Long-term adaptability is important as business requirements and data sources change. Organizations anticipating growth should plan for scalable transformation models. Cloud platforms offer elastic scaling and enable efficient processing of large data volumes, whereas on-premises solutions provide predictable performance for stable demands. The scalability of cloud infrastructure makes it ideal for rapidly growing workloads. - **Performance and latency needs.** Consider your performance requirements. On-premises setups can deliver low latency for local workloads, while cloud platforms excel at distributed, high-performance computing. Real-time use cases may benefit from hybrid approaches that optimize for both speed and scalability. - **Total cost considerations.** Evaluate both upfront costs and ongoing expenses. Cloud pricing models offer pay-as-you-go flexibility, while on-premise solutions involve higher capital expenditure but potentially lower long-term costs for steady workloads. Understanding TCO helps you make cost-effective decisions. - **Disaster recovery and backup strategy.** Cloud platforms typically include built-in disaster recovery and backup capabilities with off-site redundancy. On-premises environments require you to implement and manage your own backup and disaster recovery plans. - **Adopting a hybrid strategy.** Companies can implement a hybrid cloud migration if moving everything to the cloud is impractical. This hybrid approach lets you get the best of both worlds. Migrate in phases, keeping sensitive data on-premises while shifting analytics workloads or less critical data to the cloud. This adds cloud scalability and agility for selected tasks while keeping core systems on-premises for security and control. A hybrid cloud strategy offers flexibility to optimize each workload based on its specific requirements. ## dbt: the modern standard for data transformation - whether cloud or on-premise dbt is an open-source data transformation tool. It processes data after it's loaded into a SQL data warehouse. dbt follows the [extract, load, transform (ELT)](https://www.getdbt.com/blog/extract-load-transform) model, where data is first loaded and then transformed within the data warehouse. dbt provides data analysts with a platform to [transform](https://www.getdbt.com/blog/data-transformation), [test](https://docs.getdbt.com/docs/build/data-tests), and [document datasets](https://docs.getdbt.com/docs/build/documentation) using modular, reusable, and [version-controlled](https://docs.getdbt.com/docs/cloud/git/version-control-basics) SQL scripts. These features make collaboration easier for multiple data teams and optimize your data transformation workflow. dbt also supports automated testing to validate transformation logic and data quality. This helps in maintaining the overall reliability of the data workflow. Additionally, it generates comprehensive documentation that provides visibility into transformation processes and data lineage. ### dbt Core vs dbt: Key differences dbt provides you with both on-premises and cloud data transformation options with [dbt Core](https://github.com/dbt-labs/dbt-core) and [dbt](https://docs.getdbt.com/docs/cloud/about-cloud/dbt-cloud-features), our hosted version of dbt Core. Below is a table summarizing the key differences between them: ## Conclusion Choosing between on-premises and cloud data transformation requires careful consideration. The best choice depends on how a business wants to operate and produce outcomes. With cloud data transformation, adjusting computing resources is more scalable and flexible. Businesses can seamlessly scale resources up or down based on current needs and optimize costs through pay-as-you-go pricing models. Cloud platforms offer rapid provisioning, extensive ecosystems, and built-in automation. In contrast, on-premises data transformation is more suitable for strictly regulated and security-conscious objectives. On-prem solutions provide complete control over data security, infrastructure, and compliance. They work well for organizations with stable workloads, specific latency requirements, or regulatory constraints. The decision between dbt Core and dbt involves similar trade-offs. Teams can align their choice with their skill sets, security policies, and operational priorities. Either option helps teams run data transformations reliably and at the speed they need. Whether you choose cloud or on-premise, dbt gives you the tools to optimize your data transformation workflow. To get started with dbt in the cloud, [sign up for the dbt service](https://www.getdbt.com/signup) today. Or [download and install dbt Core](https://github.com/dbt-labs/dbt-core) to implement your own on-premises solution --- --- title: "What’s new in dbt - December 2025" description: "Catch up on the latest in dbt: from Fusion and AI updates to Core 1.11 previews and new partner integrations." url: "https://www.getdbt.com/blog/whats-new-in-dbt-december-2025" date: "2025-12-19" authors: ["Sara Gawlinski"] categories: ["Product"] --- # What’s new in dbt - December 2025 The year is wrapping up 🎁 and so is another big season of shipping at dbt. We’re coming off an action-packed Coalesce, where we shared major updates for the dbt Fusion Engine, AI, and our vision for open data infrastructure. Before we all unplug and recharge, here’s a final roundup for 2025 of what’s new in dbt, including improvements to Fusion, dbt Catalog, the dbt MCP Server, dbt Core 1.11, and new partnership updates that are already shaping how teams will work in 2026. Let’s finish the year strong. ## ICYMI: Coalesce 2025 At Coalesce this year, we showed how dbt is rewriting expectations for how data work gets done. In our [biggest keynote ever](https://www.youtube.com/watch?v=KhBsI2LQQ90), we unveiled faster development with Fusion, cost-saving pipelines through state-aware orchestration, and how governed AI experiences can be powered by the dbt MCP Server. We also announced our plans to set the standard for open data infrastructure. Together, these updates mark a major step toward faster, more intelligent, and more interoperable analytics in the AI era. Check out our [recap blog](https://www.getdbt.com/blog/coalesce-2025-rewriting-the-future) for all the details. ## Fusion A lot has landed in [Fusion](https://www.getdbt.com/product/fusion) recently: Custom materializations, Iceberg, snapshots, exposures, and big performance wins. If you want to stay in the know as new features and improvements land, be sure to follow the [Fusion Diaries newsletter](https://www.linkedin.com/newsletters/fusion-diaries-7366935294090084360/). Fusion is currently in Private Preview for eligible projects. [Review this checklist](https://docs.getdbt.com/docs/fusion/fusion-readiness) to get your projects Fusion-ready. - **Iceberg support**: You can now materialize models to Iceberg tables format as well as take advantage of dbt's catalogs.yml abstraction, making it easier to take advantage of open, performant table standards alongside Fusion’s modern architecture. - **Custom materializations**: Fusion now supports custom materializations to bring more parity to dbt Core and give teams the flexibility to define and extend how their models are built while still benefiting from Fusion’s faster parsing and richer context. Because Fusion cannot know whether a custom materialization will modify the schema of the persisted object, it will treat these nodes like introspective queries and disable static analysis. - **Exposure support**: Exposures are now supported in Fusion to enable full lineage visibility and governance across downstream dashboards, reports, and apps. ![Exposures shown in Fusion-powered lineage](https://cdn.sanity.io/images/wl0ndo6t/main/53759c112d2c9abc1a2b18bc29532e0e9c2ec002-1990x1295.png) - **[Enhancement] New selector method for exposures**: You can now select exposures directly via selectors, which makes it easier to target or test analytics assets that depend on your dbt project. - **Auto-upgrading packages with dbt-autofix**: [`dbt-autofix`](https://github.com/dbt-labs/dbt-autofix) now automatically upgrades packages for you to reduce the manual cleanup required to make your project Fusion-compatible. - **Custom snapshot support**: Fusion now supports custom snapshot strategies to allow teams to carry over advanced snapshot logic while benefiting from Fusion’s stateful execution model. - Fusion now supports **Snowflake Dynamic Tables** and **Databricks materialized views**. This means teams can use the same near-real-time materializations they use in dbt Core, but with Fusion's faster parsing and richer execution context. - **Event logging improvements**: We’ve improved event logs for Fusion runs to give teams clearer, more actionable logs that make debugging and observability easier across complex projects. - **Performance improvements**: Fusion delivers significant memory reductions in the Language Server Protocol (LSP), which makes developer experiences in VS Code and the CLI faster, lighter, and more stable, especially for large projects. ## dbt Semantic Layer + AI Here’s what’s new in the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) and [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) as we deepen our investments in AI-assisted development and governed agentic workflows. - **User PAT authentication for Semantic Layer** queries enable user-level authentication and reduce the need for sharing tokens between users. When you authenticate using PATs, queries are run using your personal development credentials. - **The dbt Semantic Layer GraphQL API now has a [`queryRecords`](https://docs.getdbt.com/docs/dbt-cloud-apis/sl-graphql#query-records) endpoint.** With this endpoint, you can view the query history both for Insights and Semantic Layer queries. - **New Admin API tools via MCP**: You can now automate key dbt workflows directly through the dbt MCP Server—list jobs, trigger runs, retrieve run details, cancel or retry runs, and manage artifacts. - **New project introspection tools:** MCP now supports a suite of introspection methods for local and remote projects, `get_model_lineage_dev`, `get_macro_details`, `get_seed_details`, `get_semantic_model_details`, `get_snapshot_details`, and `get_test_details.` This makes it easier to build intelligent agents and automated development workflows on top of dbt. ## dbt Catalog Here are the latest improvements to [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) as we expand its capabilities for multi-project exploration and governed data discovery. - **Cross-project lineage expansion**: Users can now traverse lineage across projects by opening project nodes directly in the lineage view. This lays the groundwork for better lineage exploration and more seamless exploration without navigating large, browser-crashing DAGs. ![Cross-project lineage expansion](https://cdn.sanity.io/images/wl0ndo6t/main/efa4d6329b31b3f8e87587ed82a59e8f4fcdfa02-1769x842.png) - **Favorites functionality**: Now that we have global navigation, there's much more to find compared to our previous project-only view. With favorites, users can favorite projects and assets to better organize their global navigation at the top of the file tree. - **State-aware orchestration status visibility**: Reused statuses now appear in DAGs for Fusion projects that use state-aware orchestration. These are considered in our health signals and in lineage lenses (with more improvements coming). ![State-aware orchestration status visibility](https://cdn.sanity.io/images/wl0ndo6t/main/713dbbea6a7323432d558e84aa51a8e1f4a566a7-1528x882.png) - **More granular source health signals**: Rather than simply indicating when sources are stale, we’ve expanded source health signals based on user feedback. Stale sources now break into two additional statuses—unconfigured and expired/not run recently—to give users more actionable insight into why a dependent model has a cautionary status due to its sources. ![More granular source health signals](https://cdn.sanity.io/images/wl0ndo6t/main/0d337f04b6abef32f3b8d0acadd86d8d98ebc35c-634x242.png) ## Platform We’ve also shipped some major improvements to security, connectivity, and admin configuration in dbt platform. - New [private connectivity options](https://docs.getdbt.com/docs/cloud/secure/about-private-connectivity#private-connectivity-feature-matrix) across [multi-tenant (MT) and single-tenant (ST)](https://docs.getdbt.com/docs/cloud/about-cloud/tenancy) instances. - Improved configuration for platform metadata credentials to make it easier to manage and rotate them. - Enhanced SSO configuration options for smoother enterprise administration. - All multi-tenant accounts now use static subdomains (e.g., `abc123.dbt.com`), to improve reliability and simplify network allowlisting. ## **dbt Core 1.11** Here’s a preview of what’s new in dbt Core 1.11—moving to GA this week—a release packed with quality-of-life improvements, new UDF capabilities, and several bug fixes courtesy of “Debug-cember”. Check out the [upgrade guide](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.11) to take advantage of these new capabilities. - **First-class UDFs (now more use cases!)**: Building on what we introduced in beta at Coalesce 2025, the final Core 1.11 release expands UDF functionality with support for Python UDFs, default arguments, and richer configuration options to make functions more powerful and more portable across your dbt projects. - **Deprecation warnings by default**: JSON schema validation warnings for YAML configs are now enabled out of the box to help teams catch outdated or incorrect configurations earlier in development. - **A _lot_ of bug fixes (hello, Debug-cember!)**: The community and dbt Labs team spent December hammering through long-standing issues across parsing, execution, logging, error messages, and more. (To see the full list, head over to #dbt-core-development in Slack.) Core 1.11 rolls up a ton of these fixes into one release to improve overall stability and developer experience. - **Adapter-specific improvements**: - **BigQuery**: Batched source freshness for improved performance and reduced API overhead. - **Snowflake**: Improved support to materialize Iceberg tables via a Glue catalog–linked database. - **`cluster_by` support in dynamic tables**: You can now specify the `cluster_by` configuration when working with dynamic tables to give teams more control over how data is organized and optimized. For more information, see [Dynamic table clustering](https://docs.getdbt.com/reference/resource-configs/snowflake-configs#dynamic-table-clustering) in docs. - **Spark:** [New profile configurations](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.11#spark) have been added to enhance [retry handling for PyHive connections](https://docs.getdbt.com/reference/resource-configs/spark-configs#retry-handling-for-pyhive-connections). ## Partnerships - **Microsoft Fabric:** [dbt jobs in Microsoft Fabric](https://www.getdbt.com/blog/dbt-labs-integrates-dbt-fusion-engine-in-microsoft-fabric) are now available in public preview to bring deeper performance optimizations and native integrations across the Fabric ecosystem. This first phase of the integration uses dbt Core to orchestrate transformation workflows directly within Fabric. Microsoft and dbt Labs are actively collaborating to expand this experience for the dbt Fusion engine planned for 2026. - **Databricks:** A new dbt platform Task Type is now available in beta for Databricks Lakeflow. This makes it easier for joint users to orchestrate dbt models as first-class tasks within Lakeflow pipelines. This improves reliability, observability, and end-to-end data workflow management across the Databricks Data Intelligence Platform. Learn more about this task type in the [announcement blog post](https://www.getdbt.com/blog/dbt-databricks-jobs). ## What’s next We’re excited to hear your feedback on these new features! In the meantime, don’t miss our upcoming webinar, [**Maximizing the business value of your data platform with dbt**](https://www.getdbt.com/resources/webinars/maximizing-the-business-value-of-your-data-platform-with-dbt) on January 20th & 21st. Save your spot for the live session to learn how top teams cut rework, reduce costs, and turn technical wins into real business value. Ready to get hands-on? - [Take Fusion for a spin](https://docs.getdbt.com/docs/fusion/get-started-fusion) in your dbt projects and try out the latest features. - [Join the dbt Community](https://www.getdbt.com/community/join-the-community) Slack to share feedback, swap ideas, and stay in the loop. Stay tuned for more dbt updates in the new year! ❄️✨ --- --- title: "dbt Labs expands ISO certifications" description: "dbt Labs expands ISO certifications with new standards for cloud security, privacy protection, and AI governance." url: "https://www.getdbt.com/blog/dbt-labs-expands-iso-certifications" date: "2025-12-19" authors: ["Randy Hanooman"] categories: ["Product"] --- # dbt Labs expands ISO certifications dbt Labs has expanded its security and compliance portfolio with three additional ISO certifications: ISO 27017:2015, ISO 27018:2025, and ISO 42001:2023. These certifications complement our existing ISO 27001:2022 and ISO 27701:2019 certifications to reinforce our commitment to the highest standards of security, privacy, and now artificial intelligence governance. These certifications represent a significant evolution in how dbt Labs protects your data and AI workflows. We're providing you a unified security, privacy, and AI governance framework. This means you can use dbt's cloud services and AI features knowing that specialized controls protect your data operations, ensure responsible AI use, and meet your privacy obligations; all verified by independent auditors to international standards. ## New ISO certifications **ISO 27017:2015** This certification provides guidelines for information security controls applicable to cloud services to secure your data in cloud environments. It extends the existing ISO 27001 framework with cloud-specific security controls, ensuring that our cloud service delivery meets international best practices. **ISO 27018:2025** This standard focuses specifically on protecting personally identifiable information (PII) in public cloud environments. By achieving this certification, we confirm our adherence to a comprehensive set of controls designed to protect your sensitive data and maintain privacy in cloud services. **ISO 42001:2023** ISO 42001:2023 is the world's first international standard for Artificial Intelligence Management Systems (AIMS), providing a comprehensive framework for organizations to develop, deploy, and manage AI systems responsibly. This certification demonstrates that dbt Labs has implemented: - **Structured AI governance** with clear policies, procedures, and accountability for AI systems - **Risk management frameworks** specifically designed to identify and mitigate AI-related risks - **Ethical AI principles** to ensure fairness, transparency, and explainability in our AI implementations - **Continuous monitoring** of AI systems to ensure they perform as intended and remain aligned with our values - **Stakeholder engagement** processes to consider the impact of AI on customers, employees, and communities This means dbt Copilot AI-powered features—including context-aware code generation, documentation assistance, ~~and~~ intelligent recommendations, and AI-assisted workflow automation—are developed and deployed with rigorous governance controls. We've established clear guidelines for AI development, regular audits of AI system performance, and mechanisms to address potential biases or unintended consequences. ## What this means for dbt customers These expanded certifications provide several key benefits: - **Enhanced cloud security** with specialized controls designed for cloud service environments - **Stronger privacy protections** for personal data processed in dbt’s cloud infrastructure - **Responsible AI governance** ensuring ethical and transparent use of artificial intelligence in the features you rely on daily - **Confidence in AI features** knowing that dbt AI capabilities are subject to rigorous oversight and ethical standards - **Simplified compliance** for your organization when using dbt, particularly as AI regulations evolve globally - **Independent verification** of dbt’s security, privacy, and AI governance practices by accredited third-party auditors - **Transparency and accountability** in how we develop, deploy, and monitor AI systems that impact your workflows ISO 42001 certification ensures that dbt Labs is not just innovating rapidly, but doing so responsibly. Our customers can trust that AI features in the dbt platform are built with safeguards against bias, designed with explainability in mind, and continuously monitored for performance and ethical alignment. ## Our ongoing security journey These new certifications represent important milestones in our security journey, but our commitment doesn't end here. dbt Labs continues to invest in robust security measures including: - Regular independent security assessments and penetration testing - Our active vulnerability disclosure program - Ongoing SOC2 Type II, GDPR, CCPA compliance - Continuous improvement of our security processes and technologies - Regular AI system audits and ethical reviews to maintain ISO 42001 standards We're ensuring that you can continue to use the dbt platform with confidence, knowing that your data is protected by security, privacy, and AI governance controls that meet the highest international standards. For more information about our security practices and certifications, please visit our [security page](https://www.getdbt.com/security/). --- --- title: "How Stora Enso enables autonomous data teams with dbt" description: "Stora Enso cut data delivery times from months to days by decentralizing operations with dbt" url: "https://www.getdbt.com/blog/how-stora-enso-enables-autonomous-data-teams-with-dbt" date: "2025-12-18" authors: ["Jiri Bjalek", "Hrishi Kulkarni"] categories: ["Product"] --- # How Stora Enso enables autonomous data teams with dbt Stora Enso, a Finnish-Swedish renewable materials company with 19,000 employees across Europe, Asia, and South America, is a leader in the transition from fossil-based to renewable materials. The company's five divisions each serve distinct markets with unique operational requirements. But Stora Enso's centralized data team struggled to keep pace with growing demands, which created months-long delays that slowed decision-making across the organization. By implementing dbt as the foundation for decentralized data operations, Stora Enso transformed delivery times from months to days and enabled true self-service analytics. ### **Centralized bottlenecks slow the business** "Regardless of how big the central team is, there are not enough people just to retain the business knowledge," reflects Jiri Bjalek, Data Platform Team Leader at Stora Enso. ### **Strategic shift requires the right technology foundation** Recognizing that scaling the centralized team wasn't the answer, Stora Enso made a decisive organizational change three years ago: - **Cut the central data team in half **by moving engineers directly into divisions - **Transformed the central team** from service delivery to platform provision - **Empowered divisions** to prioritize their own analytics projects "We are essentially owning just the platform," says Jiri. "We're responsible for making sure it provides all the features for divisional teams to build their data products." This organizational shift revealed a critical problem: Stora Enso's existing data transformation tool, WhereScape, wasn't designed for autonomous team operations. The platform needed to be accessible enough for diverse teams to use independently, yet robust enough to maintain enterprise standards. While exploring alternatives, Stora Enso discovered dbt. "The big benefit is that dbt uses the SQL language," says Jiri. "SQL has been the de facto standard for many data analysts over decades already, and that was the main success factor; it was so easy to spread it in the organization." Beyond SQL accessibility, the dbt platform brought built-in [data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) and [testing frameworks](https://docs.getdbt.com/docs/build/data-tests) that helped maintain quality standards across autonomous teams. ### **Self-service analytics delivers speed** With dbt as the foundation, the organizational shift delivered immediate and measurable improvements across all divisions: "They have their own teams now and can prioritize their own priority projects," explains Jiri. "It's much, much easier for them to run projects nowadays." ### **Building for governance and AI innovation** The success of the decentralized model has positioned Stora Enso to tackle more ambitious goals. The company continues to evolve its data operations with an eye toward future [compliance](https://www.getdbt.com/security) and innovation requirements: - **EU regulatory compliance** and ISO 27001 certification preparation - **AI integration exploration** built on solid data foundations - **Platform portability** through a successful Databricks proof-of-concept Stora Enso's approach demonstrates how the right platform choice can enable profound organizational change. By shifting from centralized service delivery to decentralized self-service, Stora Enso solved its fundamental challenge of scaling data operations while maintaining quality and consistency. Ready to enable autonomous data teams? [Book a demo](https://www.getdbt.com/contact) or [sign up](https://www.getdbt.com/signup) to connect your data warehouse and start building. --- --- title: "Inside Snowflake’s AI roadmap" description: "Chris Child, Snowflake's VP of Product Management, on the vision for open table formats and the future of the data engineer." url: "https://www.getdbt.com/blog/inside-snowflakes-ai-roadmap" date: "2025-12-15" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Inside Snowflake’s AI roadmap This season of The Analytics Engineering Podcast is focused on how the current data landscape is impacting the developer experience. Snowflake plays a major role in what that developer experience looks like. In this episode, Snowflake VP of Product Management Chris Child joins Tristan to unpack Snowflake’s AI roadmap and what it means for data teams. They discuss the evolution from Snowpark to [Cortex](https://docs.getdbt.com/blog/semantic-layer-cortex) and [Snowflake Intelligence](https://www.getdbt.com/blog/what-is-snowflake-intelligence-anyway), how to [govern agents](https://www.getdbt.com/blog/bring-structured-context-to-agentic-data-development-with-dbt) with row- and column-level controls, and why Snowflake is investing in [Apache Iceberg](https://www.getdbt.com/blog/iceberg-give-it-a-rest) and the [Open Semantic Interchange initiative](https://www.snowflake.com/en/blog/open-semantic-interchange-ai-standard/). dbt Labs recently open sourced [MetricsFlow](https://www.getdbt.com/blog/open-source-metricflow-governed-metrics), the technology that powers the dbt Semantic Layer, to align with the goals of OSI. Chris also shares a vision for the next five years of data engineering: fewer bespoke pipelines, more standardization and semantics, and a bigger focus on business context and data products. _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [Youtube](https://www.youtube.com/playlist?list=PL0QYlrC86xQm83Q9deiy4euEnbw8ceu3I) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) [Watch video](https://youtu.be/5Yo0chBWt2c) ## Key takeaways ### Tristan Handy: Where have you spent your time professionally? **Chris Child:** I didn’t end up in data on purpose. I found myself here through a series of hops. I was working at Redpoint Ventures and got excited by a company we invested in, RelateIQ. I left to join RelateIQ, building an intelligent CRM. We captured emails and meetings and built profiles of everyone you interacted with. We were acquired by Salesforce. Looking at what sales teams needed, I realized they also needed product usage data, marketing data, and campaign data, with a platform to pull it all together. That led me to Segment. I joined when it was about 50 people. Segment was mostly analytics.js then, loading different JavaScript on your webpage for tracking. We had just built the first warehouse connector to Redshift and got huge usage sending click and user data to Redshift. ### The original Redshift connector was a nightmare to work with. Like many startup things, one engineer built it in a week. Suddenly a ton of people used it, and enterprise customers depended on it. We had to rebuild it several times. You could see the future there. Folks I worked with went on to start companies like Census and Hightouch, thinking the CDP should be built on top of the warehouse, which Segment evolved toward. We also built a Snowflake connector because customers demanded it in addition to Redshift. ### It’s funny to think back a decade to how small Snowflake was. A couple customers demanded it; we built it, and we were sending a ton of data. That led to the realization that a customer data platform is one instance of a data warehouse, and there are others you need. Seeing how fast Snowflake was growing, I wanted to build the next layer of infrastructure. I joined Snowflake seven and a half years ago. I’ve had three key roles. First, I built areas of the product: the UI, billing, product-led growth engines and free trial infrastructure, and application capabilities for connecting into and building on Snowflake. After Sridhar became CEO, he asked me to reconnect product and sales by leading solutions engineering, reporting to the CRO. Leading a global technical seller org was very different for a product person, but it helped align teams at scale. About eight months ago, I returned to lead data engineering: how people bring data into Snowflake, how they transform it—spending a lot of time with dbt—and work around Iceberg and interoperability for worlds where not all data sits in Snowflake. ### I didn’t realize the path started in investing. Are you a finance person way back? My undergrad is in computer science. I started programming in fifth grade on an Apple IIe, learned C before high school, and followed that thread. In college I noticed business folks often made the decisions. I wanted to learn that side. After college I joined a consulting firm, then private equity, then an MBA. I realized I didn’t want to be a finance person. I moved to venture as a bridge to building products, but I wanted to build, so I jumped into operating roles. ### Tell the story of Snowflake and AI. In the 2010s there was huge demand for easier, scalable, cloud-oriented data solutions. Then 2022 happened, ChatGPT launched, and the world changed. How did Snowflake respond, and where are you today? Even pre‑2022 we saw customers putting their most important business data into Snowflake, then pulling data out for things they couldn’t do inside: training ML models and other analyses that SQL wasn’t a great fit for. Customers told us they didn’t like losing governance and lineage when data left. We invested in ways to bring more of that work to Snowflake. Snowpark was the first big step: a runtime for non‑SQL code (Python, Java, Scala) with APIs inspired by Spark, plus capabilities like forecasting. It’s great for some workloads, but most customers don’t train most ML models inside Snowflake yet. We also acquired Applica for document extraction using early LLM techniques, and Neeva for web search based on LLM approaches. When ChatGPT arrived, we saw two major influences. First, people wanted to chat with data they’d brought into Snowflake and transformed with dbt. That’s hard because LLMs are great with unstructured data and less great at turning business questions into correct SQL. Second, LLMs are very good at writing code, including Python and even dbt code. They’re not perfect for data engineering code yet, but they help. Our goal is to help customers activate important enterprise data safely in AI models, deploy agents at scale under existing governance, and keep up with exploding data volumes without 10x headcount. ### What are the key product pieces—Cortex, Snowflake Intelligence, etc.—in the Snowflake AI stack? First, you need a great data foundation. That isn’t new: get the data in one place, apply good governance and permissions, know your data, tag PII, and raise the standard of care. AI raises the bar because agents can expose sensitive data faster than dashboards. OSI (Open Semantic Interchange) work is part of this; LLMs need explicit semantics and cataloging they can consume, not tacit knowledge hidden in downstream tools. Companies with strong hygiene move faster with AI. Roles matter; if a product manager role has access to certain rows and columns, an agent acting within that role can safely answer questions. Agents can run inside or outside Snowflake, but should assume appropriate roles when querying. On the AI stack, after the data foundation, Cortex provides higher‑level APIs for unstructured processing, RAG, and structured processing. You can choose models (OpenAI, Anthropic, Mistral, Gemini, Llama, etc.), but most folks don’t want to manage prompts and GPUs. Cortex AI SQL lets you express intent like sentiment filters or fuzzy joins. It’s powerful for exploration but non‑deterministic, so you need care in production. Costs map to tokens at higher abstractions, with budgets and guardrails similar to variable compute in the cloud. At the top, Snowflake Intelligence is a UI and agent framework. You define agents with access to specific datasets and semantic models, plus gold queries and usage guidance. It looks like a chat interface over your governed data. Inside Snowflake, we’ve deployed a GTM assistant that blends product usage, Salesforce, notes, docs, and content—structured and unstructured—respecting row‑level security for every seller while giving leaders broader access. ### Let’s talk open formats and Iceberg. Why lean in when it opens up the data? Our aim isn’t to lock up data, it’s to help customers get value. Snowflake began as a reaction to Hadoop—betting on SQL at cloud scale with our own formats and catalog because they didn’t exist then. Those proprietary pieces let us evolve quickly. Iceberg is now almost as good, and we’re contributing to make it better. Openness is a win for customers and expands the universe of data Snowflake can query, run Cortex on, and power Intelligence with. The tradeoff is standards move slower. Variant type support is a good example—we contributed our approach and shepherded it into the v3 spec. Next up, the community is wrestling with fine‑grained access control beyond table‑level policies. It’s hard and will take time, but the outcome should be better for everyone. ### Give us your view on the future of data engineering. Data volume is exploding, including unstructured data that’s now usable. You can’t hand‑build every pipeline. Demand is also exploding as agents query more things in more ways. Teams must operate at a higher level: automate, standardize, and reduce bespoke pipelines. Expect more shared semantic models across consumers and packaged semantics coming from systems like SAP. You’ll also build data‑engineering agents to do work and monitor pipelines. The role looks more like architect and manager, allocating budgets, deduplicating work, and—most importantly—deeply understanding the business. The best data engineers shift from code output to data products, with clear semantics and context. ### Talk more about context. The day‑to‑day activity shifts, but the output is still data products. Great data products come with instructions, definitions, lineage, quality expectations, and how to get correct answers to common questions. We need that context captured where work happens—models, visualization, quality systems—and made available everywhere: catalogs, agents, and UIs. As you build, you should also document, and those semantics should flow consistently into tools like Snowflake Intelligence so agents can reason correctly. A big part of the challenge is selecting just‑enough context per question. ## Chapters - 00:01:50 — Chris’s path: RelateIQ, Segment, Snowflake - 00:05:40 — Roles at Snowflake: product, solutions engineering, data engineering - 00:09:00 — Snowflake and AI: foundations before ChatGPT - 00:11:40 — Why keep ML and non-SQL work closer to governed data - 00:13:40 — Applica and Neeva acquisitions, enterprise search context - 00:14:50 — Two big AI influences: chat with data and code generation - 00:16:50 — Scaling agents while preserving governance and cost controls - 00:18:40 — Why governance must live at the data layer (roles, rows, columns) - 00:22:00 — Inside vs. outside Snowflake: how agents assume roles - 00:23:02 — Cortex: higher-level APIs over many LLMs - 00:24:06 — AI SQL: joins/where by intent and the non-determinism tradeoff - 00:27:40 — Cost models, tokens, and guardrails - 00:29:10 — Snowflake Intelligence: agents over a governed foundation - 00:32:10 — Open formats and Iceberg: Why Snowflake leaned in - 00:36:00 — Standards tradeoffs: variant type and community progress - 00:38:40 — Fine-grained access control for Iceberg: thorny but necessary - 00:40:40 — The future of data engineering: scale, unstructured data, agents - 00:43:20 — No more bespoke pipelines; standardized models, and semantics - 00:44:50 — Data engineers as architects and business partners - 00:50:00 — Code vs. context: data products and shared semantics - 00:53:10 — Capturing context where work happens (models, viz, quality) - 00:55:00 — Selecting just enough context for agent reasoning - 00:56:30 — Closing --- --- title: "Scale reliable analytics in the AI era with dbt and Databricks" description: "Here’s why standards don’t matter when it comes to unlocking data siloes and fueling the agentic AI future." url: "https://www.getdbt.com/blog/reliable-analytics-dbt-databricks" date: "2025-12-11" authors: ["Ryan Segar"] categories: ["Insights"] --- # Scale reliable analytics in the AI era with dbt and Databricks If you're a data leader planning for the future, you're probably feeling the pressure. Your team's backlog could fill the next year. Leadership wants AI results yesterday. And somewhere in the back of your mind, you're wondering if your current infrastructure can actually support what comes next. Taking a look at the state of the data art, it’s easy to understand why everyone’s sweating bullets. Preparing for the AI future requires: - Accessing your data at scale, no matter how (or where) it’s stored - Migrating off of legacy systems and into more modern data formats - Making overwhelmed data engineers more productive And, somehow, we need to do all of this while also increasing stakeholders’ trust in data. What’s needed to meet these challenges isn’t necessarily more standards. What’s needed is a more open data infrastructure that, perhaps ironically, enables us to think less about infrastructure. I had a good chat recently about this with David Totten, VP of Field Engineering at Databricks. We talked about how dbt and Databricks are working to prepare companies to scale analytics reliably in the AI era by avoiding vendor lock-in, making migration painless, and leveraging AI to augment (not replace) overworked and understaffed data teams. **** ## dbt and Databricks: An expanding partnership Databricks’ core mission is to get data into a [data lakehouse architecture](https://www.databricks.com/glossary/data-lakehouse). The data lakehouse architecture gives customers the ultimate power and control over their data in the quickest, efficient, and most productive manner. Using a data lakehouse, companies can drive analytics more cheaply and efficiently than ever. dbt enables moving data into and out of the lakehouse, providing a seamless way to create and manage high-quality [data pipelines](https://www.getdbt.com/blog/data-pipelines). Once inside a Databricks data lakehouse, all of that data is governed by the Databricks [Unity Catalog](https://www.databricks.com/product/unity-catalog), which provides a single location for governing, discovering, monitoring, and sharing data across the enterprise. This tight integration between dbt and Databricks - with dbt serving as a [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) and Databricks as a single source of truth - provides a unified, vendor-agnostic approach to managing all of your data. You can put your data in any format you desire and manage it in a highly scalable and governed manner. ## How dbt and Databricks help companies move faster Almost every customer discussion I’ve had at dbt Labs since taking over as Chief Product Officer has focused on two things: AI and [Apache Iceberg](https://iceberg.apache.org/). These are the two areas where we find all of our customers have common ground. What's fascinating is that many teams are still treating these as separate problems. They want to build AI capabilities and, separately, are evaluating Iceberg for their data lakehouse. But here's what David and I both see from the field: these aren't separate problems. Iceberg is essential to getting your data out of siloes and scaling your AI strategy. This is symbolic of a bigger shift in our industry. For years, we've dealt with the consequences of closed systems and proprietary formats. We've watched teams spend months—sometimes years—on migration projects, only to find out the company has moved on to the next thing before they're done. The challenge, in other words, is not to get behind the curve while trying to stay ahead of it. Here is what we’ve seen successful customers do. ### Why standards don’t (or shouldn’t) matter anymore The modern data stack started as a kit of point solutions. You grabbed best-of-breed tools for extraction, transformation, loading, and observability, and stitched them together. Then we moved to platforms that tried to own entire boxes in that stack. Now we're in a different era. Storage is decoupled from compute. SQL can (and should) run anywhere. Your data lakehouse should work with multiple engines and storage formats. This isn't just good architecture—it's survival. The question isn't whether you're using the "right" format or the "right" platform. The question is whether you're using open standards that let you move freely as technology evolves. Because technology is evolving faster than ever, and you can't afford to be locked in. David made a great point about Databricks' journey with [Delta](https://delta.io/) and Iceberg. They started with Delta, built incredible technology, and then, when the world said it wanted Iceberg, they hired all the Iceberg people and embraced the open standard. That's the kind of flexibility you need to compete today. Open source and integration of systems are the key. You don’t want to spend 75% to 80% of your annual IT product budget and nine months figuring out basic infrastructure. That’s why Databricks and dbt have worked together to make it easy to operationalize and govern your data, no matter what standard it adheres to and what format it’s in. ### Transpilation: How dbt Fusion will make migration easier Let’s cut to it: “Transpilation” is just an overly complicated word to describe a problem we’ve wrestled with in the data space for decades. You need to migrate data from one system to another. But all your SQL is vendor-specific. This necessitates a painful migration process. Six months in, you find you’re maybe 80% done (if you’re lucky). Meanwhile, your target technology is already out of date - the industry’s moved on to the next best thing. The dream is that you write a SQL query once, and it runs in any environment. That’s the reality we should all be striving for: to ensure that we never have to use the word “migration” again. This is why, at dbt, we’ve heavily invested in [the dbt Fusion engine](https://docs.getdbt.com/docs/fusion/about-fusion), the tech we acquired from SDF Labs. We call Fusion an SQL compiler - and it is. It has a native understanding of SQL across multiple engine dialects, meaning it can compile and validate SQL against your data warehouse even as you type in your Integrated Development Environment (IDE). That accelerates development by eliminating lengthy check-in/test/fix cycles common to [CI/CD pipelines](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud). At scale, however, Fusion should be a transpiler. I.e., you shouldn’t have to worry if you have 10 or even a hundred thousand stored procedures in one data warehouse, or get hung up because of that one function you’re calling that’s specific to your platform. This is a huge blocker to companies meeting their AI goals. As David noted, most organizations have a bunch of legacy data stored offline in physical systems or in various clouds. They want - they need - to make this data available to AI agents. That’s not possible if they can’t leave these legacy systems. This is something we’re closely tracking at dbt. We believe that Iceberg and AI are both integral to unlocking that last 20% of the migration process - to inspecting code before it’s used anywhere and letting you know if it’s truly agnostic and can run anywhere. ### How AI is boosting (not replacing) data teams There’s a lot of talk these days about the human impact of AI. Particularly, people are worried about the impact that AI technology might have on their jobs. There’s no need to worry about this in the data space. There’s no data team on the planet that thinks, “We have plenty of time and resources, and no ticket backlog - we’re fine.” Every data team we’ve ever seen is at 110% capacity (or worse). AI is not coming for the data engineer, [analytics engineer](https://www.getdbt.com/blog/what-is-analytics-engineering), data scientist, or analyst. Rather, AI is poised to boost these roles, augmenting experts so they can burn down those backlogs more quickly. It’s not just the “boring” stuff - skeleton code, documentation, etc. - that AI can assist with either. We’ve got a great customer who has (for good reasons) strict internal policies on what you can or can’t use inside of data transformation code. Before AI, enforcing these policies was impossible. Inspecting code took too long. With AI-based code inspection, it’s finally possible. Around 70% of dbt users are leveraging AI to generate code, documentation, or tests. That needs to evolve so we can get out of this reactive mode of burning down Jira backlogs and back to the strategic work that drives the company forward. **** ### Transparency leads to trust Trust has always been a key issue with data. And trust is, at its core, in infrastructure problem. Trust is, in the end, about being transparent. It’s about helping a user understand, not just what data’s lineage is and where it came from, but where something might have gone wrong in the data pipeline. That’s what I’m excited about this paramount shift in infrastructure - about Iceberg, AI, and the scalability and transparency they can bring. Using AI, we can give users a rough estimate of the accuracy of the given response and help them navigate how to use it. We can have self-healing pipelines that generate their own bug fixes and file their own pull requests. This provides a new level of transparency. We don’t have to certify that something is 100% correct - because when is it ever? Things are moving too fast to guarantee 100%, which is why people lose trust in their data. Instead, we give both technical and non-technical experts the information they need to solve the problems they find in data using a common language. ## How to scale analytics well into the future That leads to the question of where to go next. Most companies and data leaders know the direction in which they need to move. It can be challenging, however, to figure out how to take the first steps. David and I recommend two basic strategies: **Shift from infrastructure to experimentation**. You simply cannot drive AI experimentation into production if you're worried about data access, format conversion, and infrastructure management. You need to shift 95% of your team's time from worrying about infrastructure to building actual use cases. AI is here, right now, and you're missing the wave if you're still figuring out how to access your data. **Ask easy questions**. Where do you start first? Simple - look at your infrastructure and make a call: red light or green light? If something’s a green light, it’s not a problem. If it’s a red light, it’s a blocker. Start tackling the red lights, one by one. The partnership between dbt and Databricks isn't just about technical integration. It's about a shared belief that customers deserve flexibility, not lock-in. Teams should spend their time solving business problems, not building infrastructure. The challenges we discussed aren't going away. Your backlog will still be full tomorrow. The pressure to deliver AI results will keep increasing. And the pace of technological change will only accelerate. But for the first time, we have the tools to meet these challenges. The AI era isn't coming—it's here. The question is whether your data foundation is ready for it. --- --- title: "Bring structured context to agentic data development with dbt" description: "Make data development with agents safe, cost-efficient, and scalable with dbt structured context and MCP server." url: "https://www.getdbt.com/blog/bring-structured-context-to-agentic-data-development-with-dbt" date: "2025-12-10" authors: ["Chakshu Mehta", "Ludwig Sewall", "Sai Maddali"] categories: ["Product"] --- # Bring structured context to agentic data development with dbt AI is writing code at Google and Microsoft. Automating deployments at startups. But when it comes to building production data pipelines, it's still mostly manual. The missing piece is the same as in [Part 1 of our series](https://www.getdbt.com/blog/bring-structured-context-to-conversational-analytics-with-dbt): for AI agents to develop analytics safely, cost-efficiently, and at scale, it needs what every data engineer depends on, **structured context**. This is **Part Two** of our series, _Bringing structured context to AI_ where we explore how **dbt** and the [**dbt Model Context Protocol (MCP) server**](https://docs.getdbt.com/docs/dbt-ai/about-mcp) make the rich metadata in your dbt project accessible, cost-efficient, and trustworthy for AI agents. In this post, we turn to the next frontier: **agentic data development.** Where AI systems don't just query dbt data, but develop, refactor, test, and migrate dbt projects safely and at scale. That means eliminating unnecessary warehouse compute and ensuring every change aligns with your org’s rules and definitions. ## AI agents accelerate app development, but stumble on data pipelines In software engineering, AI agents are already rewriting the rules. Google reports that [over 30% of its new code is now AI-generated](https://www.entrepreneur.com/business-news/ai-is-taking-over-coding-at-microsoft-google-and-meta/490896#:~:text=Key%20Takeaways,are%20%E2%80%9Cnot%20that%20great.%E2%80%9D)_._ And this isn't just autocomplete. These agents can interpret intent, map out multi-step plans, refactor codebases, open pull requests, and more. AI agents thrive in software engineering because application codebases live in highly structured environments. Clear module boundaries, type systems, tests, CI, and version control give agents a predictable world to reason about and safely automate changes in. **Data pipelines live in a very different universe.** In many organizations, analytics environments don’t look like that. Without tools like dbt, the "code" is only a tiny slice of the truth. Every model, metric, and transformation is entangled with business logic, source-system drift, undocumented constraints, and years of accumulated tribal knowledge scattered across chats, documentation, and people's memories. If you're using legacy visual ETL tools or ad-hoc SQL without structured context, agents don’t get the foundation they need to reason effectively. So yes, agents can generate SQL, but they can't understand **_why_** something exists, **how** it’s used, or **what** it’s allowed to do. This means that without structured context, these data agents produce code that looks correct but violates your organization's standards and definitions. _To demonstrate this, we asked ChatGPT to create a lifetime value model with no access to a dbt project or structured context. As expected, the output from this zero-shot prompt looks plausible but is very poor._ ![LTV model from chatgpt prompt](https://cdn.sanity.io/images/wl0ndo6t/main/8722dc2835a5756b05d10b1a9931b54cd6b7183a-2410x1320.png) When agents work over your warehouse like it’s just a pile of tables, you get a consistent set of problems: - **Missing SQL/system understanding:** If an agent only sees a single SQL file at a time, it has no way to understand how that code fits into the broader pipeline. It can’t see column lineage, compilation errors or recent job logs, it cannot understand how changes relate. So it “fixes” a query or rewrites a model in isolation and produces code that breaks downstream. - **Missing business context and logic:** Generic data agents don’t know how your company defines LTV, ARR, churn, or an “active user.” They can’t tell which table is a trusted source of truth and which is a one-off analysis from last quarter. Without access to company documentation, semantic definitions, or domain rules, it might hallucinate and return logic that appears correct but produces incorrect metrics and results leading to fast erosion of trust in anything the AI touches. - **No change-impact awareness**: Refactoring a model or adding a new one requires understanding all upstream and downstream dependencies. Without that, agents change SQL in isolation and quietly break BI dashboards, reports, and other models. They should be able to see the blast radius of a change and make decisions with that impact in mind. - **High context switching for humans in the loop:** Humans often have to chase context across multiple systems to verify AI outputs. If developers must log into the data platform UI, review dbt logs, and check the BI tool to validate outputs, they lose continuity. Iteration slows down, review becomes tedious, and AI feels like more overhead than help. - **No tight local validation loop:** In many AI experiments, the only way to know if a change “worked” is to push it all the way to the warehouse and run a heavy job. If the agent can’t compile models or validate changes locally, even simple changes require full warehouse runs, creating long feedback cycles and increased compute cost. Even with access to the local codebase, the agent will be forced to guess, because it reads your SQL as unstructured text, not as the interconnected, governed system of models, tests, lineage, and semantics your team works with every day. That’s why today’s AI-driven data development feels brittle. ## How dbt's structured context layer enables agentic data development Safe, reliable, and cost-efficient agentic data development requires giving systems a structured understanding of your analytics environment, not just access to code. **dbt’s [structured context layer](https://www.getdbt.com/blog/bring-structured-context-to-conversational-analytics-with-dbt) provides exactly that. **The [dbt MCP Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) exposes rich project metadata (semantic definitions, lineage, contracts, owners, tests, freshness information, and CI artifacts) through the same structured context and engines used by dbt Platform, Fusion, and the dbt VS Code extension. That structured context layer enables a few key capabilities: - **System-level context:** Agents can see your dbt project as a graph, not a pile of files. This means agents start behaving like engineers who understand the system, proposing changes that preserve contracts and avoid surprise breakages. - Through dbt’s [lineage](https://docs.getdbt.com/docs/explore/column-level-lineage), run [artifacts](https://docs.getdbt.com/reference/artifacts/dbt-artifacts), and [metadata](https://docs.getdbt.com/docs/explore/explore-projects#dbt-metadata) (exposed via the [dbt MCP Server](https://docs.getdbt.com/docs/dbt-ai/about-mcp)), they can inspect dependencies, see what’s upstream and downstream, and understand which changes are safe. - **Shared business semantics and logic:** Agents can reuse governed metrics and dimensions defined centrally instead of defining their own versions of metrics. This means AI-generated logic matches your existing definitions, so practitioners and leaders see consistency, not conflict, across reports and tools. - By exposing dbt docs, owners, and [the dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) via the dbt MCP Server, agents can reuse governed metrics, dimensions, and logic instead of reinventing them. - **Impact-aware planning:** Because agents see the DAG and downstream consumers, they can reason about the blast radius of a change _before_ making it. They can answer, “Which dashboards and models will this refactor affect?” and choose safer patterns, add tests, or propose follow-up changes instead of blindly editing SQL. - With dbt lineage, exposures, state from dbt artifacts, and the [**Fusion MCP tool**](https://docs.getdbt.com/docs/dbt-ai/about-mcp#fusion-tools-remote) together give agents a concrete view of upstream/downstream impact for every change. - **Less context switching, more trust:** When humans review and steer agent output in one governed environment instead of bouncing across multiple disconnected tools it results in ****higher productivity and trust. - With [dbt Fusion Engine](https://docs.getdbt.com/docs/fusion/about-fusion), the [dbt language server (LSP)](https://docs.getdbt.com/docs/about-dbt-lsp), and [the dbt VS Code extension](https://docs.getdbt.com/docs/install-dbt-extension), agent proposals, lineage, docs, and compile results all show up in the IDE where engineers already work. - **Fast, scoped validation loops:** Agents can compile locally, run targeted checks, and validate only the impacted slice of the DAG. Resulting in cheap, rapid feedback. Teams can let agents propose frequent, incremental changes while maintaining the same quality bars they expect from human-led development. - By plugging into [dbt’s compilation engine](https://docs.getdbt.com/docs/fusion/supported-features#features-and-capabilities), [tests](https://docs.getdbt.com/reference/commands/build#tests), [contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts), and Slim CI/[state-aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-about), agents can compile locally, run only the relevant checks, and validate just the impacted slice of the DAG. By operating through a structured context layer and governed interfaces (like the dbt MCP Server), SQL stops being just text and becomes part of a governed system. With this foundation, data development agents can read dependencies, reuse existing logic, run compile-time checks, validate outputs, and operate within the same guardrails as your analytics engineers. And the best part? You can build this today. dbt’s ecosystem already provides nearly everything data development agents need to reason like analytics engineers, forming the foundation for reliable, explainable, and cost-efficient agentic data development. > _Structured context is the multiplier. With dbt as our source of definitions and lineage and MCP exposing that context across Snowflake and Claude, we can add new agent skills without re-plumbing governance. We are excited for dbt Agents to bring purpose-built automation that moves us from reactive tickets to proactive, agent-driven operations and spares us the overhead of bespoke bots. > -_**Øyvind Eraker, Senior Data Engineer at NBIM** ## Building a reliable data development agent with dbt To build a trustworthy, safe, and cost-efficient data dev agent workflow, your agent needs two core capabilities: 1. Access to structured context - the complete meaning behind your models, lineage, tests, contracts, and metrics. 2. A grounded execution environment - a local, governed workspace that reflects your actual dbt project. dbt provides both through its structured context layer, exposed to agents via dbt MCP Server and executed locally through dbt Fusion Engine, which powers the dbt VS Code extension. ### How agents access structured context - **Model Context Protocol (MCP):** Think of MCP as an API for LLMs. It passes structured data from external systems (like GitHub, Notion, dbt, or any system with an MCP server) directly to the agent. This ensures agents always reason from current, complete, authoritative information, not partial context pasted into prompts. - **dbt Fusion Language Server Protocol (LSP):** The Fusion LSP powers the dbt VS Code extension, giving agents and developers: gives agents real-time, local insight into your dbt project, model structures, dependencies, refs, columns, and local compilation results. Combined with Fusion Engine, this enables fast, accurate feedback without hitting the backend and makes code changes more reliable by grounding them in the actual project structure. Together, MCP + LSP give agents the context and a **local execution environment** inside VS Code enabling safe agentic development for tools like [Cursor](https://docs.getdbt.com/docs/dbt-ai/integrate-mcp-cursor), [Claude](https://docs.getdbt.com/docs/dbt-ai/integrate-mcp-claude) Code, or any MCP-enabled IDE. ![Conceptual architecture of agentic dbt development](https://cdn.sanity.io/images/wl0ndo6t/main/f61c5176a4250d9c2879565a98b348cf830beec8-2280x1242.png) ## The three capabilities your dbt agent must implement Here’s how an agent uses dbt's structured context to develop safely. #### 1. Ingest the right context (via MCP) Instead of switching between multiple tools, it makes more sense to directly ingest this information from multiple MCP servers. For example, an MCP for your issue tracker, another for internal docs, another for GitHub, and the **dbt MCP Server** for dbt metadata. For dbt specifically, the dbt MCP Server exposes the structured context layer through dedicated tools: - the [**Admin tool**](https://docs.getdbt.com/docs/dbt-ai/about-mcp#available-tools) for reading job logs and un history - the [**Discovery tool**](https://docs.getdbt.com/docs/dbt-ai/about-mcp#available-tools) for exploring lineage and metadata - the [**Query/Semantic Layer tool**](https://docs.getdbt.com/docs/dbt-ai/about-mcp#available-tools) for querying governed metrics and dimensions - the [**CLI tool**](https://docs.getdbt.com/docs/dbt-ai/about-mcp#available-tools) for local compilation, testing, and validation - the **[Fusion tool](https://docs.getdbt.com/docs/dbt-ai/about-mcp#fusion-tools-remote) **for leveraging Fusion’s compiler, diagnostics, and project state This gives ‌the agents access to the “why” and “how” behind every model before it generates a single line of SQL. #### 2. Make safe, local code changes (via Fusion + VS Code) Once the agent has gathered context from the relevant MCP servers, it can begin making changes. This is where **dbt VS Code extension** and the language server **(LSP)** feed the agents with information from the local environment, where changes can be inspected, validated, and governed before touching the warehouse. - [**Fusion**](https://docs.getdbt.com/blog/dbt-fusion-engine) provides local SQL understanding and compilation, letting the agent analyze errors and dependencies without warehouse queries for fast, safe, and cost-efficient iteration. - [**Fusion LSP**](https://www.getdbt.com/blog/fusion-and-dbt-vs-code-extension-preview-launch) exposes your project’s full structure, models, columns, sources, refs, and configs, giving the LLM a real-time, authoritative view of your dbt graph for fewer hallucinations. Inside agent VS Code, agents (e.g. GitHub Copilot, OpenAI, Claude) can propose and apply fixes grounded in the actual project, not guesses, and never blind SQL generation. Fusion + VS Code is where structured context becomes safe, less costly, and actionable. #### 3. Validate and deploy changes safely After successful compilation, the **dbt MCP [CLI tool](https://docs.getdbt.com/docs/dbt-ai/about-mcp#available-tools)** can run tests, sampling, and model executions. Agents then uses this structured context to auto create PRs enriched with accurate model diffs, lineage impact, and downstream test results Once the PR is opened, it automatically kicks off **column-level CI** which validates only the models that directly depend on the changed code. This keeps iteration fast, cost-efficient, and production-safe, while still giving a full downstream impact analysis. This agentic workflow mirrors an analytics engineer’s workflow, but automated with complete awareness of your structured context layer. ## See it in practice Here’s a real example: a **CI Agent** that uses dbt’s structured context layer to diagnose an issue, fix the model, and open a PR. [Watch video](https://www.youtube.com/watch?v=n6T6gN63y_w) And this isn’t a one-off. Teams are already using agentic workflows for tasks like refactoring deprecated models, fixing failing tests, adding new metrics, or even migrating to a new warehouse, all powered by dbt’s structured context layer. In fact, dbt partner [Indicium helped Aura Minerals](https://indicium.ai/knowledge-hub/blog/ai-in-data-migration-aura-minerals/) use AI and dbt’s structured context layer to build a migration agent that moved their PySpark estate to dbt. The results: **400+ notebooks** and **130 workflows** migrated, pipeline time cut by **87%**, **~99%** code conformity, and **66%** less stakeholder coordination. With structured context, Aura adopted a governed dbt environment quickly and safely. ## Why this matters for you as a Data Engineer Anyone who’s maintained a major data model knows the grind: chasing lineage, debugging tests, rerunning builds. Agents powered by dbt’s structured context layer handle this repetitive work without risking your production environment. This workflow gives you: - **Faster debugging** — all relevant context is pulled automatically. - **Higher accuracy** — fixes are based on real dbt metadata, not guesses. - **Safer and cost-efficient iteration** — local compilation avoids unsafe or costly backend queries. - **Automates low-value work** — PR creation and validation can be automated. Engineers spend less time gathering context and more time architecting systems, exactly where human expertise matters most. ## **Thinking about making data development with agents possible? Start with dbt.** Before you introduce agents into your data development workflow, make sure your foundation is ready: - Do you have tests, contracts, or CI that catch issues early? - Is your business logic centralized, not scattered across dashboards, SQL files, and Slack? - Do your models have lineage, ownership, and documentation? - Can your project compile cleanly in a local environment? If so, your project is already agent-ready. You can connect your own agents through the [**dbt MCP Server**](https://www.getdbt.com/blog/dbt-agents-remote-dbt-mcp-server-trusted-ai-for-analytics) and start shipping agent powered data pipelines today! **Or, if you prefer not to build it yourself…** We’re introducing [**dbt Agents**](https://www.getdbt.com/product/dbt-agents) (coming soon), our purpose-built agents in dbt Platform that plan, generate, and validate dbt code using your existing definitions, tests, lineage, and permissions. - The **Developer Agent** (coming soon) will support refactoring, migrations, validation - The **Observability Agent** (coming soon) will monitor builds and surfaces likely root causes. All powered by structured context and governed development. It's time to bring governed, cost-efficient agentic workflows to data development! **Get a [demo](https://www.getdbt.com/contact)** of the dbt MCP Server or dbt Agents to ship reliable agentic data development today. [Watch video](https://www.youtube.com/watch?v=VMkRXWkEcKk) --- --- title: "Bring structured context to conversational analytics with dbt" description: "Make conversational analytics trustworthy with dbt as your foundation." url: "https://www.getdbt.com/blog/bring-structured-context-to-conversational-analytics-with-dbt" date: "2025-12-03" authors: ["Sai Maddali", "Chakshu Mehta"] categories: ["Product"] --- # Bring structured context to conversational analytics with dbt AI is everywhere, but reliable, production-safe conversational AI is still surprisingly rare. That’s because most AI workflows today are missing a critical ingredient: **structured context**. This is the first post in our **four-part series**, _Bringing structured context to AI. _We'll learn how dbt, along with the dbt Model Context Protocol (MCP) server, lets you see the important data and metadata inside your dbt project. This makes it easy to use, cost-effective, and safe for AI systems. Let’s start with the most immediate and visible use case: **conversational analytics.** ## The promise (and limits) of conversational analytics The dream has always been simple: Ask a question in plain language. Get answers you can trust. Text-to-SQL made that feel within reach. It taught models to translate natural language into a query. But anyone who has tried to rely on text-to-SQL in production knows its limits. Without context, AI guesses. It picks the wrong model, joins on the wrong key, applies outdated metric logic, overlooks governance rules, or runs expensive queries that blow up warehouse costs. That’s because LLMs predict likely responses probabilistically based on patterns in training data. They don’t understand your business context. When they operate without structure, they can’t reliably generate SQL, interpret business logic, or respect governance rules. And real analysis isn't just one SQL query. It’s multi-step. It involves planning, comparing, validating, and explaining. This is where the next generation of conversational analytics becomes agentic. ## Structured context layer: turning AI into an analyst Agentic analytics systems don’t just generate one query. They plan, iterate, check assumptions, apply lineage, and explain reasoning. They behave the way an analyst behaves. Conversational analytics agents can only do this if they have access to a **structured context layer.** This is a foundation that makes AI outputs predictable, governed, explainable, and cost-efficient not just in compute, but in token usage as well. A recent [Gartner](https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk) survey found that through 2026, 60% of AI projects will be abandoned if they aren’t supported by “AI-ready” data. The path to adoptable, trustworthy AI starts with a data foundation that provides structure and context. ## What is the structured context layer? **The structured context** **layer** is the maintained layer of logic and metadata that gives data models meaning. It includes: - metric and dimension logic, - lineage, - tests, - ownership, - policies, - and business rules. The structured context layer sits on top of your warehouse tables and models and turns raw data into shared understanding. In the dbt ecosystem, this layer already exists. dbt has become the standard for analytics engineering precisely because teams define meaning, quality, and relationships **once,** then reuse it everywhere. This layer is the connective tissue between your data, your business logic, and the AI systems acting on it. So as your AI stack evolves with new tools, agents, and interfaces, your structured context layer remains the same. But defining context isn’t enough. AI must be able to consume it. A structured context layer includes definitions and governance, but also the ability to expose context consistently across tools, workflows, and analytical interfaces. And it does this in a cost-efficient and permission-aware way to reduce warehouse compute and token overhead. ## What an effective structured context layer looks like (and how dbt provides it) An effective structured context layer must deliver five essential capabilities that make AI reliable, explainable, and production-ready. dbt provides these capabilities out of the box: - **Defines and optimizes business logic**: Metrics and transformations should be centrally defined and compiled into efficient SQL, to make AI outputs faster, cheaper to compute, and more reliable. - **With dbt:** The [dbt Semantic Layer](https://www.getdbt.com/blog/how-the-dbt-semantic-layer-works), powered by open-source [MetricFlow](https://www.getdbt.com/blog/open-source-metricflow-governed-metrics), provides governed definitions and compiles them into efficient SQL to [guarantee accuracy, consistency, and performance](https://www.getdbt.com/blog/why-your-ai-will-fail-without-a-semantic-layer). And [dbt Fusion](https://docs.getdbt.com/docs/fusion/about-fusion) ensures that SQL executes consistently and efficiently on your warehouse. AI retrieves both the correct meaning and the correct computation of concepts like “revenue” without the model wasting tokens trying to infer missing logic. - **Guardrails to shrink the search space:** AI needs guardrails so it only chooses from correct models, joins, and keys. - **With dbt:** dbt narrows ambiguity by giving AI a known structured schema, governed metrics, naming conventions, and artifacts like manifest.json. The model no longer has to infer structure, it chooses correctly to reduce unnecessary reasoning tokens. - **Ensures correctness through version-control, validation, and dynamism:** Your context must evolve safely as your business changes. - **With dbt:** All context lives in version-controlled code and is continuously validated through contracts, tests, CI, and `dbt build`. This makes your context layer **dynamic but stable.** It updates with the business while remaining reliable for AI. - **Provides lineage and freshness for operational grounding:** AI must understand where data comes from, what it depends on, and whether it’s trustworthy. - **With dbt:** The lineage graph, metadata freshness, and documentation give AI agents full insight into dependencies and data health. This is foundational for multi-step reasoning. - **Makes context machine-readable and interoperable:** A structured context layer must be accessible to any LLM, agent, or workflow in standardized, permission-aware formats. - **With dbt:** The [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) and artifacts expose metrics, lineage, documentation, and model metadata through [dbt’s MCP tooling](https://docs.getdbt.com/docs/dbt-ai/about-mcp#supported) (Semantic Layer tool, Discovery tool, CLI tool, Fusion tool, and Admin tool) that honors your existing access controls and makes your context layer fully interoperable across AI tools, BI interfaces, and orchestration systems. Together, these capabilities form **dbt’s structured context layer,** a single, explainable source of truth for any AI system. [Watch video](https://youtu.be/kwgWevMXaTg) With this foundation, AI systems stop guessing and start reasoning. Your AI system knows what “revenue” means. It knows how it’s calculated. It knows which upstream models feed into it. And it can use that knowledge to navigate multi-step analytical workflows, reliably and cost-efficiently. You can’t predict the full range of questions users will ask. What you _can_ control is the quality of the structure they rely on. The better your dbt project is modeled, governed, and documented, the more reliably your AI agents can perform, no matter the query. ## How customers use dbt to ship chat with their data experiences Teams across analytics, finance, marketing, and product are already using conversational interfaces to safely access and chat with governed data. They can: - **Run quick, safe analysis:** “Average order value by region?” - **Ask clarifying definitions:** “What does ‘qualified lead’ mean?” - **Trace lineage:** “What feeds into `monthly_churn_rate`?” - **Investigate anomalies:** “Did any upstream sources change before this spike?” - **Find models:** “Is there a model for trial-to-paid conversion?” Teams like **Norlys** and **LEAP Consulting** are already running this in production, grounding every chatbot answer in the same definitions and models the business already trusts. > “With dbt MCP server and dbt Semantic Layer we see a great opportunity to leapfrog our ambitions to introduce a "metrics first" model across our organisation. Our ambition is to have a few standard reports with core metrics combined with an innovative solution that enables anyone at Norlys to retrieve the insights they need simply by asking a question through a chatbot interface powered by an LLM. dbt MCP server and dbt Semantic Layer enable this.” — Søren Persson, Director of Data Engineering, Norlys > > “dbt provides the governed context with metrics, lineage, and tests, and the dbt MCP Server makes that context usable by AI systems. Together, they let us design and set up trustworthy conversational analytics quickly for customers, plugging answers into their AI chat or agent workflows with an audit trail by default. It is one of the fastest paths to reliable, production-grade AI we have seen.” — Jonas Munk, Partner, LEAP Consulting That’s not a prototype. That’s conversational AI, in production. Live, reliable, and powered by dbt’s structured context layer. ## Thinking about conversational analytics? Start with dbt as your foundation If you're exploring conversational access to data, the key question isn’t _which model_ to use, it’s _whether your foundation is ready._ Ask yourself: - Is your data structured and governed? - Are metrics defined, tested, and version-controlled? - Can AI access lineage and documentation to explain results? - Is access properly governed and cost-efficient? With dbt, you can connect any LLM (Claude, OpenAI, or your own internal assistant) and build your own multi-step conversational analytics agents via the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp). Or try the [dbt Analyst Agent](https://www.getdbt.com/product/dbt-agents) (now in beta), a purpose-built AI agent that reasons directly on your governed data. It’s time to unlock the structured data and metadata your team already trusts. Get a [demo](https://www.getdbt.com/contact) of the dbt MCP server today or join the [dbt Agents waitlist](https://www.getdbt.com/product/dbt-agents) today. [Watch video](https://www.youtube.com/watch?v=VMkRXWkEcKk) --- --- title: "Why AI raises the bar for data cleaning" description: "AI amplifies dirty data. Learn why modern systems require scalable, automated cleaning logic for trustworthy model outcomes." url: "https://www.getdbt.com/blog/ai-clean-data-requirements" date: "2025-12-01" authors: ["Joey Gault"] categories: ["Pulse"] --- # Why AI raises the bar for data cleaning Traditional analytics workflows could often tolerate certain levels of data inconsistency. An analyst reviewing quarterly sales figures might notice and mentally adjust for obvious outliers or formatting issues. Human analysts bring contextual understanding that can compensate for data quality problems, at least to some degree. AI systems operate differently. Machine learning algorithms process data at scale without the contextual awareness that human analysts provide. A single customer record with an incorrectly entered purchase amount of $100,000 instead of $100.00 doesn't just skew a quarterly report; it can fundamentally alter model training, leading to systematically incorrect predictions across thousands of future transactions. This amplification effect means that data quality issues that were merely inconvenient in traditional analytics become critical failures in AI applications. When calculating customer lifetime value using machine learning models, dirty data doesn't just produce one incorrect result: it corrupts the entire model's understanding of customer behavior patterns. The algorithm learns from these errors, embedding them into its decision-making process and propagating them across every prediction it makes. The modern [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) approach has made addressing these quality requirements more strategic than ever. Unlike traditional ETL processes where cleaning happened before loading, ELT allows teams to clean data within the warehouse environment, leveraging the full computational power of modern cloud platforms. This shift enables more sophisticated cleaning operations and makes it easier to iterate on cleaning logic as AI requirements evolve. ## Core cleaning techniques for AI readiness Effective [data cleaning](https://www.getdbt.com/blog/data-cleaning-transformation-quality) for AI applications requires a systematic approach that goes beyond traditional data quality measures. Missing value handling represents one of the most critical cleaning operations, but the stakes are higher when feeding machine learning models. Rather than simply dropping incomplete records, sophisticated cleaning processes must evaluate the context of missing data with AI consumption in mind. For customer records used in recommendation engines, missing demographic information might be acceptable if the model can rely on behavioral data. However, missing transaction timestamps would break time-series forecasting models entirely. The cleaning logic must understand these AI-specific contexts to make appropriate decisions about data completeness and imputation strategies. Duplicate detection and removal requires even more nuance when preparing data for machine learning. Customer records might have slight variations in name spelling or address formatting while representing the same entity. Advanced cleaning processes implement fuzzy matching algorithms to identify these near-duplicates and establish canonical representations. This cleaning step is crucial for training accurate customer segmentation models and prevents the algorithm from treating the same customer as multiple distinct entities. Data type standardization ensures that AI models can process inputs reliably. When date fields arrive in multiple formats (some as strings, others as timestamps with different timezone information) cleaning processes must standardize these representations. This standardization prevents model training failures and ensures that time-based features produce consistent results across different data sources. AI models are particularly sensitive to these inconsistencies because they rely on mathematical operations that require uniform data types. Format consistency extends beyond data types to business-specific standards that AI models depend on. Phone numbers might arrive as "(555) 123-4567", "555-123-4567", or "5551234567". While human analysts can recognize these as equivalent, machine learning algorithms treat them as entirely different values. Cleaning processes must establish canonical formats that AI models can depend on, ensuring that feature engineering and pattern recognition work consistently across all data sources. ## The compounding benefits of AI workflows Clean data creates a virtuous cycle throughout AI development and deployment processes. When initial cleaning operations remove inconsistencies and errors, subsequent model training steps can focus on learning genuine business patterns rather than compensating for data artifacts. This separation of concerns makes AI models more accurate and reduces the likelihood of biased or incorrect predictions. Machine learning models particularly benefit from thorough data cleaning because they learn from every data point in their training sets. When training a customer churn prediction model, dirty data can create false patterns that the algorithm incorporates into its decision-making logic. A customer record with an incorrectly entered account creation date could make the model associate longer tenure with higher churn risk, leading to systematically incorrect predictions for all established customers. Data integration becomes significantly more reliable when source data has been properly cleaned for AI consumption. Training recommendation engines requires joining customer data from CRM systems with transaction data from e-commerce platforms and behavioral data from web analytics. If the cleaning process hasn't standardized customer identifiers across these systems (removing extra whitespace, converting to lowercase, handling common typos) the integration will miss legitimate matches and create incomplete customer profiles that reduce model accuracy. The reliability gains from proper cleaning compound over time as AI systems become more sophisticated. As organizations deploy more complex models that depend on multiple data sources, the cost of data quality issues increases exponentially. A cleaning error that affects foundational customer data will impact every downstream model that uses customer features: from recommendation engines to fraud detection systems to lifetime value predictions. ## Implementing scalable cleaning with modern tools Managing data cleaning for AI applications at enterprise scale requires more than ad hoc scripts and manual processes. [Modern data transformation platforms like dbt ](https://www.getdbt.com/product/what-is-dbt)provide structured approaches to implementing and maintaining cleaning logic that can support both traditional analytics and AI workloads. By treating cleaning operations as code, teams can version control their cleaning rules, test them systematically, and deploy changes through proper CI/CD processes. dbt's approach to data cleaning emphasizes repeatability and transparency, which are crucial for AI applications that require auditable data lineage. Cleaning logic written as dbt models can be reviewed, tested, and documented alongside other transformation code. This integration ensures that cleaning operations receive the same engineering rigor as business logic transformations. When cleaning rules need to change (perhaps to handle new data sources or evolving AI model requirements) the changes can be implemented, tested, and deployed through established workflows. The [testing capabilities built into dbt ](https://www.getdbt.com/product/test-and-observe)are particularly valuable for AI-focused data cleaning operations. Teams can write tests that verify cleaning logic produces expected results for machine learning consumption: ensuring that feature scaling handles all expected input ranges, confirming that categorical encoding doesn't inadvertently create new categories, or validating that time-series data maintains proper chronological ordering. These tests run automatically as part of the transformation pipeline, catching cleaning failures before they affect model training or inference. Version control becomes crucial when cleaning logic evolves to support new AI use cases. As machine learning requirements change or new models are developed, cleaning rules must adapt accordingly. With dbt, these changes are tracked in Git, making it possible to understand why cleaning logic changed and to roll back problematic updates. This historical context is invaluable when debugging model performance issues or explaining AI system behavior to stakeholders and regulators. ## Monitoring and continuous improvement for AI applications Data cleaning for AI applications isn't a set-it-and-forget-it operation. As machine learning models evolve and new AI use cases emerge, cleaning logic must adapt accordingly. Effective monitoring helps teams understand when cleaning processes need attention and provides insights for continuous improvement that supports both current and future AI applications. Automated monitoring should track both the volume and types of cleaning operations performed, with particular attention to metrics that affect AI model performance. If the percentage of records requiring outlier detection and correction suddenly increases, it might indicate changes in upstream data collection processes that could affect model accuracy. Similarly, if feature distribution patterns shift after cleaning operations, it could signal the need to retrain existing models or adjust cleaning parameters. Performance monitoring ensures that cleaning operations scale with the data volumes required for modern AI applications. As datasets grow to support more sophisticated machine learning models, cleaning logic that worked well on smaller datasets might become bottlenecks. Monitoring query performance and resource utilization helps teams optimize cleaning operations before they impact overall pipeline performance or delay model training cycles. Quality metrics provide feedback on cleaning effectiveness specifically for AI consumption. Teams should track measures like feature stability across cleaning operations, the consistency of statistical distributions before and after cleaning, and the rate of model training failures that can be attributed to data quality issues. These metrics help quantify the business value of cleaning investments and guide prioritization of improvement efforts as AI initiatives expand. ## Strategic considerations for the AI era For data engineering leaders, the intersection of AI adoption and data cleaning represents both a technical challenge and a strategic opportunity. Organizations that invest in robust cleaning processes specifically designed for AI consumption gain competitive advantages through more accurate models, faster time-to-deployment for new AI use cases, and reduced risk of algorithmic bias or failure. The build-versus-buy decision for cleaning capabilities becomes more complex when considering AI requirements. While custom cleaning solutions offer maximum flexibility for specific machine learning use cases, they require ongoing maintenance and specialized expertise in both data engineering and AI model requirements. Modern transformation platforms like dbt provide built-in cleaning capabilities that can handle most common scenarios while allowing customization for specific AI needs. This approach often provides the best balance of capability and maintainability as organizations scale their AI initiatives. Team structure and skills development play crucial roles in cleaning success for AI applications. Data cleaning for machine learning requires understanding both technical implementation details and the specific requirements of different AI model types. Teams need members who can write efficient cleaning logic while also understanding how data quality issues propagate through model training and inference. Cross-training between data engineers, ML engineers, and data scientists often produces the best outcomes for AI-focused cleaning initiatives. Governance and compliance considerations become more complex as cleaning operations scale to support AI applications. Cleaning processes that modify or remove data must comply with regulatory requirements while also meeting the specific needs of machine learning models. Documentation becomes crucial, not just for technical maintenance, but for demonstrating compliance with AI governance standards and explaining model behavior to stakeholders. Modern transformation platforms provide audit trails and documentation capabilities that support these governance requirements while enabling the transparency that responsible AI deployment demands. ## Conclusion The rise of AI has fundamentally changed the stakes for data quality, transforming data cleaning from a best practice into a business imperative. Machine learning models amplify data quality issues in ways that traditional analytics never could, making robust cleaning processes essential for successful AI initiatives. When implemented systematically using modern tools like dbt, cleaning processes become scalable, maintainable, and auditable components of the AI-ready data infrastructure. The key to successful data cleaning in the AI era lies in treating it as an integral part of the machine learning lifecycle rather than a separate preprocessing step. By embedding cleaning logic within transformation workflows and designing it with AI consumption in mind, teams can ensure that quality improvements support both current analytics needs and future AI applications. As AI adoption accelerates and model requirements become more sophisticated, organizations with strong data cleaning foundations will be better positioned to extract value from their AI investments while maintaining the trust and reliability that stakeholders and regulators demand. ## Data cleaning FAQs **What is the difference between data cleaning and data validation?** **How can missing values be handled during data cleaning, and when should you prefer imputation over deletion?** Missing value handling requires evaluating the context of missing data, especially for AI applications. For customer records used in recommendation engines, missing demographic information might be acceptable if behavioral data is available, but missing transaction timestamps would break time-series forecasting models entirely. Imputation should be preferred over deletion when the missing data doesn't fundamentally compromise the model's ability to learn patterns, when you have sufficient related data to make reasonable estimates, and when maintaining dataset size is crucial for model performance. Deletion is more appropriate when the missing values represent critical features that cannot be reliably estimated or when imputation might introduce bias into the model. **Why stress over data cleaning?** Data cleaning has become a business imperative in the AI era because machine learning algorithms amplify data quality issues in ways traditional analytics never could. A single incorrectly entered customer purchase amount can fundamentally alter model training, leading to systematically incorrect predictions across thousands of future transactions. Unlike human analysts who can mentally adjust for obvious errors, AI systems process data at scale without contextual awareness. Dirty data doesn't just produce one incorrect result; it corrupts the entire model's understanding of patterns, with the algorithm learning from these errors and embedding them into every prediction it makes. This amplification effect means that data quality issues that were merely inconvenient in traditional analytics become critical failures in AI applications. --- --- title: "Using state-aware orchestration to slash your data costs" description: "Why run a job when there’s no new data? Here’s how state-aware orchestration in the dbt Fusion engine saves you money." url: "https://www.getdbt.com/blog/using-state-aware-orchestration-to-slash-your-data-costs" date: "2025-11-26" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Using state-aware orchestration to slash your data costs **** There are multiple factors that make it challenging to manage data pipelines at scale. One of the biggest is cost. Many teams don’t consider cost when they first begin building out their data pipelines. As the number of pipelines they manage and the volume of data increase, however, the cloud compute bill becomes too large to ignore. That sets them scrambling to find ways to enable the rapid deployment and development of data pipelines at the optimum price. [The dbt Fusion engine](https://www.getdbt.com/product/fusion), the next-generationg dbt engine, makes it easier than ever to save money on data pipelines. It does this using **[state-aware orchestration](https://www.getdbt.com/blog/announcing-state-aware-orchestration),** which uses the current state of your pipeline to make intelligent decisions around which models it ought to rebuild. Let’s look at what drives the cost of data pipelines up and how state-aware orchestration brings those costs down. We’ll also cover the other ways that Fusion drives down data prices. ## How data pipelines waste money Most data pipelines are complex. A data pipeline created using dbt, for example, can contain multiple [models](https://docs.getdbt.com/docs/build/models). These models can themselves [materialize into multiple tables or views](https://docs.getdbt.com/docs/build/materializations), each connected to one another via dependency declarations. dbt represents this in a rich [directed acyclic graph, or DAG](https://www.getdbt.com/blog/dag-use-cases-and-best-practices). This means that when you create and run [a data pipeline job](https://docs.getdbt.com/docs/deploy/jobs) in dbt, it’s aware of the dependencies between your models and knows in which order to run them. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0842aea8940295892b5de66c707bc7ce1658e92e-790x432.png) dbt builds models whenever you use the [dbt run command](https://docs.getdbt.com/reference/commands/run). By default, run rebuilds all models in a project. You can also tell dbt to run a model and all its dependencies (dbt run -m my_model+). Even with the latter option, however, dbt is running every model in the dependency chain—whether its model has changed or not, and whether it has new data or not. This means running all of those models for every dev, staging, or production run. This is an extra cost that you don’t need. It may be trivial for a single pipeline and a single developer. But this cost multiplies quickly across a company as the complexity of models, the number of models, and the number of people modifying models increase. If your teams are shipping model changes rapidly as part of [a well-tuned analytics workflow](https://www.getdbt.com/resources/the-analytics-development-lifecycle), this cost can quickly spiral. This is one of the many problems that the dbt Fusion engine solves. ## Reducing data costs with the dbt Fusion engine [Fusion](https://docs.getdbt.com/docs/fusion/about-fusion) is a new version of the dbt engine that turns dbt into a full SQL compiler, one that understands the syntax and semantics of all major data warehouses. That means it can be smarter about what needs to be rebuilt in a dbt model’s dependency chain. Let’s look at what this means in practice. dbt Core always runs with Just In Time (JIT) rendering. This means it renders a model, runs it in the data warehouse, and then moves on to the next model in the chain. By contrast, Fusion defaults to Ahead-of-Time (AOT) compiling. It renders all models in the project, producing and statically analyzing every model’s logical plan **before** it runs any models in the data warehouse. ## How state-aware orchestration in the dbt Fusion engine saves money AOT compilation means Fusion understands your data model’s dependencies. It also knows if your data or your model code has changed between runs. This enables Fusion to use [state-aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-about). With state-aware orchestration, Fusion will only run those models and any dependent models with pending changes. Fusion isn’t just aware of single-job state, either. If you’re running multiple Continuous Integration (CI) jobs that use the same models, Fusion is smart enough to only run them once across all running pipelines. ## An example of state-aware orchestration in action To understand this better, let’s walk through an example. Take the following DAG. The model takes the raw data tables (prefixed with raw.), creates staging models to transform the data, and then creates dimension and fact tables for a data warehouse from those. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/b91a790bfc94abc55647746f6da230a44165d99c-706x368.png) Let’s say that when this model runs in production, the raw.wizards model is the only source model with fresh data. A traditional data pipeline might run every single model here by default. By default, the dbt Fusion engine only rebuilds raw.wizards and its dependencies - stg_wizards, dim_wizards, and fct_orders. It will leave raw.worlds, raw.order, raw.wands, and all their downstream tables as is, since it knows these don’t require a refresh. ## How state-aware orchestration works You might be wondering how all of this operates under the hood. The answer is that Fusion is built upon a few core principles: - **Real-time shared state**: All jobs leverage a shared model-level state, which enables tracking model runs across jobs - **Model-level queuing**: If multiple jobs leverage the same model, these queue up at the model level to prevent collisions and unnecessary rebuilds - **State-aware and state-agnostic support**: A job can run in either a state-aware or state-agnostic mode; in both cases, dbt updates the shared state to accurately reflect the model’s run status ## The benefits of state-aware orchestration Using state-aware orchestration immediately brings multiple benefits to your data pipelines: **Reduced data costs**. Rebuilding fewer models results in reduced cost per run, as it reduces the amount of compute needed to run each job. **Faster data pipeline runtimes**. Fewer model runs also reduce the time it takes for a given pipeline to run. That means data engineers will spend less time waiting for job runs to complete successfully, boosting overall developer productivity. **Out of the box operation**. Fusion uses AOT by default. There’s no need to tinker with configurations to get better performance immediately. **High configurability for more demanding use cases**. That said, you can also tailor state-aware orchestration for your needs as required. For example, you can further control when models are run in jobs [by specifying source freshness intervals](https://docs.getdbt.com/reference/resource-configs/freshness) on a per-model basis. **Usable from data engineers’ favorite IDEs**. Fusion is built into the [dbt Studio IDE](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/develop-in-the-cloud). It also tightly integrates with [Visual Studio Code](https://code.visualstudio.com/) via [our official extension](https://www.getdbt.com/blog/get-to-know-the-new-dbt-fusion-engine-and-vs-code-extension). ## Getting started with state-aware orchestration The free** [dbt VS Code extension](https://docs.getdbt.com/docs/install-dbt-extension)** is the best way to develop locally with the dbt Fusion engine. Follow [our step-by-step guide to using Fusion](https://docs.getdbt.com/guides/fusion?step=1). If you’re new to dbt, you’ll want to start with [the Quickstart for your data warehouse](https://docs.getdbt.com/docs/get-started-dbt). ## How the dbt Fusion engine accelerates data development at less cost Fusion can save you both time and money in data development: - Fusion is a rewrite of the core dbt engine in Rust. This means it delivers superior performance over the previous Python-based engine—up to 30x faster per run - AOT compilation reduces data warehouse round-trips during development; the dbt Fusion engine can check SQL syntax locally, resulting in less churn in your deployment pipelines Get started on Fusion quickly with the [**dbt VS Code extension**](https://docs.getdbt.com/docs/install-dbt-extension) and [**talk with a dbt expert today**](https://www.getdbt.com/contact). --- --- title: "Reducing ETL licensing costs with the dbt Fusion engine" description: "dbt can already save you a lot on traditional ETL processing costs. Here’s how the dbt Fusion engine saves even more." url: "https://www.getdbt.com/blog/reducing-etl-licensing-costs" date: "2025-11-26" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Reducing ETL licensing costs with the dbt Fusion engine As the data landscape changes, Extract, Transform, and Load (ETL) tools are struggling to scale. A key reason? Cost. Managing ETL data pipelines often requires licensing expensive, turnkey software tools. [The licensing costs for these tools alone can be steep](https://www.intsurfing.com/blog/how-much-do-etl-systems-cost-factors-cost-breakdown/), ranging between $20,000 and $100,000 for mid- to large-scale enterprises. As data systems scale, this pricing can go from costly to prohibitive. For fixed-seat licensing systems, costs increase as companies bring on more employees to handle the growing volumes of data required for analytics and AI solutions. For usage-based systems, costs can spike dramatically as companies grow from terabyte to petabyte workloads. The fact is that traditional ETL systems are struggling to keep pace in the age of AI. Let’s dive into how traditional ETL licensing works, why costs are rising, and how dbt and the new [dbt Fusion engine](https://www.getdbt.com/product/fusion) can help you nip rising ETL costs in the bud. ## Issues with ETL licensing costs ETL is the legacy approach to data transformation, where data is transformed before it’s loaded into a data warehouse. [It’s being supplanted by the Extract, Load, and Transform (ELT) pattern](https://www.getdbt.com/blog/etl-vs-elt), where raw data is loaded into a data warehouse first and transformed multiple times, each time tailored to a specific use case. These days, most traditional ETL data pipeline tools support either ETL or ELT approaches. In what follows, we’ll use “ETL” as a shorthand for both. ETL tools use one of two pricing models: - **Fixed price**: Vendors charge companies for user seats and/or for a fixed amount of storage and compute - **Pay-as-you-go**: The cloud or “utility” model, in which companies only pay for the amount of computation and storage they actually use Both of these models have their downsides for managing ETL costs at scale. **Fixed price** is good for controlling costs and budget planning. However, “fixed price” can also mean fixed performance and fixed scale. Companies that hit their fixed-price limits may see their data processing grind to a halt, leading to service outages or processing bottlenecks. **Pay-as-you-go** has become more popular in the age of the cloud, where people have grown accustomed to treating compute as a service akin to electricity. [This is much more flexible](https://hevodata.com/learn/etl-cost/), as you won’t run into usage limits and are assured of continuous performance. The problem with pay-as-you-go pricing is that, as usage grows, costs can spiral out of control. Multiple dimensions of scaling—e.g., user base growth, service complexity, data volumes, etc.—can lead to fees growing at multiples of prior usage. Most ETL tools don’t provide a built-in way to split the difference between these two models and manage costs effectively. And that’s a problem, because figuring out how to control costs is crucial to maintaining scalable and stable data systems. ## How dbt helps with ETL costs [dbt](https://www.getdbt.com/product/what-is-dbt) is a [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) that provides a complete framework for developing, testing, deploying, and monitoring ETL/ELT workloads. With dbt, data teams can treat develop reusable code they can use to manage their data transformation logic in a vendor-agnostic manner. dbt provides numerous features that make developing and shipping ETL pipelines more efficient out of the box. It supports developing data transformation code for analytics and AI workloads in accordance with the [analytics development lifecyle (ADLC)](https://www.getdbt.com/blog/adlc-plan), which treats analytics code as software. This means that all code can be thoroughly reviewed and tested beforehand to ensure it’s written in the most performant, cost-conscious manner possible. ## The pain of JIT and round-tripping One of dbt’s main advantages is that it provides a single and consistent approach for modeling data transformations across the industry’s leading data warehouses. It provides data producers and consumers with a common platform and syntax for working with data. When it comes time to render these models, dbt Core, has always used a Just In Time (JIT) model. It connects to your data warehouse to run and verify your SQL or Python code. This is true, not only for production code, but for all code under active development. This means that developing and testing code are also contributing to your increasing data warehousing costs. Every time a data engineer creates a new data transformation model, fixes a bug, tests a fix, etc., they’re consuming more data warehouse compute. If their SQL is incorrect or malformed in any way (using an invalid column name, spelling a keyword wrong, etc.), they’ll only discover this after running the code remotely. Magnify this by dozens or hundreds of data and analytics engineers across a company, and you’re talking a significant impact on your bottom line. ## How the dbt Fusion engine cut licensing costs The new [dbt Fusion engine](https://docs.getdbt.com/docs/fusion/about-fusion) solves exactly this problem. Fusion is a complete rewrite of dbt in Rust that natively understands SQL across multiple data warehousing engine dialects. This gives dbt a powerful new set of local development capabilities. With Fusion, data engineers can: - Catch SQL errors immediately as they type in [Visual Studio Code, Cursor](https://docs.getdbt.com/docs/install-dbt-extension), or [dbt Studio](https://docs.getdbt.com/docs/dbt-versions/upgrade-dbt-version-in-cloud#dbt-fusion-engine) - Preview inline Common Table Expressions (CTEs) - Trace model and column definitions across your dbt project Besides accelerating data pipeline development times, Fusion helps companies save on costs in three ways: - Ahead-of-Time (AOT) compilation - State-aware orchestration - Super-fast performance ## Ahead-of-Time (AOT) compilation The key is that **Fusion can do all of this without ever connecting to your data warehouse**. That’s because Fusion uses [Ahead-of-Time (AOT) compilation](https://docs.getdbt.com/docs/fusion/new-concepts). Instead of running SQL directly in your data warehouse (the JIT model), AOT comprehends and analyzes your project locally, statically analyzing every model’s logic plan before running anything in your data warehouse. In other words, developers can now be assured that their models are syntactically correct before running them. That cuts down round-tripping to the data warehouse, which reduces compute costs. ## State-aware orchestration Fusion also grants dbt another superpower. It uses [state-aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-about) when running Continuous Integration (CI) jobs in a dbt environment, further reducing data warehouse costs. [State-aware orchestration is now in preview for Fusion projects](https://www.getdbt.com/blog/announcing-state-aware-orchestration). State-aware orchestration maintains a real-time shared state of every model in a project. This means it knows if, for example, Job 2 is using an upstream model that Job 1 just rebuilt a few seconds ago. In this case, it will use the results from Job 1’s run, saving you from an unnecessary rebuild. State-aware orchestration implements a model-level queue, which avoids collisions and prevents unnecessary model rebuilding. It’s also simple to configure, as it works out of the box and is Fusion’s default for model builds. You can use [advanced configuration](https://docs.getdbt.com/docs/deploy/state-aware-setup#advanced-configurations) to further tailor Fusion’s builds to meet your specific needs. ## Built for speed Finally, Fusion saves on cost and time through sheer performance. Whereas dbt Core was written in Python, the dbt Fusion engine is written in Rust. As a compiled binary written in a high-performance language, it parses even large projects up to 30x faster than dbt Core. This translates to less compute time for data pipeline jobs - as well as faster development times for data engineers. ## Conclusion ETL has been around for a while. However, while the data industry has changed significantly in the past two decades, many ETL tools have failed to keep up. Fixed-price tools are too rigid to meet the rising demands for data. On the other hand, most pay-as-you-go tools don’t provide any out-of-the-box tools to monitor and control costs. The dbt team has always strived to save our customers money by providing them with tools to keep the overall price of data transformation as low as feasible. The dbt Fusion engine builds on this legacy. Features such as Ahead-of-Time compilation and state-aware orchestration enable companies to scale data development without breaking the bank. [Sign up for an account](https://www.getdbt.com/signup) and walk through [our dbt Fusion quickstart](https://docs.getdbt.com/guides/fusion?step=1) today. --- --- title: "Building a multimodal lakehouse for AI" description: "Chang She, the CEO of LanceDB, and Tristan Handy go deep into the bridge between analytics and AI engineering." url: "https://www.getdbt.com/blog/building-a-multimodal-lakehouse-for-ai" date: "2025-11-23" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Building a multimodal lakehouse for AI Welcome back to The Analytics Engineering Podcast! Last season, we explored a host of topics on the developer experience ([something the dbt Labs crew has been pretty vocal on recently](https://www.youtube.com/watch?v=WidQLYon2_I&t=5s)). This season, we’re expanding that theme to look at how the current data landscape is impacting the developer experience. [Open data infrastructure](https://www.getdbt.com/blog/what-is-open-data-infrastructure) is on the rise; AI is pushing teams to rethink how data is modeled, governed, and scaled; and the developer experience is evolving. In this episode, Tristan Handy sits down with Chang She—a co-creator of Pandas and now CEO of LanceDB—to explore the convergence of analytics and AI engineering. The team at LanceDB is rebuilding the data lake from the ground up with AI as a first principle, starting with a new AI-native file format called Lance and building upward from there. Tristan traces Chang’s journey as one of the original contributors to the pandas library to building a new infrastructure layer for AI-native data. Learn why vector databases alone aren’t enough, why agents require new architecture, and how LanceDB is building a AI lakehouse for the future. _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ [Check out the dbt VS Code extension](https://docs.getdbt.com/docs/install-dbt-extension) **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [Youtube](https://www.youtube.com/playlist?list=PL0QYlrC86xQm83Q9deiy4euEnbw8ceu3I) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways ### Tristan Handy: You’re the founder and creator of the Lance file format and LanceDB. Before diving into vector search and vector databases, tell us about your background. **Chang She:** I love talking to analytics engineers because that’s my background. I started about 20 years ago in quantitative finance. As a junior analyst, you do a lot of data engineering and analytics, which got me into open-source Python. I became one of the co-authors of the pandas library—initially to solve my own problem of not wanting to do analytics engineering in Java or VBScript. ### You worked for a hedge fund? Yes, AQR. ### Did they know you were contributing to pandas? Hedge funds aren’t known for open source. My roommate and colleague at the time was Wes McKinney. He showed me a proprietary Python library he was working on. It was life-changing. I started using and contributing. He spent about six months convincing the fund to open-source it. This was around 2010, and they were ahead of the industry in that respect. ### I didn’t know pandas started at AQR. That’s fascinating. So much of your circa-2010 analytics work was done in early pandas? Exactly. We went through several iterations, even debated the name. Because it was a hedge fund, there was a lot of econometrics and “panel data,” so Wes named it “pandas” for panel data analysis. ### That origin story isn’t widely known. You then founded two companies, sold one to Cloudera, and were there during an interesting time. Wes and I created DataPad—cloud BI before cloud BI really took off—and sold it to Cloudera. I spent about four and a half years in the Hadoop “big data” world, where I met my co-founder. He worked on HDFS at Cloudera, and several ex-Cloudera folks are at LanceDB today. After that I moved into machine learning at Tubi TV, working on recommender systems, ML serving, and experimentation/AB testing. That exposed me to embeddings. We dealt with videos, poster art images, and synopses—data that doesn’t fit neatly into pandas or even Spark data frames. That inspired me to build better infrastructure for these data types—what we now call “classical” machine learning—which led to LanceDB. ### So that’s our bridge to vectors. You experienced these problems at Tubi, then founded the company. And Tubi used dbt? Heavily. Thank you for creating it—it was critical to our stack. ### Give us a non-technical intro: what are vectors used for? Many people focus on the latest models and techniques. My perspective: everyone has access to similar models—your differentiation comes from your data and how effectively you connect data to AI. Vectors are a way to represent any kind of data in a form models understand: high-dimensional arrays of floating-point numbers—1,500, 3,000 dimensions, etc. Early statistical models might have a few interpretable dimensions; now you can have thousands where individual dimensions aren’t necessarily interpretable, but the space captures semantics. Beyond RAG, vectors power internal model representations, recommender systems, and personalization—the original mainstream use case. ### Search is also a good use case. How is vector search different from full-text search or Command-F? Full-text search (e.g., Elasticsearch) returns documents containing the exact terms you searched. If you search for “customer,” it finds “customer/customers,” but might miss “user,” “adopter,” “organization,” etc. Vector search uses dense representations where semantically similar words and documents live near each other in high-dimensional space. Search for “customer,” and you get results that include semantically related terms. ### Would you combine vector and full-text search? Yes—hybrid search. Early RAG demos often used pure vector search for speed. Now enterprises need production-grade relevance. Many combine keyword and vector search with a re-ranking step to reach higher precision/recall. ### Early RAG pipelines often chunk text, embed, and call it done. But more thoughtful pipelines do something closer to feature engineering, right? Absolutely. Thought goes into what you feed the embedding model. For example: add a document- or section-level summary alongside each chunk before embedding; include multimodal features—artistic descriptions, literal captions, tags; create multiple embedding columns (e.g., different prompts/modalities) and search across them with re-ranking. High-quality retrieval requires feature-engineering-like decisions before embedding. ### Let’s talk vector file formats (Lance) and vector databases (LanceDB). My crude belief: a vector database is a standard database with additional indexes. True? Not wrong, but my hot take: with Lance and LanceDB, we’re building a lakehouse for multimodal data that includes vectors. Many “vector databases” are optimized only for vectors and struggle with other data types and workloads. The category needs to evolve—either toward new-generation search engines or new-generation lakehouses. We set out from day one to build the broader lakehouse, not just a vector index. ### Outline your AI-enabled data lake vision. I’m familiar with Snowflake and Databricks’ lakehouse. How do you see the world differently? We assumed everyone would use Parquet and tried for months to support AI workloads—search, training, preprocessing—on it. We couldn’t make it work well. Talking to computer-vision and ML practitioners, no one had something effective. That gave us confidence to build a new format. In AI you manage vectors, long documents, images, and videos. The first problem is storage. With Parquet, mixing wide blob columns with narrow metadata columns leads to out-of-memory issues due to row-group design. If you shrink row groups to fit blobs, read performance tanks. Even once data is in Parquet, AI needs random access and secondary indexes. Parquet doesn’t support efficient random row access: retrieving scattered rows forces reading entire row groups. With media, that’s prohibitively expensive—both for search and for training (e.g., global shuffle). Data evolution is also hard: with table formats like Iceberg, backfills often mean copying entire datasets. Copying petabytes of media is a non-starter. These issues motivated Lance. ### I have a good mental model of Parquet with structured data. With images or video, do you put them in blob columns? Yes. We use Apache Arrow types. Images/audio/video are large binary columns. Vectors are fixed-width list columns (e.g., 1,536-dimensional). But Parquet’s row-group mechanics and lack of random access make these workloads painful. ### So Lance was the first thing you built. It has solid traction on GitHub. Who uses a file format—users or vendors? Both. Frontier labs use Lance to store training data—e.g., for image/video generation—replacing stacks like TFRecords, WebDataset, Parquet, and BigQuery. Large tech companies and vendors also build on Lance: Databricks, Tencent, Alibaba, Netflix, NVIDIA, Uber, among others. ### Databricks uses Lance? For parts of their AI-specific offerings. ### You’ve raised several rounds—the format is Apache-2 licensed. How do you commercialize? Our commercial offering is a data platform for large-scale AI production: vector search, data preprocessing, training/serving cache, and an analytics engine for curation and exploration. It supports ML training workflows and AI application development, solving the hard distributed-systems problems along the path. We partner closely with big vendors; we’re generally not competitive because goals and customer bases differ. Cloud providers seek platform consumption; we focus on an AI-optimized data platform for specific workloads and users. ### The commercial product is called LanceDB, but you prefer to position it not just as a database. Right—we’re an AI-native data platform/lakehouse for multimodal data, with Lance as the common format. ### How does this space play out over the next two to three years? Two big predictions. First, multimodal will be 100× bigger—more usage and more data. Audio is exploding; video generation is resurging; robotics is next. Second, our data infrastructure isn’t ready for agents driving search and retrieval. ### Let’s unpack both. On multimodal: unlike structured analytics, where every company needs it, multimodal workloads seem concentrated. Do all enterprises really need this? I think every enterprise becomes multimodal. Take insurance: tons of documents to digitize, extract, search, and analyze; drones capturing images/video to assess risk and improvements over time. Existing businesses become more efficient; AI-native entrants gain structural advantages. Multimodal data underpins both. ### It’s a heavy lift. Will every Fortune 500 insurer build these capabilities in-house, or will vendors package them? Likely both—just like analytics engineering emerged as a role, with adjacent talent re-skilling. We see the same with AI engineering. ### What titles are hands-on with your product? AI researchers and AI engineers. Many app developers building AI features now carry the “AI engineer” title. ### On agents: how do their access patterns change platform requirements? RAG was one-shot: ask, retrieve, answer. Agents iterate: they decompose problems into sub-questions, refine queries and results, and run many steps in parallel. Load skyrockets—humans type slowly; agents can issue hundreds of queries simultaneously. Queries are more varied and selective, and agents are creative in combining modalities and sources: schemas, SQL over structured data, prior analyses and charts, document stores, image/video metadata, etc. Traditional vector databases aren’t designed for this breadth and scale. If you bolt together multiple specialized systems, your “agent stack” balloons into a maintenance nightmare. Our approach: put all data in one place with a single system that supports vector search, keyword search, filters, key-value lookups, re-ranking, analytics, and efficient random access—on top of an AI-native file format (Lance). ### For listeners whose curiosity is piqued, any resources you recommend? **Chang She:** Yes—our blog series by Weston Pace, the tech lead for Lance format. It dives into encodings, I/O, and has great reads for analytics engineers: [lancedb.com/blog](http://lancedb.com/blog) . ## Chapters - 00:00 – Intro: Analytics meets AI - 03:20 – Chang’s background and how Pandas began - 06:40 – Lessons from Cloudera and metadata - 08:30 – Multimodal data and LanceDB’s origin story - 10:00 – Why vector search matters (beyond RAG) - 12:00 – What are vectors and why do we use them? - 15:00 – Full-text vs vector search - 18:00 – Feature engineering in AI use cases - 21:15 – Lance format - 28:00 – Storage, scale, and the problem with Parquet - 35:30 – Building a business on open source - 41:00 – Two big bets: multimodal data and agents - 46:00 – Every company will become multimodal - 50:00 – Agent access patterns will redefine data - 54:00 – Why dbt-style workflows matter now more than ever --- --- title: "What is a data observability platform?" description: "A data observability platform helps teams monitor data quality, performance, and reliability across modern pipelines." url: "https://www.getdbt.com/blog/data-observability-platform" date: "2025-11-20" authors: ["Joey Gault"] categories: ["Pulse"] --- # What is a data observability platform? Modern data stacks, while powerful and flexible, have actually increased the surface area for potential issues. The proliferation of data sources, transformation tools, and consumption endpoints creates more points of failure and makes it harder to maintain end-to-end visibility. This fragmentation makes observability not just helpful, but essential for maintaining reliable data operations at scale. Consider the typical modern data architecture: data flows from hundreds of sources through various ingestion tools, gets transformed by frameworks like dbt, and ultimately feeds dozens of downstream applications and dashboards. Each step in this pipeline represents a potential failure point, and the interconnected nature of these systems means that issues can propagate quickly and unpredictably. Data engineering leaders face questions they often can't answer within reasonable timeframes: Why isn't my model up to date? Is my data accurate? Why is my model taking so long to run? How do I speed up my data pipeline? How should I materialize and provision my model? These questions highlight the gap between having data infrastructure and truly understanding how that infrastructure performs. ## Core components of data observability platforms A comprehensive data observability platform typically consists of several key components that work together to provide visibility into data systems. The foundation begins with data collection mechanisms that capture metadata about data pipelines, transformations, and quality metrics. This metadata serves as the raw material for all observability insights. Monitoring capabilities form the reactive component of observability, detecting anomalies and alerting teams when issues arise. These systems track data freshness, volume changes, schema evolution, and quality degradation across the entire data pipeline. Advanced monitoring platforms can establish baselines for normal behavior and identify deviations that warrant investigation. [Testing frameworks](https://www.getdbt.com/blog/data-testing-framework) provide the proactive element of observability, allowing teams to define expectations for data behavior and catch issues before they reach production. These tests can validate data quality, check business logic, and ensure that transformations produce expected results. When integrated properly with development workflows, testing creates a safety net that prevents many issues from ever affecting end users. Performance monitoring adds another crucial dimension, tracking execution times, resource utilization, and system bottlenecks. This capability helps teams optimize their data pipelines and make informed decisions about infrastructure scaling and query optimization. ## Integration with transformation workflows The relationship between data observability and transformation tools like dbt represents a particularly powerful combination. d[bt provides the foundation for data transformation](https://www.getdbt.com/product/what-is-dbt), offering models that clean raw data from various sources and create high-quality, usable datasets. The dbt testing framework verifies data quality through automated checks, while dbt documentation creates consistent, well-documented models that serve as single sources of truth. When observability platforms monitor [dbt models](https://docs.getdbt.com/docs/build/models) and pipelines in production, they can detect anomalies and ensure ongoing data accuracy. The real power emerges from how these tools work together rather than in isolation. For instance, when monitoring systems detect anomalies that indicate serious data quality issues, teams can create corresponding dbt test cases that prevent pipelines from proceeding if the same conditions occur again. This integration shifts responsibility for data quality upstream, enabling business users to address issues at their source rather than waiting for data engineering intervention. It also standardizes quality checks across all models, creating a consistent baseline for data quality expectations while providing automated monitoring and alerting for proactive issue notification. ## Leveraging native artifacts for observability While third-party observability tools provide valuable capabilities, teams can also build significant observability using native artifacts from their transformation tools. dbt, for example, generates detailed artifacts after every run, test, or build command, containing granular information about model execution, test results, and pipeline performance. These artifacts serve as a rich data source for custom observability solutions. The project manifest provides complete configuration information for dbt projects, while run results artifacts contain detailed execution data for models, tests, and other resources. When combined with [data warehouse](https://www.getdbt.com/discover/understanding-cloud-data-warehouses) query history, these artifacts enable deep insights into model-level performance that can inform optimization decisions. Teams have successfully built lightweight ELT systems that ingest artifact data into their data warehouses, then use dbt itself to transform this metadata into structured models that power dashboards and alerting systems. This approach leverages existing infrastructure and skills while providing customizable observability tailored to specific organizational needs. The key components of such a system include orchestration that reliably captures artifacts regardless of pipeline success or failure, storage that preserves historical artifact data for trend analysis, modeling that transforms raw artifacts into actionable insights, and alerting that notifies relevant stakeholders when issues arise. ## Effective alerting strategies Effective alerting represents one of the most critical aspects of data observability, yet it's often implemented poorly. The goal is to provide timely, actionable notifications to the right people without creating alert fatigue or overwhelming teams with false positives. Best practices for data alerting include implementing domain-specific tagging that allows alerts to be routed to appropriate team members based on model ownership. Every model in a dbt deployment might include domain tags like "growth," "finance," or "catalog," which correspond to communication channels containing relevant stakeholders. This targeted alerting ensures that model owners receive notifications about their specific models rather than broadcasting alerts to entire teams. The alerts should include sufficient context for debugging, including error messages, model names, and timestamps, enabling recipients to quickly understand and address issues. Importantly, teams should avoid introducing anomaly notifications to business users at the beginning of new data integrations. When data engineers themselves don't fully understand new data sources, it's counterproductive to alert business users about anomalies. Taking time to let pipelines stabilize and understand normal data patterns before involving business users prevents unnecessary noise and maintains alert credibility. ## Performance optimization through observability Beyond alerting and quality monitoring, observability data provides valuable insights for performance optimization. By combining transformation artifacts with data warehouse query history, teams can identify models that are candidates for different materialization strategies, clustering improvements, or warehouse sizing adjustments. Performance dashboards can surface models with high execution times, excessive data spillage, or inefficient partition scanning patterns. Time series views of individual models help identify performance degradation over time, while pipeline-level visualizations reveal bottlenecks that affect overall execution times. These insights enable data teams to make informed decisions about optimization priorities. Rather than guessing which models might benefit from incremental materialization or increased warehouse sizes, teams can use concrete performance data to guide their efforts and measure the impact of changes. Pipeline bottleneck visualization can be particularly powerful, showing thread utilization over time and helping identify models that hold up entire pipeline executions. When hourly jobs start taking longer than an hour, or nightly jobs begin affecting downstream processes, these visualizations help pinpoint specific optimization targets. ## Measuring business impact The business value of comprehensive data observability can be substantial and measurable. Organizations implementing robust observability practices often see dramatic improvements in both cost efficiency and data reliability. Performance monitoring frequently reveals opportunities for significant cost reduction through the identification of inefficient queries, unused models, and optimization opportunities. Teams have reported reductions in cloud data warehouse credit usage of 70% or more by systematically addressing performance issues identified through observability platforms. Beyond cost savings, observability frameworks enable organizations to scale while keeping costs stable. Despite adding many more models and bringing additional data sources online, teams can maintain stable job execution times while decreasing cost per unit of computation through systematic optimization. Data quality improvements are equally significant. When organizations integrate new data sources, observability platforms initially detect spikes in anomalies as teams learn to work with new data. However, the systematic conversion of anomalies into test cases leads to corresponding decreases in anomalies over time, creating self-improving systems that become more robust with experience. ## Building organizational capabilities Successful data observability requires more than just technology; it demands organizational commitment and cultural change. The most effective implementations treat observability as a core competency rather than an afterthought, investing in the tools, processes, and skills necessary to maintain visibility into data systems at scale. This includes training team members to leverage observability tools for incident routing and data discovery, creating a self-service culture that reduces the burden on data engineering teams. When business users understand how to interpret observability data and respond to alerts, they can often resolve issues without escalating to technical teams. The combination of proactive testing and reactive monitoring creates more resilient systems than either approach alone. Transformation frameworks catch many issues before they reach production, while observability tools detect the problems that slip through. When important anomalies are detected, creating test cases that prevent pipeline execution until upstream issues are resolved shifts responsibility appropriately and prevents the propagation of known data quality problems. ## Future directions and considerations As data systems continue to evolve, observability practices must adapt to new challenges and opportunities. The emergence of [data mesh architectures](https://www.getdbt.com/blog/data-mesh-architecture-explained) and domain-driven data ownership creates new requirements for observability that spans organizational boundaries while maintaining appropriate access controls and governance. Artificial intelligence and machine learning are beginning to enhance observability capabilities, from automated anomaly detection to intelligent alerting that reduces false positives. However, the fundamental principles of comprehensive monitoring, proactive testing, and effective alerting remain constant. The tools and approaches used for observability should integrate well with existing workflows and infrastructure. Solutions that require significant additional overhead or specialized expertise are less likely to be maintained effectively over time. The most successful implementations leverage existing skills and infrastructure while providing clear value to both technical and business stakeholders. ## Conclusion Data observability platforms represent a critical evolution in how organizations manage and trust their data infrastructure. By providing comprehensive visibility into data systems, these platforms enable teams to proactively identify and resolve issues, optimize performance, and maintain high levels of data quality at scale. The most successful data teams will be those that treat observability as a core competency, investing in the tools, processes, and cultural changes necessary to maintain reliable data operations. As data becomes increasingly central to business operations, organizations that master data observability will have a significant competitive advantage in their ability to make reliable, data-driven decisions at scale. Building effective observability requires commitment and investment, but the returns in terms of cost savings, improved reliability, and increased trust in data justify the effort. The combination of proactive testing, reactive monitoring, and performance optimization creates a foundation for data reliability that enables organizations to scale their data operations confidently and efficiently. ## Data observability platform FAQs **What are data observability tools?** Data observability tools are comprehensive platforms that provide visibility into data systems by monitoring, testing, and tracking data pipelines, transformations, and quality metrics. These tools consist of several key components including data collection mechanisms that capture metadata, monitoring capabilities that detect anomalies and alert teams when issues arise, testing frameworks that validate data quality and business logic, and performance monitoring that tracks execution times and resource utilization. They work together to help organizations maintain reliable data operations at scale by providing insights into data freshness, volume changes, schema evolution, and quality degradation across entire data pipelines. **How do data observability tools improve data reliability?** Data observability tools improve data reliability through a combination of proactive testing and reactive monitoring that creates more resilient systems than either approach alone. They enable teams to define expectations for data behavior and catch issues before they reach production through automated testing frameworks, while monitoring systems establish baselines for normal behavior and identify deviations that warrant investigation. When anomalies are detected, teams can create corresponding test cases that prevent pipelines from proceeding if the same conditions occur again, shifting responsibility for data quality upstream and enabling business users to address issues at their source rather than waiting for data engineering intervention. **What kinds of features should you look for in a data observability tool?** When evaluating data observability tools, look for comprehensive monitoring capabilities that track data freshness, volume changes, schema evolution, and quality degradation across your entire data pipeline. The tool should include automated anomaly detection that can establish baselines without manual configuration, proactive testing frameworks that validate data quality and business logic, and performance monitoring that tracks execution times and resource utilization. Additionally, seek tools that integrate well with your existing data stack and transformation workflows, provide effective alerting with domain-specific routing to appropriate team members, and offer insights for performance optimization including identification of bottlenecks and optimization opportunities. --- --- title: "AI readiness: How to assess and improve" description: "How to build the data foundation needed for shipping scalable, trustworthy AI workloads." url: "https://www.getdbt.com/blog/ai-readiness" date: "2025-11-19" authors: ["Daniel Poppy"] categories: ["Learn"] --- # AI readiness: How to assess and improve Artificial intelligence (AI) has become a mainstream capability, with most organizations now using AI solutions across multiple functions. However, Boston Consulting Group (BCG) reports that [74% of companies have yet to show tangible value from their AI initiatives](https://www.bcg.com/press/24october2024-ai-adoption-in-2024-74-of-companies-struggle-to-achieve-and-scale-value). AI projects that seek to scale on inconsistent, undocumented, or untrustworthy data often end in stalled efforts, lost trust, and missed opportunities. The way forward involves reliable, well-managed data backed by regular testing, documentation, and teamwork. You can only determine where you need to go next if you know where you are now. This article will discuss how organizations can conduct AI readiness assessments by evaluating their data quality, infrastructure, culture, and data governance. ## Understanding AI readiness AI readiness is an organization's ability to adopt, scale, and govern AI responsibly. It combines the right AI technologies, people, and business processes to ensure AI delivers measurable business value. Key dimensions of an AI readiness assessment include: - **Data quality and trust**. The information should be precise, verified, and regularly organized to provide sound AI system responses. High-quality testing and documentation minimize uncertainty and enhance downstream confidence. - **Infrastructure scalability**. AI use cases need infrastructure capable of managing high data volumes and automating processes. Scalable systems allow teams to transition between prototypes and production without performance concerns. - **Skills and collaboration**. Teams require a common definition of data and metrics, consistent development patterns, and aligned workflows. Close cooperation decreases the volume of rework and speeds up decision-making and analytics delivery. - **Governance and ethics**. Well-defined ownership of data, lineage, and quality criteria holds AI systems to account. Responsible AI frameworks promote accountability and transparency while addressing data privacy concerns. - **Leadership and strategy**. Leaders must set clear direction, allocate resources, and establish a roadmap aligned with long-term AI strategy and business objectives. A focused approach prevents isolated AI initiatives and strengthens organization-wide alignment. ## The data foundation of AI readiness A strong data foundation determines how reliably AI models perform. When inputs are inconsistent or unclear, machine learning models inherit those flaws. Well-structured, well-governed data reduces errors and supports stable, predictable outputs. ### Consistent structures AI systems need datasets that follow stable formats and definitions. Variations in types, fields, or schema introduce inconsistencies that affect model behavior. Consistent structures keep downstream outputs reliable. ### Automating manual steps Manual processes create logic gaps that are hard to track or reproduce. As pipelines scale, these gaps compound into larger issues. Automation of core tasks such as data validation, transformations, and scheduled refreshes keeps data flows predictable and reduces failure points. Automated workflows eliminate repetitive work and accelerate AI adoption. ### Validation, documentation, and lineage Validation ensures data meets expected rules before it moves downstream. Documentation clarifies logic so teams interpret data the same way. Lineage shows where data comes from and how it changes, supporting faster issue resolution. ### Engineering discipline Reproducible transformations depend on controlled development practices. Versioning, structured workflows, and early checks catch issues before they reach production. This discipline preserves data integrity as AI applications evolve. [A data control plane like dbt](https://www.getdbt.com/blog/data-control-plane-introduction), for example, provides the structure needed to maintain high-quality, trustworthy data. Version control brings discipline to transformation work, and [automated tests](https://docs.getdbt.com/docs/cloud/git/version-control-basics) validate assumptions before they reach production. ## Assessing your AI readiness Assessing AI readiness requires a clear view of your current data ecosystem. Organizations typically evaluate four technical areas to understand how well they can support AI implementation at scale. ### Data maturity Organizations with strong data readiness share several traits: - Clean, deduplicated datasets - Departments share definitions for key metrics and data elements in metadata or a [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction), ensuring teams use the same logic across the business - [Data testing](https://docs.getdbt.com/docs/build/data-tests) at every stage of the [analytics data lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle) to verify analytics code changes before pushing data to production - Reliable [documentation](https://docs.getdbt.com/docs/build/documentation) that is rebuilt with every push to production, so that data consumers can understand the meaning and usage of data - Traceable [data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) from data sources to production, so that consumers can verify the origin and ownership of data Teams can quantify progress using [data quality metrics](https://www.getdbt.com/blog/data-quality-metrics) such as test coverage, model reliability, and transformation cycle times. These benchmarks help measure data-driven maturity. ### Infrastructure scalability Scalable AI requires infrastructure that can: - Automate data workflows and optimize resource allocation - Support event-driven or [Continuous Integration/Continuous Delivery (CI/CD)](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud) operations - Manage rising data volume, processing demands, and workload spikes automatically without requiring manual fixes or user intervention - Maintain stable performance across workloads Cloud platforms, modern data warehouses, and orchestration AI tools are critical because they provide elastic scaling, optimized processing, and automated workflow management. Real-time processing capabilities further enhance responsiveness to business needs. ### Team alignment Data readiness is not a purely technical problem. Collaboration among stakeholders determines whether organizations move quickly or get stuck. Well-aligned teams operate with shared expectations and a common approach to how data work gets done, which shows up in practices such as: - Sharing definitions and KPIs - Reviewing transformations together - Participating in model design and validation - Using standardized processes for analytics development Shared context reduces rework and strengthens trust in the final outputs. Change management practices help teams adapt to new tools and workflows as AI capabilities mature. ### Governance and accountability [Robust governance](https://www.getdbt.com/blog/understanding-data-governance-ai) ensures data and AI systems operate safely and consistently as they scale. AI governance comes with defined regulations and protections that minimize risk and maintain reliable results. - Track AI model performance to identify drift, performance declines, or unforeseen results promptly - Apply ethical guardrails to define appropriate use cases and clarify when human oversight is required - Create incident response playbooks through risk management frameworks to address data or model crashes promptly and reduce downstream damage - Apply cross-functional supervision so that key decision-makers are aware of technical, legal, and business risks Accountability frameworks ensure that AI development is controlled, ethical, and consistent across teams. Data management protocols preserve data privacy and security throughout the AI lifecycle. dbt gives organizations measurable signals to assess their readiness. Metrics like [model test coverage](https://github.com/slidoapp/dbt-coverage), lineage completeness, and deployment cycle times reveal the strength of existing practices and highlight where AI investments will have the greatest impact. ## Accelerating AI readiness with modern data practices AI readiness does not come with standalone tools or individual efforts. It relies on the way an organization develops, sustains, and refines its data through time. Successful modern data teams use the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle). The [ADLC](https://www.getdbt.com/blog/adlc-plan) takes an idea to production through five stages, including design, build, test, deploy, and monitor. Every ADLC stage offers guardrails to minimize risk, enhance collaboration, and meet business needs. dbt puts this lifecycle on a unified platform with automatic CI/CD validation, built-in testing, and lineage updates automatically provided as part of the development process. Distributed controls and orchestration allow AI deployments to be uniform and audited. This architecture reduces technical debt, eliminates manual overhead, and allows teams to concentrate on model creation and supporting AI use cases. ### Modern data practices Beyond core engineering basics, several modern practices help teams progress from development to production faster without sacrificing quality. - **Incremental development.** Teams update only the data that changes, reducing compute costs and shortening development cycles. This approach streamlines iteration and keeps pipelines efficient as AI-powered workloads expand. - **Peer review and standardized workflows.** Consistent review processes catch logic issues early and ensure changes follow shared patterns. This alignment reduces rework and keeps transformations maintainable over time. - **Environment-based promotion.** Clear separation between development, staging, and production environments ensures changes move through controlled paths. This guards against breaking changes and supports predictable releases. - **Orchestrated, observable operations.** Centralized scheduling, visibility into run history, and lineage-aware monitoring help teams understand system health and resolve issues quickly. Observable data pipelines enable faster optimization and troubleshooting. ### Overcoming barriers to AI readiness with the ADLC and dbt Despite broad AI adoption, most organizations still struggle to generate consistent value from generative AI and other AI technologies. That's because many lack a clear, unified approach for transforming data across the organization. The good news is you can overcome these issues with a data control plane that implements the ADLC, such as dbt: - **Data silos**. Disconnected data sources create inconsistent logic and conflicting definitions. dbt [centralizes transformation](https://www.getdbt.com/blog/why-data-transformation-matters) logic into a single, version-controlled layer, ensuring teams work from shared models and consistent business rules. This unified approach supports AI integration across the entire ecosystem. - **Manual workflows**. Ad-hoc ETL and hand-built scripts slow iteration and introduce errors. dbt replaces manual steps with [automated testing](https://www.getdbt.com/product/test-and-observe), CI/CD pipelines, and scheduled runs that make transformations consistent, repeatable, and auditable. - **Unclear ownership**. Weak governance and undefined data owners produce duplicated work and untrusted outputs. dbt's [model ownership](https://www.getdbt.com/blog/the-four-principles-of-data-mesh) patterns and data contracts clarify responsibilities, enforce standards, and make quality expectations explicit across domains. Clear ownership enables effective AI governance and accountability. Together, these AI-driven capabilities create a structured, governed transformation layer that removes friction and provides the reliable data foundation AI systems require. Organizations can implement AI more confidently when their data infrastructure supports both flexibility and control. ## Conclusion When pipelines are structured, transparent, and well-governed, organizations can move from isolated experiments to meaningful, scalable outcomes. The ADLC provides the discipline needed to maintain this consistency, ensuring that data products evolve predictably as demands grow. [dbt](https://www.getdbt.com/product/dbt) brings this discipline into a single workflow, standardizing development, strengthening quality controls, and reducing the manual overhead that slows teams down. Using dbt as your data control plane, you can implement modern data practices and easily overcome the barriers that keep most organizations from achieving AI success. Whether you're beginning your AI readiness assessment or advancing an existing business strategy, having a solid data foundation is essential. The right combination of data governance, scalable infrastructure, and collaborative workflows positions your organization to capture business value from AI projects while managing risk effectively. If your organization is ready to modernize its data foundation and support scalable intelligence, [try dbt today for free](https://www.getdbt.com/lp/dbt-free-account). --- --- title: "dbt Labs Expands dbt Fusion Engine Ecosystem with Microsoft Fabric Integration" description: "dbt Labs deepens its collaboration with Microsoft and enables faster, more governed data transformations in Data Factory" url: "https://www.getdbt.com/blog/dbt-labs-integrates-dbt-fusion-engine-in-microsoft-fabric" date: "2025-11-18" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Expands dbt Fusion Engine Ecosystem with Microsoft Fabric Integration _Fusion integration will enable faster, more governed data transformation for analytics and AI workloads_ **PHILADELPHIA – **November 18, 2025 – [dbt Labs](http://getdbt.com), a leader in standards for AI-ready structured data, today announced the expansion of the dbt Fusion engine ecosystem with a new native integration in Fabric Data Factory. The integration, which builds on a longstanding collaboration with Microsoft, will allow users to build, test and orchestrate dbt transformations directly within Microsoft Fabric. The new capability will enable data teams to move faster and improve data governance while enhancing the developer experience and delivering orders of magnitude improvements in developer productivity. “The speed and complexity of modern analytics and AI projects demand that teams can transform and serve data seamlessly across their stack,” said Ryan Segar, Chief Product Officer at dbt Labs. “By bringing the dbt Fusion engine into Microsoft Fabric, users can now rely on the standard for transformation to deliver trusted, production-grade data and responsibly scale analytics for AI.” Built on Rust and equipped with deep SQL comprehension, the [dbt Fusion engine](https://www.getdbt.com/product/fusion) is a monumental evolution of the technology that powers dbt. It delivers a significantly enhanced developer experience, all while empowering organizations to operate with the highest quality, context-rich data at scale. With parse times up to 30x faster than dbt Core, Fusion executes large dbt projects in milliseconds instead of minutes. Bringing Fusion into Data Factory will allow teams to seamlessly author, test, and deploy dbt models directly in Fabric, with no Command Line Interface (CLI) setup and no external orchestration. Enterprise security controls via Fabric are automatically applied to every transformation, delivering governance at scale. “With Microsoft Fabric, teams are building the data products required to deliver trustworthy AI to their organizations,” said Faisal Mohamood, Corporate Vice President at Microsoft. “Embedding dbt into Microsoft Fabric brings together Microsoft’s cloud-scale capabilities with dbt’s transformation framework, helping customers build a robust data foundation required for AI.” Announced today at Microsoft Ignite 2025, dbt job in Microsoft Fabric is available now as public preview. The initial integration uses dbt Core and will be expanded to include the dbt Fusion engine in 2026. **About dbt Labs** Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 80,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at getdbt.com, and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "Modeling for success: Building data structures that last" description: "How to choose the correct data models for a successful, scalable data transformation architecture." url: "https://www.getdbt.com/blog/modeling-success-dbt" date: "2025-11-18" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Modeling for success: Building data structures that last Once upon a time, data transformation used to be hard. The good news is that, technically, it's now largely a solved problem. Tools like dbt have made it easy for data engineers, analytics engineers, and others to build high-quality data pipelines, no matter the size of their organization. The problem (if you can call it that) is that dbt might have made things _too_ easy. dbt's ease of use is its greatest strength - and sometimes its biggest pitfall. Teams jump in and start building, delivering value quickly. Eventually, however, they hit a wall where their approach no longer scales. The temptation is understandable. Someone asks you to whip out a quick report. You say, I have a simple tool for that! And for a while, this one-off approach works…until it doesn't. To build a successful data transformation architecture, you need to step back and consider the bigger picture. In this article, we'll look at the difference between dbt models and data models, how to choose the correct data model for a given use case, and best practices for data modeling with dbt. **** ## Understanding data models vs. dbt models First, let's start by resolving a basic confusion. Some customers confuse a **dbt model** with your **data model**. But the two are distinctly different. A [**dbt model**](https://docs.getdbt.com/docs/build/models) is a set of files that define transformation logic for a given set of data tables. These models consist of YAML files that contain the SQL or Python code laying out how source tables are converted into destination tables. A **data model**, by contrast, is the complete blueprint of how data is organized, stored, and connected across an entire system. It can span many dbt models and defines the overall data structure and relationships within the data. Think of the data model as the crafting recipe itself, with dbt models as the individual ingredients that come together to create the final product. This distinction matters. While dbt makes it easy to create individual transformation files, the real challenge lies in architecting how those pieces fit together into a coherent, scalable whole. ## The spectrum of data modeling approaches That makes it important to talk about the different types of data modeling. When we talk about data modeling, we're essentially talking about the difference between normalized and denormalized models. Think of these types of normalization as different types of vaults within _Fallout_. More normalized models like third normal form excel at write optimization, making them ideal for operational systems that process transactions and handle frequent updates. Denormalized models, on the other hand, optimize for read performance, which is exactly what analytics workloads need. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/e1a308cbb330b8b0e5cb1bcab24bfce001deba62-2892x1162.png) Four common data modeling methodologies illustrate this spectrum: **Third Normal Form** minimizes redundancy and dependency through highly structured, isolated tables. While it offers efficient storage, it requires complex queries for reporting. **Data Vault** tracks every decision with full audit history, making it ideal for high-governance use cases. It's highly scalable and flexible for data integration, but complex to implement upfront. **Dimensional Modeling** pre-aggregates data and optimizes for specific analytical use cases. It's designed for business users to understand and works exceptionally well with BI tools, though it can be challenging to change once established. **One Big Table** preserves everything in one place for simple access, but suffers from high redundancy and scaling challenges. It's important to emphasize that none of these data models is the "right" model. That depends on your use case. For example, a data vault delivers high scalability along with comprehensive auditing and efficient incremental updates. However, it's harder to engineer and query. At the other extreme, One Big Table delivers architectural simplicity and fast read performance, but doesn't scale as your business grows. The table below gives a good overview of the pros and cons of each model: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/6ce9bb6b3e14ddf3792bdf4d41eddd476e6793e6-2632x1284.png) It's possible to have all four of these approaches in your data architecture. Some companies, for example, adopt what's called a "medallion architecture," with raw data (Bronze level) as the base and other models built on top of that. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/79e8bd0491c635833c8f49decbf3f36887e6f2e0-2156x870.png) ## Why dimensional modeling wins for analytics with dbt That said, dimensional modeling has become nearly synonymous with dbt, and for good reason. dbt is predominantly a batch processing tool, and dimensional modeling is typically the best selection for batch-processed analytics. Originally developed in the 1990s by the Kimball Group, dimensional modeling was designed to solve two core problems: deliver data that's understandable to business users, and deliver it fast. While some argue that modern infrastructure has made dimensional modeling obsolete, this perspective misses the approach's most valuable benefit: making data understandable to business users. This alignment with business needs enables self-service analytics in a way that normalized transactional structures cannot. The benefits of dimensional modeling extend well beyond historical performance concerns: **Alignment with business needs:** Building data structures according to how the business thinks and operates ensures that stakeholders can actually use the data being produced. **Simplified data:** Column naming, values, and structures become user-friendly. A business user could theoretically work with a dimensional model in Excel and make sense of it, unlike transactional systems filled with cryptic IDs and bit columns. **Single source of truth:** Building each business process once eliminates redundant values replicated across multiple dbt models—a common problem in poorly architected projects. **Fast query performance:** Denormalized structures enable faster reads, crucial for analytics workloads that consume large datasets. **Modular and scalable:** Fewer dbt models need to be built, run, and maintained when following dimensional modeling principles. **Optimal for BI tools:** Industry-leading platforms like Tableau, Power BI from Microsoft, and other BI apps are built to sit on top of well-constructed dimensional models and achieve peak performance with them. ## Translating business needs into dimensional models At its core, dimensional modeling separates measurable facts from descriptive attributes: - **Facts** represent the verbs—the events and actions happening within a system - **Dimensions** represent the nouns—the people, places, and things that describe those events Facts are tables that store quantifiable data about business processes: sales transactions, shipments, website visits, payments. They contain measurable columns that can be aggregated—summed, counted, or averaged—along with foreign keys that link to dimensions. The grain of a fact table determines its level of detail. An orders fact table could exist at the total order level or drill down to the order line item level. Starting at the lowest level of granularity that might be needed makes sense, since aggregating up is always easier than disaggregating. Dimensions provide the descriptive context that enables slicing, filtering, and grouping of facts in meaningful ways. They represent business entities: customers, products, employees, locations, and time. Dimensions contain attributes that provide context. A product dimension might include category, brand, and SKU; a customer dimension might include segment, region, and contact details. The relationship between facts and dimensions creates either a star schema or a snowflake schema. **Star schemas** feature clear separation between facts and dimensions with relatively simple joins, making them ideal for analytics dashboards. **Snowflake schemas** normalize dimensions further, which can reduce duplication but adds join complexity and moves away from the benefits that make dimensional modeling valuable for analytics in the first place. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/2d9be4107f480199cf9fdbde58f4dab2149412b8-1916x1338.png) Star schemas remain the preferred approach for most use cases. ## Building scalable, analytics-ready data models with dbt Building on this, we can think of facts as your verbs. They're the events happening within your system. Your dimensions are your nouns - people, places, etc. They're values by which you filter. Starting from this, we can start to build out a star schema approach, where we have the centralized fact table. The fact table should be built around the business process you're trying to model or you're trying to solve with it. This allows us to bring everything together into a very intuitive, easy, and performant manner. What does this have to do with dbt? dbt enables us to modularize our transformations into reusable models, building them out with the proper facts and dimensions. This workflow mirrors software engineering best practices, where code is structured for maintainability and reuse. Along the way, we can make sure these models are [documented](https://docs.getdbt.com/docs/build/documentation), [well-tested](https://docs.getdbt.com/docs/build/data-tests), and clear and reliable for team members. You can use dbt to gradually build out the medallion layers we talked about above - starting with raw data sources and then building up your [staging](https://docs.getdbt.com/best-practices/how-we-structure/2-staging), [intermediate](https://docs.getdbt.com/best-practices/how-we-structure/3-intermediate), and [mart](https://docs.getdbt.com/best-practices/how-we-structure/4-marts) layers through an iterative development process. You can use tests for validation to enforce key relationships, uniqueness, and referential integrity to prevent model drift. dbt also generates dependency graphs that visualize the data flow across your entire project, making it easier to understand how models connect. ## Building with the bus matrix How do we start to put all these things together? The Kimball bus matrix provides a practical tool for visualizing dimensional models during the planning phase. This matrix maps business processes (which become fact tables) against dimensions they need to connect to. Business processes appear as rows, dimensions as columns, and intersections show which dimensions each business process requires. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/d5f86e09a60cf00d7df3760ba86a1b55f238034e-2844x700.png) Creating a bus matrix involves gathering information from stakeholder conversations and organizing it into a format that both technical teams and business users can understand. The grain of each business process should be documented—the lowest level of detail needed to serve current and anticipated use cases. This becomes a living document that grows as new dimensions and business processes are added. ### Translating business needs into dimensional models Understanding how to translate business conversations into facts and dimensions is a critical skill for data engineering teams. Business users don't speak in technical terms about tables and joins—they describe their needs in terms of processes they want to track and attributes they need for analysis. Consider a sales analytics use case. When a business stakeholder says, "We need to track every sale—who bought what, when, how much they spent, and if they got any discounts," they've essentially outlined both the fact table and several dimensions: The business process (sales) becomes the fact table. Measurable elements (amount spent, discounts) become numeric columns that can be aggregated. The descriptive elements (who, what, when) point to dimensions that need to be created or linked: customer, product, and date dimensions. When the same stakeholder adds, "It would be great to know more about our customer base—what type of customers they are and where they're from," they've identified attributes for the customer dimension: customer type and geographic details. These attributes will enable filtering and grouping of sales data by customer characteristics. A request to "see details about our products—category, brand, SKU—so we know what's selling well" similarly outlines a product dimension with specific attributes. Each additional requirement builds out more of the dimensional model, creating a comprehensive structure for analysis. Time dimensions deserve special attention. While it might seem like a date is just a single value, a well-constructed date dimension can contain hundreds of variations: weekday, month, year, quarter, fiscal values, day of week, and more. This allows users to slice data in countless ways without adding complex calculations to individual queries. Time dimensions standardize how the entire organization thinks about dates, whether using fiscal calendars or standard calendar years. ## Best practices for long-term success The impact of poor data model design often isn't realized until it's too late, and there's no easy button to fix a fundamentally flawed model. The cost to rebuild can be substantial. The following best practices help avoid that pain: **Start with business needs in mind:** Every data model should solve for a business process, enabling business users to answer their own questions and explore new ones. This data-driven approach ensures your models deliver real value. **Be clear about grain:** Building at the wrong grain creates significant technical debt. Determine the lowest level of detail needed before development begins. **Create conformed dimensions:** Master data elements like customer and product dimensions should be built once for the entire project and shared across all subject areas. This ensures data quality and consistency and makes numbers tie out across business processes. **Keep it simple:** Don't over-engineer. Build only what makes sense for the business need. Focus on clarity and reusability. **Manage slowly changing dimensions:** Understand which types of data changes need to be tracked historically versus simply overwritten. Different types of data require different handling strategies. **Govern evolution incrementally:** Establish processes for how models will evolve and how breaking changes will be handled to avoid downstream outages in your data pipelines. **Leverage SQL as your standard:** SQL remains the universal language for data transformation. Writing transformations in SQL ensures portability across platforms and enables broader team participation in the development workflow. ## Addressing data modeling from the outset dbt's accessibility makes it easy to deliver quick wins. But scalable, long-term success comes from being intentional. The lesson is clear: even with tools as user-friendly as dbt, a well-thought-out data model should be the first thing addressed in any new implementation, not an afterthought. Whether you're building data warehouses for analytics or operational pipelines for real-time decision-making, proper data modeling lays the foundation. Teams that invest in thoughtful data modeling upfront avoid the costly rebuilds that plague organizations taking shortcuts. You can either have pain now or pain later. The pain of thoughtful upfront design is significantly less than the pain of rebuilding a poorly architected system. In the analytics wasteland, survival depends on building a vault—a data model—that can withstand whatever comes next. --- --- title: "Talk to your data: AI-powered conversational analytics with the dbt MCP server" description: "Learn how one company enabled AI-powered data conversations with the dbt MCP Server and Semantic Layer." url: "https://www.getdbt.com/blog/dbt-mcp-server-conversational-analytics" date: "2025-11-13" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Talk to your data: AI-powered conversational analytics with the dbt MCP server The promise of conversational analytics has captivated data leaders for years. Imagine business users simply asking questions in natural language and receiving accurate, governed insights instantly. Yet most organizations struggle to move beyond proof-of-concepts. That’s often because of a lack of **structured context**. AI agents need to know not only what data exists, but how it's connected, what it means, and how it should be used. That spells the difference between an AI that's guessing and one that's giving accurate, trustworthy results. [Norlys](https://norlys.dk/), Denmark's largest integrated energy and telecommunications group, faced this challenge on an unprecedented scale. Following a major acquisition and organizational split, the company needed to rebuild its entire analytics infrastructure from scratch. Rather than simply migrating legacy systems, the data team saw an opportunity. Leveraging the [dbt Model Context Protocol (MCP) Server](https://www.getdbt.com/blog/mcp), they reimagined how the organization interacts with data, putting conversational analytics at the center of their vision. **** ## From 500 separate apps to one data platform Norlys serves between five and six million people in Denmark, with 800,000 customer-owners in a co-op structure. The company delivers energy, charging stations, internet, television, and mobile services while owning critical infrastructure, including Denmark's largest fiber network. When Norlys acquired the Danish operations of [Telia Mobil](http://teliacompany.com) in 2024, the merger brought together two nearly equal-sized organizations. Each had mature but fundamentally different data systems. The acquisition created both massive complexity and a unique opportunity. Telia Denmark was deeply integrated with approximately 500 applications shared across the group. Disentangling this was a complex project in its own right. Simultaneously, the Danish competition authorities mandated that Norlys separate its infrastructure businesses into distinct companies: one for its electricity business and one for its fiber-optic network. This forced additional data architecture changes across the organization. The scale of technical debt was substantial. Legacy systems had accumulated through 40 mergers over 10 years, leaving data scattered across multiple business intelligence platforms with inconsistent definitions. When different departments were asked simple questions like "how many customers do we have," the answers rarely aligned. The company had over 500 existing BI solutions that reflected outdated business structures and siloed thinking. Even _finding_ all of these solutions was challenging - when they thought they had them all, more would pop up. Rather than settle for this, company management decided that what it needed was one united workforce, working from one data platform and one single source of truth. They didn’t want 500 reports that obscured data, but a single location where users could ask data-driven questions and get accurate answers back. ## Building the foundation: Metrics first Fulfilling this bold vision meant starting fresh with a metrics-first foundation designed for modern analytics, including AI-powered conversational interfaces. The metrics-first approach centers on defining canonical business metrics once, in a centralized location, rather than scattering logic across transformation pipelines, BI tools, and application layers. Using [dbt's Semantic Layer](https://www.getdbt.com/product/semantic-layer), Norlys began systematically defining core metrics spanning strategic, tactical, and operational needs. dbt is a [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) that works across various cloud and data platform environments, providing a flexible, collaborative, and trustworthy environment for accessing data, no matter where in the organization it lives. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/b1b39f7b9a5b599fa716b6a48c3b91e56acdfe14-2318x1094.png) This meant establishing fundamental definitions, such as what constitutes a customer, how orders are measured, and how financial metrics are calculated. These definitions required extensive collaboration, bringing together energy experts, telecommunications specialists, and domain leaders to agree on common standards. [The metrics are declared in YAML files](https://docs.getdbt.com/docs/build/metrics-overview) within dbt, creating a single source of truth with several key advantages: - There’s one version of “the truth” and only one place to update logic when business rules change. - Complete [data lineage](https://www.getdbt.com/blog/what-is-data-lineage) becomes immediately visible, showing exactly which source systems and transformations feed each metric. - The approach eliminates dependencies on specific BI platforms, allowing the organization to adapt as technology evolves without rebuilding business logic. Building this foundation required significant upfront investment. The team onboarded over 100 data sources. Recognizing that they were a different business now, they carefully staged and documented each one rather than rushing to recreate legacy reports. The process took months. Ultimately, that created the structured context necessary for reliable AI applications. ## Technical implementation and partnership To implement this, Norlys selected [Snowflake](https://snowflake.com) and dbt as the core of their modern data platform, complemented by [Apache Airflow](https://airflow.apache.org/), [Qwik Talend](https://www.talend.com/), and the lightweight Python package [dlt](https://dlthub.com/) for orchestration and ingestion. This consolidated stack replaced numerous legacy platforms, giving the organization a focused foundation for scaling data operations. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/3f6f90e81597155a08eab92f00f47735378d1b57-2310x1162.png) Implementing dbt at this scale required expertise that the team was building internally. Norlys partnered with [Leap](https://leap.new/), a specialized AI and data consulting firm with deep dbt experience, to accelerate the journey. This collaboration proved essential for establishing best practices from day one, creating blueprints that enabled 60+ data engineers to develop consistently. The blueprints cover everything from code style to testing standards to documentation requirements. While this might seem rigid, the structure enables sustainability as team members change. dbt gave all data engineers and data analysts a consistent way of working together. The team set down common guidelines and best practices, which included creating [data tests](https://docs.getdbt.com/docs/build/data-tests) and thorough [documentation](https://docs.getdbt.com/docs/build/documentation) for every new data model. With this approach, any engineer can pick up work someone else started because the patterns are predictable and well-documented. It also freed the team from endless debates about implementation details, allowing them to focus on solving business problems. With that many people involved, the team feared over-generalization of roles. If everyone had a responsibility, that meant no one had it. So they created new titles and positions with clearly defined responsibilities. The partnership helped Norlys avoid common pitfalls in large-scale dbt implementations. Rather than discovering organizational patterns through trial and error over months, the team established effective project structures immediately. This meant that when they were ready to build the conversational analytics proof-of-concept, the underlying Semantic Layer was production-ready. ## From metrics to conversations with the dbt MCP server Once the Semantic Layer contained well-defined metrics with rich metadata, the path to conversational analytics became clear. The dbt MCP Server provides a standardized way for AI agents to access this structured information. With MCP Server, AI systems can understand not just what data exists but how it connects, what it means, and how it should be used. Norlys developed a proof-of-concept called Orion to demonstrate the concept, focusing initially on the finance domain where metrics were most mature. The implementation leveraged [Claude Desktop](https://www.claude.com/download) with the dbt MCP server, connecting directly to the company's defined financial metrics in dbt Cloud. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/d2bfece4996bd10f1def48e3efa8d1e51d426f25-2144x1044.png) The architecture is intentionally modular. While the proof-of-concept used Claude, the team designed the system to swap components as technology evolves. The dbt MCP server acts as an intermediary layer, allowing large language models to query metadata through APIs while maintaining governance and access controls. Users with rights to specific data domains can interact with those metrics through the conversational interface, while sensitive information remains protected. The experience differs fundamentally from traditional dashboards. When a business user asks a question, the AI agent understands the request, identifies relevant metrics from the Semantic Layer, generates appropriate queries, and returns interactive visualizations. Critically, the user can then drill deeper with follow-up questions in the same natural conversation. They can explore different angles without waiting for dashboard modifications or developer queue times. ## Demonstrating value quickly The finance domain proof-of-concept took just two to three days to build once the metrics were defined. This rapid implementation was possible because the hard work had been done upfront: staging data sources, defining metrics, writing documentation, and establishing governance. Adding the conversational layer on top required minimal additional effort. The demonstration showed business users asking questions in natural language and receiving accurate, interactive charts within seconds. Behind the scenes, the AI agent interpreted the request using the rich metadata in the Semantic Layer, identified relevant metrics, and generated visualizations. The system also provided transparency, allowing users to follow the agent's workflow and understand how it arrived at each answer. This initial success validated the metrics-first strategy and opened conversations about broader applications. The modular architecture means additional use cases can be added incrementally, each leveraging the same foundational Semantic Layer while potentially using different AI models or interfaces based on specific needs. ## Scaling the vision Norlys envisions conversational analytics not as a replacement for all traditional BI. Rather, it’s a complementary capability that addresses different use cases. Dashboards remain valuable for standardized reporting and monitoring. Conversational interfaces excel when users need to explore data, ask follow-up questions, or investigate anomalies without waiting for dashboard development cycles. The longer-term vision extends beyond asking questions about existing data. The team sees the conversational interface evolving into a "personal mission control console" where users not only gain insights but take action. For example, if an analysis reveals customers with missing contact information, an MCP-connected system could automatically trigger outreach campaigns to collect the needed data. This requires additional MCP servers beyond dbt, each exposing different capabilities while maintaining consistent governance. The modular approach allows Norlys to experiment with new AI models and tools as they emerge, swapping components without rebuilding the entire system. ## Lessons for other organizations The Norlys journey offers several lessons for data teams considering similar transformations. First, **the metrics-first approach requires significant upfront investment but pays dividends** across all analytics use cases, not just conversational AI. Having canonical metric definitions improves traditional BI, self-service analytics, and operational reporting alongside AI applications. Second, **starting small accelerates learning and builds organizational confidence**. Rather than attempting to migrate 500 legacy reports, Norlys selected one well-understood domain, defined those metrics carefully, and demonstrated value quickly. The success created momentum for expanding to additional domains. Third, **data quality matters more than ever with AI applications**. When a CFO asks about revenue through a conversational interface and receives incorrect numbers, the credibility damage is severe. The structured context from dbt's Semantic Layer helps ensure accuracy, but the underlying data must be sound. Norlys invested heavily in data staging, testing, and documentation to build that foundation. Fourth, **domain expertise must be centralized and formalized**. The days of scattered BI teams building siloed solutions are ending. Norlys brought together experts from energy, telecommunications, and other domains into a unified data and AI organization. These experts collaborated on metric definitions, ensuring consistency across the business. Focusing on domains also enables setting up your governance and security structure. For example, you can ensure that employees in HR have access to employee data, but that others outside of the department can’t see (for example) the salaries of their co-workers. Finally, **partnering with experienced practitioners accelerates time to value.** While building internal dbt expertise is essential, working with consultants who have implemented similar solutions at scale helps avoid common mistakes and establishes best practices from day one. ## The path forward Norlys continues building on this foundation, expanding metric coverage across additional domains and exploring new use cases for conversational analytics. The data team sees particular potential in customer operations, where AI-powered insights and actions could optimize processes at scale. The organization's approach demonstrates how to build an effective [Analytics Development Lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle) and [data control plane](https://www.getdbt.com/resources/whitepaper-the-control-plane-for-data-collaboration-at-scale) in practice. By centralizing business logic in dbt's Semantic Layer and exposing it through standards like the MCP server, Norlys built a flexible foundation that supports both current needs and future innovation. As AI capabilities evolve rapidly, this structured approach provides stability. The metrics, lineage, and documentation in dbt remain consistent even as the organization experiments with new large language models, interfaces, or AI agents. The investment in structured context pays returns across an expanding portfolio of AI applications, all grounded in the same trusted foundation. For organizations overwhelmed by the pace of AI innovation, the Norlys story offers a pragmatic path. Focus first on building structured, well-governed data foundations using proven tools like dbt Cloud. Once that groundwork is solid, AI applications become faster to build, more reliable in production, and easier to expand over time. The future of analytics may be conversational, but it's built on a foundation of carefully defined metrics and trusted data. --- --- title: "Automating data transformations for scalable analytics" description: "Automate your data transformations to speed delivery, ensure quality and scale with confidence." url: "https://www.getdbt.com/blog/automating-data-transformation" date: "2025-11-11" authors: ["Joey Gault"] categories: ["Pulse"] --- # Automating data transformations for scalable analytics Data is growing faster, business needs are evolving more rapidly, and the pressure to turn raw data into reliable insights is higher than ever. Manual transformation processes — writing SQL, managing dependencies, cleaning data by hand — can no longer keep up. Automation is not just a nice‑to‑have, it’s essential. In this article, we’ll walk through how automating the data transformation layer helps teams scale, improve data quality, and focus on strategic insights instead of tedious plumbing. ## The foundation of modern data operations [Data transformation](https://www.getdbt.com/blog/data-transformation) represents the systematic process of converting raw data from various sources into structured, analysis-ready formats. This process involves cleaning inconsistencies, standardizing formats, applying business logic, and creating reliable datasets that serve as the foundation for decision-making across the organization. The transformation layer sits at the heart of the modern data stack, bridging the gap between raw operational data and meaningful business insights. Without proper transformation processes, organizations find themselves trapped in cycles of manual data preparation, inconsistent metrics, and fragmented analytics efforts that undermine confidence in data-driven decisions. Modern data transformation follows the [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) paradigm, where data is first extracted from source systems and loaded into a central warehouse before transformation occurs. This approach leverages the computational power and scalability of cloud data warehouses, enabling more flexible and efficient processing compared to [traditional ETL methods](https://www.getdbt.com/blog/extract-transform-load). ## The imperative for automation Manual data transformation processes create significant bottlenecks in analytics workflows. Data teams spend considerable time writing repetitive SQL queries, managing dependencies between transformations, and ensuring consistency across different datasets. These manual processes are not only time-intensive but also prone to errors that can propagate throughout downstream analytics. Automation addresses these challenges by establishing systematic workflows that handle routine transformation tasks without human intervention. Automated systems can process new data as it arrives, apply consistent business rules, and maintain data quality standards across all transformations. This shift allows data teams to focus on higher-value activities such as developing new analytical capabilities and supporting strategic business initiatives. The benefits of automation extend beyond efficiency gains. Automated transformation processes provide better auditability, as all changes are tracked and versioned. They also enable more reliable testing and validation, ensuring that data quality issues are caught early in the pipeline rather than discovered in production dashboards. ## Building scalable transformation architectures Effective automation requires a well-designed transformation architecture that can handle growing data volumes and increasing complexity. This architecture must support [modular development](https://www.getdbt.com/blog/modular-data-modeling-techniques), where individual transformations can be developed, tested, and deployed independently while maintaining proper dependencies and relationships. A robust transformation layer incorporates several key components. [Version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) systems track changes to transformation logic, enabling teams to collaborate effectively and roll back problematic changes when necessary. [Automated testing frameworks ](https://docs.getdbt.com/docs/build/data-tests)validate data quality at multiple stages, from individual transformation steps to final output datasets. Documentation systems maintain current information about data lineage, business logic, and usage patterns. The architecture must also support different development environments, allowing teams to test changes against representative data before deploying to production. This capability is essential for maintaining system stability while enabling continuous improvement of transformation processes. ## Implementing automated workflows [Modern data transformation tools like dbt](https://www.getdbt.com/product/what-is-dbt) have revolutionized how organizations approach automation. dbt enables teams to define transformations as code using SQL, creating a development workflow that mirrors software engineering best practices. This approach brings version control, testing, and documentation directly into the transformation process. Automated workflows in dbt begin with modular transformation logic that can be reused across different models and projects. These transformations are defined in SQL files that reference other models, creating a dependency graph that the system can execute in the proper order. The framework automatically handles complex dependency resolution, ensuring that upstream models complete successfully before downstream transformations begin. Testing automation represents another critical component of modern transformation workflows. dbt includes built-in testing capabilities that can validate data quality assumptions automatically. These tests run as part of the transformation process, catching issues such as null values in required fields, duplicate records, or unexpected data distributions before they impact downstream users. Documentation automation eliminates the traditional burden of maintaining separate documentation systems. dbt automatically generates documentation from the transformation code itself, including data lineage diagrams that show how data flows through the system. This automated documentation stays current with the actual implementation, providing reliable reference material for both technical and business users. ## Orchestration and scheduling Automated transformation systems require sophisticated orchestration capabilities to manage complex workflows efficiently. Modern orchestration platforms can trigger transformations based on data availability, schedule regular updates, and handle error recovery automatically. These systems monitor upstream data sources and initiate transformation processes when new data becomes available, ensuring that analytical datasets remain current without manual intervention. Intelligent scheduling algorithms optimize resource utilization by running transformations during periods of lower system demand. They can also prioritize critical transformations during peak business hours while deferring less urgent processes to off-peak periods. This [optimization reduces infrastructure costs](https://www.getdbt.com/product/cost-optimization) while maintaining service levels for business-critical analytics. Error handling and recovery mechanisms are essential components of automated orchestration. When transformations fail, the system can automatically retry operations, send notifications to appropriate team members, and implement fallback procedures to maintain system availability. These capabilities ensure that temporary issues don't cascade into broader system failures. ## Quality assurance through automation Automated quality assurance processes are fundamental to reliable transformation systems. These processes go beyond basic data validation to include comprehensive testing of business logic, performance monitoring, and consistency checks across related datasets. Automated testing frameworks can validate that transformations produce expected results under various conditions, catching logic errors that might not be apparent during initial development. Data quality monitoring systems continuously assess the health of transformation outputs, tracking metrics such as record counts, value distributions, and freshness indicators. When these metrics deviate from expected ranges, automated alerts notify relevant team members, enabling rapid response to potential issues. This proactive monitoring prevents data quality problems from impacting business operations. Automated regression testing ensures that changes to transformation logic don't inadvertently break existing functionality. These tests compare outputs from modified transformations against baseline results, flagging unexpected differences for review. This capability is particularly valuable in complex transformation environments where changes to one model might have subtle effects on downstream processes. ## Performance optimization and cost management Automated transformation systems must balance performance requirements with cost considerations, particularly in cloud environments where compute resources are billed based on usage. Modern transformation tools include optimization features that can automatically improve query performance and reduce resource consumption. Query optimization algorithms analyze transformation logic and suggest improvements such as more efficient join strategies, better indexing approaches, or opportunities to reduce data scanning. Some systems can automatically implement these optimizations, while others provide recommendations for manual review and implementation. Resource management automation adjusts compute capacity based on workload demands, scaling up during peak processing periods and scaling down during quieter times. This dynamic scaling ensures that transformation jobs complete within acceptable timeframes while minimizing unnecessary infrastructure costs. ## Governance and compliance automation Automated governance processes ensure that transformation systems comply with organizational policies and regulatory requirements. These processes can automatically apply data classification rules, implement access controls, and maintain audit trails of all transformation activities. Automated compliance checking validates that transformations follow established data handling procedures and flag potential violations for review. [Data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) tracking automation maintains comprehensive records of how data flows through transformation processes, supporting both operational troubleshooting and regulatory compliance requirements. These systems can automatically generate lineage documentation and impact analysis reports, showing which downstream systems might be affected by changes to specific data sources or transformations. ## Integration with AI and machine learning The integration of artificial intelligence into transformation workflows represents a significant advancement in automation capabilities. AI-powered tools can automatically generate transformation code from natural language descriptions, reducing the time required to implement new analytical requirements. These tools understand existing data structures and business context, producing code that follows established patterns and conventions. Machine learning algorithms can optimize transformation performance by analyzing historical execution patterns and predicting optimal resource allocation strategies. They can also identify anomalies in data patterns that might indicate quality issues or changes in upstream systems, enabling proactive response to potential problems. Automated code generation and optimization capabilities are becoming increasingly sophisticated, with tools like dbt Copilot providing context-aware assistance for transformation development. These AI-powered assistants can suggest improvements to existing code, generate test cases, and create documentation, further accelerating the development process while maintaining quality standards. ## Future directions in transformation automation The evolution of automated data transformation continues to accelerate, driven by advances in cloud computing, artificial intelligence, and data processing technologies. Emerging capabilities include more sophisticated automated optimization, intelligent error recovery, and adaptive scheduling that responds to changing business priorities. Real-time transformation automation is becoming increasingly important as organizations seek to reduce the latency between data generation and analytical insights. Stream processing capabilities enable transformations to occur as data flows through the system, rather than in batch processes that introduce delays. The integration of automated transformation systems with broader data platform capabilities creates opportunities for more comprehensive automation. These integrated platforms can automatically provision resources, configure security settings, and optimize performance across the entire data pipeline, reducing the operational overhead associated with managing complex data environments. As organizations continue to recognize the strategic value of automated data transformation, investment in these capabilities will likely accelerate. The organizations that successfully implement comprehensive automation will gain significant competitive advantages through faster time-to-insight, improved data quality, and more efficient resource utilization. The foundation for this success lies in thoughtful architecture design, appropriate tool selection, and a commitment to engineering best practices that ensure automated systems remain reliable, maintainable, and aligned with business objectives. ## Automating data transformation FAQs **What tools and methods can you use to automate data extraction, transformation, and loading (ETL) processes?** Modern data transformation automation relies on several key tools and methods. dbt (data build tool) has revolutionized the space by enabling teams to define transformations as code using SQL, incorporating version control, testing, and documentation directly into the transformation process. Modern orchestration platforms provide sophisticated scheduling and workflow management capabilities, triggering transformations based on data availability and handling error recovery automatically. Cloud data warehouses support the ELT paradigm, where data is extracted, loaded, and then transformed using the computational power of scalable cloud infrastructure. Additionally, AI-powered tools are emerging that can automatically generate transformation code from natural language descriptions and optimize performance through machine learning algorithms. **How does a data transformation layer automate processing to reduce manual data preparation and ensure consistency?** A data transformation layer automates processing by establishing systematic workflows that handle routine transformation tasks without human intervention. The system automatically processes new data as it arrives, applies consistent business rules across all datasets, and maintains data quality standards throughout all transformations. Automated workflows use modular transformation logic that can be reused across different models and projects, with dependency graphs that execute transformations in the proper order. Built-in testing capabilities validate data quality assumptions automatically, catching issues like null values, duplicates, or unexpected data distributions before they impact downstream users. This automation eliminates the time-intensive manual processes of writing repetitive SQL queries and managing dependencies, while providing better auditability through tracked and versioned changes. **What governance, monitoring, and version control capabilities are needed to safely automate data transformation at scale?** Safe automation at scale requires comprehensive governance and monitoring capabilities. Version control systems must track all changes to transformation logic, enabling effective team collaboration and the ability to roll back problematic changes. Automated testing frameworks should validate data quality at multiple stages, from individual transformation steps to final output datasets. Data quality monitoring systems need to continuously assess transformation outputs, tracking metrics like record counts, value distributions, and freshness indicators, with automated alerts when metrics deviate from expected ranges. Documentation systems must maintain current information about data lineage, business logic, and usage patterns, ideally generated automatically from the transformation code itself. Additionally, automated governance processes should apply data classification rules, implement access controls, maintain comprehensive audit trails, and ensure compliance with organizational policies and regulatory requirements through automated compliance checking and lineage tracking. --- --- title: "What is data pipeline observability?" description: "Discover how data pipeline observability improves system trust, performance, and troubleshooting in modern ELT stacks." url: "https://www.getdbt.com/blog/data-pipeline-observability" date: "2025-11-10" authors: ["Joey Gault"] categories: ["Pulse"] --- # What is data pipeline observability? The shift toward cloud-native data architectures has fundamentally changed how organizations process and manage data. [Modern ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) pipelines leverage the processing power of cloud data warehouses like Snowflake, BigQuery, and Databricks to transform data after it's loaded, enabling parallel processing and more flexible data handling. However, this architectural evolution has also introduced new complexities that traditional monitoring approaches struggle to address. Today's data pipelines consist of multiple interconnected components: ingestion systems that pull data from various sources, loading processes that move raw data into warehouses, transformation layers that clean and model data, orchestration systems that manage execution timing, and storage systems that maintain processed data for analysis. Each component represents a potential point of failure, and the interactions between components can create cascading effects that are difficult to trace without proper observability. The fragmentation of modern data stacks compounds these challenges. Organizations typically use multiple tools across their data pipeline (different systems for ingestion, transformation, orchestration, and consumption). This fragmentation makes it difficult to maintain visibility across the entire data estate, particularly when issues span multiple systems or when the root cause of a problem lies upstream from where symptoms appear. ## Understanding the scope of pipeline observability Data pipeline observability encompasses several key dimensions that work together to provide comprehensive visibility into data operations. Performance monitoring tracks how long transformations take to execute, identifies bottlenecks in pipeline execution, and surfaces opportunities for optimization. This includes monitoring query performance, warehouse utilization, and resource consumption patterns that can inform decisions about materialization strategies, clustering, and infrastructure sizing. Data quality monitoring goes beyond simple accuracy checks to include freshness monitoring that ensures data arrives when expected, completeness validation that verifies all expected data is present, and consistency checks that ensure data conforms to expected formats and business rules. These monitoring capabilities must operate continuously and provide early warning when data quality issues emerge. Lineage tracking provides visibility into how data flows through the pipeline, from source systems through various transformation stages to final consumption points. Column-level lineage enables teams to understand the journey of individual data elements, making it easier to trace the impact of changes and identify the root cause of issues when they occur. Error detection and alerting systems must be sophisticated enough to distinguish between minor anomalies and critical issues that require immediate attention. Effective alerting routes notifications to appropriate stakeholders based on model ownership and domain expertise, providing sufficient context for rapid diagnosis and resolution. ## The business impact of observability gaps The consequences of inadequate pipeline observability extend far beyond technical inconvenience. When data teams cannot quickly identify and resolve issues, the impact cascades through the organization. Executive dashboards may display incorrect metrics, automated marketing campaigns may target the wrong customers, and financial reporting may be delayed or inaccurate. These incidents erode trust in data systems and can lead to hesitation in making data-driven decisions. Organizations with poor observability often experience what industry practitioners call "data downtime": periods when data is partial, erroneous, missing, or inaccurate. During these periods, data consumers lose confidence in the systems they depend on, and data teams spend disproportionate time firefighting rather than building new capabilities. The cost of poor data quality has been shown to impact significant portions of company revenue, making observability not just a technical necessity but a business imperative. The complexity of modern data operations means that issues can remain hidden for extended periods before being discovered. Without proactive monitoring, teams may only learn about problems when business users report discrepancies in reports or dashboards. By this time, the issue may have affected multiple downstream systems and require extensive investigation to identify and resolve. ## Building observability with dbt artifacts [dbt](https://www.getdbt.com/product/what-is-dbt) provides a foundation for pipeline observability through its comprehensive artifact system. Every time dbt executes a run, test, or build command, it generates detailed artifacts containing granular information about model execution, test results, and pipeline performance. These artifacts serve as a rich data source for building custom observability solutions that can be tailored to specific organizational needs. The project manifest provides complete configuration information for the dbt project, including model definitions, dependencies, and metadata. Run results artifacts contain detailed execution data for models, tests, and other resources, including execution times, success or failure status, and error messages when issues occur. When combined with data warehouse query history, these artifacts enable deep insights into model-level performance that can inform optimization decisions. Teams have successfully built lightweight ELT systems that ingest artifact data into their data warehouses, then use dbt itself to transform this metadata into structured models that power dashboards and alerting systems. This approach leverages existing infrastructure and skills while providing customizable observability tailored to specific organizational requirements. The key to effective artifact-based observability lies in reliable collection and processing of this metadata. Systems must capture artifacts regardless of pipeline success or failure, since understanding what went wrong is often more important than tracking successful executions. Automated processes should upload artifacts to external storage immediately after dbt execution, ensuring that metadata is preserved even when pipeline failures occur. ## Implementing effective alerting strategies Effective alerting represents one of the most critical aspects of data pipeline observability, yet it's frequently implemented poorly. The goal is to provide timely, actionable notifications to the right people without creating alert fatigue or overwhelming teams with false positives. This requires careful consideration of alert routing, content, and timing. Domain-specific alerting ensures that notifications reach the people best positioned to address issues. By tagging [dbt models](https://docs.getdbt.com/docs/build/models) with domain identifiers like "finance," "marketing," or "product," teams can route alerts to appropriate stakeholders rather than broadcasting notifications to entire data teams. This targeted approach reduces noise while ensuring that model owners receive timely notifications about issues affecting their specific areas of responsibility. Alert content must provide sufficient context for rapid diagnosis and resolution. Effective alerts include model names, error messages, execution timestamps, and links to relevant documentation or dashboards. This information enables recipients to quickly understand the scope and nature of issues without requiring additional investigation to gather basic facts. Timing considerations are equally important. Alerts should be triggered quickly enough to enable rapid response, but not so aggressively that temporary issues generate unnecessary notifications. Implementing appropriate delays and thresholds helps distinguish between transient problems that resolve themselves and persistent issues that require intervention. ## Performance optimization through observability Pipeline observability data provides valuable insights for performance optimization that extend beyond simple monitoring. By combining [dbt artifacts](https://docs.getdbt.com/reference/artifacts/dbt-artifacts) with data warehouse query history, teams can identify models that would benefit from different materialization strategies, clustering improvements, or warehouse sizing adjustments. Performance dashboards can surface models with consistently high execution times, excessive data spillage, or inefficient partition scanning patterns. Time series views of individual models help identify performance degradation over time, while pipeline-level visualizations reveal bottlenecks that affect overall execution times. These insights enable data teams to make informed decisions about optimization priorities rather than guessing which changes might improve performance. Observability data also supports capacity planning and cost management. Understanding which models consume the most resources, when peak usage occurs, and how performance changes over time helps teams make informed decisions about infrastructure sizing and scheduling. This data-driven approach to resource management can result in significant cost savings while maintaining or improving performance. ## Integration with broader data quality initiatives Pipeline observability works most effectively when integrated with broader data quality initiatives rather than implemented in isolation. The combination of proactive testing through dbt and reactive monitoring through observability tools creates more resilient systems than either approach alone. dbt tests catch many issues before they reach production, while observability tools detect problems that slip through initial validation. Converting observability alerts into proactive test cases creates a self-improving system that becomes more robust over time. When monitoring systems detect anomalies that indicate serious data quality issues, teams can create corresponding dbt tests that prevent pipelines from proceeding if the same conditions occur again. This approach shifts responsibility for data quality upstream, enabling business users to address issues at their source rather than waiting for data engineering intervention. This integration also standardizes quality expectations across all models. Every dbt model becomes subject to consistent testing requirements, creating a baseline for data quality that applies throughout the organization. Regular performance reviews ensure that models don't degrade over time, while automated monitoring provides ongoing validation of data quality assumptions. ## The strategic value of observability Data pipeline observability represents more than a technical capability: it's a strategic enabler that allows organizations to scale their data operations while maintaining reliability and trust. Teams with comprehensive observability can confidently make changes to their pipelines, knowing that issues will be detected and addressed quickly. This confidence enables more rapid iteration and innovation in data products and analytics. Observability also supports the transition to more distributed data ownership models. As organizations adopt data mesh architectures and push data ownership closer to business domains, observability becomes essential for maintaining quality and reliability across decentralized teams. Clear visibility into data lineage, quality metrics, and performance characteristics enables domain teams to take ownership of their data products while maintaining organizational standards. The investment in pipeline observability pays dividends as organizations grow and data complexity increases. Rather than constantly fighting fires and rebuilding fragile systems, teams with solid observability foundations can focus on delivering business value through innovative data products and insights that drive competitive advantage. This shift from reactive maintenance to proactive development represents a fundamental transformation in how data teams operate and deliver value to their organizations. As data becomes increasingly central to business operations, the organizations that master data pipeline observability will have significant advantages in their ability to make reliable, data-driven decisions at scale. The combination of comprehensive monitoring, proactive testing, and effective alerting creates the foundation for trustworthy data systems that can support ambitious analytics and AI initiatives while maintaining the reliability that business stakeholders require. ## Data observability FAQs **How does pipeline observability monitor and alert on the health and performance of CI/CD pipelines across platforms** Pipeline observability monitors health and performance through comprehensive tracking of multiple key dimensions including performance monitoring that tracks execution times and identifies bottlenecks, data quality monitoring that ensures freshness and completeness, and lineage tracking that provides visibility into data flows. Effective alerting systems distinguish between minor anomalies and critical issues, routing notifications to appropriate stakeholders based on domain expertise and providing sufficient context for rapid diagnosis and resolution. **How do the five pillars (freshness, distribution, volume, schema, and lineage) work together to ensure real-time observability and reliability of data pipelines?** The five pillars work together to create comprehensive pipeline visibility by addressing different aspects of data quality and flow. Freshness monitoring ensures data arrives when expected, distribution and volume tracking detect anomalies in data patterns, schema validation ensures data conforms to expected formats and business rules, and lineage provides visibility into how data flows through transformation stages. These pillars operate continuously to provide early warning when issues emerge and enable teams to trace the impact of changes across the entire data pipeline. **What approaches help reduce false positives in anomaly detection while scaling pipeline observability across complex, distributed, and legacy-integrated systems?** Reducing false positives requires implementing appropriate delays and thresholds to distinguish between transient problems that resolve themselves and persistent issues requiring intervention. Effective approaches include domain-specific alerting that routes notifications to the right stakeholders rather than broadcasting to entire teams, converting observability alerts into proactive test cases that create self-improving systems, and integrating observability with broader data quality initiatives. This combination of targeted alerting, intelligent thresholds, and proactive testing creates more resilient systems that scale effectively across complex architectures. --- --- title: "What is data infrastructure and how to design it" description: "Use dbt to design and modernize your data infrastructure for AI-powered analytics." url: "https://www.getdbt.com/blog/data-infrastructure" date: "2025-11-10" authors: ["Daniel Poppy"] categories: ["Learn"] --- # What is data infrastructure and how to design it In a world that demands agility and adaptability, a company can move forward with the right data infrastructure. Companies need effective systems that maintain low costs and high performance while supporting big data initiatives. These days, many companies are looking at AI to make their data infrastructure leaner and more efficient. Gartner reports that [over 54% of operational leaders use AI in their infrastructure](https://www.gartner.com/en/newsroom/press-releases/2025-10-29-gartner-survey-54-percent-of-infrastructure-and-operations-leaders-are-adopting-artificial-intelligence-to-cut-costs) to automate processes and optimize costs. However, simply adding AI isn't enough. Teams struggle with fragmented data pipelines and inconsistent transformation logic that slow analytics workflows. Poor data quality and siloed data sources create bottlenecks that prevent organizations from becoming truly data-driven. dbt resolves these challenges by automating transformations and standardizing data modeling, no matter where your data lives. In this article, we'll explore how to design modern data infrastructure with dbt to streamline transformations and modeling workflows while ensuring scalability and data governance. ## What is data infrastructure? Data infrastructure is the system of tools and processes that businesses use for data management. This ecosystem may involve data warehouses, data lakes, and cloud platforms. Its key components include data ingestion, data storage, data processing, transformation, and secure access controls. Robust data infrastructures support data analytics, enabling businesses to make more insightful, data-driven decisions. A solid data infrastructure also addresses data quality concerns through automated testing and validation. It's important to design data infrastructure with data security features that protect sensitive information from unauthorized access. This approach ensures data privacy and helps companies remain compliant with regulations like GDPR. ## How to design a modern data infrastructure Designing your data infrastructure must begin with careful planning. Best practices in data ingestion, storage, and transformation help build efficient systems for all data operations. The right data architecture balances business needs with technical functionality. ### Designing your data ingestion layer Businesses obtain data from multiple data sources, including CRMs, APIs, websites, and spreadsheets. This ingestion can be automated with tools like [Airbyte](https://airbyte.com/) or using custom scripts. These tools provide connectors for standard data sources and support various data types and formats. However, for internal APIs and unique systems, custom scripts are often more effective, as they offer greater customization and control over data flow. ETL (Extract, Transform, Load) processes handle the movement of raw data from source systems into your data infrastructure. It's essential to enforce data schemas during ingestion. This prevents structural inconsistencies that could negatively impact downstream data analytics. Frameworks like [AWS Glue](https://docs.getdbt.com/docs/core/connect-data-platform/glue-setup) help properly format data to reduce downstream errors. This proactive step ensures reliable datasets for all data analysis and reporting that follows, supporting both business intelligence and machine learning use cases. ### Choosing the right data storage architecture Data should be stored in centralized storage systems that are reliable and secure. Data storage architectures define how data is structured and integrated within an organization. The right storage solution depends on your data volumes and workloads. Comparing data storage architectures: **Data warehouses** like [Redshift](https://aws.amazon.com/redshift/) and [BigQuery](https://cloud.google.com/bigquery), as well as cloud-based platforms, store structured data and support complex queries. They enable fast data retrieval and reporting, providing timely insights for decision-making. Data warehouses excel at handling analytics tools and business intelligence workflows. **Data lakes** can store unstructured data and semi-structured data. They retain the original data form, making it reusable for both analytics and AI use cases. Data lakes also offer flexible storage for diverse datasets and support large-scale data processing. Technologies like [Hadoop](https://hadoop.apache.org/) enable processing of big data within data lake environments. **Data lakehouses** are hybrid storage solutions that combine the reliability of data warehouses with the flexibility of data lakes. This data architecture provides the best of both worlds for modern data infrastructure. Both on-premises and cloud-based storage systems have their place, though cloud solutions offer better scalability for growing data volumes. Consider data integration requirements when choosing between these options. ### Transforming and modeling your data Data must be cleaned and adequately structured after centralization so it's usable. Removing errors and inconsistencies makes the dataset reliable enough for analysis and data modeling. [The data transformation process](https://www.getdbt.com/blog/data-transformation) involves preparing raw data for analysis by cleaning and validating it. Naming rules, joins, and logical data models are applied to present information in a meaningful way. This stage is critical for ensuring data quality and data integrity. Tools like dbt simplify data transformation tasks and enable [modular transformations](https://www.getdbt.com/blog/modular-data-modeling-techniques) with automated testing. Data engineers and analytics engineers use SQL to build reusable data models. dbt also integrates with orchestration tools like Airflow to keep transformations consistent when new data arrives. Reliable and well-modeled data enhances the accuracy of predictive models and AI-powered systems. Teams can build dashboards and generate insights that accurately reflect real business conditions. This data-driven approach gives organizations a competitive advantage. ### Making data accessible for insights Data modeling and validating data are only useful if teams can effectively access it. The way data flows into dashboards and reports reveals patterns and risks that might otherwise remain hidden. Data accessibility is crucial for enabling self-service analytics. **Dashboards** highlight critical metrics where decisions happen. When users can quickly identify trends through visualization, they adjust strategies before small issues become big problems. Modern dashboards provide real-time data updates for faster decision-making. **Automated reports** deliver the right data at the right time and prevent teams from relying on outdated data. Automation frees analysts to focus on data analysis instead of manual data compilation. **Analytics tools** can be embedded into existing workflows to deliver actionable insights. Users explore data within their apps, so decisions are made contextually without switching tools. This embedded approach supports various use cases across the organization. When teams access real-time data, they use it confidently and make better decisions. Strong data governance ensures that users work with trusted, accurate information. ## How dbt helps teams make smarter decisions Data transformation often becomes complex due to inconsistent logic and manual workflows across different teams. dbt makes the transformation step practical for analysts by using standardized SQL models. Data and analytics engineers define business logic and transformation steps directly in code. These SQL models become the modular components of subsequent data pipelines. Designing efficient data infrastructure with dbt: **Modular transformation.** Data engineers use dbt to build modular [SQL models](https://docs.getdbt.com/docs/build/sql-models) that depend on each other. This modular approach creates a clear and maintainable data pipeline from raw to curated layers. The functionality extends across your entire data engineering ecosystem. **Version control and data governance.** [Version-controlled data models](https://docs.getdbt.com/docs/mesh/govern/model-versions) using Git track every change transparently. This enhances data governance and ensures consistent transformation logic across environments. Teams maintain clear ownership and accountability for data assets. **Automated testing.** [Automated tests](https://www.getdbt.com/blog/build-trust-through-data-testing) in data models maintain unique and complete records and verify freshness with each run. These validations detect data quality issues early, before they affect dashboards or AI outputs and lead to incorrect decision-making. Testing is essential for building trust in your data infrastructure. **Cost and speed optimization.** [Incremental materializations](https://docs.getdbt.com/docs/build/materializations) handle only new data to cut warehouse load and runtime costs. This optimization keeps queries fast and data pipelines more efficient as data volumes increase. Smart resource allocation delivers better performance at lower cost. **Enhanced transparency.** The automatically generated [data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) graph displays every dependency from data source to report. This visual lineage adds trust by making ownership and data flow visible. Understanding data lineage is crucial for troubleshooting and optimization. **Accessible documentation.** dbt generates [documentation](https://docs.getdbt.com/docs/build/documentation) for data models, columns, and data sources within the project. Documentation enables all users to understand data logic without relying on data engineers. This democratizes data access across the organization. **Semantic layer.** dbt's [semantic layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) enables codifying metrics as declarative assets. Teams then store these definitions centrally and share them with BI tools, machine learning pipelines, and APIs. This shared logic keeps dashboards and AI models aligned on the same trusted definitions, eliminating silos. **APIs.** With [dbt Cloud APIs](https://docs.getdbt.com/dbt-cloud/api-v2#/), you can trigger downstream pipelines after completing the transformation. This orchestration ensures AI workflows always use the latest validated data. APIs enable data integration with various analytical tools and platforms. **Fast, low-cost development.** dbt's [Fusion engine](https://docs.getdbt.com/docs/fusion/about-fusion) understands the syntax of all major data warehouses and can validate changes locally, before ever checking in a single line of code. It delivers stateful performance, only re-running models when needed, resulting in significant cost savings. This scalable approach works regardless of data platform. With centralized logic, different business units operate using the same metric definitions. This helps business users make confident, data-driven decisions based on high-quality data. ## How your business can benefit from a modern data infrastructure A well-designed data infrastructure provides stability and ensures data flows reliably from source to dashboard. This stability removes manual effort and immediately reduces operational costs. Solid data infrastructure creates a foundation for data strategy and initiatives. A streamlined data infrastructure adds several business benefits: **Accelerated decision-making.** Immediate data access enables businesses to absorb insights and act in real-time. Teams quickly adjust their strategies to respond to market changes. Real-time analytics support faster, more informed decisions. **Enhanced efficiency.** Automation in data ingestion and data processing minimizes manual labor time. This frees human experts to focus on complex data analysis and strategic initiatives rather than routine data management tasks. **Scalable growth.** The scalability and elasticity of cloud solutions enable seamless scaling of data volumes and workloads. This capability powers both current and future analytics and AI initiatives across your data centers and cloud infrastructure. **Personalized experience.** Real-time data processing yields faster customer insights. These insights enable AI-powered personalization services, such as intelligent recommendations. Better data quality directly improves customer experiences. Organizations that invest in modern data infrastructure gain a competitive advantage through faster insights, better data governance, and more efficient data management across their entire ecosystem. ## Conclusion Data infrastructure isn't a one-and-done affair. It's a strategic asset. Teams avoid data silos and respond faster to market changes when data infrastructure is efficiently designed. Strong data architecture supports scalability and adaptability as business needs evolve. Data transformation becomes more straightforward when using dbt. With dbt, you can manage your data via a centralized data control plane, bringing consistency and data quality to your data no matter where it lives—from data warehouses to data lakes to hybrid environments. A data infrastructure designed with intention creates a solid foundation for your data management and data strategy. To try it for yourself, [sign up for free today](https://www.getdbt.com/signup) and start automating your data infrastructure for high-performance, AI-powered analytics that deliver real business value. --- --- title: "How to find balance in data work (and prevent burnout before it finds you)" description: "The View on Data hosts share real stories about burnout, boundaries, and finding balance in data careers." url: "https://www.getdbt.com/blog/how-to-find-balance-in-data-work" date: "2025-11-07" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # How to find balance in data work (and prevent burnout before it finds you) In the latest episode of _The View on Data_, hosts Erica “Ric” Louie, Faith McKenna, Paige Berry, and Jerrie Kenney get real about one of the hardest parts of working in data: **finding balance** and recognizing burnout before it finds you. From late-night debugging sessions to the pressure to learn faster and ship more, the team talks about how they set boundaries, recharge creatively, and redefine what success looks like as their lives and careers evolve. 🎧 Listen & subscribe: [Spotify](https://open.spotify.com/show/0Itxg3BoF6zaQdORfVuY6I?si=6133f88257624765) | [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-view-on-data/id1835755886) | [Amazon Music](https://music.amazon.com/podcasts/84c47dbe-adf7-4d33-a371-a1fe55255495/the-view-on-data) | [YouTube](https://youtu.be/x4pzq7HMbJ4) ## What balance looks like when you work in data The group started with a simple but impossible question: _Does balance even exist?_ For Jerrie, a mom of three, balance means being fully present in whatever is in front of her. “There’s no perfect balance,” she said. “It’s always shifting based on what needs my attention most right now.” Having hard boundaries on start and finish times—along with being fully in “work mode” or “family mode”—helps her stay grounded. Faith described balance as something seasonal. “Coalesce is the perfect example,” she said. “During the run-up, there’s pretty much no work-life balance, and that’s okay. But once it’s over, I rebuild those boundaries. I start my day at nine, and if I’m ready early, I use that time for something personal instead of jumping back into work.” Paige and Ric both rely on structure to create separation. Paige starts every morning with meditation and a treadmill walk before checking Slack. Ric plans her week every Sunday and has a hard 7:00 p.m. cutoff to avoid letting work spill into her evenings. “If I work past seven, I know I won’t feel good,” Ric shared. “It’s about being honest with myself about what’s sustainable.” ## How to recognize burnout before it sneaks up on you Burnout can show up quietly, and often by the time you notice, it’s already taken hold. For data practitioners, it often hides behind productivity. You might be shipping dashboards, answering questions, or debugging pipelines, but something feels off. For Jerrie, the red flag is when work follows her into her dreams. “When I start dreaming about data problems, that’s when I know I need to step away,” she said. Her fix is simple: she writes the problem down, gets it out of her head, and gives herself permission to think about it later. Paige recognizes burnout when she stops feeling joy outside of work. “If I’m anxious all the time or overthinking something from the day, I write it all down,” she said. “Once I can see everything in front of me, I realize not everything needs solving right now.” Faith described how burnout often stems from an imbalance between effort and reward. “Burnout happens when you put in a lot of work but don’t feel any payoff,” she said. Ric agreed, adding that when she starts oversleeping or dreading work, she knows it’s time to pause. “Sometimes that means taking a vacation. Sometimes it means asking, ‘What’s this all for?’” ## Why boundaries matter in data work Setting boundaries is hard in any field, but especially in data, where work is naturally reactive. There’s always another request, another metric to check, another dashboard to fix. Faith noted that boundaries shift depending on your stage of life or career. “It’s something we have to keep re-learning,” she said. “What works when you’re early in your career might not work later, especially when life outside of work changes too.” Jerrie shared that the hardest part isn’t saying no, it’s giving herself permission to pause. “I’ve always operated on go, go, go,” she said. “Now I block recharge time in advance. I pick an afternoon or a random day off that doesn’t even have to make sense, because if I wait until I feel burnt out, it’s already too late.” Paige talked about learning to accept what comes with setting boundaries. “If I miss a conversation at work because I’m at the vet with my cat, that’s okay,” she said. “That moment mattered more.” Her morning routine has become non-negotiable: 12 minutes of meditation, 30 minutes on the treadmill, and no Slack until both are done. “It makes the chaos of the day manageable,” she said. Ric added that boundaries don’t have to be hard lines. They can also be questions. “I keep a sticky note on my desktop that says, ‘Does this need to be done by me? Does it need to be done now? Can it wait until next week?’” she said. “It helps me remember that not everything needs to happen right away.” ## Redefining what success means in your career As the conversation unfolded, the group got candid about ambition and how their definitions of success have changed over time. Paige shared that early in her career, she worked nonstop to prove herself. “A lot of it was about showing my parents that I was successful,” she said. “But at some point, I realized their opinion didn’t matter as much as my own. I’m proud of where I am and what I do, and maybe that’s enough.” Faith talked about the anxiety that comes with being a household breadwinner. “There’s no safety net,” she said. “I sometimes overwork because I’m afraid of what would happen if I lost my job.” Her way to reset? Spreadsheets. “I look at my finances and remind myself I’m fine. It’s grounding to have something tangible.” Jerrie shared how motherhood and experience shifted her mindset. “I’ve had to accept some seasons of being mediocre,” she said with a laugh. “There’s nothing wrong with just being okay. I still care deeply about my work, but I don’t have to be at 110% all the time.” Ric agreed, emphasizing that growth isn’t mandatory. “There’s so much pressure to keep climbing,” she said. “But if you love what you do and it’s enough for you, that’s valid. You can step back without it meaning less.” ## Making space for joy outside of work The episode closed with a simple truth: rest and creativity are part of the job. For Jerrie, that looks like drawing and dreaming of one day writing a graphic novel. “It’s my way of staying playful,” she said. “I’ll even draw with my kids. It’s a creative recharge.” Faith finds energy in improv and writing. “When I’m not doing improv, I can feel it,” she said. “It makes me see everything differently. It’s like a mental reset.” Paige signs up for one creative class each term at her local community college—acting, stop-motion animation, and next, screenwriting. “I love being a beginner again,” she said. “It gets me out of my head.” Ric’s hobbies are tactile. “I build mechanical keyboards and do woodworking,” she said. “Anything that gets me off the screen. Rest is work too. It’s what helps me show up better on Monday.” ## Episode takeaways Finding balance isn’t a one-time fix. It’s a practice. It’s checking in with yourself when things start to feel off. It’s giving yourself permission to pause. It’s remembering that “enough” can be exactly that. And it’s making space for the version of you that exists outside of your job title. As Ric summed it up, “Work is a marathon, not a sprint. But it’s also not the whole marathon.” --- --- title: "How AI is changing the analytics stack" description: "AI is transforming how enterprises analyze data to reduce silos, boost access, and enable agentic analytics." url: "https://www.getdbt.com/blog/how-ai-is-changing-the-analytics-stack" date: "2025-11-05" authors: ["Daniel Poppy"] categories: ["Insights"] --- # How AI is changing the analytics stack Today’s cloud-native data processing architectures have enabled cost-efficient data warehousing, real-time data processing, and a new level of self-service analytics. But most enterprises find themselves running up against a data ceiling. Driving insights from today’s data stack takes a huge productivity toll on all involved. AI is about to change everything for the analytics stack. **** ## Our fragmented reality For years, extracting insights from enterprise data has meant wrestling with a patchwork of tools. BI dashboards, SQL editors, data warehouses, ETL pipelines, and governance systems. Each has its own interface and learning curve. Most business users lack the technical skills to use these tools independently. That forces a dependency on analysts or engineers. The result is reports that arrive days or weeks later, often with insights that have already gone stale. The numbers paint a stark picture. Enterprise analytics teams [now work across an average of 400 data sources](https://www.cdpinstitute.org/news/average-enterprise-analyzes-400-data-sources-idg-report/). Nearly one in five enterprises juggle more than 1,000. A 2024 industry survey revealed that [more than 70% of data teams rely on 5 to 7 different tools](https://www.datavisualizationsociety.org/soti-report-2024) just to get through daily workflows. About 10% juggle more than 10. The productivity toll is measurable. In the 2025 State of Analytics Engineering Report, 57% of analytics and data professionals [said they spent most of their time maintaining or organizing datasets](https://www.getdbt.com/resources/state-of-analytics-engineering-2025). That’s the same level as the prior year, despite 70% using AI to help write code and documentation. [Data scientists spend about 60% of their time cleaning and organizing data](https://www.dataversity.net/survey-shows-data-scientists-spend-time-cleaning-data/), and another 19% gathering datasets. This leaves roughly 20% for actual analysis and insight generation. [Between 60% and 73% of all enterprise data never gets used for analytics](https://www.forrester.com/blogs/hadoop-is-datas-darling-for-a-reason/). 83% of organizations [suffer from data silos](https://www.kenan-flagler.unc.edu/perspectives/breaking-barriers-how-to-free-your-organization-from-the-silo-mentality/), and 97% believe [those silos hurt performance](https://www.kenan-flagler.unc.edu/perspectives/breaking-barriers-how-to-free-your-organization-from-the-silo-mentality/). Many users don't even know what data exists within their organization, let alone how to access or apply it. This fragmented reality is where most enterprises find themselves today. And this is where AI can help. ## **From dashboards to dialogue** For much of the modern data era, business intelligence has been defined by static dashboards and the technical expertise required to navigate them. Pulling meaningful insights often meant waiting for scarce data analysts to run SQL queries or create tailored reports. This system empowered only those with the tools and training to interpret raw datasets. This bottleneck is now easing. AI-powered conversational analytics lets users ask questions in plain English and iterate in real time while preserving definitions and controls behind the scenes. [The global conversational AI market hit $13.2 billion in 2024](https://www.marketsandmarkets.com/Market-Reports/conversational-ai-market-49043506.html) and is projected to reach $49.9 billion by 2031, growing at nearly 25% annually. Gartner predicts that [by 2025, natural language will be the main way people interact with data systems](https://www.wire19.com/gartner-top-data-and-analytics-predictions-for-2024/). That change alone is expected to drive a 100x surge in data usage across organizations. As access improves, the value of data doesn't just rise—it multiplies. ## **Beyond conversational analytics** Beyond one-off questions, agentic AI coordinates multi-step work—planning, writing code or SQL, running checks, and proposing changes. Research from Capgemini found that [50% of enterprises plan to implement AI agents in 2025](https://www.adamsstreetpartners.com/insights/the-next-frontier-the-rise-of-agentic-ai/), with adoption expected to reach 82% in 2028. User expectations have shifted with tools like [ChatGPT](https://chatgpt.com/), [Claude](https://claude.ai), [Gemini](https://gemini.google.com/app), and [Microsoft Copilot](https://copilot.microsoft.com/). The natural language interface has become not only a standard for AI systems but also a key feature for many traditional applications. As major AI developers roll out new capabilities to enormous user bases, expectations are normalizing around systems that don't just answer—they act within guardrails. **** ## **A glimpse into the future with agentic development** Modern AI IDEs point to where interfaces are heading. Consider what this could look like for data engineering. You've been assigned to build a weekly [ETL](https://www.getdbt.com/blog/extract-transform-load) pipeline to aggregate customer activity, enforce data quality standards, calculate summary metrics, and push the final output into production. In an agentic system, you define your goal in natural language: "Create a weekly customer activity ETL pipeline. Include data quality checks for nulls and duplicates, calculate weekly active users, and push summary tables to the analytics warehouse." From that point, the AI agent gets to work. It scans your project, considering schema definitions, naming conventions, current pipeline structures, and warehouse configuration. It outlines a detailed plan—creating a new model, drafting [Python](https://python.org) scripts for anomaly detection, and preparing orchestration configs aligned with your tech stack. Once you approve, the agent transitions into automated development. It writes SQL aggregation logic, test scripts, [Continuous Integration/Continuous Development (CI/CD)](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud) configurations tailored to your stack, and optional README updates. Everything is formatted to match your project's style guidelines. Next comes execution and validation. The agent runs the full pipeline in a [sandboxed environment](https://docs.getdbt.com/docs/environments-in-dbt), executes the SQL, initiates data quality scripts, and runs your test suite. This real-time feedback loop ensures problems are caught before human review. Finally, when the pipeline passes all checks, the agent handles deployment—opening a [pull request](https://www.getdbt.com/blog/how-to-review-an-analytics-pull-request) or directly pushing updates to your orchestration layer. ## **Why structured context is critical** For agentic systems to operate safely and autonomously in complex environments, you need more than instructions and advanced AI. You need structured data and metadata—the schemas, semantics, relationships, permissions, and lineage that describe how your data works. [Generative AI (GenAI)](https://www.techtarget.com/searchenterpriseai/definition/generative-AI) has earned its spotlight largely for what it can do with unstructured data. But the emergence of [agentic AI](https://www.ibm.com/think/topics/agentic-ai) has pushed enterprise companies to rethink their entire data foundation. In this shift, structured context moves from "nice to have" to "non-negotiable." Enterprise automation depends on structured context. Without it, agents can't function safely or effectively. ## **The two critical integrations** An AI system must integrate with: **Structured data:** The rows in your data warehouse or lakehouse from customer relationship management (CRM), enterprise resource planning (ERP), human capital management (HCM), and other enterprise systems. **Structured metadata:** Data about models, sources, lineage, dependencies, tags, and governance tags—essentially a map of your data ecosystem. With structured data and metadata, AI agents can discover tables, understand relationships, check quality, and plan safe actions. This enables powerful conversational analytics powered by planning, reasoning, and autonomous decision-making. But it also includes governance—policies and permissions embedded in metadata that determine who can access which datasets, which transformations are permitted, and how sensitive fields must be handled. Structured context equips agents with three key capabilities: 1. Memory via metadata, so they know what assets exist and how they relate 2. Boundaries via clear definitions, permissions, and rules so they don't wander outside guardrails 3. The ability to take useful actions with validated tools to read and write safely When you combine these, agents evolve from chatbots into reliable teammates. They can plan, reason, and execute tasks at scale. We're already seeing agents autonomously modify data pipelines, fix errors, manage migrations, and spin up new data products—all driven by structured inputs and aligned with business logic. ## **When AI gets it wrong** Weak governance and poor inputs have predictable consequences. Studies cite that [between 70% and 80% of AI projects don't succeed](https://www.rand.org/pubs/research_reports/RRA2680-1.html). That’s nearly double the failure rate of traditional IT projects. In most cases, it comes down to bad data. Gartner puts the average cost of poor data quality [at around $12.8 million per year per organization](https://www.gartner.com/smarterwithgartner/how-to-improve-your-data-quality). Some companies lose as much as [6% of annual revenue from flawed AI outputs](https://www.fivetran.com/blog/new-ai-survey-poor-data-quality-leads-to-406-million-in-losses). High-profile examples show the impact. A major airline was taken to court after its chatbot [promised a bereavement fare refund](https://www.forbes.com/sites/marisagarcia/2024/02/19/what-air-canada-lost-in-remarkable-lying-ai-chatbot-case/) that didn't exist. The airline had to pay the customer over $600. Lawyers have faced sanctions for submitting briefs citing fictional cases. News sites have published AI travel content directing readers to unsafe destinations. OpenAI's newest models [hallucinate at higher rates than predecessors](https://techcrunch.com/2025/04/18/openais-new-reasoning-ai-models-hallucinate-more/), with error rates hitting 48% and 33%. ## **Regulatory considerations** Regulators are making guardrails explicit. Under [the EU AI Act](https://artificialintelligenceact.eu/), especially Articles 10 and 27, organizations face serious compliance risks. [Article 10](https://artificialintelligenceact.eu/article/10/) mandates that high-risk AI systems use datasets that are complete, accurate, representative, and error-free. Organizations must document everything—data sources, annotation methods, quality checks, and bias mitigation. [Article 27](https://artificialintelligenceact.eu/article/27/) requires Fundamental Rights Impact Assessments that examine fairness, dignity, and non-discrimination. Companies must map data flows, retention policies, oversight mechanisms, and risk mitigation steps, sometimes reporting to regulators. The EU AI Act makes clear that data quality and governance are legal obligations that need to be baked into every layer of the AI pipeline. Real compliance means engineering a governance framework that is automated, auditable, and built to scale. ## **The industry response** The technology industry is actively working to fix what's broken. Vendors are converging on a context-first, governance-forward model. Quality and policy controls are moving closer to where queries run and transformations execute. Data quality and observability are receiving an AI boost. Vendors are adding AI and ML capabilities to detect anomalies or quality issues in real time. New features in cloud data platforms can automatically monitor freshness, null spikes, or other data health metrics and alert teams before issues propagate downstream. **** ## **The changing role of data engineers** This deeper integration of intelligence signals dramatic changes for data engineers. Tasks that once defined the profession—building ingestion pipelines, wrangling schemas, writing ETL logic, monitoring data flows—will increasingly be handled by intelligent systems. Yet this evolution doesn't mean data engineers are becoming obsolete. As routine responsibilities fade, engineers are shifting focus to more strategic concerns. They're designing resilient and adaptive data architectures, validating the integrity and semantics of AI-driven pipelines, maintaining rigorous standards for data quality, and embedding ethical and compliance principles into core infrastructure. Data engineers are evolving from system operators to system stewards. The work now demands fluency in AI-native tools, semantic data modeling, governance strategy, and the supervision of autonomous agents. The shift isn't about doing less—it's about doing more of what truly matters. For this new era to become reality, there must be a solid foundation where LLMs effectively interact with structured data and metadata. Otherwise, ‌capabilities remain shallow, not grounded in the relevant data, processes, and workflows of the enterprise. **** --- --- title: "What are the benefits of data observability?" description: "Learn how data observability reduces costs, boosts trust, and enables faster decisions with resilient data systems." url: "https://www.getdbt.com/blog/benefits-of-data-observability" date: "2025-11-05" authors: ["Joey Gault"] categories: ["Pulse"] --- # What are the benefits of data observability? One of the primary benefits of data observability is the dramatic improvement in system reliability and organizational trust in data. When teams implement comprehensive observability frameworks, they gain the ability to detect and address issues before they impact business operations. This proactive approach transforms data teams from reactive firefighters into strategic partners who can guarantee data reliability. The [experience at SurveyMonkey illustrates this benefit clearly](https://www.getdbt.com/blog/surveymonkey-monte-carlo-dbt). When their data engineering team conducted an internal survey to understand data challenges across the organization, they discovered that 53% of respondents cited data quality issues as a primary concern, while 50% identified data processing issues as equally significant challenges. By implementing a systematic observability approach that combined reactive monitoring through Monte Carlo with proactive data management through dbt, they created a comprehensive framework that addressed both concerns simultaneously. This integration enabled SurveyMonkey to convert Monte Carlo anomalies into dbt test cases, creating a self-improving system that became more robust with experience. When the monitoring system detected an anomaly indicating a serious data quality issue, the team would create a corresponding dbt test that would prevent the pipeline from proceeding if the same condition occurred again. This approach shifts responsibility for data quality upstream, enabling business users to address issues at their source rather than waiting for data engineering intervention. The result was a marked increase in data quality over time. When SurveyMonkey integrated a third-party marketing analytics platform, Monte Carlo initially detected a spike in data anomalies as the team learned to work with the new data sources. However, the systematic conversion of anomalies into dbt test cases led to a corresponding decrease in anomalies over time, demonstrating how effective observability creates resilient, self-healing systems. ## Significant cost optimization and performance improvements [Data observability](https://www.getdbt.com/blog/data-observability) delivers substantial cost benefits through performance monitoring and optimization capabilities. By providing visibility into query performance, resource utilization, and pipeline efficiency, observability tools enable teams to identify and eliminate waste while optimizing for performance. SurveyMonkey's implementation demonstrates the potential scale of these benefits. The most striking outcome was a 73% reduction in Snowflake credit usage across nearly 10,000 credit jobs. This dramatic improvement came primarily from performance monitoring that identified inefficient queries, unused models, and optimization opportunities. Performance monitoring revealed long-running jobs performing unnecessary cross-joins, enabling the team to simplify SQL statements and merge redundant queries. The systematic identification and removal of unused models and tables further contributed to cost reduction. In one particularly impressive case, obsoleting unused models and revamping inefficient queries resulted in a 94% reduction in pipeline runtime and a 97% reduction in Snowflake credit usage. These improvements weren't one-time gains; the observability framework enabled SurveyMonkey to scale while keeping costs stable. Despite adding many more dbt models, including bringing a marketing analytics platform in-house, job execution times remained stable while cost per credit steadily decreased. This demonstrates how effective observability can support growth without proportional increases in operational overhead. Teams can make informed decisions about optimization priorities rather than guessing which models might benefit from incremental materialization or increased warehouse sizes. Concrete performance data guides efforts and measures the impact of changes, creating a virtuous cycle of continuous improvement. ## Proactive issue detection and faster resolution Traditional monitoring approaches often rely on business users or downstream consumers to report data issues, creating delays between when problems occur and when they're addressed. Data observability fundamentally changes this dynamic by enabling proactive detection and automated alerting for a wide range of potential issues. Effective observability systems monitor standard metrics including run statuses, record counts, schema versions, and latency, establishing SLA thresholds and statistical alerts that can detect anomalies before they impact business operations. This capability extends beyond simple threshold monitoring to include sophisticated anomaly detection that can identify subtle patterns indicating emerging problems. The alerting capabilities enabled by observability create accountability and urgency around data quality issues. When important anomalies are detected and converted into test cases that prevent pipeline execution until upstream issues are resolved, it gets everyone's attention and shifts responsibility appropriately. This prevents the propagation of known data quality problems while ensuring that the right people are notified to address issues quickly. Domain-specific alerting, as implemented at SurveyMonkey, ensures that model owners receive notifications about their specific models rather than broadcasting alerts to entire teams. Every model in their dbt deployment includes domain tags like "growth," "finance," or "catalog," which correspond to Slack user groups containing relevant stakeholders. This targeted approach includes sufficient context for debugging, including error messages, model names, and timestamps, enabling recipients to quickly understand and address issues. ## Improved collaboration and data democratization Data observability breaks down silos between data engineering teams and business users by providing shared visibility into data systems and creating self-service capabilities for data discovery and troubleshooting. This democratization of data insights enables more effective collaboration and reduces the burden on data engineering teams. When business users have access to observability tools and documentation, they can independently investigate data issues, understand data lineage, and make informed decisions about data usage. This self-service capability reduces the number of support requests to data engineering teams while empowering business users to work more effectively with data. The integration of observability tools with documentation systems, such as combining dbt documentation with Monte Carlo Asset discovery, creates a single pane of glass for data assets. Users can search and filter available tables based on comprehensive documentation, understanding not just what data is available but how it was created, when it was last updated, and what quality checks are in place. This transparency builds trust and confidence in data systems while enabling more sophisticated use cases. When business users understand data lineage and quality measures, they can make more informed decisions about which datasets to use for different purposes and how to interpret results appropriately. ## Scalable governance and compliance As organizations grow and data systems become more complex, maintaining governance and compliance becomes increasingly challenging. Data observability provides the foundation for scalable governance by creating audit trails, monitoring data access patterns, and ensuring that quality standards are consistently applied across all data assets. Observability systems automatically track data lineage, transformation logic, and quality metrics, creating comprehensive audit trails that support compliance requirements. This automated documentation reduces the manual effort required to maintain governance standards while providing more complete and accurate records than manual processes. The systematic application of quality checks and monitoring across all data assets ensures that governance standards are consistently enforced rather than applied ad hoc. When every dbt model is subject to mandatory testing and monitoring, it creates a consistent baseline for data quality expectations that scales with the organization. Regular performance reviews and automated monitoring ensure that models don't degrade over time, maintaining quality standards even as teams grow and change. This systematic approach to governance scales more effectively than manual processes while providing better outcomes. ## Enhanced decision-making capabilities Perhaps the most important benefit of data observability is its impact on organizational decision-making capabilities. When stakeholders have confidence in data quality and understand how data flows through systems, they're more likely to rely on data for important decisions. This increased trust in data systems enables more sophisticated analytics use cases and supports data-driven decision making at scale. Observability provides the context necessary for interpreting data correctly. When business users understand data freshness, quality measures, and transformation logic, they can make more informed decisions about how to use data and how to interpret results. This contextual understanding prevents misuse of data while enabling more sophisticated analysis. The real-time nature of modern observability systems enables faster decision-making by ensuring that stakeholders have access to current, high-quality data when they need it. Rather than waiting for data engineering teams to investigate and resolve issues, business users can quickly understand data status and make decisions accordingly. ## Building organizational resilience Data observability contributes to organizational resilience by creating systems that can adapt to change and recover quickly from disruptions. When teams have comprehensive visibility into data systems and automated processes for detecting and addressing issues, they can respond more effectively to unexpected challenges. The self-improving nature of well-designed observability systems means that organizations become more resilient over time. Each incident that's detected and resolved strengthens the system's ability to prevent similar issues in the future. This creates a virtuous cycle where observability investments compound over time, delivering increasing returns. The combination of proactive testing and reactive monitoring creates more resilient systems than either approach alone. dbt tests catch many issues before they reach production, while observability tools detect the problems that slip through. This layered approach provides multiple opportunities to catch and address issues before they impact business operations. As data becomes increasingly central to business operations, organizations that master data observability will have a significant competitive advantage in their ability to make reliable, data-driven decisions at scale. The investment in observability tools, processes, and skills pays dividends in terms of cost savings, improved reliability, and increased trust in data systems. The future of data observability will likely include enhanced artificial intelligence and machine learning capabilities for automated anomaly detection and intelligent alerting that reduces false positives. However, the fundamental principles of comprehensive monitoring, proactive testing, and effective alerting will remain constant. The most successful data teams will be those that treat observability as a core competency rather than an afterthought, building the foundation for reliable, scalable data operations that support business growth and innovation. ## Data observability FAQs **Why is data observability important?** Data observability is crucial because it dramatically improves system reliability and organizational trust in data. It enables teams to detect and address issues before they impact business operations, transforming data teams from reactive firefighters into strategic partners who can guarantee data reliability. By providing comprehensive monitoring, proactive testing, and effective alerting, data observability creates resilient, self-healing systems that become more robust over time and support reliable, data-driven decision making at scale. **How does observability reduce downtime and improve mean time to resolution (MTTR)?** Observability reduces downtime by enabling proactive issue detection through automated monitoring of metrics like run statuses, record counts, schema versions, and latency. Instead of waiting for business users to report problems, sophisticated anomaly detection identifies subtle patterns indicating emerging issues before they impact operations. When problems are detected, domain-specific alerting ensures the right people are immediately notified with sufficient context for debugging, including error messages, model names, and timestamps, enabling rapid resolution and preventing the propagation of data quality issues. **What measurable business value such as ROI, increased uptime, or better customer conversion can organizations gain from implementing observability?** Organizations can achieve substantial measurable benefits from data observability implementation. Cost optimization alone can deliver dramatic results, with examples showing up to 73% reduction in cloud computing credits across thousands of jobs, 94% reduction in pipeline runtime, and 97% reduction in resource usage through performance monitoring and optimization. Beyond cost savings, observability enables organizations to scale operations while keeping costs stable, maintain consistent quality standards across growing data assets, and build organizational resilience through self-improving systems that become more reliable over time. --- --- title: "What is Snowflake Intelligence anyway?" description: "Conversational AI is only as good as your data. Here’s why dbt is foundational to Snowflake Intelligence." url: "https://www.getdbt.com/blog/what-is-snowflake-intelligence-anyway" date: "2025-11-04" authors: ["Luis Leon"] categories: ["Partnerships"] --- # What is Snowflake Intelligence anyway? The promise of being able to 'talk with your data' is currently in its second decade of being delivered. But we're here now. Instead of it feeling like you were just asking questions using verbal sql, the advent of new tools, LLM and open semantics means you can chat with your data better than you can with 'that guy' at the meat raffle _[Editor’s note: This is a very niche reference from the U.S. Midwest contributors to this blog post]_. Snowflake Intelligence is a new general-purpose agentic platform from Snowflake. It lets anyone in your organization build and interact with AI agents to ask questions of your data and get instant, accurate answers without writing SQL, building dashboards, or sending tickets to your data teams. It's designed to close the gap between having data and‌ using it. What sets Snowflake Intelligence apart is that it is built directly into the Snowflake platform on top of proven Snowflake technologies like Cortex Analyst, Cortex Search & Document AI. It understands your data, semantic definitions, and through Cortex Knowledge Extensions can even access and analyze information from public sources like The Associated Press, Stack Overflow, or even information directly from Snowflake’s own documentation. Leveraging existing Snowflake AI tools also means Snowflake Intelligence can interpret both structured (curated dbt models) and unstructured data (PDFs, emails, Slack conversations). This provides a more complete picture of your business and more accurate answers to your business questions. Most importantly, and as you would expect from a Snowflake solution, it has all the governance, access, and security built in to work at any scale. Equally exciting, the release of Snowflake Intelligence and the continued refinement and proliferation of AI solutions creates an even larger importance on the role of the analytics engineer and the need for dbt. Let’s get into why. ## What role does dbt play in Snowflake Intelligence? Solid foundations of data quality and governance are critical for success in any AI project. dbt is the standard for data transformation in the cloud data warehouse; it provides reliable, consistent, and well-controlled data that is needed for any AI project to work well at scale. dbt is also a main source for the types of metadata needed to provide the appropriate context for AI to answer questions as accurately and credibly as possible; think data quality, data freshness, and crucially, semantic definitions. Semantics are becoming increasingly more critical for successful AI. The dbt Semantic Layer makes it possible to define business metrics as part of your dbt pipelines. This allows you to create quality and contextual data. Defining metrics in dbt not only ensures your metrics are consistent and well-governed but are also connected to the contextual metadata of your dbt project. Snowflake Intelligence uses semantic definitions to generate more accurate responses, while the metadata generated by dbt provides context to allow both the AI and the person asking to understand the reasoning of the response. Due to this importance of semantic definitions, Snowflake and dbt Labs are excited to be two of the founding members of the [Open Semantic Interchange](https://www.snowflake.com/en/blog/open-semantic-interchange-ai-standard/). While basic functionality exists today—via the ability to define Snowflake Semantic Views with [dbt](https://hub.getdbt.com/Snowflake-Labs/dbt_semantic_view/latest/)—we are excited to partner more closely to promote the open interoperability of semantic definitions regardless of where these measures are defined or consumed. ## The dbt MCP Server and Snowflake Intelligence Another exciting integration point between dbt and Snowflake Intelligence is the [dbt MCP Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server), which opens new possibilities for extending Snowflake Intelligence capabilities. The dbt MCP Server provides a standardized interface that enables AI agents to interact directly with dbt projects, exposing development tools, discovery capabilities, and semantic layer access through the MCP protocol. Users can ask questions like "Which dbt models reference the customers table?" or “What’s the overall health of my dbt projects and what opportunities are there for improvement?” and get answers directly. They can even execute dbt specific operations like compiling models and running tests or executing pipelines. What’s even more exciting is pairing this with dbt’s new [Fusion engine](https://www.getdbt.com/product/fusion), and its ability to parse and validate SQL code. With this workflow, Snowflake Intelligence can use the dbt MCP Server to compile dbt projects, the dbt Fusion engine to catch any errors in the code, and Snowflake Cortex to fix these errors—all before any code is ever executed in Snowflake. While native integration for Snowflake Intelligence and external MCP Servers like dbt’s is on the roadmap, teams can build working prototypes today by creating a Streamlit application within Snowflake that combines the Cortex Agents API (which powers Snowflake Intelligence) with the remote dbt MCP Server. My colleagues at dbt have published [this example repo](https://github.com/dbt-labs/streamlit_mcp_cortex), which is designed to let organizations experiment with the integration pattern, validate use cases, and prove the initial value of adding the dbt MCP Server to Snowflake Intelligence. Fundamentally, a conversational AI is only as good as the data it queries. If your underlying datasets are inconsistent, poorly documented, or scattered across dozens of schemas with unclear ownership, even the smartest AI will give you unreliable answers. This is where [dbt's data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) becomes critical. ## What benefits will you see with Snowflake Intelligence? With so many new AI tools and agents available, who should use Snowflake Intelligence and what use cases does it help us solve? Snowflake Intelligence is flexible and easy to use. It can help anyone who uses Snowflake, from business leaders and executives who use data to the engineers and analysts who create it. **Business leaders and executives**: Business leaders can ask complex questions in plain English—for example, "Show me win rates by region and deal size for Q3." They can get answers, including data, charts, and most importantly, the reasoning behind the answer, all in seconds. Now let's consider a more complex, real-world example showcasing the power of combining Document AI, Cortex Search & Cortex Analyst. With Snowflake Intelligence, you can analyze an uploaded PDF containing the text of a doctor’s clinical note, parse the text for diagnostic, and treatment details from the patient's visit. Next, Cortex Analyst can query relevant clinical information from the patient's record, for example, the medications prescribed and doses. Combining the analysis of structured and unstructured data means answering complex, valuable, real-world questions. **Data analysts and engineers**: Freeing data analysts and engineers from pulling data and answering follow-on questions related to routine requests allows them to shift from being ticket-takers to strategic partners. They can focus on high-value tasks like onboarding new data sets and use cases or tackling the complex analytical challenges that‌ move the business forward. Snowflake Intelligence can find all the data and metadata in Snowflake, including all dbt metadata. This makes it a great way to find data quickly and easily. Asking questions like ‘what is the best source for certified up-to date customer data?’ leads analysts to the right data products quickly. This data discovery flow increases developer productivity and ensures the data products they build can be deployed faster and with higher accuracy. Finally, by integrating Snowflake Intelligence with [dbt’s MCP Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server), analysts and engineers can allow Snowflake Intelligence to take dbt specific actions like assessing dbt project health, pulling dbt specific metadata and finding and remediating issues in pipelines. Even more powerfully, a workflow with Snowflake Intelligence, the dbt MCP Server, and the dbt Fusion Engine can be used to fix pipeline issues before they run against your Snowflake warehouse. ## What benefit does dbt provide beyond using Snowflake Intelligence? Snowflake Intelligence is a great solution for interacting with data inside of Snowflake. But before anyone can meaningfully interact with this data, it must be transformed, validated, and governed. As the standard for data transformation in the warehouse, dbt is the natural tool for ensuring data quality, usefulness, and governance for use by Snowflake Intelligence. Additionally, the same context provided to the AI is made available to the people consuming the data. This means that people who use these answers can review and check all the same information. This makes Snowflake Intelligence's response not only accurate, but also trustworthy and actionable. Semantics are becoming more important for AI accuracy. dbt plays three different roles in providing the semantic definitions that make Snowflake Intelligence more accurate: 1. Teams can define Snowflake Semantic Views directly in dbt, ensuring metric definitions are version-controlled and tested alongside transformations. 2. The dbt Semantic Layer integrates with Snowflake Intelligence through the dbt MCP Server, allowing conversational queries to leverage the same certified business metrics that power BI dashboards and reports. 3. Both dbt Labs and Snowflake are founding members of the [Open Semantic Interchange (OSI)](https://www.getdbt.com/blog/dbt-labs-affirms-commitment-to-open-semantic-interchange-by-open-sourcing-metricflow), committing to enhance semantic interoperability industry-wide so metrics remain consistent regardless of where they're defined or consumed. Integrating Snowflake Intelligence with the dbt MCP Server represents an exciting area of innovation beyond pure analytics. Through the MCP Server, Snowflake Intelligence can answer dbt-specific questions and execute dbt operations like test, run, and compile. When paired with dbt's Fusion engine—which can parse and validate SQL code—this creates opportunities for Snowflake Intelligence Agents to compile dbt code and fix errors before any code executes in your warehouse. Beyond these specific capabilities, dbt provides the contextual metadata and workflow—data quality metrics, freshness indicators, lineage information, and semantic definitions—that help both the AI and the people asking questions understand the reasoning behind answers and trust the responses. This combination of reliable data foundations, semantic precision, and operational intelligence makes dbt essential for organizations that want Snowflake Intelligence to deliver accurate, trustworthy insights at enterprise scale. ## What do you need to make it all work? Getting started with dbt and Snowflake Intelligence is straightforward, especially if you're already using the dbt platform. All AI projects are data projects. When working with Snowflake Intelligence, you should follow all the best practices for building and scaling a successful data project. ### 1. Ensure your dbt project is well-documented Snowflake Intelligence relies on documentation to understand your data semantics. Take the time to: - Add meaningful descriptions to your models explaining what business entities they represent - Document key columns, especially dimension keys, dates, and metrics - Define metrics in your dbt Semantic Layer for consistency - Use clear naming conventions that make intent obvious You don't need perfect documentation to start, but the better documented your models are, the more accurate Snowflake Intelligence's answers will be. If you’re using the dbt platform, [dbt Copilot](https://www.getdbt.com/blog/introducing-dbt-copilot) makes documenting your models a breeze. ### 2. Implement data quality tests Data quality issues that slip through to Snowflake Intelligence will erode trust in the system: - Add schema tests for uniqueness, not null, and relationships - Write custom data quality tests for business-specific rules - Set up alerts for test failures so issues get addressed quickly - Use dbt's store_failures configuration to investigate quality issues Ensuring that primary key fields are tested for uniqueness and non-nullability makes a tremendous difference in quality. dbt Co-Pilot writes test for you too. ### 3. Structure your project with dbt Mesh (for larger organizations) If you have multiple teams building data products, [dbt Mesh helps maintain quality at scale](https://www.getdbt.com/product/dbt-mesh): - Break your monolithic dbt project into smaller, domain-specific projects - Define clear ownership and data contracts between projects - Use groups and access controls to manage dependencies - Publish stable models for consumption by other teams This modular structure makes it easier to evolve your data products without breaking downstream dependencies. This is critical when those dependencies include conversational AI queries. ### 4. Configure Snowflake Intelligence While Snowflake Intelligence does not require specific set-up, it relies upon other Snowflake services that must be configured before use: - Set-up [Cortex Search](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-search/overview-tutorials) - Set-up [Cortex Analyst](https://quickstarts.snowflake.com/guide/getting_started_with_cortex_analyst/index.html#0) - Set-up [Cortex Knowledge Extension](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-knowledge-extensions/cke-overview) (optional) - Set-up [Cortex Agents](https://quickstarts.snowflake.com/guide/getting_started_with_cortex_agents/index.html#0) (optional) - Access Snowflake Intelligence (ai.snowflake.com) This approach makes getting started quick and easy, while still providing flexibility for highly tailored agents and workflows. ### 5. Monitor and iterate Like any AI system, Snowflake Intelligence improves with feedback: - Monitor what questions users are asking and what answers they're getting - Identify gaps in data coverage or documentation - Iterate on your dbt models based on usage patterns - Gather user feedback and adjust your data products accordingly The beauty of this architecture is that improvements to your dbt project immediately make Snowflake Intelligence more valuable. Better documentation leads to more accurate answers. New models expand what questions can be answered. Better data quality increases trust. ## Looking ahead with Snowflake Intelligence and dbt Snowflake Intelligence represents a fundamental shift in how organizations interact with data. But this shift is only possible because of the reliable, well-governed data foundation that tools like dbt provide. The combination of dbt's transformation and governance capabilities with Snowflake Intelligence's conversational interface creates something powerful: enterprise data that's both trustworthy and accessible to everyone. Your carefully curated dbt models, your tested transformations, your documented metrics—all of these become instantly queryable by anyone in your organization, in plain English, with enterprise-grade security. We've just seen how this partnership unlocks new value from data investments you've already made. The dbt project you've been building becomes the knowledge base for conversational AI. The governance you've implemented ensures security at scale. The semantic layer you've defined provides metric consistency. And the data quality tests you've written build trust in every answer. The future of analytics isn't just faster queries or prettier dashboards—it's data that anyone can talk to, learn from, and act on. With dbt and Snowflake Intelligence working together, that future is here. Ready to get started? Check out the [Snowflake Intelligence documentation](https://docs.snowflake.com/en/user-guide/snowflake-cortex/snowflake-intelligence) and work with your Snowflake account team to request preview access. **If you're not yet using dbt, [install the dbt VS code extension](https://docs.getdbt.com/docs/install-dbt-extension) and start building the reliable data foundation that makes conversational AI possible.** Your data has answers. Now everyone in your organization can ask the questions. --- --- title: "Speed, simplicity, cost savings: Experience the dbt Fusion engine" description: "How dbt Fusion can save your organization time—and shave up to 29% off your data spend." url: "https://www.getdbt.com/blog/dbt-fusion-experience" date: "2025-10-30" authors: ["Kathryn Chubb"] categories: ["Product"] --- # Speed, simplicity, cost savings: Experience the dbt Fusion engine Data teams today face three daunting challenges: improving data quality, improving velocity, and managing costs. As an industry, we’ve made significant progress in building pipelines and accessing compute resources. However, organizations struggle with data quality issues, ambiguous ownership, and stakeholder literacy. Overcoming these challenges requires more than incremental improvements to existing tools. The solution lies in rethinking the foundation of how data transformation works. [The dbt Fusion engine](https://www.getdbt.com/product/fusion), a complete rewrite of the dbt engine built in [Rust](https://rust-lang.org/) with SQL comprehension capabilities, represents a fundamental shift in analytics engineering. By understanding SQL at a deeper level and introducing state management, Fusion delivers dramatic performance improvements, substantial cost savings, and a developer experience that fundamentally changes how teams build and maintain data pipelines. **** ## Why we rewrote dbt from scratch While dbt's [authoring layer](https://docs.getdbt.com/docs/build/models) has evolved to provide developers with increasingly sophisticated functionality, the original dbt Core engine has remained built on the same technology and design principles from 2016. This created two fundamental problems that couldn't be solved iteratively. First, dbt Core's [Python](https://www.python.org/)-based architecture became a performance bottleneck. For larger projects, parse times and compilation could become unworkable. Even smaller projects needed step-change improvements to power truly excellent developer experiences. Second, dbt Core renders SQL but doesn't comprehend it. The engine treats SQL as text to be templated and passed to the warehouse, but lacks understanding of SQL semantics. This meant any functionality requiring deep knowledge of SQL code structure—like impact analysis or intelligent optimization—was impossible to build. The solution required rebuilding the engine entirely. The [dbt Fusion](https://www.getdbt.com/product/fusion) engine is fully rewritten in [Rust](https://rust-lang.org/) and incorporates SQL compiler technology from SDF (recently acquired by dbt). Not a single line of code is shared between dbt Core and Fusion, aside from adapter macros. The result is an engine built for speed, one that truly understands code, and one that powers next-generation developer experiences. ## Performance that transforms workflows The speed improvements in Fusion are immediately noticeable. Parse times are up to 30 times faster than dbt Core, and full-project compilation is twice as quick. [The Visual Studio Code extension](https://docs.getdbt.com/docs/install-dbt-extension) can recompile entire projects in the background with near-instant turnaround for individual files. This performance leap changes how developers work. Instead of running a model to check for errors, developers get real-time feedback as they type. There's no more cycle of writing code, running dbt build, discovering a typo, fixing it, and running again. The development loop tightens dramatically. The Rust architecture enables this performance without requiring any cloud connectivity or compute charges during development. Everything happens locally, instantly, without a round-trip to the warehouse. ## SQL comprehension unlocks developer experience innovations Fusion's SQL compiler represents a fundamental advancement. Unlike dbt Core, which treats SQL as templated text, Fusion parses and understands SQL syntax and semantics across data platforms. It knows about columns, functions, type signatures, and how transformations propagate through data lineage. The VS Code extension powered by Fusion introduces capabilities that fundamentally change the development experience. At its most basic level, Fusion catches simple errors without round-tripping. For example, without Fusion, if you use a SQL window function in an incorrect context (i.e., outside of SELECT, QUALIFY, or ORDER BY clauses), you’d have to check it into source control, run your [Continuous Integration (CI) pipeline](https://docs.getdbt.com/docs/deploy/continuous-integration), wait for the failure, fix it, and then check in a new change. By contrast, the VS Code extension catches both simple and complex SQL errors immediately, so you can fix them before ever typing git add. That shaves significant time off of your development cycles. ![Fusion demo 1](https://cdn.sanity.io/images/wl0ndo6t/main/0655185cd44db2e2acaf12570be8d0bce4208ff9-2630x1652.png) But that’s not all. Common Table Expression (CTE) Preview lets developers preview the output of CTEs without commenting out downstream code. Fusion understands the CTE lineage and runs only those models and downstream models that need to be run. ![Fusion demo 2](https://cdn.sanity.io/images/wl0ndo6t/main/f9ab164042eb9b551f26ab5634f6afc7ee4e9603-2630x1650.png) The Compare Changes feature, currently in alpha testing, performs data diffs during development. It computes the exact impact of code changes on production data before deployment. Compare Changes shows removed records, modified values at the column level, and changes to column metadata. This enables developers to detect changes that might compile and seem harmless but that could have a significant impact on the data. For example, removing the LEFT keyword from the JOIN below might seem innocuous. However, Compare Changes can reveal that it eliminates 99.7% of rows. ![Fusion demo 3](https://cdn.sanity.io/images/wl0ndo6t/main/794c1224769fa357994759e316b02880e645b548-2628x1644.png) Local compilation in Fusion does more than just look at one file. Impact analysis runs continuously across dbt projects as devs make changes. Fusion highlights all downstream models that break when upstream changes occur. This keeps the mental load manageable—developers don't need to remember every dependency chain in their heads. This also means you can leverage Fusion to perform refactoring. For example, if you rename a column in one model, Fusion will rename it automatically in all downstream models. Cross-platform compatibility checking leverages SQL comprehension in a unique way. Developers can switch the target warehouse in their project configuration, and Fusion immediately identifies models that use platform-specific functions. For organizations maintaining vendor flexibility or planning migrations, this makes multi-warehouse support practical. ## State-aware orchestration drives 29% cost savings One of Fusion's most significant innovations is state-aware orchestration. The engine understands the current state of data sources and can optimize execution based on what's actually changed. When sources have no new data, Fusion skips building downstream models that would simply recompute the same results. Take the example of a data model, represented here by a [directed acyclic graph (DAG)](https://www.getdbt.com/blog/dag-use-cases-and-best-practices), with seven tables. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/065098da757ed5f7bd56e15c7496e4c38265044c-2664x1038.png) src_orders has no new data since the last run. It’s stale. src_customers has new data - it’s fresh. The models that depend on src_customers - e.g., stg_customers - are also out of date, since they’re dependent on src_orders, and need to be run. However, all of the models depending on src_orders don’t need to be recalculated. They would just be computing the same result unnecessarily. This is where state-aware orchestration comes in. Fusion can detect this state and reuse the models for everything dependent on src_orders instead of re-running those models. That’s money in your pocket. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/edd78131544e3cafc1daa16c3f5cff9c9570be56-2624x1048.png) This aspect of Fusion simplifies scheduling. Now, you don’t have to worry about timing your data pipeline runs around data arrival times. You can run your pipelines on a set schedule, confident that Fusion will only process what needs to be changed. Early testing shows that this basic capability alone delivers an average 10% cost reduction. But Fusion goes further with advanced configurations that enable fine-grained control. In practice, this delivers substantial cost savings. Simply enabling state-aware orchestration provides approximately 10 percent cost reduction on average. Adding advanced configurations like updates_on and build_after adds another 15 percent. These configurations let teams specify precisely when models should update—whether when all upstream dependencies have fresh data, any upstream has fresh data, or specific sources reach defined freshness thresholds. Test optimization provides additional savings. Fusion can aggregate multiple tests per column into a single query, reducing testing costs significantly. It also performs intelligent column-level lineage analysis to skip redundant tests. If a column is simply copied from one model to another without transformation, Fusion recognizes that testing it in both locations is unnecessary and skips the downstream test. Combined, these optimizations can deliver up to a **29% reduction in deployment costs** - 10% from state-aware orchestration, 15% from advanced configurations, and 4% from testing savings. ## Rich context for AI-ready data Fusion's SQL comprehension and metadata capabilities position it as an ideal foundation for AI-assisted development. Traditional approaches to AI coding assistance lack context about SQL semantics, leading to lower-quality outputs and requiring expensive query execution to validate changes. Fusion provides AI agents with rich metadata about column lineage, type signatures, and impact analysis. An AI renaming a column can immediately understand which downstream assets are affected without executing queries. This reduces both the cost of AI-assisted development and increases the fidelity of generated code. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/355a17b0f8410bdcea866cdc1af929a00a0af407-2772x1106.png) The speed and rigor of Fusion help produce higher-quality AI-generated code. Developers can vet, validate, and comprehend AI suggestions more efficiently when the development environment provides instant feedback about correctness and impact. ## Open data infrastructure and vendor flexibility By abstracting SQL from specific dialects and understanding it semantically, Fusion enables truly platform-independent data workloads. Combined with open table formats like [Apache Iceberg](https://iceberg.apache.org/), this creates an ideal architecture for organizations that want to avoid vendor lock-in while maintaining sophisticated data transformation capabilities. Fusion works identically whether deployed locally on the CLI or used [on the dbt platform](https://www.getdbt.com/product/dbt). Core users can adopt Fusion without any contracts or cloud requirements. The engine runs anywhere, on any supported platform, maintaining the open-source principles that have defined dbt since its inception. ## Real-world impact at DPG Media The benefits of Fusion extend beyond technical specifications. DPG Media, a major media company in Belgium and the Netherlands with more than 80 print publications, has been a dbt customer for years. DPG analytics engineer Sonja Strempel says dbt reduced data quality issues, consolidated repetitive data preparation tasks, and enabled users to self-service ad hoc questions. Strempel says the company has been using Fusion in production for several months. Their analytics teams work across multiple dbt projects, each with its own repository, all centrally managed through [Terraform](https://developer.hashicorp.com/terraform). The most immediate impact was the elimination of the constant context switching between VS Code and dbt. Previously, developers would write code locally, then switch to dbt to run and validate changes. With Fusion's VS Code extension, everything happens in a single environment. Live error detection proved particularly valuable. Small mistakes like typos or incorrect column references get caught immediately, rather than after running dbt build and waiting for results. This saved significant time across both small tasks and large refactoring projects. For schema redesigns and table refactoring, the combination of impact analysis and column lineage visualization made complex changes much safer. Developers could see exactly which downstream models would be affected by changes and ensure all references were updated correctly before creating pull requests. [The column-level lineage feature](https://docs.getdbt.com/docs/explore/column-level-lineage) transformed how teams onboard new members and collaborate on existing code. When explaining models to colleagues, developers can interactively show where columns originate and how they transform through the DAG. The ability to command-click through to source definitions makes exploring unfamiliar codebases intuitive. From a business perspective, these developer experience improvements translate directly to faster time-to-insights. When stakeholders request new metrics or analysis, analytics teams can deliver more efficiently. The reduced error rate and better impact analysis also increase trust in data quality. ## Getting started with Fusion The dbt Fusion engine is currently in preview, with expanding availability across hosted dbt projects. dbt Core users can install Fusion locally and begin using it immediately. [The VS Code extension](https://docs.getdbt.com/docs/install-dbt-extension) is available in the VS Code extension marketplace and automatically installs the Fusion-powered CLI. For organizations with existing dbt projects, migration tooling helps address the changes required. The dbt-autofix helper automatically handles many common migration tasks. While any major version upgrade requires some code changes, the goal is maximal compatibility with existing dbt code. Fusion represents the future of dbt—faster, smarter, and built for the AI-driven analytics landscape ahead. Early adopters are already seeing significant benefits in developer productivity, cost optimization, and data quality. The technology that seemed impossible with dbt Core is now becoming a reality, one deployment at a time. --- --- title: "Key components of data governance" description: "A practical breakdown of the key components of data governance and how they support trust, scale, and AI‑ready data." url: "https://www.getdbt.com/blog/key-components-of-data-governance" date: "2025-10-29" authors: ["Joey Gault"] categories: ["Pulse"] --- # Key components of data governance Modern data governance isn’t just about compliance — it’s the backbone of any scalable, trustworthy analytics or AI initiative. But what does effective governance actually require? It goes far beyond assigning ownership or enforcing policies. A mature governance program rests on a clear framework: a set of structural pillars, technical capabilities, and collaborative processes that together ensure data is accurate, secure, discoverable, and ready for use. In this article, we break down the key components that make governance actionable — and sustainable — at scale. ## Essential organizational roles Successful data governance requires clearly defined roles and responsibilities that span both technical and business functions. These roles work collaboratively to ensure that governance becomes embedded in how organizations naturally work with data rather than existing as a separate, parallel process. At the executive level, organizations need someone who serves as the strategic champion for data governance initiatives. Whether this is a dedicated Chief Data Officer or another senior leader, this role is responsible for securing resources, removing organizational barriers, and ensuring that data governance aligns with business strategy. This executive sponsor translates technical governance concepts into business value and establishes governance as an organization-wide priority. Data stewards represent the operational backbone of any governance program, serving as the front line for governance initiatives. They are collectively responsible for defining and documenting the organization's data assets, ensuring data quality, and promoting effective data sharing across teams. Data stewards act as liaisons between different teams, helping to bridge the gap between technical and business stakeholders when data problems need resolution. Data owners typically represent the individuals or teams closest to where data is created and initially managed. They have primary accountability for specific datasets and are responsible for making decisions about data access, usage policies, and quality standards for their domain. Their domain expertise is essential for making informed decisions about data governance policies and procedures. Technical governance roles, including data engineers and architects, are responsible for implementing the systems that enable governance. They build and maintain the infrastructure that supports governance, including data pipelines, quality monitoring systems, and access controls. These roles implement automated data quality checks, establish monitoring and alerting systems, and create the technical infrastructure that enables data lineage tracking and metadata management. **** ## Modern governance strategies Traditional enterprise data governance has been largely static and top-down, with central authorities laying out standards and policies for all teams. This manual approach worked for a while but doesn't scale in the age of AI and rapidly changing business requirements. Modern data governance requires strategies that are dynamic, continuous, automated, and responsive to fast-changing regulatory environments. **Enabling High-Quality Dataset Creation** across all teams represents a fundamental shift from centralized control to distributed responsibility. Rather than having a single team responsible for all data quality, modern governance empowers every team to create and publish high-quality datasets using standardized tools and processes. This approach requires adopting a single data control plane that enables consistent data modeling and transformation across the enterprise. Tools like dbt support built-in testing frameworks, enabling data engineers to build out test suites that verify all changes before release, ensuring that data transformations generate correct outputs before data is made available to downstream consumers. **Emphasizing Data Collaboration** addresses the challenges created by do-it-yourself approaches to data quality that make it difficult for teams to work together or share results. When everyone uses different tools for data transformation, it becomes impossible for teams to collaborate on common problems or share data transformation code. Modern governance strategies establish common toolsets that allow anyone who knows SQL to understand and contribute to data models, enabling data engineering teams to share common transformation code across projects and reducing the need to start from scratch with each new dataset. **Defining Continuous Release Processes Centered on Quality** ensures that only high-quality code makes its way to production. This involves creating continuous integration release processes that include peer review, automated testing against non-production databases, and continuous monitoring of production data. These processes use role-based access control to define who's authorized to make changes to specific models, ensuring appropriate oversight while enabling efficient development workflows. ## Technical implementation components The technical infrastructure supporting data governance must provide comprehensive capabilities for managing data throughout its lifecycle while enabling collaboration and maintaining security. Modern governance frameworks require several key technical components working together seamlessly. **Data Cataloging and Discovery** capabilities ensure that data assets remain findable and understandable across the organization. This includes comprehensive metadata management, automated documentation generation, and searchable catalogs that enable self-service data discovery. Tools like dbt Catalog make data models discoverable to anyone with appropriate permissions, allowing them to read accompanying documentation, understand data lineage, and leverage datasets for their own work. **Data Lineage and Impact Analysis** provide visibility into how data flows through systems and how changes might affect downstream consumers. This capability is essential for understanding dependencies, assessing the impact of changes, and troubleshooting data quality issues. Comprehensive lineage tracking enables teams to trace data from its source through all transformations to its final consumption points. **Automated Quality Monitoring** moves beyond manual testing to provide continuous validation of data quality. This includes automated data quality checks, anomaly detection, and alerting systems that notify relevant stakeholders when issues arise. Modern governance frameworks integrate quality monitoring directly into data transformation workflows, making quality assurance a natural part of the development process. **Access Control and Security** mechanisms ensure that sensitive data is protected while enabling appropriate access for legitimate business needs. This includes role-based access controls, data masking capabilities, and audit logging that tracks who accessed what data when. These controls must be flexible enough to support diverse use cases while maintaining security and compliance requirements. ## Governance in the age of AI The emergence of AI and machine learning has created new governance challenges that traditional approaches weren't designed to handle. AI systems introduce unique considerations around bias, explainability, and model governance that require evolved governance strategies. Large language models and other AI systems are trained on vast amounts of data to generate probabilistic outputs, making their responses difficult to predict and creating concerns around transparency and auditability. This leads to issues such as bias in underlying data that results in biased responses, lack of transparency in how models arrive at outputs, and inability to explain or control model decision-making processes. AI systems are also susceptible to unique threats including data poisoning, prompt injection, model inversion, and private leakage attacks. These risks require governance frameworks that can address both traditional data governance concerns and AI-specific challenges. Modern governance strategies address these challenges by emphasizing data products: structured and curated assets designed to solve specific business problems. Data products are particularly well-suited for AI workloads because they make it easier to verify the origin and quality of data that comprises AI solutions. This approach enables organizations to maintain the high data quality standards necessary for reliable, unbiased AI outcomes. **** ## Building sustainable governance programs Successful governance programs recognize that components and responsibilities must evolve as organizations grow and data landscapes become more complex. The most effective approach starts with core components and adapts based on organizational maturity and requirements. The key to sustainability lies in embedding governance into natural workflows rather than creating separate, parallel processes. This requires selecting governance frameworks that integrate seamlessly with existing data platforms and development workflows. Modern tools like dbt enable this integration by providing governance capabilities directly within data transformation processes, making governance a natural part of development rather than an external constraint. Governance frameworks should be vendor-agnostic and integrate with major data cloud platforms like Snowflake, Databricks, and BigQuery. They should utilize centralized, reusable models that foster collaboration, reduce duplication, and ensure consistent data definitions across teams. Robust audit logging and access control features help safeguard data integrity while supporting software development best practices including portability, CI/CD, observability, and documentation. The ultimate goal is creating an environment where everyone in the organization can extract value from data assets while managing cost and complexity. When governance frameworks automate data flow traceability and process transparency, organizations can focus on optimizing operations, improving performance, and achieving strategic goals while minimizing data security and privacy risks. Effective data governance transforms data from a potential liability into a strategic asset. By implementing comprehensive governance components that address quality, stewardship, protection, and management while supporting modern collaborative workflows, data engineering leaders can build systems that enable faster, more confident decision-making across their entire organization. The investment in robust governance components pays dividends through enhanced data quality, reduced management costs, accelerated insights from trusted data sources, and the foundation necessary for successful AI initiatives. ## Components of data governance FAQs **What are the 10 core components of a data governance program outlined in the article?** The article focuses on four foundational pillars rather than ten specific components. These four pillars are: Data Quality and Trust (maintaining accuracy, completeness, and consistency across organizational data assets), Data Stewardship (creating clear roles and responsibilities for individuals who manage and monitor data quality), Data Protection and Compliance (encompassing security measures, privacy protections, and regulatory compliance processes), and Data Management Visibility (covering processes for storing, accessing, and manipulating data including metadata management and data lifecycle management). **What roles and responsibilities should be defined for the steering committee, senior management, data stewards, data owners, and data users in a data governance program?** At the executive level, organizations need a strategic champion (such as a Chief Data Officer) who secures resources, removes organizational barriers, and aligns data governance with business strategy. Data stewards serve as the operational backbone, defining and documenting data assets, ensuring data quality, and acting as liaisons between teams. Data owners are individuals or teams closest to where data is created, having primary accountability for specific datasets and making decisions about data access and usage policies. Technical governance roles, including data engineers and architects, implement the systems that enable governance by building infrastructure, data pipelines, quality monitoring systems, and access controls. **What tools and technologies are essential to support monitoring, validation, protection, and traceability?** Essential technical components include Data Cataloging and Discovery capabilities for comprehensive metadata management and searchable catalogs that enable self-service data discovery. Data Lineage and Impact Analysis tools provide visibility into data flows and dependencies throughout systems. Automated Quality Monitoring moves beyond manual testing to provide continuous validation through automated data quality checks, anomaly detection, and alerting systems. Access Control and Security mechanisms include role-based access controls, data masking capabilities, and audit logging. Modern tools like dbt integrate these capabilities directly into data transformation workflows, making governance a natural part of the development process. --- --- title: "AI unlock: Empowering future-ready analysts" description: "AI isn’t replacing analysts—it’s elevating them into strategic, future-ready decision accelerators." url: "https://www.getdbt.com/blog/ai-unlock-empowering-future-ready-analysts" date: "2025-10-27" authors: ["Daniel Poppy"] categories: ["Insights"] --- # AI unlock: Empowering future-ready analysts Artificial intelligence is reshaping conversations about the future of work. In analytics, the prevailing narrative often drifts toward replacement: machines taking over tasks once done by humans. **But the reality is different.** Analysts don’t see AI as a threat. **They see it as a force multiplier**—expanding their ability to drive strategy, retain top talent, and accelerate innovation, [according to a survey](https://www.getdbt.com/resources/the-analyst-revolution-unlocking-tomorrows-ai-initiatives) of analysts conducted by dbt Labs in conjunction with The Harris Poll. Far from eliminating the role, AI is creating the conditions for analysts to step into more visible, more strategic, and more impactful positions inside their organizations. **** ## Future-ready analysts The new era of analytics will be **conversational.** Prompting is already replacing many manual tasks, multiplying analyst impact instead of diminishing it. Instead of spending hours writing SQL queries or building dashboards manually, analysts are increasingly guiding AI systems through natural language, defining questions, setting boundaries, and ensuring outputs align with business needs. This shift is already reshaping collaboration. Analysts are no longer order-takers. They are becoming **co-pilots in decision-making**—sitting with stakeholders to explore scenarios in real time, pressure-testing assumptions, and helping leadership see risks and opportunities more clearly. They’re also emerging as **storytellers**. With AI-assisted tools, analysts can craft narratives that combine structured data, qualitative signals, and forward-looking insights. This positions them not just as technical specialists, but as strategic advisors who bridge the gap between raw information and business outcomes. **In short, the modern analyst is a decision accelerator—someone who ensures that insight travels faster, with more accuracy and more context than ever before. **But to reach this future, organizations must first remove the inefficiencies that dominate analysts' work today. ## The AI unlock: From busywork to business impact Despite these possibilities, the current reality for many analysts is less inspiring. Too often, they spend more time battling inefficiencies than delivering insight. **Data validation, formatting, and error-checking consume valuable hours that could be spent on higher-order strategy.** The result is unnecessary risk, slower delivery, and team burnout. The appetite for change is overwhelming. According to a survey conducted by dbt Labs and The Harris Poll, over 90% of analysts: - Want governed self-service platforms; - Believe all-in-one solutions would increase productivity; and - Say they’re more likely to stay with employers who invest in workflow optimization. At the same time, **85% would consider leaving an employer that forces them to use outdated tools.** These statistics go beyond preference. They signal a profession on the cusp of transformation. Analysts want speed, but not at the expense of trust. They’re asking for structured, **AI-powered environments **where they can self-serve safely, collaborate effectively, and focus on strategy, not maintenance. With capabilities like automated visualization, intelligent data cleaning, and real-time data quality detection, **AI removes repetitive friction points**. Instead of spending hours validating a dataset, analysts can guide decisions, monitor outcomes, and ensure insights are both explainable and actionable. This shift reclaims capacity that would otherwise be lost, redirecting it toward innovation and forward-looking analysis. ## Balanced autonomy and governance Outdated tooling isn’t just a workflow problem, it’s a **retention risk. Ninety-four percent of analysts say access to self-service tools is a critical factor when evaluating employers. **Losing top analysts means more than the cost of recruitment. It also erases institutional knowledge, complicates ongoing projects, and jeopardizes momentum on strategic initiatives. Leaders often face a false choice: restrict access and slow analysts down, or open the floodgates and risk chaos. Neither approach is sustainable. The solution is balance: **governed self-service** that combines autonomy with oversight. In practice, this means systems that provide audit trails, standardized definitions, role-based permissions, and clear lineage tracking. With these guardrails in place, analysts can explore data freely while leaders remain confident in the accuracy, security, and compliance of the insights being shared. **The benefits are tangible.** Governed self-service eliminates tool sprawl, reduces redundant work, and ensures consistency across business functions. It accelerates collaboration, minimizes governance bottlenecks, and allows analysts to spend more time shaping outcomes. Organizations that fail to invest in this balance risk more than inefficiency. They risk falling behind competitors who empower their analysts to move with speed, trust, and scale. ## A C-suite imperative For executives, the implications are clear: investing in AI and modern data platforms isn’t simply a technology decision. **It’s a talent decision, a retention strategy, and a competitive imperative.** Companies that give their analysts modern, AI-powered environments enjoy **higher morale, stronger retention, and more consistent insights. **They also build organizations that can pivot quickly, seize opportunities sooner, and respond to market changes with greater confidence. On the other hand, organizations that fail to evolve will see their most strategic minds trapped in manual workflows, unable to contribute at the pace the business demands. They’ll also face higher turnover, as analysts choose employers who empower them to do their best work. **The mandate for leaders is simple: empower analysts with AI-powered, governed workflows, or risk losing both talent and competitive ground.** AI-powered analytics doesn’t replace analysts—it elevates them. By embedding AI into governed workflows, organizations unlock the full potential of their data talent. The result is an analytics function that’s faster, more trusted, and more strategically central to the business than ever before. These findings are just the beginning. Explore the full research in [The Analyst Revolution](https://www.getdbt.com/resources/the-analyst-revolution-unlocking-tomorrows-ai-initiatives). **** --- --- title: "The governance gap: How shadow AI is already reshaping analytics" description: "Shadow AI is on the rise, and it's reshaping how analytics teams work, govern, and take risks with data." url: "https://www.getdbt.com/blog/the-governance-gap-how-shadow-ai-is-already-reshaping-analytics" date: "2025-10-20" authors: ["Daniel Poppy"] categories: ["Insights"] --- # The governance gap: How shadow AI is already reshaping analytics The promise of self-service analytics was simple: give analysts faster access to data, and they would deliver insights at the speed of business. Instead, fragmented tool adoption and rigid governance processes have created a dangerous middle ground—one where analysts, under pressure to move quickly, **bypass approved systems and turn to unapproved AI tools.** This governance gap is fueling the rise of shadow AI. Workarounds may help analysts meet short-term deadlines, but they expose organizations to **compliance violations, data leaks, and delayed projects.** ## The scale of the shadow AI problem Rising demand for rapid insights amid haphazard tool adoption and uneven governance policies has created a dangerous environment for analysts. Many are forced to work at the margins of governance, using tactics that create systemic risk. While executives are demanding faster, AI-powered insights, the vast majority of analysts feel that they lack the tools they need. [According to a survey](https://www.getdbt.com/resources/the-analyst-revolution-unlocking-tomorrows-ai-initiatives) conducted by dbt Labs in conjunction with The Harris Poll, **nearly all (90%) say that their organization needs more efficient tools to deliver insights, **and most report underinvestment in AI-powered platforms at the organizational level. **** Many analysts use** tools like ChatGPT, personal API keys, or free online tools **outside approved systems to analyze company data. And almost one-third (32%) admit to going even further—not just taking chances in their work, but actively **creating workarounds to bypass governance processes.** Unapproved AI tools for data analysis introduce uncertainty into workflows, fragment organizational knowledge, and force teams to retroactively validate results before they can be trusted. The cost of this workaround economy is steep, and analysts know it. Instead of accelerating insights, **shadow AI introduces friction and risk** into already overburdened workflows. A majority of analysts (63%) agree, reporting that **working outside governed systems further delays projects.** But in context, it's understandable why so many analysts feel pushed to "go rogue" and adopt shadow AI tools and strategies. Faced with rigid controls, tight timelines, and outdated systems, analysts turn to whatever gets the job done. The lesson is clear: **governance without enablement doesn’t work.** Analysts either wait on bottlenecked data teams, or they move ahead with unapproved solutions. **Both paths erode trust and slow down the business.** This is the governance gap: over-governed analysts face delays, while under-governed analysts create risk. **The organizations that succeed will be those that balance autonomy and governance, empowering analysts with modern tools inside governed workflows.** ## The risks of shadow AI and the governance gap For operational team leads, the dangers of shadow AI are immediate and tangible. **Compliance violations and legal exposure** can arise when sensitive data is processed through unapproved tools. **Personal API keys and free online platforms expand the attack surface for data breaches.** Shadow AI tactics also increase operational inefficiency and fragmentation. Retroactive validation wastes time, forcing projects into rework cycles instead of accelerating decisions. And analysts working in silos produce **inconsistent metrics and conflicting insights, undermining confidence across the organization.** These risks are particularly challenging because they remain **mostly invisible**. Shadow AI often operates outside IT’s line of sight, leaving leaders unaware of the vulnerabilities until something breaks. To close the governance gap, organizations must tackle a few key challenges: - **Integrate AI into sanctioned workflows**: By embedding AI into governed platforms, companies can deliver secure, high-functioning environments for the work and preserve both speed and compliance. - **Simplify tool sprawl:** On average, analysts juggle more than five platforms daily. Consolidating fragmented environments into a single governed control plane reduces context switching, minimizes risk, and speeds insight delivery. - **Empower analysts without bypasses:** To avoid the problem of analysts going rogue, leaders should ensure members of their team can access trusted data directly within governed systems. This reduces reliance on data engineering backlogs and keeps outputs reliable. - **Evolve governance models:** Governance shouldn’t be static compliance—it should be a dynamic, future-ready framework that balances autonomy with guardrails. In governed self-service environments, analysts can move fast without breaking trust. Solving this challenge doesn’t mean locking analysts out of AI. Instead, it means bringing AI into governed workflows. **The future of AI governance isn’t restriction, it’s enablement.** And that shift is already underway. ## The way forward The rise of shadow AI is a warning sign. **Analysts are already adopting new ways of working, but organizations are not keeping pace with the governance structures** needed to support them. Team leaders face a choice: either continue firefighting compliance breaches and project delays, or invest in governance models that align analyst empowerment with organizational trust. **The era of AI and analytics depends on threading this needle. Equip analysts with governed access to AI tools, and they will shift from risky workarounds to high-impact insights. Leave the governance gap unaddressed, and shadow AI will continue to grow—undermining trust, compliance, and competitiveness.** The findings are just the beginning. Explore the full research in [The Analyst Revolution](https://www.getdbt.com/resources/the-analyst-revolution-unlocking-tomorrows-ai-initiatives). **** --- --- title: "Data transformation in the data warehouse" description: "Why transformation matters in the warehouse—and how dbt makes it modular, testable, and scalable." url: "https://www.getdbt.com/blog/data-transformation-in-data-warehouse" date: "2025-10-16" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data transformation in the data warehouse The market for data warehousing is worth over[ USD 11 billion in 2025](https://www.mordorintelligence.com/industry-reports/global-active-data-warehousing-market-industry) due to the rising demand for state-of-the-art business intelligence (BI) solutions and data for AI. Thanks to the cloud, modern warehouses can easily scale to the size and speed needed for modern business. The blocker is data quality. Data transformation pipelines struggle with schema drift, duplicated logic, and opaque lineage. The result is inconsistent metrics, fragile dashboards, and delays as upstream changes ripple through dependent models. The challenge here isn’t warehouse compute. It’s the lack of structure in managing data transformation code. Ad hoc SQL scripts and stored procedures scattered across data warehouses make it hard to enforce standards, test assumptions, or track dependencies across teams. Addressing this complexity requires more than raw compute power. It calls for a framework that applies engineering discipline to the transformation layer. This is where [dbt](https://docs.getdbt.com/docs/get-started-dbt) comes in. dbt is a data control plane that acts as a single point for managed data workflows across your enterprise, turning fragile SQL workflows into reproducible and governed pipelines. In this article, we’ll look at why data transformation in the warehouse is critical to deliver value and how you can use dbt to deliver reliable, scalable pipelines. ## What is data transformation in the warehouse? Transformation is the “T” in [ELT (Extract, Transform, Load)](https://www.getdbt.com/blog/extract-load-transform). It happens after raw data has been loaded into the warehouse. Unlike the traditional ETL process, where data is extracted, transformed, and loaded, modern systems load raw data first. Transformations are then performed directly within the warehouse as needed for various use cases. This ELT model uses cloud data warehouse architecture for efficient, cost-effective transformations, using features like columnar storage, massively parallel processing (MPP), and elastic scaling. Many cloud warehouses utilize features such as [micro-partitioning (Snowflake)](https://blog.devgenius.io/snowflake-micro-partitions-clustering-keys-dbt-b6cb1212dcbe) or distributed query execution ([BigQuery](https://cloud.google.com/bigquery/docs)) to scale transformations efficiently. For instance, Snowflake automatically segments tables into micro-partitions, which supports pruning and clustering. ## Why data transformation matters Transformation is more than just a technical step in the pipeline; it determines whether data can be trusted, scaled, and used effectively. Let’s look at why this stage matters so much in the warehouse. ### Data quality and trust Source systems often produce mismatched data types, null values, and duplicate keys that disrupt joins and inflate metrics. Transformations ensure uniqueness, referential integrity, and standardized schemas for consistent downstream queries. ### Speed and agility Transformations materialize cleaned and aggregated tables that analysts can query directly instead of applying fixes to raw data. Techniques like incremental processing and partition pruning reduce execution time by avoiding full table scans, allowing rapid iteration cycles. ### Scalability and maintainability Ad hoc SQL scripts break down when data grows large or new sources are added. Modular transformation layers in dbt decouple staging, business logic, and marts, making pipelines testable and easier to refactor without system-wide failures. ### Observability and governance Transformation pipelines generate artifacts that document the flow of data from raw ingestion to analytics-ready outputs. Lineage graphs reveal dependencies between tables, while schema validations and anomaly checks ensure inputs match expected contracts. ### Analytics and AI performance Analytics and BI platforms use star schemas and pre-aggregated tables for performance, while ML pipelines need engineered features like rolling averages and cohorts. Transformations deliver these artifacts directly to the warehouse. ### Cost efficiency External ETL pipelines often duplicate workloads and transfer data unnecessarily, resulting in increased I/O and storage costs. In-warehouse transformations minimize data movement and apply optimizations such as clustered storage and materialized views. ## How dbt makes data transformation easier Transformation is central to how data becomes reliable and reusable in the warehouse. Understanding its importance also shows why dbt is built to strengthen this stage. - **Standardization:** Transformations clean, type, and align raw data before it flows into downstream models. This standardization allows dbt’s [schema tests](https://docs.getdbt.com/docs/dbt-cloud-apis/discovery-schema-job-tests) and documentation features to enforce consistency and prevent schema drift across projects. - **Lineage:** Every transformation adds a node to the dependency graph that dbt builds. The resulting [directed acyclic graph (DAG)](https://www.getdbt.com/blog/dag-use-cases-and-best-practices) makes data flows transparent and supports impact analysis when upstream schemas or data evolve. - **Modularity: **Arranging transformations into staging, intermediate, and marts layers creates a clear modular structure. dbt expands on this by supporting macros and packages, enabling teams to centralize logic and reuse it across models instead of duplicating SQL code. - **Optimization:** dbt’s [incremental materialization](https://docs.getdbt.com/docs/build/incremental-models) feature ensures that only new or changed records are processed after the initial model build. This approach reduces unnecessary recomputation, shortens runtimes, and improves overall warehouse efficiency. - **Consistency: **Transformations define business metrics once in central models that feed all BI tools and reports. dbt’s model architecture enables consistent metric reuse across dashboards and reports. **Safe evolution**: Layered transformations combined with lineage make it possible to detect and manage schema or source changes. dbt’s [version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) and testing catch issues early, enabling safe pipeline evolution without disrupting downstream analytics. ### How dbt makes common data transformations easier Using dbt, you can easily implement and manage the most common types of data transformations: - **Cleaning:** Removes duplicates, standardizes formats, and handles null values. In dbt, these checks are expressed as tests for [uniqueness](https://www.getdbt.com/blog/data-pipelines), validity, and completeness, ensuring consistent results across datasets. - **Enrichment: **Integrates data from multiple systems to create unified views, such as combining sales and support data into a customer profile. dbt’s [modular models](https://www.getdbt.com/blog/data-integration) make these joins and integrations reusable and traceable. - **Aggregation: **Summarizes granular records into higher-level metrics such as daily revenue or churn rates. Aggregations in dbt can be version-controlled and referenced across teams, reducing duplication of metric logic. - **Modeling: **Structures data into schemas optimized for analysis, including star and snowflake designs. dbt uses dependency graphs to enforce structures, and tests verify referential integrity and historical accuracy with slowly changing dimensions. ### Layers of transformation Transformations are structured into layers, with each stage refining data as it moves from raw input to analysis-ready output. dbt provides built-in support for managing these layers: - **Staging layer:** Holds raw data in a simple, lightly processed form while keeping it close to the original source. In dbt, this corresponds to [staging models](https://docs.getdbt.com/best-practices/how-we-structure/2-staging) that provide a consistent base for downstream logic. - **Intermediate layer: **Applies business rules and integrates dimensions across domains. dbt [intermediate models](https://docs.getdbt.com/best-practices/how-we-structure/3-intermediate) capture this logic centrally, reducing metric drift and aligning definitions across teams. - **Analytics or marts layer: **Provides business-ready datasets, optimized with summary tables, wide schemas, and incremental builds for efficient use in analytics, machine learning, and applications. This layered approach is reflected directly in dbt projects. Structured directories and dependency graphs ensure transformations run correctly, can be debugged efficiently, and align with the warehouse’s optimizations. ## Best practices with dbt Establishing best practices for data transformation keeps projects reliable, scalable, and easy to maintain as they grow. dbt supports these best practices directly, simplifying the creation, management, and maintenance of data pipelines at scale across the enterprise. - **Layered architecture:** A layered model design ensures that transformations remain modular and maintain their integrity. In dbt, this structure makes dependency graphs easier to interpret, separates raw data handling from business rules, and provides clarity when scaling pipelines. - **Version control:** As code-first projects, dbt [models](https://docs.getdbt.com/best-practices/best-practice-workflows) are typically managed in Git. Branching, pull requests, and peer reviews ensure traceability and help prevent regressions. This process aligns well with CI/CD pipelines, where automated tests are executed before deployment. - **Reusable code:** [Macros](https://docs.getdbt.com/docs/build/jinja-macros) and [packages](https://docs.getdbt.com/docs/build/packages) bring common logic into reusable units. This approach limits duplication, keeps SQL concise, and enables teams to enhance pipelines without introducing inconsistencies. - **Testing discipline:** Quality checks are embedded directly in the transformation layer. With dbt, [data tests](https://docs.getdbt.com/docs/build/data-tests) run alongside builds, so errors are caught before they reach production systems, reducing downstream rework and improving confidence in data outputs. - **Documentation:** Because dbt couples [documentation](https://docs.getdbt.com/docs/build/documentation) with models, projects become self-describing. Teams can navigate a living catalog of models and columns without depending on external wikis or one-off documentation efforts. - **Performance tuning:** Warehouse efficiency is shaped by how models are materialized. dbt gives developers control over [materializations](https://docs.getdbt.com/docs/build/materializations), [clustering](https://docs.getdbt.com/reference/resource-configs/snowflake-configs), and partitioning, aligning transformation performance with the cost and scale of the underlying platform. - **Metric consistency:** The dbt [Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) defines KPIs once and exposes them across tools. Using a single revenue definition across dashboards prevents metric drift and boosts confidence in reports. ## Case study: JetBlue and dbt ### Challenge JetBlue’s centralized [data engineering](https://www.getdbt.com/case-studies/jetblue) team faced significant bottlenecks. Legacy ETL pipelines built with [SQL Server Integration Services** (**SSIS)](https://learn.microsoft.com/en-us/sql/integration-services/sql-server-integration-services?view=sql-server-ver17) were slow and contributed to periods when the data warehouse was unavailable, reportedly resulting in it being live only 65% of the time. ### Solution with dbt JetBlue modernized its stack by migrating to Snowflake and adopting dbt for transformations. Over three months, it onboarded 26 data sources and created 1,200+ dbt models. Transformations were shifted closer to analysts to distribute ownership, while engineers focused on governance and reliability. JetBlue used dbt’s testing for data quality, lineage for dependencies, and auto-generated docs to standardize metrics. The team also implemented lambda views to union historical data with real-time streams, enabling faster operational insight. ### Impact - Pipeline uptime rose to 99.9% (from 65%), as measured by dbt Labs. - Analysts gained faster access to clean, trusted data. - Metric inconsistencies dropped, and reporting confusion decreased. - Documentation and testing enhanced transparency and trust. - Scalability improved without a rise in total cost of ownership. By integrating dbt into its warehouse transformation layer, JetBlue transformed its data stack into a more reliable, governed, and democratized system. The shift empowered analysts, improved pipeline stability, and scaled insights across the organization. ## Conclusion Data transformation converts raw warehouse data into consistent, analysis-ready datasets. Without it, data warehouses function as costly storage rather than engines of insight. dbt enhances this stage by making transformations modular, testable, and scalable. Version control, automated testing, documentation, and lineage bring structure and reliability, ensuring pipelines can grow without losing trust or transparency. The combination of modern warehouses and [dbt](https://www.getdbt.com/product/what-is-dbt) creates a foundation where data is not just stored, but consistently transformed into value. Try [dbt for free](https://www.getdbt.com/signup) today [using one of our quickstarts](https://docs.getdbt.com/docs/get-started-dbt) to see for yourself how it simplifies managing data transformation. --- --- title: "Main uses of data integration tools" description: "Explore how data integration tools consolidate data, power real‑time and governance use cases, and enable analytics across teams." url: "https://www.getdbt.com/blog/main-uses-of-data-integration-tools" date: "2025-10-14" authors: ["Joey Gault"] categories: ["Pulse"] --- # Main uses of data integration tools The primary use case for data integration tools is consolidating data from various systems into a single, queryable environment. Organizations typically operate dozens of different platforms, from CRM systems like Salesforce to advertising platforms like Google Ads, backend databases, and SaaS applications. Each system generates valuable data, but analyzing it in isolation provides limited insights. Data integration tools address this fragmentation by extracting data from these diverse sources and loading it into centralized platforms such as cloud warehouses or lakehouses. This consolidation enables cross-functional analysis that would be impossible when data remains siloed. For example, a retail organization might combine e-commerce transaction data, customer service interactions, and marketing campaign performance to build comprehensive customer profiles that drive personalization strategies. The extraction process itself varies depending on the data source. Well-supported platforms often provide prebuilt connectors that handle the technical complexity of API calls, authentication, and data formatting. For custom or legacy systems, data integration tools may require more sophisticated extraction logic, but they abstract much of the complexity that would otherwise require custom scripting. ## Enabling real-time and near-real-time analytics Modern business operations increasingly demand fresh data for decision-making. Data integration tools support this requirement through change data capture (CDC) and streaming capabilities that detect and synchronize source system changes as they occur. This real-time data movement is particularly valuable for use cases where delayed information can lead to poor decisions or missed opportunities. Financial services organizations use CDC to power fraud detection systems that must evaluate transactions within milliseconds. Similarly, e-commerce platforms rely on real-time inventory updates to prevent overselling and optimize pricing strategies. The key advantage of using data integration tools for these scenarios is their ability to handle the complexity of exactly-once delivery, schema evolution, and error handling that real-time pipelines require. Log-based CDC, which reads directly from database transaction logs, offers the lowest latency and overhead for these applications. When direct log access isn't available, trigger-based CDC provides an alternative approach by emitting change events from within applications. Data integration tools manage these technical complexities while providing the reliability and monitoring capabilities that production systems require. ## Supporting compliance and data governance In regulated industries such as finance and healthcare, data integration tools play a critical role in maintaining compliance while enabling analytics. These tools can apply data masking, encryption, and access controls during the integration process, ensuring that sensitive information is properly protected before it reaches analytical systems. The ability to transform data during extraction (a hallmark of traditional ETL approaches) remains valuable for scenarios where personally identifiable information (PII) or protected health information (PHI) must be anonymized before storage. Data integration tools can hash sensitive fields, remove identifying information, or apply other transformations that meet regulatory requirements while preserving the analytical value of the data. Beyond privacy protection, these tools support governance through comprehensive audit trails and lineage tracking. Data engineering leaders can trace exactly how data flows from source systems through transformations to final outputs, which is essential for regulatory reporting and internal compliance processes. This visibility becomes increasingly important as organizations scale their data operations and need to demonstrate control over their data handling practices. ## Handling diverse data types and formats Modern organizations work with structured data from traditional databases, semi-structured data like JSON from APIs and applications, and unstructured data from logs and documents. Data integration tools excel at handling this variety, providing the flexibility to ingest data in its native format and transform it as needed for specific use cases. Cloud warehouses now natively support semi-structured formats, which means data integration tools can load JSON, XML, and other formats directly without requiring upfront schema definition. This capability is particularly valuable for organizations dealing with rapidly evolving data sources where schema changes are frequent. Rather than breaking pipelines when new fields are added or data types change, modern data integration tools can adapt to these variations automatically. The ability to work with diverse data types also supports exploratory analytics and data science workflows. Data scientists often need access to raw, unprocessed data to identify patterns or build models. Data integration tools can provide this access while simultaneously supporting the structured, cleaned datasets that business analysts require for reporting and dashboards. ## Optimizing performance and cost Data integration tools provide significant performance advantages over custom-built solutions, particularly when working with cloud-based data platforms. By leveraging the native compute capabilities of warehouses like Snowflake, BigQuery, and Databricks, these tools can process transformations at scale without requiring separate infrastructure. The shift from ETL to ELT architectures exemplifies this optimization. Rather than transforming data on separate servers before loading, ELT approaches load raw data first and perform transformations within the warehouse using its elastic compute power. This approach reduces infrastructure costs, improves scalability, and enables faster iteration on analytical models. Incremental processing capabilities further enhance performance and cost efficiency. Instead of reprocessing entire datasets with each update, data integration tools can identify and process only new or changed records. This approach dramatically reduces compute costs and processing time, especially for large datasets where only a small percentage of records change between updates. ## Enabling self-service analytics Data integration tools democratize data access by providing business users with reliable, well-documented datasets they can analyze independently. Rather than requiring technical expertise to extract and prepare data from source systems, business analysts can work with pre-integrated datasets that are already cleaned, standardized, and optimized for analysis. This self-service capability is enhanced by semantic layers that provide consistent definitions of key business metrics across different tools and use cases. When data integration tools work in conjunction with semantic layers, organizations can ensure that metrics like "customer churn" or "monthly recurring revenue" are calculated consistently whether they're accessed through dashboards, notebooks, or AI applications. The documentation and metadata capabilities of modern data integration tools also support self-service analytics by providing context about data sources, transformation logic, and data quality. Business users can understand the provenance and reliability of their data without needing to consult with engineering teams for every question. ## Facilitating data science and machine learning Data science and machine learning workflows have specific requirements that data integration tools are increasingly designed to support. These workflows often require access to both historical data for model training and real-time data for inference. Data integration tools can provide both through batch processing for historical analysis and streaming capabilities for real-time model serving. Feature engineering, a critical component of machine learning pipelines, benefits from the transformation capabilities of data integration tools. Data scientists can define feature calculations as part of the integration process, ensuring that the same logic is applied consistently across training and production environments. This consistency is essential for model performance and reduces the risk of training-serving skew. The ability to handle large volumes of data efficiently also makes data integration tools valuable for machine learning applications. Training modern machine learning models often requires processing massive datasets that would be impractical to handle with custom scripts or manual processes. Data integration tools provide the scalability and reliability needed for these demanding workloads. ## Supporting operational analytics and reverse ETL Beyond traditional analytical use cases, data integration tools increasingly support operational analytics through reverse ETL capabilities. This involves taking insights generated from integrated data and pushing them back to operational systems where they can drive business processes. For example, customer segmentation models built from integrated data might be used to personalize marketing campaigns in email platforms or CRM systems. Product recommendation engines might feed suggestions back to e-commerce platforms in real-time. These operational use cases require data integration tools that can not only bring data in but also push processed insights back to where they can create business value. The reliability and monitoring capabilities of data integration tools become particularly important for operational use cases where data quality issues can directly impact customer experience. Unlike analytical applications where bad data might lead to incorrect insights, operational applications can affect customer-facing processes in real-time. ## Conclusion Data integration tools have evolved far beyond simple data movement to become comprehensive platforms that enable modern data-driven organizations. They address the fundamental challenge of turning diverse, distributed data sources into unified, reliable datasets that power everything from executive dashboards to machine learning models. For data engineering leaders, the key is selecting tools that align with their organization's specific requirements around data sources, scalability, governance, and use cases. The most effective implementations combine robust data integration capabilities with transformation frameworks like dbt that bring software engineering best practices to the analytical workflow. As organizations continue to generate more data from more sources, and as real-time requirements become more demanding, data integration tools will remain essential infrastructure for any serious data operation. The organizations that invest in building reliable, scalable data integration capabilities will be best positioned to extract value from their data assets and maintain competitive advantages in increasingly data-driven markets. ## Data integration tools FAQs **What are the different types of data integration tools?** **Do you need data to sync in real-time or due to a particular action?** Real-time data synchronization is essential for use cases where delayed information can lead to poor decisions or missed opportunities. Financial services organizations use change data capture (CDC) for fraud detection systems that must evaluate transactions within milliseconds, while e-commerce platforms rely on real-time inventory updates to prevent overselling. Log-based CDC offers the lowest latency by reading directly from database transaction logs, while trigger-based CDC provides an alternative by emitting change events from within applications when direct log access isn't available. **Does the tool support data transformation, enrichment, deduplication, and quality checks?** Modern data integration tools provide comprehensive transformation capabilities including data masking, encryption, and access controls during the integration process. They can hash sensitive fields, remove identifying information, and apply transformations that meet regulatory requirements while preserving analytical value. These tools also handle diverse data types from structured databases to semi-structured JSON and unstructured logs, with automatic schema adaptation capabilities. Additionally, they provide incremental processing to identify and process only new or changed records, reducing compute costs and processing time. --- --- title: "Coalesce 2025: Rewriting the future of data, analytics, and AI" description: "Coalesce 2025 highlights how Fusion, AI, and state-aware orchestration are rewriting data workflows for speed, trust, and savings." url: "https://www.getdbt.com/blog/coalesce-2025-rewriting-the-future" date: "2025-10-14" authors: ["David Tishgart"] categories: ["Company", "Product"] --- # Coalesce 2025: Rewriting the future of data, analytics, and AI Coalesce 2025 is off and running with exciting keynotes, lively hallway conversations, and more than 100 breakout sessions on AI, Fusion, Iceberg, and everything in between. More than 2,000 people packed the general session this morning at Resorts World in Las Vegas, with another 10,000 joining online from around the world, making this the largest Coalesce event to date. In a jam-packed 90 minute keynote, dbt executives, product leaders, and some of our customers shared the latest developer experience features of the dbt Fusion engine, demonstrated how to reduce your cloud bill with state-aware orchestration, and showed off the dbt platform’s new governed agentic AI experiences. But today wasn’t just about new platform features. It was about giving practitioners and leaders the tools they need to move faster, cut costs, and trust AI in ways that weren’t possible before. In other words, we are **_rewriting your expectations for how data work gets done_.** ## **The era of open data infrastructure** On Monday, we announced a merger with Fivetran to set the standard for [**open data infrastructure**](https://www.getdbt.com/blog/what-is-open-data-infrastructure) . ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/078229e2d49590add65df6a8f263f54c6f0ca308-1920x1280.png) The topic of open data and the merger of two complementary technologies led today’s keynote. While there is still lots to sort out over the coming months, dbt Labs and Fivetran have committed to the following: - dbt Core and Fusion will both continue to be shipped under their current licenses. - We will continue to maintain dbt Core indefinitely. - We will continue to support and foster the dbt Community Be sure to check out a replay of the opening keynote to hear from both Tristan and George about why we're excited for this next step, together. Now, on to the product announcements! 👇 ## Rewrite the developer experience with Fusion _Faster dev cycles, fewer errors._ We’ve rebuilt dbt from the ground up. The new **Fusion engine is available in Preview in the dbt platform for eligible projects**. Back in August, [we launched Fusion into Preview for _local_ development](https://www.getdbt.com/blog/fusion-and-dbt-vs-code-extension-preview-launch) in the CLI and VS Code extension. With today’s release, you can choose to run your projects on Fusion whether you develop locally or in the cloud. Built in Rust, Fusion parses 30x faster than dbt Core and introduces a compiler and stateful architecture that understands SQL deeply across platforms. That foundation unlocks key features that data developers have been asking for including: - Intellisense, hover insights, go-to-definitions, and instant refactoring - Live previews of CTEs without leaving your flow - `dbt compare`, a brand-new way to see data diffs directly in VS Code ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/7bc98f4dc86cd9ef20af2a0154d886c74a6895a0-2940x1832.gif) Fusion is also compatible with dbt Core and the projects you’ve built in the platform. Your existing dbt code will largely _just work_ with it, and we’ve built [Autofix tools](https://github.com/dbt-labs/dbt-autofix) to help smooth any upgrade issues. dbt Core isn’t going away—it’ll continue to be supported indefinitely—but Fusion is where innovation is headed. > Preventing human errors with live error detection saves DPG Media valuable time. This feature in Fusion through the dbt VS Code extension has been a game-changer. > - Sonja Strempel, Analytics Engineer at DPG Media Adoption is strong, with over 1,500 weekly average projects running on Fusion and over 8,000 weekly average users of the [dbt VS Code extension](https://docs.getdbt.com/docs/about-dbt-extension). With today’s release, many more teams will get the chance to run their projects on Fusion for the first time. [Click here](https://docs.getdbt.com/guides/fusion) to take Fusion for a test drive. ## Rewrite the rules with state-aware orchestration _Run when it matters. Reuse when it doesn’t._ Also available in Preview for projects running on Fusion in the dbt platform is [**state-aware orchestration**](https://www.getdbt.com/blog/announcing-state-aware-orchestration), a new way to avoid unnecessary costs associated with executing data pipelines while still meeting your data business requirements. Here’s how it works: instead of rebuilding every model in your DAG, Fusion has the stateful intelligence to automatically skip the ones that don’t need to be refreshed—because either no data has changed upstream, or you’ve defined certain rules for how often you want your models refreshed—reducing wasted compute and delivering faster pipelines. ### Automatically reuse unchanged models, save 10% Turning state-aware orchestration on is simple. Once your projects are running on Fusion, just flip the toggle in the dbt platform, and dbt automatically pinpoints which models should be refreshed, based on whether or not new data is produced upstream. Models without any upstream changes are automatically reused. ![Models built vs. reused chart](https://cdn.sanity.io/images/wl0ndo6t/main/f280e2c557dba38c828840f0b913092fc5c66952-2336x1124.png) Results from our beta cohorts showed this simple change resulted in an average of **10%** cost reduction in data platform costs, with no required changes to existing projects. ## Tuned configurations for more savings This is just the start. Instead of wrestling with schedules, we have made it easy to fine-tune your configurations: you can declare your freshness requirements, and then dbt will detect exactly which upstream tables have new or changed data and update models accordingly based on the freshness targets you’ve set across your project or on each model. For example, you can set your maximum staleness window before a model needs to rebuild (say, six hours) with a `build_after=6h` configuration. Or you can define whether all sources need to be refreshed before updating a downstream model (`updates_on=all`). ![state aware orchestration tuned configuration](https://cdn.sanity.io/images/wl0ndo6t/main/6bd181f7e884604ccff599fc74582cc7821f8856-960x540.jpg) In our beta cohorts, tuned configurations reduced annual cloud costs by _at least_ an additional **15%+**. ### More efficient testing Testing also gets smarter with Fusion. New **column-aware** and **aggregated testing** features eliminate redundant checks so you’re not validating the same data over and over. Early testing shows these improvements reduce CI and testing costs by an average of 4%. ![Column aware testing](https://cdn.sanity.io/images/wl0ndo6t/main/048851c02251ffb45c2c446bd232fec3c905c30d-960x540.jpg) ### Run Fusion in production and save 30% Put together, these optimizations can drive a potential **29%+ reduction** in annual data costs: - 10% from just turning on state-aware orchestration - 15%+ from tuning the configuration with business SLAs - 4% from more efficient testing and CI Intelligence is now the default in dbt. You don’t need to rebuild your project: just run it on Fusion. Learn more about state-aware orchestration [here](https://www.getdbt.com/blog/announcing-state-aware-orchestration) and see how dbt Labs saved 64% on our own projects! ## Rewrite how you work with data with AI _Governed, auditable AI in the tools you already use._ Before we get into how AI works with your data, let’s talk about how _you_ interact with your data, particularly as it relates to development and query. **[dbt Canvas](https://docs.getdbt.com/docs/cloud/canvas),** generally available (GA) in May 2025, makes it easy for analysts to contribute safely to production pipelines in a visual way. New to dbt Canvas is the ability to **upload a CSV file** and drag it into your project. > The drag-and-drop interface in dbt Canvas makes it easier and faster to update existing models without breaking anything. > - Tony Mayer, Senior LOB Reporting & Analytics Manager, Fifth Third Bank Another popular platform feature, **[dbt Insights](https://docs.getdbt.com/docs/explore/dbt-insights),** is now generally available today, giving anyone governed, fast analysis with confidence. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/65e23d80ae7d67e064845cfd38f7c75ed1902105-2940x1592.png) ### AI-powered analytics is here AI has moved fast in the past two years. What started with basic SQL generation has become end-to-end project automation, and dbt sits at the center by powering your AI with a structured context layer—models, metrics, tests, and lineage—so automation is trustworthy and results are reliable. Today at Coalesce, we [announced several major updates](https://www.getdbt.com/blog/dbt-agents-remote-dbt-mcp-server-trusted-ai-for-analytics) to the [**dbt MCP server**](https://docs.getdbt.com/docs/dbt-ai/about-mcp), the universal bridge between AI tools and dbt structured context layer: - Local MCP now supports **OAuth**, so you can securely log in with your dbt credentials. - The **remote MCP server is now generally available** and exposes one secure endpoint per environment, making it simple for tools like OpenAI, Anthropic, and Cursor to connect to your team’s dbt projects. - **Fusion MCP tools are now integrated with the dbt MCP server**, giving agents richer SQL comprehension and lower-cost execution. On top of this foundation, we introduced [**dbt Agents**](https://www.getdbt.com/product/dbt-agents): out-of-the-box, auditable agents that extend governance into everyday AI workflows. The following are coming soon to dbt: - **Developer agent (coming soon) :** Explains model logic, predicts downstream impact, validates before merge, and can draft or refactor from prompts. It will run in VS Code or dbt Studio and is powered by dbt context, so every change can be shipped quickly and safely. - **Observability agent (coming soon):** Helps you monitor jobs, pinpoint likely root causes, and cut resolution time. Designed to reduce noise and cut investigation time dramatically. - **Discovery agent** (beta): Helps you find the right dataset or metric in plain language, along with clear definitions and why it’s trustworthy. Surfaces governed sources and dbt lineage so anyone can explore data with confidence. - **Analyst agent** in dbt Insights (beta) answers natural language questions with governance intact. Answers come with definitions, lineage, and tests—what we call “answers with receipts.” - Sign up for the [dbt Agent waitlist today](https://www.getdbt.com/product/dbt-agents). Developer agent (coming soon) Analyst agent (beta) The impact is already visible. At Norway’s sovereign wealth fund, NBIM, any employee can now use conversational analytics powered by the dbt MCP server and Anthropic’s Claude to ask questions and get verified, accurate results. > Chat with your data’ works reliably when every answer is governed and explainable. dbt is our governance backbone, and MCP exposes that structured context to our chat experience so any NBIM employee can ask questions and get trusted answers. With dbt, Claude, and Snowflake powering our chat experience, adoption is 10× our previous catalog and tickets to core data teams are down. The same governed foundation now powers agentic workflows that flag anomalies, deliver morning briefs, and open small PRs. > — Øyvind Eraker, Senior Data Engineer, NBIM With dbt MCP and dbt Agents, AI is no longer a side experiment. It’s how governed data gets built, managed, and consumed at scale. Finally, we announced that [**MetricFlow is now open source**](https://www.getdbt.com/blog/open-source-metricflow-governed-metrics) under Apache 2.0, with co-maintenance from partners like **Snowflake** and **Salesforce,** and aligned to the Open Semantic Interchange (OSI) ecosystem efforts . This makes MetricFlow a reliable engine and critical AI infrastructure that ensures metrics like “revenue” or “active users” are consistent across dashboards, notebooks, and AI agents. Open sourcing MetricFlow enables partners and the community to build together on a a common, open engine to accelerate AI adoption and preserve trust in every result. ## Rewriting how data work gets done Day one of Coalesce 2025 was about rewriting your expectations for how data work gets done. - **Rewrite the developer experience** with Fusion, now available in Preview on the dbt platform for eligible projects, as well as in Preview for local development. - **Rewrite the rules for your cloud costs** with state-aware orchestration, also available in Preview for dbt platform projects running on Fusion. - **Rewrite how you build, manage, and consume data** in the age of AI with dbt Insights and the dbt Remote MCP Server, both GA today and with new AI agents in beta and coming soon. This is more than a roadmap. It’s already here. Thousands of teams are adopting Fusion today, and the results are clear: better performance, lower costs, and AI they can trust. If you’re ready for a deeper dive into dbt Fusion, check out our upcoming webinar, [Speed, simplicity, cost savings: Experience the dbt Fusion engine](https://www.getdbt.com/resources/webinars/speed-simplicity-cost-savings-experience-the-dbt-fusion-engine). --- --- title: "dbt Labs Affirms Commitment to Open Semantic Interchange by Open Sourcing MetricFlow" description: "dbt Labs is open-sourcing MetricFlow with an Apache 2.0 license, a significant step towards advancing trustworthy AI." url: "https://www.getdbt.com/blog/dbt-labs-affirms-commitment-to-open-semantic-interchange-by-open-sourcing-metricflow" date: "2025-10-14" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Affirms Commitment to Open Semantic Interchange by Open Sourcing MetricFlow _Shift to Apache 2.0 license model fuels trustworthy AI, closely following launch of OSI initiative_ **PHILADELPHIA, October 14, 2025** – [dbt Labs](http://www.getdbt.com), the leader in standards for AI-ready structured data, today announced at its annual conference, [Coalesce 2025](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/summary), that it is open sourcing [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) with an Apache 2.0 licence. This marks a significant step towards advancing trustworthy AI across enterprises and comes as the company has committed to the Open Semantic Interchange (OSI), a joint initiative led by industry leaders aimed at creating vendor-neutral standards for semantic data exchange across analytics platforms and AI tools. Inconsistent metrics and fragmented definitions have become a persistent problem as enterprises rush to deploy AI, leading to diminished trust and slow adoption. Because of this, there is a massive need for a single standard for metrics that every tool and agent can rely on along with an engine that renders a metric into its provably-correct calculation. MetricFlow is the core engine that compiles metric definitions into the code that computes them, and, unlike text-to-SQL methods, that computation is explainable and reliable every time. **Driving governance for AI-ready data** MetricFlow, which has powered the dbt Semantic Layer following the company’s [acquisition of Transform in 2023](https://www.getdbt.com/blog/dbt-acquisition-transform), uses information from semantic model and metric YAML configurations to construct and run SQL in a user’s data platform, providing governed metrics. By open sourcing MetricFlow under the Apache 2.0 license, dbt Labs is providing the community with a transparent, extensible engine that will enable AI agents to leverage trusted metric definitions for governed conversational analytics, ensuring teams get consistent results across tools and clouds for scale. "dbt is rooted in our open source DNA. This transition to open sourcing MetricFlow will unlock new opportunities for data practitioners to deliver huge value for their organizations," said Ryan Segar, Chief Product Officer of dbt Labs. "Metrics drift across tools and trust erodes if there isn’t a single source of truth, and with [90% of analysts agreeing](https://www.getdbt.com/resources/the-analyst-revolution-unlocking-tomorrows-ai-initiatives) that they desperately need more efficient tools to meet business demands, open sourcing this engine greatly diminishes the cadence of re-work while improving organizational trust in data.” **Empowering Organizations to Standardize Fragmented Data** dbt Labs’ [recent commitment](https://www.snowflake.com/en/news/press-releases/snowflake-salesforce-dbt-labs-and-more-revolutionize-data-readiness-for-ai-with-open-semantic-interchange-initiative/) to the OSI initiative, led by [Snowflake](http://www.snowflake.com), the AI Data Cloud Company, and other industry leaders including Salesforce and Sigma, focuses on solving the costly impact of nonstandardized data definitions. “The OSI initiative aims to standardize how semantic metadata is defined and shared, and dbt Labs’ change of the MetricFlow license to Apache 2.0 is an important step in this mission,” said Josh Klahr, Director of Analytics Product Management at Snowflake. “Fragmented data definitions are one of the largest barriers to AI adoption, and we believe that MetricFlow can serve as a key component in helping us achieve our vision of providing a shared set of analytic metadata that will propel trusted data forward.” The collaboration acknowledges the pitfalls that come with proprietary semantic standards and offers a solution to directly target the bottleneck holding organizations back from achieving lofty AI-driven goals. MetricFlow will be a central component in realizing the vision of the OSI initiative. “Defining metrics in MetricFlow is crucial for a unified source of truth,” said Rob Vicker, Data Analytics Architecture Director at EMC Insurance. “However, BI and AI tools often interpret metrics in their own ways. By leveraging open-source MetricFlow, OSI can ensure every tool is consistent in the consumption of metrics. This saves analysts time, streamlines audits, and gives users the flexibility to access data as they prefer, without constant code updates." For more information, visit https://www.getdbt.com/blog/open-source-metricflow-governed-metrics. **About dbt Labs** Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 80,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at [getdbt.com](http://getdbt.com), and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "dbt Labs Delivers Significant Cost Optimization Results and Agentic AI Features, Powered by Fusion" description: "Double-digit compute spend reductions and a new suite of goverened AI agents among new features announced at Coalesce 2025" url: "https://www.getdbt.com/blog/dbt-labs-cost-optimization-agentic-ai-product-announcements" date: "2025-10-14" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Delivers Significant Cost Optimization Results and Agentic AI Features, Powered by Fusion _New capabilities accelerate developer workflows, automatically reduce compute spend, and bring intelligent AI assistance to structured data_ **LAS VEGAS** – October 14, 2025 — On the keynote stage at Coalesce 2025, [dbt Labs](http://getdbt.com) – the leader in standards for AI-ready structured data – showcased the continued evolution of the dbt platform, unveiling cost optimization outcomes and performance enhancements powered by the dbt Fusion engine. dbt Labs also introduced dbt Agents, a suite of intelligent AI assistants built into dbt and made accessible via the remote dbt MCP server, to supercharge development, improve governance and deliver trustworthy AI outcomes. Collectively, these platform updates accelerate development, support cross-platform portability and lay an important foundation for new analytics use cases.  “Open standards and AI are fueling the next era of analytics, and the dbt Fusion engine is the bridge that data teams need to move toward that future,” said Tristan Handy, founder and CEO at dbt Labs. “Fusion delivers robust context, tools and error-correction mechanisms for both humans and agents. It is the enabler of next generation, AI-powered data infrastructure.” Fusion, now in Preview for eligible projects on BigQuery, Databricks, Snowflake and Redshift, is building on its supercharged developer experience capabilities and enabling customers to dramatically optimize compute spend, eliminate wasted cycles, and focus teams on innovation and faster insights delivery. **State-aware orchestration**, in Preview for projects running Fusion, instantly allows teams to reduce compute spend by approximately 10% simply by activating the feature. It ensures pipelines only run models that have changed, allowing organizations to reduce unnecessary compute costs, all without rewriting projects or restructuring jobs. Teams can further tune their pipelines by providing specific data freshness requirements, and state-aware orchestration determines the most efficient job execution path to meet those needs. Organizations can expect an additional estimated 15%+ data platform cost savings with these tuned configurations, and in early testing, some organizations have realized over 50% in total savings. **This represents a meaningful reduction in customers’ spend on data infrastructure.** Fusion is powering more than just smarter orchestration; it is supporting evolving analytics use cases. dbt-powered pipelines can now **create and manage Apache Iceberg tables in Snowflake and Databricks**, laying the groundwork for easier adoption of open table formats and cross-platform portability. In addition, the **dbt VS Code Extension**, now in Preview, lets developers run Fusion locally for tighter inner loops and parity with production environments, while **dbt Insights**, now Generally Available, uses Fusion’s language server bringing definitions, lineage, cost, performance, and reliability into one place for faster, smarter decisions. “The improvements the dbt Fusion engine delivers address many of the pain points we currently face in our dbt development cycle,” said James Dorodo, VP of Data Analytics at Bilt Rewards. “It will provide a step function increase in our velocity.” **Agentic AI Supports the Evolution of Data Analytics** As the leader in standards for AI-ready structured data, dbt is introducing governed AI agents powered by dbt’s uniquely powerful context, to make analytics faster and smarter while preserving quality, trust, and governance. dbt Agents is built directly into the platform and includes: - **Developer agent**: Explains logic, flags duplicates, validates, and authors/refactors from prompts in VS Code or dbt Studio for faster, safer shipping. - **Discovery agent**: Finds the right datasets and definitions, highlighting trusted sources for faster exploration. - **Observability agent**: Monitors jobs, identifies likely root causes, and proposes fixes to reduce manual remediation work. - **Analyst agent**: Built into dbt Insights, this agent answers questions about models, jobs, and metrics, dramatically accelerating the insight-generation process. These agents bring AI into the heart of the Analytics Development Lifecycle, helping teams accelerate outcomes, improve quality, and maintain governance, ensuring AI delivers tangible business impact. This structured context is now universally accessible to AI systems through the **remote dbt MCP server**, now Generally Available. It runs in the cloud, connecting AI tools to projects in dbt without local setup. Now, dbt’s context, tooling, and error-correction is accessible to model providers and IDEs like OpenAI, Anthropic, and Cursor for safer, more reliable AI systems. “dbt is our governance backbone, and MCP exposes that structured context to our chat experience so any NBIM employee can ask questions and get trusted answers. With dbt, Claude, and Snowflake powering our chat experience, adoption is 10× our previous catalog and tickets to core data teams are down,” said Øyvind Eraker, Senior Data Engineer at NBIM. “AI only works at our scale when metadata, quality, and governance come first. dbt gives us that foundation.” **dbt Labs Deepens OSI Commitment with MetricFlow License Change** dbt Labs also announced that [MetricFlow is now fully open source](https://www.getdbt.com/blog/dbt-labs-affirms-commitment-to-open-semantic-interchange-by-open-sourcing-metricflow), with an Apache 2.0 license. This reaffirms dbt Labs’ commitment to the [Open Semantic Interchange](https://www.snowflake.com/en/blog/open-semantic-interchange-ai-standard/) (OSI) initiative. By standardizing metrics and semantics across tools, organizations can ensure consistent, trusted outcomes across analytics and AI workflows. “We're thrilled to see the introduction of dbt Agents at this year's Coalesce,” said Southard Jones, Chief Product Officer at Tableau. “Tableau and Salesforce believe deeply in making agents as successful as possible for our customers. That is why, along with dbt and others, we joined the OSI to deliver consistent definitions to the agents our customers are building to drive tangible business results.” To learn more about the product innovations unveiled at Coalesce 2025, register now to watch keynotes and select sessions virtually at [coalesce.getdbt.com](http://coalesce.getdbt.com) or join the [Oct. 28 and 29 live virtual event](https://www.getdbt.com/resources/webinars/speed-simplicity-cost-savings-experience-the-dbt-fusion-engine), focused on how the dbt Fusion engine transforms analytics workflows across dbt Core and the dbt platform. **About dbt Labs** Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 80,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at getdbt.com, and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). **** --- --- title: "Announcing dbt Agents and the remote dbt MCP Server: Trusted AI for analytics" description: "Introducing new AI agent capabilities to speed up your analytics lifecycle." url: "https://www.getdbt.com/blog/dbt-agents-remote-dbt-mcp-server-trusted-ai-for-analytics" date: "2025-10-14" authors: ["Chakshu Mehta", "Tom Grabowski", "Sai Maddali"] categories: ["Product"] --- # Announcing dbt Agents and the remote dbt MCP Server: Trusted AI for analytics Today, at [Coalesce 2025](https://www.getdbt.com/blog/coalesce-2025-rewriting-the-future), we announced the general availability of the **remote dbt MCP server** and introduced [**dbt Agents**](https://www.getdbt.com/product/dbt-agents), a new family of governed, task-specific AI agents built on the dbt platform, now available in beta. The **remote dbt MCP server** runs in the cloud and exposes one secure endpoint per environment so AI tools can connect to your dbt project without local setup or custom connectors. **dbt Agents** operate inside dbt guardrails. They read the same structured context your team already trusts, so development moves faster, self-service is safer, and data quality goes up. Together, they make dbt the standard context layer for agentic analytics. Whether you're using AI tools like Claude or Cursor, or you want a more dbt-native experience, we want to enable that choice while moving fast and keeping governance intact. Get started with the **remote or local [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) today** to build your own copilots and agents, or [request access to dbt Agents](https://www.getdbt.com/product/dbt-agents) (beta/coming soon) and be sure to check out the [Coalesce keynote recap](https://www.getdbt.com/blog/coalesce-2025-rewriting-the-future) to learn more about our latest launches. ## The agent era: The next step for trusted analytics The world of analytics is shifting. Data leaders are already rolling out AI that speeds how teams build, manage, and consume data, and many deployments are live, not experimental. As those efforts scale from pilots to production, one pattern is clear: agents are only as good as the context they run on. Without shared context, an agent guesses. It may pick the wrong joins, apply the wrong filters or time grains, or use out-of-date logic. With shared context, an agent knows. It can read definitions for metrics and dimensions, follow lineage to see dependencies, check tests and freshness, respect policies and roles, and explain exactly what it did. That's where dbt comes in. dbt has long provided data teams this **structured context layer** (definitions, lineage, tests, and semantics) that codifies how data should behave. That foundation now unlocks the **next step:** an agent-first era where AI plans work, executes tasks end to end, and checks results against the same definitions your team already trusts. > As Øyvind Eraker, Senior Data Engineer at NBIM put it, “Structured context is the multiplier. With dbt as our source of definitions and lineage and MCP exposing that context across Snowflake and Claude, we can add new agent skills without re-plumbing governance.” Today, we’re taking that step with the launch of **dbt Agents**, powered by structured context and made accessible through the [remote dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/setup-remote-mcp). Together, they redefine how you build, manage, and consume data, so development speeds up, self-service expands, and data quality improves. > "In five years, your most reliable data developer will be an agent that commits code, passes tests, and explains its work. **It will do that because it stands on dbt**. With agents running on the dbt structured context layer and exposed through the remote dbt MCP server, this is how analytics will be built, managed, and consumed in this new era.” _- Tristan Handy, CEO, dbt Labs_ ## A standards-based foundation for AI: the remote dbt MCP server is GA Powering this launch is the structured context that already lives in dbt: the definitions, lineage, tests, and semantics that teams already use to make analytics trustworthy. This structured context is now universally accessible to AI systems through our version 1.0 remote and local [dbt MCP Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server). It turns dbt’s structured context into a standard interface that AI tools can safely call, across your warehouse. This means your governed models, metrics, tests, and lineage are available to clients like [OpenAI](https://github.com/dbt-labs/dbt-mcp/tree/main/examples), [Anthropic](https://docs.getdbt.com/docs/dbt-ai/integrate-mcp-claude), and [Cursor](https://docs.getdbt.com/docs/dbt-ai/integrate-mcp-cursor) for safer, more reliable AI. We are also bringing the [**dbt Fusion engine**](https://docs.getdbt.com/docs/fusion/about-fusion) to MCP through our new [Fusion MCP tools](https://docs.getdbt.com/docs/dbt-ai/about-mcp#fusion-tools-remote), so clients can use Fusion’s compiler, diagnostics, and metadata, giving agents precise awareness of how transformations work and what might break before they act. Fusion enables analytics agents to be reliable and accurate. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/e768386817c81317327ea765caf4c6f7a2c20580-1164x2346.png) And because strong execution only matters if it is secure, we are expanding authentication and control. The [**local dbt MCP server**](https://docs.getdbt.com/docs/dbt-ai/setup-local-mcp) now supports [**OAuth**](https://docs.getdbt.com/docs/dbt-ai/setup-local-mcp#dbt-platform-authentication), so teams can use their dbt login for secure access with a simple setup. **OAuth for the remote dbt MCP server is on the way**, bringing centralized authentication and auditability to your single endpoint per environment. With the MCP server standardizing access to the structured context layer that lives in dbt, and an expanding suite of MCP-native tools, you get a universal, governed bridge between AI and trustworthy data. ### Why dbt is the natural standard for agentic analytics In analytics, trust is critical. Without a structured context layer, AI tools can make mistakes that damage confidence. dbt solves this problem by providing clear rules for how data should be organized, tested, and used. We provide this context layer through several key tools: the dbt MCP server creates one secure gateway for AI to access properly structured context in any warehouse; dbt Fusion delivers fast data mapping, code checking, and validation to ensure work is reliable; and [MetricFlow (now open source)](https://www.getdbt.com/blog/open-source-metricflow-governed-metrics) ensures consistent measurements across all your tools. We're committed to continuing to strengthen this foundation and make it even more AI-friendly over time. And the momentum is real. More than 900 data teams and partners have adopted the dbt MCP server to prototype conversational analytics and agentic workflows across their stacks. > “Chat with your data’ works reliably when every answer is governed and explainable. dbt is our governance backbone, and MCP exposes that structured context to our chat experience so any NBIM employee can ask questions and get trusted answers. With dbt, Claude, and Snowflake powering our chat experience, adoption is 10× our previous catalog and tickets to core data teams are down. The same governed foundation now powers agentic workflows that flag anomalies, deliver morning briefs, and open small PRs.” — Øyvind Eraker, Senior Data Engineer, NBIM Because the remote server exposes your project through one endpoint, teams are moving from chat into agents that run end-to-end tasks across the analytics development lifecycle. They are deploying changes faster, rethinking BI with governed answers, and using the warehouse as durable memory for longer, multi-step agent work. With the dbt MCP server as the standard for structured data in AI, you can start building agents today. Simply connect the remote server, plug in your preferred model, and ship agents that deliver faster development, safer changes, and higher quality. But we want to make autonomous, AI-powered analytics truly seamless for data teams. ## Introducing native agents to dbt That’s why, on this foundation, we’re introducing [**dbt Agents**](https://www.getdbt.com/product/dbt-agents) in the dbt platform. These out-of-the-box agents help you build, manage, and consume data faster within dbt guardrails. As routine work shifts to agents, analytics engineers can focus on tuning agent behavior to accelerate their workflows and unlock safe self-service for the business. These agents use the same structured context layer your team already trusts so they act confidently without sacrificing governance. Today we’re introducing the following agents: - **Analyst agent (available in beta):** Available inside [dbt Insights (GA)](https://docs.getdbt.com/docs/explore/dbt-insights) - ask questions in plain English and get governed answers. The agent uses your dbt project context to generate SQL directly from your models, then executes it in your warehouse and returns verified results with definitions, tests, and lineage. When metrics are defined in the dbt Semantic Layer, the agent automatically resolves them for higher accuracy. Otherwise, it reasons from your dbt context to generate the right query. - **Discovery agent (in beta).** Find the right dataset or metric in plain language, along with clear definitions and why it’s trustworthy. Surfaces governed sources and dbt lineage so anyone can explore data with confidence. - **Observability agent (coming soon).** Helps you monitor jobs, pinpoint likely root causes, and cut resolution time. Designed to reduce noise and cut investigation time dramatically. - **Developer agent (coming soon).** Explains model logic, predicts downstream impact, flags duplicate logic, and validates changes before merge. Runs directly in VS Code or dbt Studio, powered by dbt’s context, so every change can be shipped quickly and safely. > “We are excited for dbt Agents to bring purpose-built automation that moves us from reactive tickets to proactive, agent-driven operations and spares us the overhead of bespoke bots.” - Øyvind Eraker, Senior Data Engineer, NBIM The momentum does not stop here. We will keep expanding agent capabilities, integrations, and multi-agent workflows so more of your analytics development lifecycle can run with confidence on the same governed foundation. ## What’s next The first generation of dbt Agents turns structured context into action, helping teams ship faster, catch issues earlier, and expand safe self-service. But this is only the start. These agents act as powerful force multipliers for data teams, automating routine tasks, enhancing collaboration, and allowing data engineers to focus on higher-value strategic work instead of repetitive operations. As these agents mature, they will fundamentally transform how data engineering teams operate. Rather than handling every data request, engineers will orchestrate and govern agent-driven workflows that scale their expertise across the organization. This means smaller teams can support larger data ecosystems while maintaining quality and governance, truly democratizing data without sacrificing reliability. In the future, we'll continue to deliver experiences that accelerate more stages of the analytics lifecycle and introduce multi-agent workflows that help you work faster. We'll keep deepening dbt's structured context layer through richer metadata intelligence in dbt Fusion, so agents become more aware of how data behaves across production systems. And we'll continue expanding through open standards - MCP, MetricFlow, metadata to ensure dbt integrates cleanly with the broader AI ecosystem. We envision a world where data engineers spend more time on innovation and less on maintenance, as agents handle increasingly complex tasks with minimal supervision. Start today by enabling the remote [dbt MCP Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server), try the [Fusion MCP tool](https://docs.getdbt.com/docs/dbt-ai/about-mcp#fusion-tools-remote), and request access to [dbt Agents](https://www.getdbt.com/product/dbt-agents). This is the beginning of the agentic era for analytics and we can't wait to see what you build with us! --- --- title: "Announcing open source MetricFlow: Governed metrics to power trustworthy AI and agents" description: "Open source MetricFlow delivers governed, portable metrics to power accurate AI, analytics, and agent workflows." url: "https://www.getdbt.com/blog/open-source-metricflow-governed-metrics" date: "2025-10-14" authors: ["Ryan Segar"] categories: ["Product"] --- # Announcing open source MetricFlow: Governed metrics to power trustworthy AI and agents Today, at [Coalesce 2025](https://www.getdbt.com/blog/coalesce-2025-rewriting-the-future), we announced our clear commitment to open, portable semantics for everyone. - We are [**open sourcing** MetricFlow](https://github.com/dbt-labs/metricflow), the technology that powers the dbt Semantic Layer, by moving it to the Apache 2.0 license. - We are building it in public with partners like [Snowflake and Salesforce](https://www.getdbt.com/blog/dbt-labs-affirms-commitment-to-open-semantic-interchange-by-open-sourcing-metricflow), so metric interoperability becomes the default across tools and clouds. - We are aligning **MetricFlow** with the goals of [**Open Semantic Interchange (OSI)**](https://www.snowflake.com/en/blog/open-semantic-interchange-ai-standard/) so any vendor can participate and any team can adopt without lock-in. We believe that open standards for semantic technology will power trusted AI systems for many years to come. MetricFlow is our contribution to that future. ## Why open semantics now? The semantic layer, a layer that defines business logic and its associated computations, has felt important to us for many years. We announced in 2021 that we were making investments in building our own semantic layer. In 2023, [we acquired Transform](https://www.getdbt.com/blog/dbt-acquisition-transform), the authors of MetricFlow, the leading technology in the space. In the years since then, we’ve seen consistent adoption of MetricFlow for BI and embedded analytics use cases. But as useful as semantic technology has been as a part of the BI stack, what no one saw coming was how critical it was about to become for AI. As it turns out, the semantic layer is [_the critical component_](https://www.getdbt.com/blog/why-your-ai-will-fail-without-a-semantic-layer) to build a bridge between AI and structured data. It allows AI and agents to apply correct business logic and [return reliable results](https://roundup.getdbt.com/p/semantic-layer-as-the-data-interface) every time. And the [recent](https://www.databricks.com/blog/whats-new-databricks-unity-catalog-data-ai-summit-2025) [investments](https://docs.snowflake.com/en/user-guide/views-semantic/overview) being made by major players in the space make it clear that this is now widely understood. Now in 2025, companies are racing to put trustworthy “chat with your data” in the hands of every employee. Semantic layers make that possible. Without them, AI emits raw SQL, guesses at joins, filters, time grains, and windows, and each model guesses differently. Numbers don’t align. Trust erodes. Adoption slows. Metrics should not be probabilistic or depend on an LLM guessing each calculation. **They should be deterministic.** MetricFlow makes this possible. It's the single most advanced way to take natural language business concepts and map them to code. This enables every LLM call—agent or a human, GPT5 or Claude 4—to measure everything exactly the same way. We feel so strongly about the industry standardizing on a shared approach that we’ve made some significant changes to the way that MetricFlow is both licensed and governed. 1. **Apache 2.0:** As of today, MetricFlow is now [fully available under an Apache 2.0 license](https://github.com/dbt-labs/metricflow). This includes all of the code required to define metrics and calculate them in multi-dialect SQL, as well as its metadata representation of semantic constructs. Other vendors can now build MetricFlow into their products in a first class way, building a high-quality bridge from business metrics to models and agents. 2. **Aligned with OSI:** MetricFlow will be governed and [maintained with OSI partner](https://www.getdbt.com/blog/dbt-labs-affirms-commitment-to-open-semantic-interchange-by-open-sourcing-metricflow) organizations like Snowflake and Salesforce. Our shared goal is for MetricFlow to power semantic interoperability between partner platforms. We will work collectively with OSI partners and with the community to evolve MetricFlow’s metadata representation of semantics to serve the needs of this initiative. We hope to eventually donate this technology to a leading open source foundation to ensure that semantic interoperability be powered by a truly open standard, maintained by the community in perpetuity. > “Defining metrics in MetricFlow is crucial for a unified source of truth. However, BI and AI tools often interpret metrics in their own ways. By leveraging open-source MetricFlow, OSI can ensure every tool is consistent in the consumption of metrics. This saves analysts time, streamlines audits, and gives users the flexibility to access data as they prefer, without constant code updates.” > — Rob Vicker, Data Analytics Architecture Director, EMC Insurance ## An open standard for the AI age [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) gives the community an open standard for semantic metadata and an extensible engine that turns semantic intent into fast, warehouse-specific SQL for your AI systems. How MetricFlow powers trustworthy AI: - **Built with the ecosystem.** Vendors and open source contributors co-maintain the metadata spec under OSI and in public to power metric + semantic interoperability across tools and clouds. - **Plan once, run anywhere.** You define your metrics and dimensions once. When a tool integrated with MetricFlow asks for a metric, the engine plans from that spec and compiles optimized, dialect-specific SQL. - **Explainability by design.** Every query is inspectable. You can see how joins, filters, time grains, and policies are applied, which makes reviews and audits faster. - **Performance aware.** The engine optimizes queries so teams do not trade correctness for speed. - **Made for complex calculations.** MetricFlow models joins, windows, cohorts, and semi-additive measures so hard problems are correct by default. Opening the engine means agents can ask for a metric by name and receive proven SQL. Then the [dbt Semantic Layer](https://www.getdbt.com/blog/build-centralize-and-deliver-consistent-metrics-with-the-dbt-semantic-layer) governs how definitions are authored, versioned, and accessed, so agents return the same answer everywhere. This is a step toward a world where AI and BI systems share one language for metrics so teams get the same answer everywhere and AI can be trusted at scale. **Have a complex calculation? Here's an example of how MetricFlow works:** Ask your AI: “Gross margin % by month for North America last quarter (net of discounts and returns, on the fiscal calendar.)” MetricFlow pulls the right joins across orders/discounts/returns and COGS, applies the region filter, aligns to fiscal months, and computes the ratio with matching numerator/denominator populations. It compiles readable, warehouse-specific SQL and returns results with lineage; change the definition, and every connected tool picks it up, no rewrites. ## Unlocking the next chapter of data and AI Open source MetricFlow and the dbt Semantic Layer turn shared semantics into everyday wins. Here’s what this means for your business: - **Less rework, fewer tickets.** A single shared definition cuts duplicate dashboarding and ticket escalations. In a dbt Labs replication, AI answered [83 percent](https://www.getdbt.com/blog/semantic-layer-as-the-data-interface-for-llms) of addressable natural language questions correctly via the dbt Semantic Layer, with several answered at 100% accuracy. - **Portability savings.** Interoperable semantics mean tool or warehouse changes without rebuilding metric logic, reducing migration costs and vendor lock-in premiums. - **Production-ready AI.** Replace prompt-only guesses with governed queries that respect joins, filters, time grains, and policies. [Findings](https://www.getdbt.com/blog/semantic-layer-as-the-data-interface-for-llms) reinforce that structured semantics materially improve accuracy over prompting alone. - **Faster resolution time.** Clear lineage and definitions shorten investigations, which cuts on-call and incident costs. - **Reduce compute.** MetricFlow plans efficient, warehouse-specific SQL, which reduces wasted scans and long-running queries. ## What comes next We will continue to invest in MetricFlow with the community and our partners to add optimizations, expand warehouse coverage, and improve explainability so every team can trust and verify each answer. In parallel, we are deepening our investment in enterprise-ready metric consumption through the dbt Semantic Layer. In the future, this includes stronger governance and security, role-based access, versioned changes with review, auditability, lineage you can trace, and reliable APIs and connectors so metrics flow safely into the tools your teams use. Together, MetricFlow and the dbt Semantic Layer accelerate AI adoption, ensuring every answer is governed, explainable, and consistent across tools and clouds. We are grateful to our customers, partners, and community members who pushed for a simpler path. This is one step on a longer journey, and we invite you to build it with us. - [Learn more](https://github.com/dbt-labs/metricflow) and stay up to date on MetricFlow. - [Learn more](https://www.snowflake.com/en/blog/open-semantic-interchange-ai-standard/) about the Open Semantic Interchange initiative. - [Learn how the dbt Semantic Layer](https://www.getdbt.com/blog/how-the-dbt-semantic-layer-works) brings these definitions to AI systems and tools. --- --- title: "State-aware orchestration now in Preview for Fusion projects" description: "Intelligent pipelines that save time and money—state-aware orchestration now in Preview for dbt Fusion projects." url: "https://www.getdbt.com/blog/announcing-state-aware-orchestration" date: "2025-10-14" authors: ["Ernesto Ongaro", "Reuben McCreanor"] categories: ["Product"] --- # State-aware orchestration now in Preview for Fusion projects **** As your data estate grows, costs and complexity grow with it. Across the dbt ecosystem, teams built over 33 billion models in the past year—more than four per person on Earth. Scale brings real costs in compute and cognitive load. Data growth does not have to mean complexity. Today, we're excited to announce that **state-aware orchestration** is now available in Preview for dbt platform projects running on Fusion. ## What is state-aware orchestration? State-aware orchestration is a Fusion-powered innovation that fundamentally rewrites how data pipelines work—by only running what actually needs to run. In a traditional dbt run, every model in your DAG is rebuilt regardless of whether its inputs have changed. ![a traditional dbt dag](https://cdn.sanity.io/images/wl0ndo6t/main/e3949f3808073f936ef53ccc47fe9de080ce1e73-1296x522.png) With state-aware orchestration enabled, dbt moves from being stateless to stateful. Instead of simply running exactly what was specified in a job, dbt maintains a real-time fingerprint of both model code and data state. The system pinpoints exactly which models need refreshing based on when new data is produced upstream or when code changes occur. Models without any upstream changes are automatically reused, eliminating unnecessary processing and only building models that will actually produce a different result. ![state-aware orchestration](https://cdn.sanity.io/images/wl0ndo6t/main/642dd8e79d5afa7a7742c6ab0a44572146cdf962-1296x522.png) ## Advanced configuration for fine-tuned control With state-aware orchestration, you get immediate savings by simply enabling it, but the real power comes with the advanced configuration options. Instead of rigidly scheduling jobs, you can now express your data freshness requirements directly in your dbt project. These [intent-based configurations](https://docs.getdbt.com/docs/deploy/state-aware-setup#advanced-configurations) allow you to specify exactly how fresh each model needs to be and under what conditions it should rebuild. For example, you can tell a model to: - **Wait for all upstream dependencies** with `updates_on: all` - only rebuilding when all source data is fresh - **Define maximum staleness windows** with `build_after: 6h` - ensuring models rebuild at least every six hours regardless of upstream changes - **Prioritize critical models** with varying freshness requirements based on business importance The system intelligently orchestrates your pipeline based on these declared intents rather than rigid schedules. This means models will only run when they'll actually produce different results, delivering immediate cost savings without compromising data freshness. ![tuned SAO ](https://cdn.sanity.io/images/wl0ndo6t/main/2d7984474f24c17fa46b427ade96519979bdb6b8-1296x522.png) ## Smarter testing, same standards With column‑aware testing, checks only run when the underlying data actually changes, eliminating unnecessary compute. State‑of‑the‑art static analysis identifies tests that are guaranteed to pass and safely skips them, while intelligent batching combines multiple checks into fewer, more efficient operations. The result is the same rigorous quality bar with with far fewer redundant table scans and compute, and no loss in accuracy. ## The business impact: Cost savings without compromise For data leaders and platform owners who need to control costs while delivering high-quality insights, state-aware orchestration delivers measurable bottom-line benefits: - **Immediate ROI with zero-implementation effort:** Simply flip a switch to enable state-aware orchestration. No rewrites or restructuring: just **immediate 10% reduction in compute costs based** on our beta results. - **Data freshness on your terms:** [Fine-tune configurations](https://docs.getdbt.com/docs/deploy/state-aware-setup#advanced-configurations) to meet your business SLAs without over-scheduling. Declare your freshness requirements, and Fusion handles the rest: **delivering an additional 15%+ in cost savings.** - **Smarter testing:** Fusion intelligently skips tests when they do not need to run. The result is **an estimated 4% additional annual cost savings.** - **Total impact: 29%+ reduction in annual dbt-related compute costs** requiring only minor changes. ![estimated cost savings](https://cdn.sanity.io/images/wl0ndo6t/main/fc04e411d90785052e5d5bb77d6e4ed7e8dc16c4-1920x1080.png) ## Real-world results from dbt Labs At dbt Labs, we've implemented state-aware orchestration in our own internal analytics project with impressive results. With a data estate containing around 1,500 models and many jobs on various schedules, we saw immediate impact. After enabling Fusion with state-aware orchestration, we immediately saw **around 35% of our models being reused daily, resulting in 9% cost savings on our compute bill,** just from turning it on, with no tuning required. Then things got interesting. When we added tuned configurations to optimize our data freshness requirements, we achieved even more dramatic results. > What we ended up with was a much simpler configuration with shockingly impressive results: with fine-tuned configurations built on state-aware orchestration, we hit an additional 55% cost savings on our dbt workloads, in addition to the 9% we achieved when we turned on Fusion, for a total of 64% cost reduction on our compute bill. This is an incredible result. > — Ken Oster, VP of Data, dbt Labs ![savings for dbt labs](https://cdn.sanity.io/images/wl0ndo6t/main/3700c48ee0a0661cf877877cc5cb7ee0ba7c7716-1920x1080.png) When we implemented tuned configurations to optimize our data freshness requirements, the results were impressive and surprisingly easy to achieve. Here's what we did: - **Simplified freshness configurations:** We created just two freshness tiers - daily (our global default for analytical purposes) and hourly (for operational needs that require more frequent updates). - **Streamlined job management:** We reduced our job complexity from numerous over-scheduled jobs to just two: one hourly job that refreshes both daily and hourly data appropriately, and a weekend job for cleanup and targeted backfills. Many teams will ask how difficult this transformation was to implement. The answer surprised even us: just a few lines of YAML at the project definition level, some model-specific overrides, and configuring the two jobs. Here's an example of a model-specific override: ```yaml -- fct_dbt_invocations.sql {{ config( materialized = 'incremental', unique_key = 'invocation_id', freshness = { 'build_after': { 'updates_on': 'all' } } ) }} ``` And here's what our daily grain configuration looks like today: ```yaml --dbt_project.yml models: +freshness: build_after: count: 1 period: day updates_on: any ``` With this simple setup, we dramatically reduced both our cognitive overhead and our compute bill. ## Developer experience benefits Beyond the business value, state-aware orchestration transforms how data teams work: - **Faster development cycles:** Reduced run times mean quicker iterations and faster time-to-insight. - **Simplified job management:** No more wrestling with complex scheduling. Define freshness requirements once, and Fusion handles the orchestration. Jobs are also aware of each other, preventing concurrent writes to the same table, eliminating a significant pain point for many data teams. - **Fewer false alarms:** With more precise orchestration comes less noise: reducing alert fatigue and helping teams focus on actual issues. - **More time for new initiatives:** When your team spends less time managing jobs and monitoring costs, they can focus on delivering business value through data. ## Get started State-aware orchestration is available in Preview for Enterprise customers running their projects on Fusion. Enabling it is as simple as flipping a switch in your project settings. With state-aware orchestration, your data infrastructure can now be as intelligent as the insights it delivers: running smarter, not harder. Ready to rewrite the rules for your data platform? [Learn more in our documentation](https://docs.getdbt.com/docs/deploy/state-aware-about) or contact your account team to get started. --- --- title: "Reverse ETL vs ETL: What's the real difference?" description: "Reverse ETL syncs data back into tools for action. It’s not just ETL in reverse—here’s how it’s different and when to use it." url: "https://www.getdbt.com/blog/reverse-etl-vs-etl" date: "2025-10-13" authors: ["Joey Gault"] categories: ["Pulse"] --- # Reverse ETL vs ETL: What's the real difference? [Traditional ETL](https://www.getdbt.com/blog/extract-load-transform) follows a well-established pattern: extract data from operational systems, transform it for analytical purposes, and load it into a data warehouse or data lake. This process consolidates disparate data sources into a centralized repository optimized for reporting and analysis. The transformations typically involve cleaning, standardizing, and aggregating data to support business intelligence use cases. Reverse ETL operates in the opposite direction. It takes transformed data that already exists in your data warehouse and syncs it back to operational systems where business users can act on it. Rather than consolidating data for analysis, reverse ETL distributes analytical insights to drive operational workflows. This directional difference reflects a fundamental shift in how organizations think about data architecture. Traditional ETL treats the data warehouse as the final destination: a place where data goes to be analyzed. Reverse ETL recognizes that the warehouse should be a hub that not only receives and processes data but also distributes insights back to the systems where work actually happens. ## The transformation layer distinction The nature of transformations in reverse ETL differs substantially from traditional ETL. In standard ETL processes, transformations focus on data quality, standardization, and analytical modeling. You're typically cleaning messy source data, resolving schema conflicts, and creating dimensional models suitable for reporting. Reverse ETL transformations serve a different purpose. The heavy lifting of data cleaning and modeling has already been completed by [tools like dbt ](https://www.getdbt.com/product/what-is-dbt)during the initial transformation process. Instead, reverse ETL transformations focus on adapting already-clean data to meet the specific requirements of destination systems. These adaptations might involve renaming fields to match API expectations, casting data types to conform to external system requirements, or creating derived fields that make sense in the context of the destination tool. For example, when syncing customer data to an email marketing platform, you might need to create a `buyer_type` segment based on purchase history or extract date fields to month-level granularity for campaign targeting. The key insight is that reverse ETL transformations are typically lightweight and purpose-specific, building on the solid foundation of data quality and business logic established during the initial transformation phase. This is why dbt's approach to reverse ETL emphasizes creating dedicated export models that supplement existing fact and dimensional tables rather than rebuilding transformation logic from scratch. ## Architectural and operational differences The nature of transformations in reverse ETL differs substantially from traditional ETL. In standard ETL processes, transformations focus on data quality, standardization, and analytical modeling. You're typically cleaning messy source data, resolving schema conflicts, and creating dimensional models suitable for reporting. Reverse ETL transformations serve a different purpose. The heavy lifting of data cleaning and modeling has already been completed by tools like dbt during the initial transformation process. Instead, reverse ETL transformations focus on adapting already-clean data to meet the specific requirements of destination systems. These adaptations might involve renaming fields to match API expectations, casting data types to conform to external system requirements, or creating derived fields that make sense in the context of the destination tool. For example, when syncing customer data to an email marketing platform, you might need to create a "buyer_type" segment based on purchase history or extract date fields to month-level granularity for campaign targeting. The key insight is that reverse ETL transformations are typically lightweight and purpose-specific, building on the solid foundation of data quality and business logic established during the initial transformation phase. This is why dbt's approach to reverse ETL emphasizes creating dedicated export models that supplement existing fact and dimensional tables rather than rebuilding transformation logic from scratch. ## Business context and use cases The business context surrounding reverse ETL fundamentally differs from traditional ETL. [Traditional ETL](https://www.getdbt.com/blog/extract-transform-load) supports analytical use cases: generating reports, building dashboards, and enabling data exploration. The end users are typically analysts, data scientists, and business intelligence professionals who are comfortable working with data in its analytical form. Reverse ETL serves operational use cases. It enables marketing teams to create personalized campaigns based on customer lifetime value calculations, allows sales teams to prioritize leads using predictive scoring models, and helps customer success teams identify at-risk accounts using churn prediction algorithms. The end users are operational teams who need data integrated into their daily workflows rather than presented in separate analytical tools. This operational focus creates different requirements for data freshness, reliability, and usability. Marketing campaigns might need customer segments updated daily or even hourly. Sales teams require lead scores that reflect the most recent behavioral data. Customer success teams need real-time visibility into account health metrics. The self-service aspect of reverse ETL also distinguishes it from traditional ETL. While traditional ETL typically requires technical expertise to implement and modify, reverse ETL tools are designed to enable business users to configure their own data syncs once the underlying data models are established. This democratization of data access represents a significant shift from the centralized, IT-controlled approach of traditional ETL. ## The role of existing transformations One of the most significant differences between reverse ETL and traditional ETL lies in how they relate to existing data transformations. Traditional ETL creates the initial transformed datasets, implementing business logic, data quality rules, and analytical models from scratch. Reverse ETL leverages transformations that have already been completed, tested, and validated. When using dbt for [data transformation](https://www.getdbt.com/blog/data-transformation-vs-etl), reverse ETL processes can build directly on the fact and dimensional models that have undergone rigorous testing, peer review, and documentation. This creates a different risk profile and development approach. The export models created for reverse ETL should not duplicate the heavy transformation logic found in core analytical models. Instead, they should focus on the specific formatting and field requirements needed for destination systems. This architectural principle ensures that business logic remains centralized and version-controlled while enabling flexible distribution to operational systems. This relationship to existing transformations also affects governance and lineage tracking. Reverse ETL processes inherit the data quality and business logic validation from upstream transformations, but they also create new dependencies and exposure points that must be documented and monitored. ## Integration complexity and tooling The tooling ecosystem around reverse ETL reflects its distinct requirements and challenges. While traditional ETL tools focus on extracting data from various sources and loading it into analytical systems, reverse ETL tools must navigate the complex landscape of operational system APIs and integration requirements. Modern reverse ETL platforms like Hightouch, Census, and Rudderstack provide pre-built connectors for dozens of operational systems, handling the intricacies of authentication, rate limiting, and data formatting for each destination. This specialization reflects the unique challenges of operational system integration that differ significantly from the data warehouse loading patterns of traditional ETL. The integration with transformation tools like dbt also creates new architectural patterns. Rather than replacing ETL processes, reverse ETL extends them, creating a bidirectional data flow that serves both analytical and operational use cases. This requires careful coordination between transformation schedules, data freshness requirements, and operational system constraints. ## Conclusion While reverse ETL and traditional ETL share some surface-level similarities in terms of data movement and transformation, they represent fundamentally different approaches to data architecture and business value creation. Traditional ETL consolidates data for analysis, while reverse ETL distributes insights for action. Traditional ETL creates analytical datasets from raw sources, while reverse ETL adapts analytical datasets for operational use. The emergence of reverse ETL reflects the maturation of data architecture thinking. Organizations are moving beyond viewing the data warehouse as a final destination and instead treating it as a central hub that both receives and distributes data. This shift requires new tools, new processes, and new ways of thinking about data transformation and governance. For data engineering leaders, understanding these distinctions is crucial for building effective data architectures that serve both analytical and operational needs. Rather than viewing reverse ETL as simply "ETL in reverse," it's more accurate to see it as a complementary capability that extends the value of existing data transformations into operational workflows. The question isn't whether reverse ETL is just ETL: it's how to effectively integrate both approaches into a cohesive data architecture that maximizes the value of your organization's data investments. This integration requires careful consideration of transformation patterns, governance frameworks, and operational requirements that reflect the unique characteristics of each approach. ## Reverse ETL FAQs **What is reverse ETL?** Reverse ETL is a process that takes transformed data from your data warehouse and syncs it back to operational systems where business users can act on it. Unlike traditional ETL which consolidates data for analysis, reverse ETL distributes analytical insights to drive operational workflows, treating the data warehouse as a hub that both receives and distributes data rather than just a final destination. **How does reverse ETL work?** Reverse ETL works by leveraging data transformations that have already been completed, tested, and validated in your data warehouse. It creates lightweight, purpose-specific transformations that adapt clean analytical data to meet the requirements of destination systems, such as renaming fields to match API expectations, casting data types, or creating derived fields. The process focuses on distributing fresh insights to operational systems through pre-built connectors that handle authentication, rate limiting, and data formatting for each destination. **What's the difference between ETL & reverse ETL?** The key differences lie in direction, purpose, and transformations. Traditional ETL extracts data from operational systems, transforms it for analytical purposes, and loads it into a data warehouse for reporting and analysis. Reverse ETL operates in the opposite direction, taking already-transformed warehouse data and syncing it to operational systems for business action. While traditional ETL focuses on data cleaning and analytical modeling, reverse ETL performs lightweight adaptations of clean data to meet specific destination system requirements. --- --- title: "dbt Labs + Fivetran: Open data infrastructure for analytics and AI" description: "Why we’re joining forces with Fivetran to build the future" url: "https://www.getdbt.com/blog/dbt-labs-and-fivetran-merge-announcement" date: "2025-10-13" authors: ["Tristan Handy"] categories: ["Company"] --- # dbt Labs + Fivetran: Open data infrastructure for analytics and AI This is unlike any anything I’ve ever written. As I sit down to write, there are just so many emotions. I’ll reflect on some of this later in this post, but for now, I won’t bury the headline. #### Today we are announcing that we are merging with Fivetran. The merged company combines for ~$600 million in ARR with well north of 10,000 customers. The majority of companies who use the cloud for their data infrastructure use one or both of these products already. This kind of scale and reach is surreal to think about, and it brings me much closer to my own personal goal of the past 6 years: to create the organization that can invest in and steward dbt and its community over the very long term. Behind the numbers though, the real story here is what this means for the dbt Community, for our customers, and for the future of the industry. That’s what I want to spend most of this post on. But I want to start with the motivation for the merger and why I’m personally so excited about it. ## The backstory I’ve known [George](https://www.linkedin.com/in/george-fraser-a0219230/) and [Taylor](https://www.linkedin.com/in/taylorwcbrown/), the two co-founders of Fivetran, for a decade. Starting Fishtown Analytics back in 2016, Fivetran was one of our best partners. Our stack at the outset was pretty much always Redshift + Fivetran/Stitch + Looker/Mode. We brought Fivetran dozens of customers from 2016-2019. For a while, Taylor was our partner manager. I filed support tickets and asked for new connectors. As a partner, Fivetran was never anything less than fantastic. Then came the modern data stack buzz era. Both companies competed to out-fundraise each other for a couple of years ;) George and I were on a million panels together, and every single conference I went to I would always wander over to the Fivetran booth to catch up with Taylor and George. I think it’s fair to say that no two companies in the modern data stack have been more joined at the hip. The products were literally built to be used together. The companies have grown up together, the relationships at all levels of both companies are longstanding, and the respect is mutual. It has even seemed obvious for a while that it made sense to merge—the questions were really more about when and how to do it, and whether we could figure out how to make it work. I remember a conversation about this with Martin Casado at a16z, who is on the board of both companies, a few years ago. His response: “You and George have to figure that out for yourselves, but from my perspective it is just such an obvious win for everyone involved.” I agree—this has been a good idea for a while. And the answer to “why now” is really about maturity. Both companies have reached a stage where we have a level of predictability around our existing businesses, and we were both thinking about what was next. We decided that, as we traveled that road ahead, we’d be far better positioned to do it together. ## What we can build—together Right out of the gate, let me just say something basic, but important. **dbt will still be dbt. Fivetran will still be Fivetran.** We’re not planning to rename either product, not planning any disruptive product changes, we’ll continue to provide the same types of support for the dbt Community, and are aggressively executing against our respective product roadmaps. In short: _if you use and love either product today, this merger will be non-disruptive_. The more interesting question, really, is: what can we now build together that we could not build alone? This is where I get really excited. And it really all comes down to two principles that both products hold dearly: simplicity and openness. ### Simplicity From the beginning, dbt and Fivetran have both attempted to make the sometimes-arcane practice of data engineering _simpler_. Fivetran’s integrations have always been push-button: flip a switch and get data, reliably, with low latency. dbt’s opinionated programming model allows less technical data practitioners to author production-grade data pipelines. George likes to think about this like electricity: flip a switch, get data. I like to think about it like the Apple ecosystem: it all just works together, no duct tape. We both agree that far too much time is wasted in data engineering toil that creates zero business value, and we’ve built products that make the work of data engineering simpler, more automated, and more accessible. ### Openness Both products have also always been focused on being open and interoperable. That is to say, allowing users to build data pipelines that work with any underlying data platform. Both dbt and Fivetran abstract away the details of the underlying data platform and the cloud it runs on, allowing users greater strategic flexibility with one of your most core IT assets. There was a time when this open pattern was almost assumed, but as each of the data platforms has begun to invest more heavily in building their own first-party tooling, this trait becomes ever-more-important. ## Open data infrastructure Many users think of Fivetran as ingestion and dbt as transformation, but both products have grown meaningfully over the last several years and at this point cover a very significant footprint. The combination will form the most complete, most widely-deployed open data infrastructure platform on the market. Here’s what it looks like when both products come together: ![dbt and Fivetran together for te future of data](https://cdn.sanity.io/images/wl0ndo6t/main/982bb0f03c77be352577141e5554d00757132063-1974x1152.png) This architecture is what we are calling “open data infrastructure.” Open data infrastructure describes an infrastructure that is **pluggable**, relies on **integration via standards**, does not assume the usage of any one particular compute engine, and does not assume that solutions will be duct taped together from many individual products and vendors. It is more integrated than the modern data stack, and it allows for greater user choice than the all-in-one data platforms. I’ve written a whole lot more about open data infrastructure and the journey that our ecosystem has been on over the past several years [here](https://www.getdbt.com/blog/what-is-open-data-infrastructure). And if you’re interested in hearing more about why I think this is so critical to the future of analytics and AI, [tune into my Coalesce keynote tomorrow](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/regProcessStep1:ca2ac50c-de10-4af1-ab25-94141a9484f3?rp=8916244e-96f3-4641-bbce-411935d4bd29). ## Open source If you are a dbt user, what you will likely want to know is: how does this impact dbt open source projects—most importantly, Core and Fusion? You can read [George’s own post](https://www.fivetran.com/blog/the-data-platform-for-the-next-10-years) on this topic, but part of what made this merger such an obviously good idea to me was Fivetran’s complete alignment in how it sees this. Our collective commitments are as follows: - dbt Core and Fusion will both continue to be shipped under their current licenses. - We will continue to actively maintain dbt Core. - We will continue to support the dbt Community, via Slack, meetups, etc. To give credit where credit is absolutely due: Fivetran has been a long-time contributor to the dbt open source ecosystem, authoring over 100 packages with OSS licenses that are used by thousands of teams. They have 5 full time staff now dedicated to this work. I think that the extent of their contributions here over the years has gone under-appreciated; it’s likely that they have contributed more OSS dbt code than any other organization outside of dbt Labs. I know that at least some of you reading this will be concerned that Fivetran’s historical focus on building proprietary software might make its way into the ethos of dbt, but we anticipate quite the opposite. Instead, George and I are excited to bring _more openness to the Fivetran ecosystem_ with the combined company. We are already thinking about what existing software can be shipped under an OSS license (connectors? connector development kit?) and what standards we can invest in (or create). While we don’t know all of the answers yet, we absolutely believe that open source is a win-win for the community and the company, and _we anticipate building more, not less, of it_. Figuring out shared OSS strategy will be one of my biggest charters in the combined company. My title will be Co-Founder and President, I’ll have a seat on the board, and I’ll be responsible for our community and open source strategy. So: I have every intention of continuing our decade-long trajectory of innovation in the open. ## So much more to do I mentioned at the beginning that there are a lot of feelings swirling around as I write. This is true. A lot of gratitude for a lot of people. Some natural stress associated with any big change. Maybe a little bit of nostalgia hiding in there somewhere :) And a few things that I’ll probably have to work out with my coach! But this is not a freakin’ Oscar speech. This thing isn’t done. We have more work to do, I have more work to do. Open standards and AI both completely rewrite the script of how analytics is done, and I’m excited to step into this new future. The combined company is well-positioned for this future—fine—but for the moment I just want to say that _I am excited about this moment as a data practitioner_. There is so much to figure out, so much to play with. In the 2010’s we got fast analytical databases in the cloud and we got to work a lot more like software engineers. In the 2020’s we’re getting heterogeneous analytical compute and _computers that can think_. The world of the data practitioner will just be _so much different in 2030_. **The best practices for this next era have not been figured out.** We don’t yet know what’s possible. _And that’s when it’s fun._ Today feels like things felt back in 2015-2017; we get to play with these cool new toys that are starting to work and figuring out what we can do with them. I am excited to get my hands dirty and figure it out right alongside of you. And I plan on sharing, in public, throughout that journey. As always, do not hesitate to ping me on dbt Slack—I’m @tristan. --- --- title: "Fivetran and dbt Labs Unite to Set the Standard for Open Data Infrastructure" description: "Together, Fivetran and dbt are simplifying enterprise data management with a unified foundation." url: "https://www.getdbt.com/blog/dbt-labs-and-fivetran-sign-definitive-agreement-to-merge" date: "2025-10-13" authors: ["Elaine Green"] categories: ["Press"] --- # Fivetran and dbt Labs Unite to Set the Standard for Open Data Infrastructure **OAKLAND, Calif., October 13, 2025 — **Fivetran, the global leader in automated data movement, signed a definitive agreement today to merge in an all-stock deal with dbt Labs, the pioneer of modern data transformation. Under the agreement, George Fraser will serve as CEO of the unified company, and dbt Labs CEO Tristan Handy will serve as co-founder and President. Following close of the transaction, the combined company will be approaching $600M in annual recurring revenue (ARR). Fivetran and dbt share a long-standing belief that data infrastructure should be open, automated, and effortless. The combination brings complementary strengths together to deliver open data infrastructure — which unifies data movement, transformation, metadata, and activation while preserving freedom of choice for analytic compute and AI. As part of this transaction, the company is committed to keeping dbt Core open under its current license and maintaining it with and for the community, ensuring its development remains vibrant. “This is a refounding moment for Fivetran and the broader data ecosystem,” said George Fraser, CEO of Fivetran. “As AI reshapes every industry, organizations need a foundation they can trust — one that is open, interoperable, and built to scale with their ambitions. Our admiration for dbt and its remarkable community runs deep — this is about bringing together the best of both worlds to accelerate innovation and create lasting impact across the data community.” The unification of Fivetran and dbt sets the standard for open data infrastructure, a new approach that reduces engineering complexity by automating data management end to end. It works across any compute engine, catalog, BI tool, or AI model, is built on open standards like SQL and Iceberg, and remains flexible so organizations avoid lock-in and can scale with future workloads. “dbt has always stood for openness and practitioner choice,” said Tristan Handy, founder of dbt Labs. “For nearly a decade, I’ve worked to build data infrastructure that supports every engine, every format, every model, every tool, that acts as an abstraction layer across an entire ecosystem. By merging with Fivetran, we can accelerate that mission and deliver the open data infrastructure that practitioners and enterprises need in the AI era.” The transaction has been approved by the Boards of Directors and shareholders of both companies. Finalization of the merger remains subject to customary closing conditions, including regulatory approvals. Until then, Fivetran and dbt Labs will continue to operate as separate, independent companies. Qatalyst Partners served as exclusive financial advisor to Fivetran, Deloitte & Touche LLP as due diligence advisors, Wilson Sonsini Rosati & Goodrich P.C. served as lead legal advisor, and DLA Piper LLP also served as Fivetran’s legal advisor. Morgan Stanley served as the exclusive financial advisor to dbt Labs and Latham & Watkins LLP served as its legal advisor. **About Fivetran** Fivetran, the global leader in automated data movement, is trusted by companies like OpenAI, LVMH, Pfizer, Verizon, and Spotify to centralize data from SaaS applications, databases, files, and other sources into cloud destinations, including data lakes. With high-performance pipelines, seamless interoperability, and enterprise-grade security, Fivetran empowers organizations to modernize their data infrastructure, power analytics and AI, ensure compliance, and achieve transformative business outcomes. Learn more at [Fivetran.com](http://fivetran.com). **About dbt Labs** Since 2016, dbt has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 80,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at [getdbt.com](http://getdbt.com), and follow dbt on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "The era of open data infrastructure" description: "Why the dbt + Fivetran merger is the inflection point for AI, Iceberg, and the enterprise" url: "https://www.getdbt.com/blog/dbt-labs-and-fivetran-product-vision" date: "2025-10-13" authors: ["Ryan Segar"] categories: ["Company"] --- # The era of open data infrastructure For years, we have treated data like a promise that somehow never quite arrives. We built warehouses, lakes, and then a “modern data stack,” each step giving us new power but also exposing new seams. Then came the big platforms, offering simplicity but introducing lock-in. The net result is an uncomfortable truth: [enterprises still put only around a third of their data to work](https://hbr.org/sponsored/2025/09/is-your-enterprise-data-strategy-ready-for-the-age-of-intelligence), while roughly half remains totally dark; collected, stored, and never used. That is not a rounding error. It is the ceiling on what your models can know, the drag on every digital initiative, and the reason [AI pilots stall before they scale](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/). Two forces make this the right moment to break that ceiling. First, AI has moved from novelty to necessity. It now sits in decision paths where error is costly and transparency is non-negotiable. Second, Apache Iceberg has matured from a great idea into the neutral substrate enterprises needed all along. Iceberg lets data live once and work everywhere by decoupling storage from compute, enabling safe schema evolution and ACID reliability across engines and catalogs. This is not hypothetical alignment: [Snowflake open-sourced the Polaris Catalog for Iceberg under Apache 2.0](https://www.snowflake.com/en/blog/polaris-catalog-open-source/), and [Databricks announced full Iceberg support governed in Unity Catalog](https://www.databricks.com/blog/announcing-full-apache-iceberg-support-databricks). The industry has, in effect, agreed on the table standard AI can trust. **** ## **Why this is the moment for open data infrastructure** **This is the context for the [merger of dbt Labs and Fivetran](https://www.getdbt.com/blog/dbt-labs-and-fivetran-sign-definitive-agreement-to-merge).** We are not combining to create another all-in-one platform. We are unifying the backbone of a new category: open data infrastructure. The mandate is straightforward: movement, transformation, metadata, and activation must operate as a managed continuum. The fabric must be open by design, across engines, catalogs, BI tools, and LLMs, and it must deliver reliability you can measure. When that fabric is built on Iceberg, AI finally sits on a governed, portable foundation. Consider what changes, concretely, when you put these pieces together. A global retailer uses Fivetran to land operational events from hundreds of SaaS apps and OLTP systems into Iceberg tables with service-level guarantees on delivery and uptime. dbt orchestrates contracts, tests, and transformations that compile into governed models and a semantic layer. The same definitions feed Tableau and Power BI, but also AI copilots that answer “What is net revenue by cohort for the last two weeks?” and use lineage to show exactly which inputs were read and when. If a source schema shifts at noon, Fivetran’s CDC and state-aware syncs update only what changed, dbt’s tests fail fast, and the copilot offers a patch PR with a clear diff. The answer at 12:05 is both fresh and explainable. None of this requires a re-platform when you decide to run the same workloads on Snowflake today and Databricks tomorrow, or to share a governed slice via Delta Sharing with a partner while your internal teams query the same Iceberg tables from Trino. This is portability without penalty, and it rewires what “time to trustworthy answer” means inside any organization. Or take a bank that must prove how a risk model was trained. Iceberg snapshots give you time-consistent data slices, dbt’s lineage and documentation preserve feature provenance, activation from Fivetran (Census) pushes validated aggregates to downstream apps. When an auditor asks “Which version of the ‘exposure’ metric did the model ingest on March 31?” the system can reproduce it, because governance travels with the data. Unity and Polaris interoperate with the same Iceberg tables so the bank’s mixed compute estate is a feature, not a liability. The same architecture supports low-latency agents that need concurrency spikes, because you can scale compute independently of storage while keeping semantics stable. This is also a moment of standards. AI needs a clean, universal way to connect to governed data, tools, and workflows. The Model Context Protocol (MCP) is emerging as a practical standard for that as a “USB-C for AI apps,” and it is already gathering real traction in the ecosystem. A semantic layer that exports contract-backed metrics and lineage into MCP gives agents the context to be right and the paper trail to be trusted. By combining dbt and Fivetran’s complementary solutions, the transaction will create a more complete solution for data workflows that will better serve customer needs.Fivetran created the standard for automated data movement at scale, now with a published 99.9% uptime SLA for core services and data delivery. It pairs high-volume database replication methods including log-based CDC from its HVR heritage with hundreds of managed connectors so operational change is captured with low toil and predictable freshness. dbt created the standard for analytics engineering and made transformation reproducible, testable, and explainable; its semantic layer turns metric drift from an inevitability into an anti-pattern. Together, after closing, these will become one fabric that is provably reliable and measurably portable. ## Open by design Because clarity matters during transitions, let me state our commitments in unambiguous terms. dbt Core remains open source under its current license and will continue to be supported indefinitely. The dbt Fusion engine remains source-available under its current license and will continue to be supported indefinitely. [We published these commitments publicly](https://www.getdbt.com/blog/new-code-new-license-understanding-the-new-license-for-the-dbt-fusion-engine), and we are standing by them. These technologies are not just components in our architecture; they are flagships of the open movement and essential to the portability customers demand. Why insist on openness when a single platform could make the diagrams look neat? Because the market math is unforgiving. AI evolves faster than any one vendor’s roadmap. Your workloads will span warehouses, lakehouses, and specialized engines. Your models will include hosted LLMs and open-weight models that you fine-tune. Your governance surface will extend across Unity, Polaris, Glue, and whatever comes next. Open data infrastructure assumes this heterogeneity and turns it into a strength. Iceberg makes the storage layer common. Catalogs and sharing protocols make discovery and access universal. Movement, transformation, and semantics make the data verifiably right, no matter which engine reads it. If you want a simple yardstick to hold us to, use these. How long from source change to governed, queryable metric or feature, with lineage intact? How often are freshness SLOs met without brute-force recompute? How consistently does a metric return the same value in BI and in an AI copilot? How easily can you validate the same Iceberg tables across two engines and two catalogs without copying data? When these numbers move in the right direction, your 32 percent data utilization becomes 40, then 60, and the dark half of your estate starts to light up. Some will ask whether we are replacing one center of gravity with another. The answer is no. We are providing the roads: a fully managed, open, reliable infrastructure that makes choice practical and safe. Until our transaction closes, it is business as usual for customers, no changes to your contracts or support, and the work of unifying this fabric will proceed in the open, with the community that got us here. After close, our responsibility is to keep the roads smooth, the standards open, and the service levels high so that AI can scale on something worthy of your ambitions. The era of open data infrastructure has arrived. AI finally has the foundation it needs. Iceberg is the neutral substrate. dbt and Fivetran are the living system on top that moves, shapes, explains, and activates your data with the reliability an enterprise can sign its name to. Most of the world’s data has been asleep. It is time to wake it, and to do it in a way that you control. --- --- title: "What is open data infrastructure? How is it different from the modern data stack?" description: "And why we think it’s the right evolution for our industry in the age of Iceberg and AI" url: "https://www.getdbt.com/blog/what-is-open-data-infrastructure" date: "2025-10-13" authors: ["Tristan Handy"] categories: ["Company"] --- # What is open data infrastructure? How is it different from the modern data stack? Today, we announced that we are merging with Fivetran. [In my blog](https://www.getdbt.com/blog/dbt-labs-and-fivetran-merge-announcement), I shared the motivation behind this merger and our shared vision for building an **open data infrastructure**. But what exactly does that term mean…and why does it matter? In this post, I’ll dive deeper into what an open data infrastructure is, why it’s the right evolution of the “modern data stack” in the age of Iceberg and AI, and how dbt Labs and Fivetran together plan to make that vision real for data teams everywhere. ## The modern data stack solved some problems, and created new ones I hate neologisms for the sake of neologisms. No one needs a tech company to introduce new terms of art purely for marketing. If we’re going to use this phrase it’s because it actually means something. So let’s put it to the test. And let’s start with the term “modern data stack” (MDS): what it meant and how it’s failing to scale in the new era. Data practitioners are widely familiar with the term “modern data stack” at this point. The cloud, and the MDS, transformed data over the past decade. Previously, data had been slow, clunky, and expensive. Its success stories made for great headlines, but in the trenches things moved very slowly, cost too much, and working in the field was not particularly…fun. The MDS changed this. This practitioner-focused tooling was lightweight but production-grade. It allowed users to move fast, with SQL as the primary standard, and bring the best practices of software engineering and DevOps into data at scale for the first time. A once-sleepy profession, the data engineer, and a brand new one, the analytics engineer, became the focus of innovation for a fast-moving ecosystem. As the MDS gathered steam, vendors popped up to solve every conceivable problem, and customers wrestled with constructing end-to-end solutions from a dozen or more tools. More time was spent debating what tools to use and how to integrate them than was spent in actually working towards business goals. While standards _helped_ with interoperability, they couldn’t solve everything. In particular, complex problems like data quality, governance, and metadata management never seemed to quite get solved. ## The rise of “all-in-one” data platforms As a result, customers became frustrated with the tool-integration challenges and the inability to solve the larger, cross-domain problems. Customers began demanding more integrated solutions—asking their existing vendors to “do more” and leave in-house teams to solve fewer integration challenges themselves. Vendors saw this as an opportunity to grow into new areas and extend their footprints into new categories. This is neither inherently good nor bad. End-to-end solutions can drive cleaner integration, better user experience, and lower cost. But they can also limit user choice, create vendor lock-in, and drive up costs. The devil is in the details. In particular, the data industry has, during the cloud era, been dominated by five huge players, each with well over $1 billion in annual revenue: Databricks, Snowflake, Google Cloud, AWS, and Microsoft Azure. Each of these five players started out by building an analytical compute engine, storage, and a metadata catalog. But over the last five years as the MDS story has played out, each of their customers has asked them to “do more.” And they have responded. Each of these five players now includes solutions across the entire stack: ingestion, transformation, notebooks and BI, orchestration, and more. They have now effectively become “all-in-one data platforms”—bring data, and do everything within their ecosystem. These platforms often do a decent job of delivering on the integrated vision—there is typically less duct tape required than in a traditional do-it-yourself MDS solution. But there are tradeoffs: - Costs can be high and are difficult to control because all negotiations are with a single vendor who is both authoring workloads and charging for compute. - Customer choice is restricted. Different compute engines are actually good at different things, and going “off platform” is hard when you’ve made the intentional decision to bring everything to one of these tools. - Internal collaboration is harmed. Data organizations become walled gardens, where some teams are on Platform A while some are on Platform B and they have a hard time collaborating. So: customers are faced with a hard choice. The fragmentation and duct tape that often existed with the MDS? Or the high cost and restricted choice that come with the “all-in-one data platforms”? ## Open data infrastructure as the path forward This is the context in which we introduce the term “open data infrastructure”. Open data infrastructure describes an infrastructure that is **pluggable**, relies on **integration via standards**, does not assume the usage of any one particular compute engine, and does not assume that solutions will be duct taped together from many individual products and vendors. Here’s what open data infrastructure will look like: ![dbt and Fivetran together to define the future of data](https://cdn.sanity.io/images/wl0ndo6t/main/982bb0f03c77be352577141e5554d00757132063-1974x1152.png) At dbt Labs and at Fivetran, we’ve both been building towards this vision for a long time. But in order to deliver a complete open data infrastructure, it will require both of us, together. This is what we have heard customers ask us for, over and over again, over the past several years: - Help make it easy for me to deliver best-in-class data infrastructure. - Help me save costs on my compute bills. - Deliver end-to-end capabilities like governance that are hard to do piecemeal. - Help me preserve optionality with my choice of analytical compute. These are the promises of open data infrastructure. Much of this we can deliver today, but of course, there is work still to do. The growth of AI over the last few years makes delivering on open data infrastructure only more critical. For customers building new AI capabilities: - You need to have reliable, high quality data, but also _centralized and high-quality metadata_ that describes everything about it. - You need to make sure that you can access your data directly from any model you choose, without having your access to your own data locked up behind what partnerships your all-in-one data platform vendor has at the moment. - You need to have your data exposed by open standards like MCP to quickly integrate with the fast-moving AI tooling ecosystem. - You need the ability to seamlessly use AI-native analytical compute engines that deliver the types of latency and concurrency that AI systems require. Open data infrastructure accomplishes all of this without asking customers to do it all themselves. This is why we’re excited about it, and why we hope you will be too. --- --- title: "From Informatica to dbt: A migration path to an AI-ready data control plane" description: "How to migrate off the legacy ETL data pipelines that are bogging down your business and stifling your AI initiatives." url: "https://www.getdbt.com/blog/informatica-dbt-data-control-plane-ai" date: "2025-10-09" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # From Informatica to dbt: A migration path to an AI-ready data control plane Legacy ETL tools like [Informatica](https://www.informatica.com/) served their purpose for decades. The demands of AI, however, are making the limitations of this legacy tooling impossible to ignore. The old era of ETL tooling—from [Oracle](https://www.oracle.com/) and Informatica to products like [IBM DataStage](https://www.ibm.com/products/datastage) and Matillion—did a good job of solving yesterday’s problems. But they were born in a different time when storage was expensive and everything was locked in on-premise data centers. They operated on timelines of weeks or quarters. Today’s reality—one that involves cloud data warehousing, elastic compute, domain team ownership, AI initiatives—is drastically different. Today, the companies that win ship solutions in hours or days, not weeks. Legacy ETL tools can’t deliver this. And the technical debt inherent in their existing data pipelines is dragging your business down. Solving this requires more than just swapping out tools. It requires rethinking how data teams deliver value and how everyone—not just engineers—works with data. It also requires a carefully considered and executed migration strategy to succeed. In this article, we’ll dig into what that means, how some companies have tried solving this problem, and how moving to an AI-ready data control plane positions you to compete in today’s modern business landscape. **** ## Why legacy ETL tools can't keep pace with modern demands The data landscape has transformed dramatically over the past decade. Legacy pipelines and stored procedures that once powered analytics workflows now represent more than technical debt. They actively slow organizational change, obscure risk, and inflate TCO. When business logic is buried in black-box ETL jobs, every enhancement becomes a risk. Every incident takes longer to debug. Every audit becomes a complex undertaking. The teams that succeed today are those that can ship insights in hours or days rather than weeks or quarters. They’re the teams that build trust through transparent and testable code. The architectural center of gravity has shifted from heavy standalone ETL tools to in-database transformation. Cloud data warehouses provide elastic compute, data products are owned by domain teams across the organization, and AI initiatives demand governed, explainable, high-quality data. The tools and requirements have both evolved. But the requirements have changed even more than the tools have. Boards and regulators now ask where numbers come from and demand full lineage and testing. Business units run weekly experiments and expect data to move at that cadence. Costs must map to ROI with precision. Business participation in data workflows is no longer optional. Data teams aren't order-taking factories. They’re enablers helping stakeholders build and act safely. ## The hidden cost of tool sprawl As organizations grow, different teams inevitably adopt different tools, each working in isolation: - Architects might use WhereScape - Data engineers work in Informatica or Matillion - Analytics engineers prefer Alteryx or Talend - BI developers rely on Tableau or Power BI - Analysts still turn to Excel Each tool has its own workflow, terminology, and implementation patterns. And the result is, predictably, chaos: - The same key performance indicator gets implemented five different ways across the organization - There's no single place to review logic, no single source of truth - Trust erodes as teams argue over which monthly revenue number is correct—and according to their own calculations, they're all right ## The modernization opportunity However, this fragmentation also represents a critical opportunity. Modernization isn't just about swapping tools. It's about redesigning how teams deliver value and how the business participates in data workflows. Organizations that successfully modernize unlock measurable advantages across four key dimensions. **Time to value.** Time to value accelerates dramatically when work becomes modular, tested, and reviewed. Projects that once took quarters can be completed in weeks. **Cost reduction**. Cost and complexity decrease substantially—many organizations see 50 to 80 percent lower transformation costs by consolidating on in-database processing, using compute efficiently, and eliminating duplicate workflows running the same calculations across different tools. **Resilient operations**. Operations become more resilient as incident numbers drop and recovery happens faster. When lineage, tests, and logs show exactly what changed, when it changed, and where it changed, troubleshooting transforms from guesswork to precision. **Strategic reinvestment**. Perhaps most importantly, the savings from modernization fund growth initiatives like AI. AI only works when data is high-quality, governed, documented, and explainable. Trusted data models result in safer and more accurate [retrieval-augmented generation (RAG) supplementation](https://aws.amazon.com/what-is/retrieval-augmented-generation/), AI copilots that understand your data based on your [DAG](https://www.getdbt.com/blog/dag-use-cases-and-best-practices) and your [data test suites](https://docs.getdbt.com/docs/build/data-tests), and AI agents that can take action because the underlying semantics and policies for your data are codified and clear. ## Why dbt has become the modernization standard dbt has been around since before the AI boom. That’s because we saw the need years ago to transition from traditional ETL systems like Informatica to a more modern method of data processing. dbt does this by bringing software engineering best practices to analytics. It enables companies to implement an [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle), similar to the Software Engineering Lifecycle, where data producers, analysts, and stakeholders collaborate iteratively on building and shipping high-quality data pipelines. With dbt, teams build robust transformations using [modular SQL or Python models](https://docs.getdbt.com/docs/build/models), using a cross-vendor data model syntax that transforms data wherever it lives in your organization. Data engineers manage production pipelines with built-in [Continuous Integration and Deployment (CI/CD)](https://docs.getdbt.com/docs/deploy/continuous-integration), testing, and governance to ensure ongoing data quality and safety. The [dbt Fusion](https://www.getdbt.com/product/fusion) engine enables context-aware development that accelerates delivering data faster and with higher quality. For data engineers transitioning from Informatica, this represents a fundamental unlock: - Teams gain shared context, so downstream partners aren't blocking progress - Engineers spend less time refactoring or rewriting logic from fragmented tools - High-performing data teams build like software teams: modular code, test-driven development, governance, and iteration built into every workflow. dbt is different because, unlike proprietary ETL systems, it treats data like code. That includes [version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics), CI/CD, tests shipped alongside pipelines rather than maintained separately, and automatically-generated [documentation](https://docs.getdbt.com/docs/build/documentation). The result is more innovation with the confidence to ship often, because guardrails catch issues early. Trust isn't an afterthought—it's engineered into every step. ### dbt: A data control plane that defeats tool sprawl dbt isn’t a new kind of ETL system. Rather, it’s a fundamentally different way of managing data - one that is flexible, cross-platform, collaborative, and focused on producing trustworthy outputs. Think of dbt as a [one-stop data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) for your data. Capabilities that once required gluing together multiple tools—orchestration, observability, cost management, catalog, and semantic layer—are integrated in a single platform with deeply connected metadata that drives results. dbt is the data control plane for **everybody** - not just data engineers. It integrates with the most popular BI tools and AI systems, enabling easy analysis. The built-in [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) centralizes metric definitions using common business language, ensuring both consistency across the organization and accuracy for AI-derived answers. To encourage data democratization, dbt supports multiple tools for finding and transforming data. These include [dbt Studio](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/develop-in-the-cloud) and [Visual Studio Code](https://docs.getdbt.com/docs/install-dbt-extension) for developers, [dbt Canvas](https://docs.getdbt.com/docs/cloud/canvas) for analysts and tech-savvy stakeholders, and [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) for decision-makers who need to find the right data fast. Finally, dbt uses AI to build for AI. [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot) leverages AI to facilitate working with data - from constructing complex SQL queries for data models to helping stakeholders analyze datasets with natural language queries. ## De-risking the migration path That brings us to the million-dollar question: how to migrate from traditional ETL systems to a modern solution like dbt. Migration has a reputation. And it…isn’t great. When you hear “migration,” you likely think of drawn-out projects that finish over time and over budget. If they finish at all. Fortunately, these days, your data teams no longer need to disappear for years to make the jump to a new data platform. dbt and partners such as [Infinite Lambda](https://infinitelambda.com/) have developed proven playbooks for iterative migrations that deliver value quickly. Modern tech also gives you more tooling for migration than we’ve ever had previously. AI-assisted migration tooling accelerates pattern discovery and code translation. With dbt, which is SQL-versant, you can reuse valuable SQL and business logic rather than discarding it. This means that institutional knowledge becomes visible and testable rather than lost. It’s also easier to find the resources you need. Because dbt is an industry standard, organizations tap into a large talent pool, making hiring and scaling easier. That all spells less risk, faster wins, and a runway to transformative AI projects. ### Choosing the right migration approach Three primary migration approaches exist, each with distinct tradeoffs: **Lift-and-shift. **A lift-and-shift code migration attempts to translate code with automation tools. This offers perceived speed and simplicity with relatively low risk, since business logic doesn't change. However, this approach can't take full advantage of the benefits of the cloud or reduce existing technical debt. Perceived speed can be misleading once you factor in validation time. **Full rewrite**. At the opposite end, a **full rewrite** creates truly cloud-native solutions but requires enormous investment of time, money, and effort. For anyone who’s attempted a complete platform rewrite, alarm bells ring immediately. Completely rewriting logic requires understanding code sometimes written over decades by people no longer with the organization. That makes the risk here substantial. **Replatforming and refactoring**. This is the middle path and often the most viable. This lifts and shifts data and logic, then incorporates best practices and looks for improvements. A human-led approach uses code automation for straightforward translations while investing time in thoughtful refactoring, whether AI-assisted or purely human-powered. ### The critical importance of assessing the three Vs Before attempting any migration, a comprehensive assessment is essential. This means understanding exactly what exists in legacy pipelines—i.e., taking skeletons out of the closet. Teams examine data sources and destinations, analyze pipeline complexity, and score each pipeline with complexity points to estimate the human effort required. Converting a representative pipeline helps translate complexity scores into real human terms. Understanding [**the three Vs**](https://www.techtarget.com/whatis/definition/3Vs)—**volume**, **velocity**, and **variety**—enables appropriate sizing of the new data architecture and accurate cost profiling. Identifying optimization opportunities is critical. These include finding code duplication where [dbt macros](https://docs.getdbt.com/docs/build/jinja-macros) can substantially reduce the application footprint, providing long-term maintenance benefits. ## Flowline + dbt for successful migration Even with the best planning, however, success isn’t guaranteed. Statistics say that [only one in three cloud migrations succeeds](https://www.unisys.com/news-release/one-in-three-cloud-migrations-fail-unisys-cloud-success-barometer/). Recognizing this, Infinite Lambda developed [Flowline](https://infinitelambda.com/flowline-legacy-data-migration/). Flowline is a packaged software and service solution designed to help enterprises modernize legacy infrastructure from [Talend](https://www.talend.com/), Informatica, and [SQL Server SSIS](https://learn.microsoft.com/en-us/sql/integration-services/sql-server-integration-services?view=sql-server-ver17) to an AI-ready platform in weeks rather than months or years. Flowline's approach centers on a human-led methodology heavily assisted by automation and AI, distinguishing it from both purely manual migrations and fully automated approaches. The solution follows a four-step process that balances speed with quality and reduces risk at every stage: 1. **Deterministic code conversion** Flowline extracts legacy pipelines and converts them using deterministic code—notably not AI—to dbt. This deliberate choice stems from practical considerations: deterministic code is cheaper to run and produces identical results every time, allowing the system to bake in best practices consistently. This step achieves approximately 95 percent automatic code conversion, representing a 10x improvement over manual rewrites. 1. **Validation and reconciliation** This is arguably the most critical phase. Flowline compares the resulting dbt models against production-quality data through an automation-assisted, human-led process. Some organizations underestimate this phase. Infinite Lambda’s learned through years of experience that conversion is the easy part. It’s testing, validating data, understanding differences, and securing stakeholder sign-off that consume the most time. At the end of this step, pipelines are nearly 100 percent reconciled. This phase accounts for any differences introduced by moving to new data warehouse platforms. 1. **Refactoring** With a stable baseline established, Flowline applies proprietary AI with human oversight to refactor code for improved performance, reduced code size, and lower costs. This refactoring happens confidently because the previous step has already validated correctness. 1. **Onboarding** The process concludes with rapid onboarding that trains both technology teams and business users while providing comprehensive change management to ensure complete adoption rather than leaving a migrated platform isolated. ### Real-world results from modernization Multiple companies have used Flowline to streamline their migration process from Informatica to dbt. The result is substantial cost savings with better data quality. [Macif](https://www.macif.fr/), a French insurer, saved substantial licensing costs compared to their traditional ETL systems and dramatically improved operational efficiency. Operations teams now go home on time because pipelines run substantially faster. One pipeline saw its runtime reduced from over two hours to under five minutes through migration and refactoring. [AstraZeneca](https://www.astrazeneca.com/) is positioned to save USD 40 million in total cost of ownership across personnel costs and licensing fees. Their AI-ready platform now powers better data experiences and helps deliver on their core mission of discovering new drugs. ## Addressing common migration concerns Three questions frequently arise when initiating migration conversations: ### Isn't migration just a rewrite? Moving to dbt is more than refactoring SQL code from one product to another. The complete platform acts as a data control plane leveraged across teams, consolidating existing tooling and enabling safe, governed data consumption across BI tools and AI systems. ### Will we just get locked in again? Platform flexibility is essential. dbt is designed to avoid lock-in. The tool sprawl era scattered data across multiple vendors, with analytics logic duplicated across tools. This led to tight coupling and high switching costs. The SQL standardization era consolidated analytics in dbt and SQL with centralized transformation logic. While this created scalable workflows, it locked organizations into specific cloud providers. Today, over half of dbt customers work across multiple data platforms. 81 percent of enterprises use more than one cloud provider. With native hosting, organizations use dbt on their cloud provider of choice—AWS, Azure, or GCP—running close to their data without compromising governance, security, or latency. Support for Apache Iceberg and catalog integrations for select adapters provides flexibility across compute engines. This isn't about checking boxes—it's about giving teams choice without penalty. ### What about ROI? dbt drives value across tooling and maintenance costs, efficiency gains, and direct business benefits. Organizations reduce current tooling costs, achieve cloud platform savings, and improve operational efficiency by freeing engineer and analyst time for core business work. That accelerates innovation around your company’s strategic objectives. ## Moving forward Modernization from legacy ETL tools to an AI-ready data control plane represents more than a technology upgrade. It’s a shift in the way your company interacts with data. That leap forward can be daunting. Even frightening. However, with proven migration approaches, comprehensive assessments, and platforms purpose-built for modern analytics workflows, the path forward is clearer than ever. The journey begins by figuring out where you’re already at. To get started, take the Infinite Lambda [online migration readiness assessment](https://migrationassessment.infinitelambda.com/) and plan out the first steps of your journey to AI-ready data. --- --- title: "How governed self-service helps analysts move faster without losing trust" description: "Empower analysts with governed self-service to ship insights faster without breaking data quality, trust, or compliance." url: "https://www.getdbt.com/blog/how-governed-self-service-helps-analysts-move-faster-without-losing-trust" date: "2025-10-07" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # How governed self-service helps analysts move faster without losing trust Self-service analytics has become the default expectation. But in practice? Most analysts are stuck jumping between tools, tracing broken models, and re-explaining logic across disconnected workflows. Speed without structure leads to problems: broken dashboards, stale metrics, and a loss of trust in the data team. The fix isn’t more meetings or manual checks. It’s **governed self-service**, which is a way for analysts to move fast _without_ breaking governance. With the right systems in place, analysts can safely build, validate, and ship data products without waiting on engineers or sacrificing quality. Here's how to do it, and why more data teams are making this shift today. ## Why traditional self-service breaks down at scale Many self-service efforts fall short because they rely on ad hoc processes and tribal knowledge. Analysts are often expected to “just figure it out” in a maze of notebooks, staging tables, and Slack threads. According to Chris Fiore, senior data analyst at dbt Labs: _“You’re wearing several different hats. You’re in the downstream conversation, building the data viz, then jumping upstream to debug lineage and do validation. It’s unsustainable.”_ This lack of role clarity, documentation, and tooling leads to: - Conflicting dashboards - Repeated work and context switching - Long turnaround times for simple requests - Broken definitions and compliance risks It’s not that analysts _can’t_ do technical work. It’s that they need the **right structure** to do it safely, quickly, and at scale. **** ## What governed self-service means for analysts Governed self-service [enables analysts](https://www.getdbt.com/product/analyst) to own their workflows—modeling, testing, and documenting data—inside a framework that enforces best practices automatically. This is the approach we use at dbt Labs. Analysts work in dbt Canvas and Git, validate changes with CI, document work in the dbt Catalog, and explore models in dbt Insights. Governance is embedded into the process, not enforced through tickets. As Paige Berry, lead data analyst at dbt Labs, puts it: _“Before dbt Insights, I’d drop SQL snippets in Slack. Now I can send a single Insights link with everything in one place that’s clean, traceable, and ready to ship.”_ With governed self-service: - Analysts trace lineage and health signals before making changes - CI runs automatically on every PR - Roles and ownership are clear - Governance doesn’t block progress. It supports it ## A modern analyst workflow, powered by dbt Here’s what governed self-service looks like in action: ### Develop in version-controlled branches Every analyst works in a dedicated dev schema with Git-backed version control. Changes are scoped, tested, and reviewed through CI/CD pipelines before hitting production. ### Validate models and freshness in dbt Catalog and Insights Analysts use dbt Catalog to view lineage, field-level metadata, owners, and test coverage. Then, they use dbt Insights to preview data, validate freshness, and debug errors before asking engineering for help. _“If I see a dashboard failing, I can trace it back myself—lineage, freshness, the job status—and often fix it before looping anyone else in.”_ — Rachael Gilbert, staff data analyst at dbt Labs ### Leverage shared metrics and business logic With the dbt Semantic Layer, analysts don’t have to rewrite logic across tools. Metrics are defined once and used everywhere, from BI to AI. ### Use dbt Insights for reproducible ad hoc work With dbt Insights, analysts explore data, build analyses, and share results all within the same governed workspace, so there’s no more copying SQL across Notion docs or Slack messages. ### Build visually with [dbt Canvas](https://www.youtube.com/watch?v=pO_TUnCt6es) dbt Canvas gives analysts a drag-and-drop interface to model data visually, no SQL required. It’s an intuitive way to build, iterate, and understand how data flows without leaving the governed environment of dbt. ### Get context-aware AI help with [dbt Copilot](https://www.youtube.com/watch?v=Vg7Nu6SKcXE) dbt Copilot gives analysts AI assistance that understands your project’s structure, naming conventions, and documentation. Ask dbt Copilot to write or refactor models, explain logic, or suggest tests based on your existing codebase. ### Chat with your data using dbt [MCP](https://www.youtube.com/watch?v=fJ-72qOA7BE) The dbt MCP server exposes your dbt project's structured context to any AI system. That means analysts can talk to their data from tools like Slack, ChatGPT, or Claude with full trust in definitions, lineage, and freshness. ## Your analyst enablement checklist To know if your team is truly ready for governed self-service, you should be able to answer “yes” to the following: - Can every analyst trace model lineage and trust signals _without asking an engineer_? - Can they make changes safely using Git, CI, and temporary dev environments through a visual editing tool? - Are metrics centrally defined in a semantic layer, not buried in LookML or spreadsheets? - Do they have tools to preview, validate, and debug models independently? - Are governance and RBAC policies enforced _by the system_, not by gatekeeping? - Are AI tools embedded in the workflow to help analysts build models and chat with context-aware data? If not, you’re not doing self-service. You’re doing ungoverned guesswork. ## A better way to work for analysts and engineers Governed self-service doesn’t just help analysts. It frees up engineers to focus on scale, telemetry, infrastructure, and the AI projects of tomorrow rather than debugging logic or answering repeat questions. As Zach Brown puts it: _“My focus now isn’t enabling analysts directly. It’s about building the infrastructure so they don’t need me in the loop for every change.”_ The result is a system that works for everyone: - Analysts move faster, with more autonomy - Engineers spend less time unblocking tickets - The business gets insights faster with trust and consistency Explore how the dbt Labs team puts these principles into action in the whitepaper, [**An analyst’s guide to working with data engineering**](https://www.getdbt.com/resources/analysts-guide-to-working-with-data-engineering), featuring real examples and workflows from dbt’s own analysts and engineers. [**Request a demo**](https://www.getdbt.com/contact) or [**start your free trial**](https://www.getdbt.com/signup) of dbt to bring governed collaboration to your own team. --- --- title: "Why governed collaboration is the key to modern analytics workflows" description: "Unlock faster analytics with governed collaboration. Streamline data workflows without sacrificing trust or governance." url: "https://www.getdbt.com/blog/why-governed-collaboration-is-the-key-to-modern-analytics-workflows" date: "2025-10-06" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Why governed collaboration is the key to modern analytics workflows Data is growing faster than ever, and so are the demands on your team. Stakeholders want insights _now_, but without a clear structure and ownership, your analysts are chasing tickets while engineers are firefighting pipelines. You can move fast. Or you can protect quality. But doing both? That’s where governed collaboration comes in. Governed collaboration is a modern approach to how [analysts](https://www.getdbt.com/product/analyst) and engineers work together. It balances autonomy with accountability, enabling faster development without sacrificing data quality or governance. When done right, it reduces ticket backlogs, increases iteration speed, and builds trust across your team and the business. Here’s how to make it work. ## What is governed collaboration in data teams? Governed collaboration is a structured partnership between data analysts and data engineers who co-own the data lifecycle. Analysts move fast. Engineers provide guardrails. Everyone works in the same system with clear roles, workflows, and accountability. Instead of shadow pipelines and Slack messages, collaboration happens inside version-controlled projects, with testing, lineage, and documentation baked into the process. It’s a workflow, not a workaround. As Zach Brown, senior software engineer at dbt Labs, puts it: _“The analysts are the ones who know what they need. They just don’t know how to make it happen.”_ Governed collaboration closes that gap and turns requests into repeatable systems. **** ## Why data teams need governed collaboration now Let’s be honest: data chaos doesn’t scale. - AI raises the stakes for quality, compliance, and lineage. - Data volume is exploding. So are the questions. - And the old way of working (tickets, tribal knowledge, disconnected tools) is slow and risky. Poor data quality doesn’t just hurt your dashboards. It erodes trust across your org. According to our [2025 State of Analytics Engineering Report](https://www.getdbt.com/resources/state-of-analytics-engineering-2025), **over 56%** of respondents cited poor data as a top challenge. You can’t ticket your way out of this. You need a shared system with built-in governance that supports speed, not stops it. ## 5 principles of governed collaboration for analytics teams You don’t need more meetings. You need shared workflows. Here’s how governed collaboration works in practice: #### **1. Shared development environments** Analysts and engineers work in the same repo, with version-controlled development and automated CI to catch issues before they reach prod. #### **2. Continuous testing and validation** Every change, whether from an analyst or engineer, runs through the same gauntlet: tests, contracts, freshness checks. No shortcuts or surprises. #### **3. Built-in lineage and metadata** Changes are traceable. Owners are known. Stakeholders can validate where data comes from without asking in Slack. #### **4. Clear roles and defined handoffs** Specialization matters. Engineers own infrastructure and telemetry. Analysts own business logic and stakeholder alignment. Everyone knows where their work starts and stops. #### **5. Guardrails that enable self-service analytics** Analysts can ship models, tests, and docs safely, without waiting on engineers. But every change is scoped, tested, reviewed, and governed by the system. _“Having a tool like dbt makes it a lot easier... I can self-serve on tracing and put together a picture before I ask for help.”_ — Rachael Gilbert, staff data analyst at dbt Labs ## How dbt enables governed collaboration at scale dbt is built to make governed collaboration the default without requiring extra tools or overhead. You get: - **Version-controlled development** with Git + CI for every PR - **Lineage and ownership** with dbt Catalog - **Run ad-hoc analysis, validation, and previews** with dbt Insights - **Centralized metrics and logic** in the dbt Semantic Layer - **Visual, AI-powered editing and modeling** with [dbt Canvas](https://www.youtube.com/watch?v=pO_TUnCt6es), making it easier to drag-and-drop changes inside branches - **AI that’s aware of your project context** with [dbt Copilot](https://www.youtube.com/watch?v=Vg7Nu6SKcXE), helping you write, refactor, and reason about code more effectively, all while staying grounded in your team’s standards - **Role-based access control** to maintain governance at scale **** ## Why governed collaboration is the future of data work Governed collaboration is about removing the blockers that stop analysts from doing their best work. When engineers enable repeatable workflows and analysts contribute safely within guardrails, the entire org moves faster and trusts the data. Explore how the dbt Labs team puts these principles into action in the whitepaper, [**An analyst’s guide to working with data engineering**](https://www.getdbt.com/resources/analysts-guide-to-working-with-data-engineering), featuring real examples and workflows from dbt’s own analysts and engineers. [**Request a demo**](https://www.getdbt.com/contact) or [**start your free trial**](https://www.getdbt.com/signup) of dbt to bring governed collaboration to your own team. --- --- title: "The hidden tax on analysts" description: "A new report with the Harris Poll shows analysts lose $21K/year to inefficiencies. Learn how AI workflows reverse the trend." url: "https://www.getdbt.com/blog/the-hidden-tax-on-analysts" date: "2025-10-03" authors: ["Daniel Poppy"] categories: ["Insights"] --- # The hidden tax on analysts Analysts are among modern enterprises’ most valuable assets. Their insights guide strategy, shape investments, and determine how organizations compete. But the way most companies structure analyst work is fundamentally broken. Instead of fueling innovation, outdated workflows bury analysts in busywork and repetitive tasks caused by tool sprawl. **Imagine paying an extra $21,000 per analyst every year; not in salary, but in wasted time.** That's the reality for many organizations today. [According to a survey](https://www.getdbt.com/resources/the-analyst-revolution-unlocking-tomorrows-ai-initiatives) conducted by dbt Labs with The Harris Poll, analysts lose an average of **9.1 hours per week** to inefficiencies, translating into **$21,613 in lost productivity per analyst each year**. For a business with 1,000 analysts, that’s an **annual hidden tax of $21.6 million**—and that’s without accounting for the intangible losses to speed, agility, and quality of decision-making. This is not just an employee experience issue. **It’s a productivity crisis that touches every aspect of business, from growth to competitiveness.** **** ## Tool sprawl: A drag on business growth Analysts aren’t disengaged because they dislike their jobs. They love the challenge of turning data into insights. What burns them out is being forced into endless busywork by outdated systems. In fact, our survey discovered that **77.4% of analyst time is lost to mundane tasks** like data prep, validation, and repetitive queries—leaving only **22.6% for generating insights**. Executives can’t afford to ignore this drag. Every day analysts spend reconciling spreadsheets or switching between mismatched tools is a day of delayed insights and missed opportunities. Analysts don't lack talent or motivation; workflow design is the problem. As a quick fix, many organizations have responded to the rising demand for insights by layering on more self-service tools. But adding more and more tools hasn’t solved the problem, it’s compounded it. Instead of accelerating decision-making, tool fragmentation creates friction, context-switching, and errors, forcing the very people the C-suite relies on to unlock competitive advantage into low-value administrative tasks. Given that our survey found that analysts find themselves juggling 5.4 platforms daily—switching between them nearly six times a day—62% report feeling overwhelmed by tool sprawl. The result isn’t just analyst frustration. **It's measurable business loss.** ## The cost of doing nothing Inefficiency at this scale is more than a workflow annoyance.** It’s a high-stakes performance crisis with severe financial consequences.** When analysts spend three-quarters of their time validating data rather than analyzing it, leadership is flying blind. Decisions are delayed. Strategy execution slows. Responsiveness suffers. In a market where speed and accuracy define winners, inefficiency is a competitive liability. Why AI-powered automation changes everything Analysts don’t need more tools. They need structured workflows that** automate the low-value work draining their time**. The path forward lies in **AI-powered automation**, designed not as a bolt-on, but as the backbone of the [analytics development lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle). Analysts are ready for this shift. Ninety percent want more AI tools integrated into their workflows. AI-powered workflow automation, from real-time data quality checks to automated visualization, ranks at the top of their wish lists. And nearly all analysts (97%) express interest in governed self-service platforms that blend autonomy with compliance. In other words, they want AI tools that take mundane, repetitive tasks off their plates and make their work more secure and efficient. This isn’t about replacing analysts with technology. **It’s about empowering them to focus on what they were hired to do: generate insights that move the business forward.** ## Optimized workflows can power the analyst revolution The path forward is clear. Organizations that invest in workflow automation and governed AI platforms aren’t just reducing costs, they’re **unlocking growth.** Almost all analysts (96%) say they’re more likely to stay with employers who invest in workflow optimization. But this is more than a retention play. **It’s the future of AI and analytics: shifting analysts from data prep to decision acceleration.** When analysts are free from busywork, they become** strategic catalysts**—delivering faster insights, enabling bolder decisions, and strengthening organizational agility. However, failing to act equates to lost productivity, delayed insights, and unnecessary attrition. Adopting centralized, AI-powered workflows now means redirecting millions of wasted dollars into innovation and growth. This hidden tax is costly, but not inevitable. **Smarter investment in automation and governance can eliminate it altogether, and leaders who act now will gain unmatched speed, efficiency, and resilience.** These findings are just the beginning. Explore the full research in [The Analyst Revolution](https://www.getdbt.com/resources/the-analyst-revolution-unlocking-tomorrows-ai-initiatives). **** --- --- title: "How to make data-driven decisions (without being a perfectionist)" description: "The View on Data explores how to make confident, data-driven decisions, even when your data isn’t perfect." url: "https://www.getdbt.com/blog/how-to-make-data-driven-decisions" date: "2025-10-03" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # How to make data-driven decisions (without being a perfectionist) In the fourth episode of _The View on Data_, hosts Faith McKenna, Erica “Ric” Louie, Paige Berry, and new co-host Jerrie Kenney talk about one of the most persistent myths in data work: that data-driven decisions always require perfect data. Spoiler alert: they don’t. From civic engagement indexes built on grocery store data to dashboard decisions driven by intuition, the group explores what it really means to work with “enough” data and how to build confidence in your analysis, even when the dataset is messy, incomplete, or inconsistent. This episode is for anyone who’s ever agonized over missing values, hesitated to share a chart, or wondered if their work was “good enough” to make the call. Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions. 🎧 Listen & subscribe: [Spotify](https://open.spotify.com/show/0Itxg3BoF6zaQdORfVuY6I?si=6133f88257624765) | [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-view-on-data/id1835755886) | [Amazon Music](https://music.amazon.com/podcasts/84c47dbe-adf7-4d33-a371-a1fe55255495/the-view-on-data) | [YouTube](https://youtu.be/bETrfOok1Fw) ## What does it really mean to be data-driven? For many folks entering the field, “data-driven” can feel like code for _statistically significant, p-value certified, decision-changing insight_. But in practice? It’s often more about building directional confidence. “I think people conflate data-driven with 'I’m going to change the company strategy,’” Ric shared. “But sometimes the data-driven decision is: don’t change anything.” The hosts discussed how business decisions often rely on imperfect or incomplete data, and how it’s more important to ask good questions upfront than to get bogged down in chasing precision. ## What’s “good enough” when it comes to data quality? Jerrie shared a story about helping a client rethink their entire approach to data quality testing, starting with the “why.” “They were so focused on thresholds and configs, they forgot who they were doing it for,” she said. “We ended up having to zoom out and ask: what decisions are we trying to support here? And how perfect does the data really need to be?” Paige echoed that point, describing the diminishing returns of chasing exactness when the goal is clarity, not control. “Sometimes good enough is better than perfect,” she said. “Especially when what your stakeholders actually need is just confidence to move forward.” ## Using incomplete data to drive impact Throughout the episode, the group shared examples of when “directionally correct” was more than enough: - **Jerrie** talked about using social media signals and grocery store data to identify civic engagement hotspots. This helped a health equity team decide where to host in-person town halls. - **Faith** described making curriculum decisions based on a small number of user requests in course feedback: “What’s the worst that could happen? People don’t like it, and I try something else.” - **Paige** highlighted how institutional knowledge and intuition helped her decide whether a metric spike was internal testing or external usage without spending hours chasing a rabbit hole. - **Ric** emphasized the value of pairing data with strategic context: “Not every weird number needs fixing. Ask what decisions are riding on this, and how much we’re willing to invest in fixing it.” The theme? Use what you have. Pair it with experience. And don’t let perfection block progress. ## How to build confidence when your data isn’t perfect For data folks, especially those who are self-taught or newer in their careers, it can feel risky to present work that isn’t 100% airtight. The team talked about the role of impostor syndrome in driving overwork and perfectionism. “I used to feel like I had to make an ironclad case for every recommendation,” Faith shared. “But I’ve had to learn to trust my experience and say: this is enough.” Ric and Jerrie both emphasized the power of up-front conversations with stakeholders to set expectations: - What is the data for? - What will the decision be? - What are the consequences if it’s wrong? “If it’s for a board report or compliance, sure, go for precision,” Jerrie said. “But if the cost of being wrong is low, then directionally accurate is often all you need.” ## Communicating clearly when data is messy The episode wrapped with a conversation about how to present data clearly, especially when your audience ranges from execs to engineers. Paige shared her approach: “I try to understand the experience of the person reading the chart. How much time do they have? What are they trying to do with this?” Ric emphasized using context-first messaging in Slack and reports: leading with the TL;DR, followed by charts and only the most relevant details. And everyone agreed on the power of simple visuals: Favorite chart types? - Ric: Bar chart with line overlay - Paige: Sankey (for the drama) - Jerrie: Mini line charts (“spark lines”) - Faith: Big number with small number (and triangle indicators) ## Episode takeaways This conversation is a practical reminder that in real-world data work, perfection isn’t the goal, progress is. Here are a few takeaways from the episode: - Data-driven doesn’t mean statistically significant. It means directionally useful. - Ask what decision the data supports before obsessing over accuracy. - Stakeholder context matters more than fancy charts. - Show your work, but only share the summary. - Confidence in your experience is part of the job. --- --- title: "Data engineering best practices: What's new?" description: "Four big shifts in data engineering that are reshaping how teams build trusted, scalable pipelines for AI and analytics." url: "https://www.getdbt.com/blog/modern-data-engineering-best-practices" date: "2025-10-02" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data engineering best practices: What's new? Data engineering is at a turning point. We live in an era where data underpins everything from customer experiences to strategic decisions. AI-driven analytics, real-time demands, and distributed data ownership are reshaping how teams build, manage, and deliver data. Data engineering has always been about building reliable, scalable pipelines — but the landscape is changing. In recent years, three significant developments have reshaped how teams design and deliver modern data systems: - The explosion of data volume and diversity - The near-ubiquity of cloud-native infrastructure - The advent and adoption of AI across the enterprise These changes have created new challenges — and opportunities — in data engineering. This article addresses these challenges and discusses how teams can adopt best practices that improve collaboration, increase agility, and deliver trusted data at scale. ## Today's data engineering challenges (and solutions) Data engineering must continuously adapt to unprecedented growth in scale, speed, and complexity while delivering reliable, self-service analytics to drive business value. Four critical challenges define today’s landscape: - **Too much data to process** - **Data engineering teams are doing too much** - **Development and debugging are too slow** - **Data governance can’t be manual** ### Challenge 1: Too much data to process In 2024, [industry analysts estimated](https://www.domo.com/learn/infographic/data-never-sleeps-12) that the amount of data created, captured, and consumed globally was 149 zettabytes, with projections surpassing 394 zettabytes by 2028. Businesses are drowning in data — [64% of organizations](https://cdn.avepoint.com/pdfs/en/shifthappens/AI-IM-Whitepaper-v4.pdf) already manage at least one petabyte of data, and 41% manage at least 500 petabytes worth. Data can be unstructured, structured, or semi-structured. And it all needs to be analyzed to generate business value from insights. Traditional, manual approaches to data discovery and pipeline creation simply can’t keep pace. This yields backlogs, bottlenecks, and missed opportunities. The scale of today’s data landscape demands a new layer of acceleration. ### The solution: AI-powered acceleration with dbt Copilot dbt has long been the industry leader in data transformation platforms. Using dbt as your [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction), your data teams can more quickly create trustworthy data outputs with data workflows that are flexible, cross-vendor, and collaborative. Now, [dbt Copilot](https://www.getdbt.com/product/dbt-copilot) brings generative AI directly into your dbt environment, giving you a way to move faster without sacrificing quality or control. Copilot lets you harness your data's full context, including relationships, metadata, and lineage, to automate routine tasks using natural language prompts. With Copilot’s AI capabilities, you can: - **Accelerate model creation**: Generate SQL models from natural‑language prompts, using your project’s metadata (relationships, lineage, and context) to ensure relevance and accuracy. - **Auto‑suggest transformations, joins, and model structures: **Leverage warehouse context to recommend the right building blocks for your models. - **Speed up data discovery: **Use context‑aware AI that understands your dbt models’ metadata, lineage, and relationships to recommend the most relevant assets and transformations. - **Accelerate data testing:** dbt Copilot uses the context of your dbt models to suggest context-aware validation tests. With one click, it adds the corresponding test code directly to your project, ready to run during your builds. - **Embed governance from the start**: Every Copilot‑generated asset is version‑controlled, testable, and documented like any other dbt model, ensuring speed never comes at the expense of trust. Copilot is like a dedicated data intern who standardizes legacy documentation, improves query optimization, checks for SQL syntax errors, enhances metadata compliance, and speeds up migrations — all within your dbt workflows! ### Challenge 2: Data engineering teams are doing too much Today’s engineers build pipelines that feed self-service analytics, power AI models, and enable real-time decision making. They also enforce governance, maintain data quality, and scale across an expanding ecosystem of tools and data sources. As a result, engineers can become bottlenecks for basic data access, documentation, and model creation. ### The solution: Empower stakeholders with self-service tools Data self-service is now commonplace in data analytics, with modern BI platforms that enable business users to explore and analyze information without relying on IT. Data democratization has been shown to [shorten time‑to‑insight](https://www.castordoc.com/data-strategy/what-is-self-service-bi), boost productivity, and drive innovation as teams can act on trusted data in minutes instead of days. dbt has recently introduced self‑service tools for analysts (and engineers!) to accelerate model development, streamline transformations, and boost productivity without sacrificing trust. The first is [dbt Canvas](https://www.getdbt.com/blog/dbt-canvas-is-ga), a visual, drag-and-drop interface that lets teams create and edit dbt models without starting from a blank SQL file. In the dbt Canvas no-/low-code modeling environment, your analysts can: - **Build and edit models** without hand-coding every step. Simply drag and drop operators, such as input, join, select, aggregate, and formula, onto a canvas, then connect them visually. - **Generate valid SQL** code. Canvas-generated code is version‑controlled, testable, and deployable like any other model. - **Find and explore data without SQL knowledge.** Canvas includes robust search and discovery features and always-on data profiling capabilities to help analysts (and engineers) gain a deeper understanding of the data. - **Facilitate iterative collaboration.** Each transformation is visually represented as a node, making it easy to trace logic, understand relationships, and work with teammates. Step‑by‑step previews at every stage build confidence, reduce errors, and accelerate iteration. - **Harness AI‑assisted code generation.** Canvas integrates with dbt Copilot to suggest transformations and generate SQL, accelerating model development. We’ve also introduced the self-service [dbt Catalog](https://docs.getdbt.com/docs/explore/dbt-explorer-faqs) tool to help engineers and analysts alike quickly understand your entire lineage from data source to the reporting layer. Using Catalog, data teams can troubleshoot, improve, and optimize your data workflows faster. In the dbt Catalog self‑service data discovery environment, your team can: - **Search and explore data assets without writing SQL.** Analysts can instantly find models, sources, and metrics with rich metadata, descriptions, and tags to speed up analysis. - **Understand lineage and dependencies.** Interactive, column‑level lineage maps show how data flows from source to dashboard, helping teams assess impact before making changes. - **Assess data quality at a glance.** View test coverage, documentation completeness, and performance insights to ensure trusted, production‑ready datasets. - **Troubleshoot and optimize faster.** Identify slow‑running models, missing documentation, or failing tests directly from the Catalog interface. - **Enable governed self‑service.** All assets are tied to the dbt project, so definitions, lineage, and quality checks stay in sync with production, empowering analysts without sacrificing control. We believe the future of analytics engineering is collaborative, governed, and accessible to every data practitioner. Canvas and Copilot enable analysts to work with well‑governed data earlier in the pipeline in a self-service manner. That helps your engineers to focus on what matters most: designing scalable, high‑performance data models, enforcing rigorous quality and testing standards. Instead of answering simple questions about data, they can return to solving complex transformation challenges that unlock faster, more reliable insights for the business. ### Challenge 3: Development and debugging are too slow In many data engineering environments, development remains slow. Writing and editing transformation code can mean waiting minutes for parsing and compilation just to validate changes. Errors often surface only after running against the warehouse. Once changes are made, long feedback loops, redundant runs, and the challenge of pinpointing root causes further stall delivery and inflate compute costs. That leads to longer development cycles, slower debugging, and higher compute costs. ### The solution: High-velocity development with dbt Fusion [dbt Fusion](https://www.getdbt.com/blog/dbt-labs-redefines-dbt-with-new-fusion-engine) attacks these bottlenecks end-to-end in your pipeline. With lightning-fast parse times, instant error detection during coding, and state-aware orchestration that runs only what’s changed in the pipeline, teams can iterate, debug, and deliver at top speed. With** dbt Fusion’s state-aware orchestration**, teams can: - **Develop at record speed.** 30x faster parse times mean instant feedback and near‑real‑time iteration. - **Catch errors early.** Live error detection flags issues before running code against the warehouse. - **Run only what’s changed. **[State‑aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-setup) pinpoints modified models and their dependencies, avoiding costly full‑project rebuilds. - **Iterate faster.** Shorter feedback loops enable engineers to test, refine, and ship changes in minutes, rather than hours. - **Debug with context.** Built‑in dependency awareness makes it easier to isolate issues and understand downstream impacts before deploying. - **Control costs.** Early adopters report significant savings, with visibility into spend via the cost management dashboard. ### Challenge 4: Data Governance can't be manual In fast-moving analytics environments, governance processes that rely on manual checks, ad-hoc documentation, or after-the-fact reviews can’t keep pace with the speed of modern data delivery. Fragmented documentation and inconsistent metadata standards undermine trust and self-service, leaving analysts to hunt for definitions and lineage. ### Solution: Automated governance inside the pipeline dbt embeds governance directly into the transformation layer. It enforces quality, consistency, and compliance as part of the workflow, not as a bolted-on afterthought. - [dbt Catalog](https://docs.getdbt.com/docs/explore/dbt-explorer-faqs) provides interactive, column‑level lineage across your entire data estate, making it easy to audit changes, trace data flows from source to dashboard, and assess downstream impacts before deploying. - [dbt’s Semantic Layer](https://www.getdbt.com/blog/how-the-dbt-semantic-layer-works) centralizes metric definitions using MetricFlow. Every stakeholder — whether they’re working in BI tools, spreadsheets, or embedded analytics — starts from the same, governed business logic. - [dbt Fusion](https://www.getdbt.com/blog/dbt-labs-redefines-dbt-with-new-fusion-engine)’s governance‑aware orchestration automatically enforces policies in‑flight, running only the models that have changed and ensuring compliance rules are applied consistently at build time. With **dbt’s built-in governance capabilities**, teams can: - **Automate lineage tracking.** Capture end‑to‑end data flow for complete auditability and faster impact analysis. - **Standardize metrics**. Use the dbt Semantic Layer to define and enforce consistent business logic across teams and tools. - **Enforce policies in‑flight.** Apply governance rules automatically during orchestration with Fusion’s governance‑aware orchestration, not after the fact. - **Accelerate compliant delivery.** Ship governed, high‑quality data at the same speed as ungoverned pipelines. - **Build trust at scale.** Ensure every dataset meets quality, compliance, and consistency standards before it’s consumed. ## dbt and the next era of data engineering Modern data teams can’t afford to ignore these four challenges facing data engineering. If they do, they risk higher costs, lower data quality, missed opportunities, and a growing gap between strategic business needs and the insights delivered. dbt brings together the solutions to all four challenges — accelerating development, empowering stakeholders, speeding up debugging, and automating governance — in a single, integrated workflow: - **Copilot **speeds development with AI-assisted code generation. - **Fusion** eliminates bottlenecks in parsing, debugging, and orchestration. - **Catalog** provides complete lineage for auditability. - **Semantic Layer** enforces consistent, governed metrics across every downstream tool. Together, they enable teams to move faster, catch issues earlier, enforce governance automatically, and control costs — all within the same platform they already use to build and manage transformations. dbt makes this possible by enabling true collaboration, giving every role shared context, consistent definitions, and built‑in guardrails inside the pipeline. dbt equips data engineering teams to move faster and with more confidence. To learn more about how dbt can help you future-proof your data engineering practices, [request a demo](https://www.getdbt.com/contact). --- --- title: "Getting started with an ELT pipeline" description: "Learn how to build ELT pipelines that combine extraction, loading and transformation using cloud warehouses and dbt." url: "https://www.getdbt.com/blog/elt-pipeline" date: "2025-10-02" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Getting started with an ELT pipeline As businesses grow, they get insights from a range of data sources to make informed decisions. Ad hoc reports, in-app reporting, and spreadsheets can get at pieces of the data. But often, you need to bring data together from multiple sources and reshape it to yield true business insights. An [extract, load, transform (ELT)](https://www.getdbt.com/blog/extract-load-transform) data pipeline combines data from multiple sources to provide accurate and reliable business insights. Transforming large volumes of unprocessed data into a format that’s suitable for analytics isn’t easy, though. It requires breaking down siloed transformation workflows and a high degree of visibility into the data transformation processes. This article explores ELT pipelines and how to design one that scales efficiently. You’ll also learn how dbt drives modern ELT pipelines, simplifying collaboration and data transformation. ## What is an ELT pipeline? An ELT pipeline is a modern data integration process following the Extract, Load, Transform sequence. Analysts first extract data from various sources and load it into a centralized data warehouse. After that, they transform the data within the data warehouse to extract insights. ELT pipelines consist of four core components: 1. A **cloud data warehouse** for storing and transforming data. 2. A **data integration tool** for extracting and loading data. 3. A **transformation framework** for data modeling and preparation for analysis. 4. A **business intelligence (BI) tool** for insights and visualization. The choice of tools determines how efficiently data moves and transforms in your ELT pipeline. ### Why modern data teams choose ELT over ETL [Extract, transform, load (ETL)](https://www.getdbt.com/blog/extract-transform-load) pipelines transform data before loading it into the warehouse. This requires analysts to anticipate data models and reporting needs in advance. Any change in business requirements often means redesigning complex workflows and schema mappings. ETL was all the rage back when compute and storage cost multiples of what they do today. While ETL works, it’s generally an inflexible approach that doesn’t scale to meet today’s data transformation demands. ELT flips this process. It uses the scalability of cloud compute and storage to enable faster, more flexible transformations. Data teams ensure consistent access and simplify governance by centralizing data in a warehouse. Automated tools for extraction and loading keep pipelines updated and resilient against schema changes. In ETL, transformation happens before analysts even touch the data. In ELT, it occurs closer to the data so analysts can quickly iterate and build accurate models for analysis. ## How to design your ELT pipeline ![Diagram showing the four stages of an ELT pipeline: Data Warehouse (scalability, governance, cloud storage), Integration Tool (schema updates, error handling, incremental syncs), Transformation Framework (modular SQL, testing, version control, documentation, lineage), and BI Tools (dashboards, real-time insights, stakeholder decision-making, customizable visualizations).](https://cdn.sanity.io/images/wl0ndo6t/main/3330189316953615f2c171496521aef3988ea084-1600x934.jpg) Designing an ELT pipeline begins with a scalable data warehouse that acts as the main hub for all your data. When planning the warehouse: - Focus on scalability and elasticity to handle changing workloads. - Ensure governance to stay compliant and keep access secure. - Use cloud storage for better flexibility and performance. The next step is to choose a data integration tool that automates data extraction and loading. [Fivetran](https://www.fivetran.com/) and similar tools set up connectors automatically and manage scheduling in the background. Look for integration features such as: - Automated schema updates and error handling. - Broad support for data sources and destinations. - Incremental syncs for faster and efficient updates. A transformation framework turns raw data into structured models for analysis within the warehouse. Key features to consider include: - **Modular SQL-based development.** Build reusable and maintainable data models using SQL. - **Automated testing.** Validate data models to ensure their accuracy and reliability. - **Version control integration.** Facilitate collaboration and enable tracking of changes in data models. - **Built-in documentation.** Provide clear explanations of data models to enhance understanding and usability. - **[Data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) visualization.** Show how data moves and transforms across models, allowing data consumers to show where data came from so they can validate its authenticity. Modern [Business Intelligence (BI) tools](https://www.coursera.org/articles/bi-tools) help teams share real-time insights on business performance. Stakeholders use shared dashboards to make decisions using the updated data. ## Common use cases of ELT pipelines ![Diagram showing ELT pipeline use cases: • Real-time analytics (stream data from apps, live dashboards, operational metrics) • E-commerce optimization (combine sales and clickstream data, identify buying patterns, personalized recommendations) • Marketing intelligence (analyze campaign data, measure sentiment trends, cross-channel campaigns) • Healthcare outcomes (stream patient data, standardize historical records, predictive diagnostics)](https://cdn.sanity.io/images/wl0ndo6t/main/c72614214e6bcc698c4deb47b52322ce945c7f5f-1600x706.jpg) ELT pipelines enable teams to process and transform massive data streams directly within cloud data warehouses. They support transformations on demand, real-time analytics, and [machine learning (ML)](https://hub.getdbt.com/kristeligt-dagblad/dbt_ml/latest/) on continuously updated datasets. ### Real-time analytics and reporting Companies stream raw data from apps and IoT devices into data warehouses like [Snowflake](https://snowflake.com). ELT transformations then clean and model this data to power live dashboards and operational metrics. ### Retail and e-commerce optimization E-commerce platforms use ELT pipelines to combine clickstream, sales, and inventory data from multiple systems. The pipeline then applies transformations to help identify buying patterns and enhance personalized recommendations. ### Sentiment analysis and marketing intelligence Marketing teams pull campaign, CRM, and social data into a data warehouse using ELT automation. They transform it to analyze sentiment trends and measure cross-channel campaign effectiveness in real time. ### Healthcare and patient outcomes Hospitals stream patient vitals and historical records into a data warehouse through ELT pipelines. Data is then standardized to enable predictive models that improve diagnostic speed and patient care outcomes. ## Key challenges in implementing ELT pipelines Scalable ELT pipelines demand tight control over security, compliance, and cost to ensure reliability and efficiency. Here are some of the challenges in implementing ELT pipelines: **Data quality and consistency.** Loading raw data without validation can introduce duplicates, missing values, or inconsistent formats. Teams need to implement validation and anomaly detection during the loading and transformation phases. Automated testing and data lineage tracking help identify errors early and maintain consistent data quality. **Security and access control.** Transferring petabytes of data between applications and warehouses can expose pipelines to corruption or unauthorized access. Implementing encryption and [role-based access control (RBAC)](https://www.ibm.com/think/topics/rbac) limit who can view sensitive datasets. Integrating security measures at each pipeline stage ensures data remains protected across the warehouse. **Regulatory compliance.** Regulations like [HIPAA and GDPR](https://www.puredome.com/blog/intro-to-key-cybersecurity-compliance-standards) require ongoing audits and privacy by design transformations. Compliance should be integrated into ELT design with logging and documentation tools. Regular monitoring of data processes helps teams remain audit-ready and ensures clear governance. **Resource bloating.** Cloud data warehouses expand endlessly without retention strategies, which drives up compute and storage costs. Teams should implement data lifecycle policies to regularly archive or delete unused data. Automatic partitioning and tiered storage can help control costs and improve query performance. **Integration complexity.** Diverse data formats and systems require adaptable connectors and consistent schema management. Tools with broad connector support, such as Fivetran or Rivery, simplify multi-source integration. Modular transformations make pipelines easier to manage and scale over time. ## How dbt powers ELT pipelines dbt functions as the transformation layer in an ELT pipeline, orchestrating transformations within the data warehouse. dbt serves as a team’s [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction), using modular SQL models with testing, documentation, and version control to manage data in the organization in a flexible, cross-platform manner. dbt handles the following ELT tasks: **Transforms data directly within your data warehouse.** dbt runs transformations within cloud warehouses, keeping data in place within the ELT pipeline. dbt compiles [SQL models](https://docs.getdbt.com/docs/build/sql-models) into warehouse-native queries rather than exporting data for transformation. This approach enhances both performance and scalability. **Brings software engineering discipline to analytics.** Transformations are maintained in [version-controlled repositories.](https://docs.getdbt.com/docs/cloud/git/version-control-basics) Teams branch, review, and merge code just as they would with an application. This approach ensures reproducibility because everyone knows precisely which logic created a table and why. **Automates testing and quality checks.** dbt supports creating data tests, with [built-in support for common data quality checks](https://www.getdbt.com/blog/data-quality-testing) such as null values, uniqueness, and relationships. Teams can define custom tests using YAML, and dbt runs them automatically after every model build. This tightens the feedback loop and adds trust in the downstream data pipeline by flagging errors immediately. **Builds documentation and lineage as you go.** Every dbt model automatically generates [documentation and dependency graphs](https://docs.getdbt.com/docs/build/documentation). You don’t need separate data catalog tools to trace the origin of your data. Data lineage becomes an integral part of your transformation process. **Schedules and orchestrates transformations reliably.** dbt Cloud or dbt Core [integrations with tools like Airflow](https://www.getdbt.com/blog/dbt-airflow) and Fivetran handle workflow orchestration. Data pipelines can refresh incrementally or run on a set schedule. This ensures transformed models remain synchronized across dashboards and ML workflows within the ELT pipeline. ## Conclusion ELT pipelines have moved beyond mere technical workflows. They function as data systems that accelerate business growth. This means that business leaders, IT professionals, and sales teams are influenced by ELT pipelines and other types of data pipelines. These pipelines influence decision-making across the entire organization, for better or worse. In big data environments, the key advantage lies in how efficiently organizations model and transform data. dbt provides a framework for building reliable, scalable data models. Designing your ELT pipeline with the right tools ensures that your organization can make informed decisions. Start managing your data transformations effectively and take control of your ELT pipeline with [dbt](https://www.getdbt.com/product/what-is-dbt). --- --- title: "What AI data engineers actually do" description: "See how data engineers are shifting from pipelines to platforms, leveraging AI for testing, documentation, and enablement." url: "https://www.getdbt.com/blog/ai-data-engineer-tasks" date: "2025-09-25" authors: ["Joey Gault"] categories: ["Pulse"] --- # What AI data engineers actually do ## The evolution of core data engineering tasks AI data engineers continue to handle the foundational responsibilities of traditional data engineering while leveraging artificial intelligence to enhance their effectiveness. The most significant change lies not in what tasks they perform, but in how they execute them and where they focus their strategic attention. ### Creating and managing technical artifacts The creation of technical artifacts remains central to AI data engineering work, though the approach has evolved considerably. Data ingestion pipeline development now benefits from AI assistance, where engineers can generate working pipelines from virtually any data source with publicly available APIs. Using tools like [Cursor](https://cursor.com/), engineers can rapidly prototype ingestion solutions that handle pagination, edge cases, and instrumentation requirements. However, the strategic focus has shifted toward leveraging existing frameworks and vendor solutions rather than building custom ingestion code from scratch. [Data transformation](https://www.getdbt.com/blog/data-transformation) work represents perhaps the most AI-enhanced area of data engineering. Within [dbt](https://www.getdbt.com/product/what-is-dbt), practitioners can use [dbt Copilot](https://www.getdbt.com/product/dbt-copilot) to generate or refine SQL, documentation, data tests, metrics, and semantic models, all within governed workflows and subject to human review. Multi-file refactors remain possible with assistance, but results depend on code quality and conventions; changes should land via CI/CD and code review to maintain reliability. ### Automated incident resolution and monitoring AI data engineers increasingly focus on building systems that can diagnose and resolve pipeline failures autonomously. Modern AI systems can analyze complete log outputs from failed pipeline runs, examine associated project code, and generate both diagnoses and proposed resolutions. This capability extends to creating pull requests with fixes that can be automatically tested through continuous integration systems, dramatically reducing the time engineers spend on break-fix activities. The monitoring and optimization of data infrastructure costs has also become more sophisticated. AI assists in identifying performance bottlenecks, suggesting code optimizations, and recommending infrastructure adjustments based on usage patterns and cost analysis. ## Stakeholder collaboration and self-service enablement AI data engineers spend considerable time developing systems that reduce the friction between data teams and business stakeholders. Traditional data engineering often created bottlenecks where business users needed to request data access or analysis through the engineering team. AI-powered solutions are changing this dynamic significantly. ### Context-aware data discovery A major focus area involves implementing context protocols that allow AI systems to understand and provide access to organizational data assets. Engineers work on integrating metadata about data sources, quality indicators, and usage guidelines into systems that can respond intelligently to business user queries. This involves implementing standards like [Model Context Protocol (MCP)](https://en.wikipedia.org/wiki/Model_Context_Protocol) or similar frameworks that enable AI assistants to access comprehensive information about available datasets, their trustworthiness, and their suitability for specific analytical purposes. The [dbt Model Context Protocol (MCP)](https://docs.getdbt.com/docs/dbt-ai/about-mcp) server provides a standard way to expose dbt-managed metadata and execution context to AI applications and agents, enabling governed discovery and action without bypassing controls. ### Natural language data interaction AI data engineers increasingly build and maintain systems that allow business stakeholders to query data using natural language rather than SQL or other technical interfaces. This work involves implementing semantic layers that translate business terminology into appropriate database queries, ensuring accuracy and consistency in results. The engineering challenge lies in creating systems that can understand business context while maintaining data governance and security requirements. ## Framework integration and standardization The importance of frameworks in AI-enabled data engineering cannot be overstated. AI data engineers focus heavily on implementing and maintaining consistent frameworks that provide the standardization necessary for effective AI assistance. ### Leveraging established frameworks Engineers working with AI prioritize using well-documented, widely-adopted frameworks like dbt, Spark, and Airbyte. These frameworks provide the consistency and documentation that AI systems need to generate reliable, maintainable code. The homogeneous nature of framework-based development allows AI to understand patterns, generate appropriate code, and maintain consistency across projects. ### Code quality and consistency AI data engineers spend significant time establishing and maintaining coding standards, documentation practices, and testing protocols that work effectively with AI assistance. This includes implementing consistent CI/CD pipelines, standardized logging and observability practices, and well-documented best practices that AI systems can follow when generating or modifying code. ## Strategic platform development As AI automates more routine tasks, data engineers increasingly focus on higher-level platform and infrastructure concerns. This strategic shift represents one of the most significant changes in the profession. ### Data platform engineering Many AI data engineers evolve toward data platform engineering roles, focusing on the infrastructure that supports data pipelines rather than building individual pipelines. This work involves ensuring performance, quality, governance, and uptime across the entire data ecosystem. Platform engineers design and maintain the systems that enable other team members (both human and AI) to work effectively. ### Automation and business integration Another emerging focus area involves building automation systems that translate data insights into business actions. Rather than simply providing reports or dashboards, AI data engineers create systems that can automatically trigger business processes based on data analysis. This represents a shift from insight generation to action enablement. ## Quality assurance and governance Despite AI's capabilities, data quality remains a paramount concern for AI data engineers. In fact, the stakes for data quality have increased as organizations rely more heavily on AI systems that require high-quality inputs to produce reliable outputs. ### Testing and validation AI data engineers focus extensively on building comprehensive testing frameworks that validate both the data and the AI-generated code that processes it. This includes unit tests for individual transformation logic, data tests that ensure output quality, and integration tests that validate entire pipeline functionality. AI assists in generating these tests, but engineers must design the overall testing strategy and ensure comprehensive coverage. ### Documentation and metadata management Maintaining comprehensive documentation and metadata becomes even more critical in AI-enhanced environments. Engineers focus on creating and maintaining documentation that serves both human users and AI systems, ensuring that context and business logic are clearly captured and accessible. ## Looking forward The tasks that AI data engineers focus on reflect a profession in transition. While core responsibilities around data ingestion, transformation, and quality remain constant, the methods and strategic focus continue to evolve. Engineers who successfully adapt to this new paradigm combine traditional data engineering expertise with an understanding of AI capabilities and limitations. The most successful AI data engineers focus on building robust, well-governed systems that can effectively leverage AI assistance while maintaining the reliability and trustworthiness that organizations require from their data infrastructure. This involves not just technical implementation, but also strategic thinking about how AI can best serve business objectives while maintaining appropriate oversight and control. As AI capabilities continue to advance, these focus areas will likely continue evolving, but the fundamental principle remains: AI data engineers must balance automation and efficiency gains with the governance and quality requirements that make data systems truly valuable to their organizations. ## AI Data engineer FAQs **Will AI replace data engineers?** AI will not replace data engineers but will transform how they work. AI data engineers continue to handle foundational responsibilities like data ingestion, transformation, and quality assurance, but with enhanced capabilities. The most significant change is in execution methods and strategic focus rather than core responsibilities. Engineers now leverage AI to automate routine tasks like pipeline generation, documentation creation, and incident resolution, while shifting their attention to higher-level concerns like platform development, governance, and building systems that enable business stakeholders to interact with data more effectively **How should data teams standardize frameworks, coding conventions, and observability to make AI most effective at building and maintaining pipelines?** Data teams should prioritize well-documented, widely-adopted frameworks like dbt, Spark, and Airbyte, as these provide the consistency and documentation that AI systems need to generate reliable, maintainable code. Teams must establish consistent coding standards, documentation practices, and testing protocols that work effectively with AI assistance. This includes implementing standardized CI/CD pipelines, consistent logging and observability practices, and well-documented best practices that AI systems can follow when generating or modifying code. The homogeneous nature of framework-based development allows AI to understand patterns and maintain consistency across projects. **To what extent can AI automate multi-file refactoring, metric definitions, testing, and incident resolution across dbt-based DAGs?** AI can significantly automate these complex tasks in dbt environments. For multi-file refactoring, engineers can issue natural language prompts to restructure code across multiple parts of a data pipeline, minimize duplication, or propagate new fields throughout the entire transformation graph. AI assists in authoring new transformation assets, generating comprehensive documentation, and creating robust test suites. For incident resolution, AI systems can analyze complete log outputs from failed pipeline runs, examine associated project code, and generate both diagnoses and proposed resolutions, even creating pull requests with fixes that can be automatically tested through continuous integration systems. --- --- title: "Building an AI-ready data platform that supports generative AI" description: "Here’s how to build a data architecture that scales to meet the needs of modern AI-powered applications." url: "https://www.getdbt.com/blog/ai-ready-platform-generative-ai" date: "2025-09-24" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Building an AI-ready data platform that supports generative AI Every company wants to get started with implementing generative AI (GenAI) use cases. You probably have multiple ideas and initiatives ready to roll, ranging from customer support to marketing to easier employee onboarding. The challenge, as always, is the data. AI use cases require massive amounts of data. That’s true whether you’re building your own [Large Language Model (LLM)](https://www.ibm.com/think/topics/large-language-models), fine-tuning an existing one, or adding context to calls against a commercial LLM. Unfortunately, many companies are finding that the legacy [Extract, Transform, and Load (ETL) systems](https://www.getdbt.com/blog/extract-transform-load) they’ve hobbled along on for years can’t scale to the speed and volume demanded by modern AI apps. Meanwhile, companies that have gotten by with homegrown data pipelines are realizing they’re not getting the data quality, consistency, and governance they need to make their AI projects successful. In this article, we’ll look at the new challenges that AI presents for data. We’ll talk about why existing systems tend to miss the mark, and how to evolve your current data architecture into an AI-ready data platform. ## How GenAI use cases differ from analytics use cases First, some good news. At the most basic level, data for analytics use cases and data for AI use cases share a lot in common. Both require access to high-quality data that’s secure and well-governed. In other words, preparing data for AI doesn’t mean starting from scratch. Everything you know about [data quality](https://www.getdbt.com/blog/data-quality-dimensions) applies to AI and analytics data equally. There’s a key difference, however, in how the data’s used: - Analytics applications are **deterministic**. The results are numeric calculations derived using various mathematical formulas directly from the underlying data, and the results are directly connected to the data. - AI applications, such as LLMs, are **probabilistic**. They’re based on neural networks that are trained on large amounts of data containing billions or even trillions of parameters. Questions are interpreted and answered using a transformer architecture, and the outputs aren’t directly relatable to the inputs. If you ask an analytics application a question, you’ll always get the same answer. If you ask an AI application a question, you may get a different answer each time. This has multiple implications for how we handle data. It means that GenAI apps need: - Rich context to return accurate, up-to-date results. That requires bringing in high-quality data from multiple sources across the enterprise. - Access to a wide range of data in various formats. This includes not just the structured data used by analytics but semi-structured and unstructured data (i.e., data that was difficult to mine for value prior to the advent of GenAI). - Rigorous testing (of both inputs and outputs) prior to production use. While this is true for analytics, the importance of data testing increases with AI, as the connection between inputs and outputs is more opaque. - Additional security and governance precautions to avoid [security issues unique to LLMs](https://www.paloaltonetworks.com/cyberpedia/generative-ai-security-risks). ## The challenges involved in enabling AI use cases The problem is getting the right data at the right level of quality, subject to the right governance. Existing ETL and home-grown solutions tend to underperform here for four key reasons: - A lot of data remains siloed - Legacy ETL systems can’t scale to meet demand - Data quality is inconsistent - Metrics definitions vary from team to team ### A lot of data remains siloed Many ETL and home-grown data pipelines rely on data centralization to make production datasets discoverable across the company. This almost always requires that a centralized data team take on new data pipeline projects as they have capacity. This common chokepoint means a lot of key data remains locked away in [data silos](https://www.techtarget.com/searchdatamanagement/definition/data-silo). Much of this data is exactly the semi-structured and unstructured data from which GenAI applications could most benefit. Siloed data is nearly impossible to find. It’s also hard to work with, as it’s often of low quality, out of date, or out of compliance with corporate governance standards. Data silos are a significant drag on businesses. A McKinsey study found that [companies collectively are losing around $3.1T every year](https://www.jpmorgan.com/kinexys/content-hub/collective-intelligence-from-data-silos) due to the lost opportunities locked away in siloed data. Breaking down these data silos isn’t just a technical problem. It’s a business culture problem. Unlocking this data requires that it be discovered, transformed, documented, and made available for use and collaboration. This is a team effort. It requires a set of data tools that are accessible, not just to technically-savvy data producers, but to data stakeholders with varying levels of technical acumen. ### Legacy ETL systems can’t scale to meet demand Many companies are still dependent on legacy ETL systems. These systems come to us from the pre-cloud world and were built primarily to work in resource-constrained environments. Many legacy ETL tools are trying to adapt to modern times. For example, a number of tools now support [Extract, Load, Transform (ELT) pipelines](https://www.getdbt.com/blog/extract-load-transform) as well. Whereas ETL pipelines would transform the data before loading it into a data warehouse, ELT pipelines load the raw data and transform it later. This enables better data traceability, prevents data loss, and allows different teams to transform the same base data for different use cases. However, even when an ETL system has been retrofitted to support ELT workflows, the fact remains that they weren’t originally designed to handle the volume, speed, and variety of data that companies manage today. That leaves them struggling to keep up with the massive demand for data created by today’s AI solutions. ### Data quality is inconsistent Your data doesn’t live in one place. It’s scattered across dozens or hundreds of file systems, object stores, relational databases, data warehouses, and data lakes/lakehouses. This architectural heterogeneity means there often isn’t “one true way” to create data transformation pipelines in a company: - Some teams will use legacy ETL systems or dedicated data orchestration solutions like [Apache Airflow](https://airflow.apache.org/) or [Prefect](https://www.prefect.io/). - Others will program transformations directly into their [Snowflake](https://www.snowflake.com/) data warehouses. - Still others may have an ad hoc collection of Python scripts running as scheduled cron jobs scattered across their cloud infrastructure. This variability means there isn’t a single place to manage or monitor data quality. As a result, data quality differs wildly from org to org, and even team to team. This makes it hard to determine when data is reliable and “AI-ready.” ### Semantic definitions vary from team to team Another downside of this scattershot approach to data transformation is that teams may have totally different definitions of common business concepts. For example, one team might use a different formula to calculate adjusted revenue compared to another. This lack of semantic consistency is a roadblock to data sharing and collaboration. The Finance and Sales teams can’t coordinate their work easily if they’re both using different definitions of basic concepts. AI systems require consistency. Faced with two different definitions of revenue, the system will be at a loss as to which one to use. In the worst case, it may return different answers to different users. To overcome this issue, AI systems require a [semantic layer](https://www.getdbt.com/blog/why-your-ai-will-fail-without-a-semantic-layer)—a single, centralized framework that defines key metrics and business logic. Without a semantic layer, your AI initiatives are doomed to fail. ## dbt as your AI-ready data control plane The solution to these challenges isn’t to move everything into a centralized data store. That’s often an impossible and ill-advised effort. A better solution is to have a single **[data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) **for all your AI data. A data control plane is an architectural layer that sits over all of your end-to-end data activities. It enables integration, access, governance, and protection so you can manage the behavior of people and processes in a distributed and dynamic data environment. [**dbt**](https://www.getdbt.com/product/what-is-dbt) is a data control plane that provides a standardized and cost-efficient way to build, test, deploy, discover, and monitor data for both analytics and AI. It’s flexible, vendor-independent, and collaborative, making it a perfect solution for your GenAI data needs. ### Leverage AI to transform data in a uniform manner Using dbt, you can connect to a wide ecosystem of [data warehouses and other tools across your data stack](https://www.getdbt.com/product/integrations). Teams can use dbt to represent data transformations using [dbt data models](https://docs.getdbt.com/docs/build/models), a combination of YAML code and SQL or Python. Engineers and analysts alike can leverage [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot) to accelerate analytics and AI workflows with our AI-powered assistant. Teams can check all their data transformation code into source control. That makes it easy for data producers to track changes, discover existing code, and collaborate on data workflows. ### Test and monitor for data quality Data testing is often one-off and ad hoc. dbt supports creating [data tests](https://docs.getdbt.com/docs/build/data-tests) alongside your data models so that your teams can verify the quality of their data transformation code before deploying it to production. [Data health signals](https://docs.getdbt.com/docs/explore/data-health-signals) enable data teams and governance specialists to track data quality across all teams using dbt. ### Document data for users (and AI) Most data is lightly documented, if it’s documented at all. With dbt, you can create rich documentation for every single data model, so that data consumers know where data comes from and what it means. You can also supply these docs to AI as context—e.g., using the [dbt MCP (Model Context Protocol) Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server))—to give GenAI invaluable context for understanding your data. ### Define global metrics for AI-powered apps dbt also supports AI data quality through the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl), which you can leverage to represent key business metrics in one central location. You can supply these metrics directly to your LLMs and other AI solutions, eliminating confusion between conflicting metrics definitions. ### Implement governance to protect sensitive data Using both dbt models and the dbt Semantic Layer, [you can ensure proper governance](https://www.getdbt.com/product/governance) of all dbt-managed data: - Govern workflows by standardizing on a single platform and a single source of truth for metrics - Manage access to data models via role-based access control (RBAC) - Leverage auto-generate docs, version control, and visual data lineage to reduce errors and make auditing easy ## Create your AI-ready data platform today Building an AI-ready data platform isn’t an overnight task. The good news is that it doesn’t require tearing everything you’ve built down and starting from scratch, either. dbt connects seamlessly to the existing elements of your data stack. That enables onboarding the critical datasets you need for AI workloads today while leaving your existing data infrastructure in place. Over time, you can move more data pipelines over to leverage the improvements that dbt brings in data quality, scalability, observability, and data discovery. To learn more about how to leverage dbt as the data control plane for your AI journey, [book a demo with a dbt expert today](https://www.getdbt.com/contact). --- --- title: "Rewrite your career at Coalesce 2025" description: "Rewrite your data story at Coalesce 2025, where the dbt community builds the future of analytics together." url: "https://www.getdbt.com/blog/rewrite-your-career-at-coalesce-2025" date: "2025-09-19" authors: ["Daniel Poppy"] categories: ["Community"] --- # Rewrite your career at Coalesce 2025 The analytics world is at a tipping point. The old rules don’t work anymore, and at [Coalesce 2025](https://coalesce.getdbt.com/), we’re rewriting the playbook with you. This is where the dbt community comes alive: practitioners teaching hard-earned lessons, data leaders charting new directions, and thousands of your peers sharing ideas to make data work better. Coalesce 2025 is where you can accelerate your career, no matter your role. Register now before it’s tool. And if you still need proof, keep reading. ## Rewrite what’s possible for your role Whether you’re just getting started or leading a global data organization, Coalesce 2025 will give you skills, tools, and strategies to take your next step forward. **For practitioners**: Get hands-on with the latest and greatest in dbt—including the dbt Fusion engine and dbt Canvas—and become the architect of your organization’s analytics foundation. [Join the dbt MCP hackathon](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda?session=955912a0-ef15-4326-9c65-6a8527a664dd&shareLink=true) and take home working code you can adapt to your projects. [Demonstrate to your organization how your technical dbt work ties to quantifiable business value](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda?session=4d072b4d-f778-4aad-ab0c-0fa26145bd1e&shareLink=true). You’ll learn the tools, calculations, and techniques to be able to talk to your boss — and your boss’s boss — about the business value of your analytics stack. [Step into an immersive experience of a data transformation journey](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda?session=990220f5-319a-4da6-8269-c56e4953cb4e&shareLink=true) from multiple angles: data engineer, manager, analytics VP, and chief data officer. Walk away with a clear picture of the capabilities your team needs (without naming products) and a roadmap for building champions across your organization so that you can be the leader of your data journey. **For data leaders**: Connect with the top minds in the industry at exclusive networking events, and learn from peers who have built trusted, cost-optimized data programs at scale. **** ## Rewrite your network—and your influence Coalesce isn’t just about what you learn. It’s about who you learn with. Thousands of data professionals come together every year to exchange ideas. From hallway chats to structured meetups and workshops, you’ll build relationships that open doors. There are countless opportunities to connect, like the [Women in Data session](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda?session=54251ec6-b598-40a6-b786-227f825d72f4&shareLink=true), and swap stories with folks on their data, analytics, and AI journeys. Hiring managers are here. dbt Labs’ recruitment team is here. Future collaborators are here. Your next mentor—or mentee—is here. Being part of the dbt community doesn’t just help you keep up with the changes in the industry. The dbt community helps you shape the industry. ## Rewrite how you work. Unlock where you go next. [The Coalesce 2025 program](https://www.getdbt.com/blog/what-to-expect-from-sessions-at-coalesce-2025) is designed to accelerate both your day-to-day work and your long-term career. Breakout tracks include: - **Analytics development best practices** — Establish robust analytics processes, improve team workflows, and turn analytics into a strategic advantage. - **Embracing AI** — Build reliable data foundations and accelerate time to value with dbt as a key part of your AI workflows. - **Data modernization** — Tackle complex architectures and achieve tangible results while future-proofing your analytics stack. - **dbt at scale** — Learn how to grow your data environment as your organization grows. - **Empowering self-service** — Enable business users to explore and use data without compromising on governance, performance, or trust. These aren’t vendor pitches. They’re real stories from real teams who’ve built careers and impact through this community. ## Rewrite your impact with training At Coalesce 2025, you’ll gain new skills, new frameworks, and the confidence to bring change back to your organization. [**Sign up for training**](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/training-certification) with dbt experts, and leave with skills that you and your team can apply immediately. Join hands-on labs with dbt resident architects, solutions architects and technical instructors: - Getting started with dbt - Becoming a dbt Architect - Developing with dbt Canvas - Cross-platform Mesh with Iceberg Tables - Managing costs with dbt - Building a data quality framework with dbt - Upgrading to Fusion ## This is your moment. Be the one who rewrites it Coalesce 2025 is the launchpad for the next generation of analytics leadership. It’s the place to rewrite your career story with a community that’s building the future of data work together. --- --- title: "Common data transformations used in ETL processes" description: "Discover the major transformation types used in analytics pipelines and how each contributes to trusted, scalable data products." url: "https://www.getdbt.com/blog/data-transformation-types" date: "2025-09-19" authors: ["Joey Gault"] categories: ["Pulse"] --- # Common data transformations used in ETL processes In today’s data‑driven organizations, one thing is constant: raw data arrives in all shapes, sizes and formats. But before it can drive decisions, it must be cleansed, standardized, modeled and enriched. That’s where transformation comes in. In this article, we’ll walk through the core types of transformations you’ll encounter—such as cleaning, normalization, aggregation and enrichment—and show how each plays a unique role in turning raw inputs into trusted analytics assets. ## The foundation: data cleaning and quality assurance [Data cleaning](https://www.getdbt.com/blog/data-cleaning-transformation-quality) forms the cornerstone of any transformation process, addressing the fundamental quality issues that plague raw datasets. This transformation type focuses on identifying and correcting inaccuracies, filling missing values, and removing duplicate records that can compromise analytical integrity. In practice, data cleaning involves multiple validation layers. Teams implement format validation to ensure data types conform to expected standards—verifying that phone numbers contain the correct number of digits, email addresses follow proper syntax, and dates align with standardized formats. Completeness checks identify null or empty critical fields, while uniqueness constraints prevent duplicate identifiers from corrupting downstream analyses. The impact of inadequate data cleaning extends far beyond technical inconvenience. Poor-quality data can cost organizations up to 30% of their annual revenue through flawed decision-making, broken pipelines, and regulatory compliance failures. This makes cleaning not just a technical necessity but a business imperative that directly affects organizational performance and risk management. ## Normalization: creating consistency across data sources Normalization transforms data into standardized ranges and formats, enabling meaningful comparisons across different sources and time periods. This transformation proves particularly valuable in global organizations where data arrives in various currencies, time zones, measurement units, and formatting conventions. Consider a multinational retail company collecting transaction data from multiple regions. Raw data might include prices in euros, dollars, and yen, alongside varying date formats and measurement systems. Normalization processes convert all currencies to a single standard (such as USD), standardize date formats, and align measurement units, creating a unified foundation for cross-regional analysis. The normalization process often works in tandem with data cleaning, as both focus on establishing consistency and reliability. However, normalization specifically addresses structural and format variations rather than data quality issues, ensuring that mathematically equivalent values can be properly compared and aggregated regardless of their original representation. ## Aggregation: rolling up data for performance and insight Aggregation transforms granular data into summarized formats that improve query performance and reveal higher-level patterns. This transformation type proves essential when dealing with large datasets where individual transaction-level analysis would be computationally expensive or analytically overwhelming. Common aggregation patterns include temporal rollups (daily sales summarized to monthly totals), categorical groupings (customer behavior segmented by demographics), and hierarchical summaries (individual product sales aggregated to category and department levels). These transformations not only improve system performance by reducing data volume but also create the foundation for executive dashboards and strategic reporting. The strategic value of aggregation extends beyond performance optimization. Pre-aggregated datasets enable real-time decision-making by providing instant access to key metrics without requiring complex calculations at query time. This capability becomes particularly important in customer-facing applications where response time directly impacts user experience. ## Generalization and discretization: creating analytical hierarchies Generalization breaks down complex data elements into hierarchical structures that support different levels of analysis. Address data provides a clear example: a single address field can be generalized into separate components for street, city, state, and country, enabling analysis at multiple geographic levels. Discretization complements generalization by converting continuous data into categorical ranges that facilitate targeted analysis and decision-making. Age data transformed into demographic segments (18-29, 30-44, 45-60, 60+) enables marketing teams to develop age-specific campaigns, while income data discretized into brackets supports pricing strategy development. These transformations prove particularly valuable in machine learning applications, where categorical features often perform better than continuous variables in certain algorithms. They also support regulatory compliance requirements that mandate data anonymization through generalization techniques. ## Validation: ensuring data integrity throughout the pipeline Validation transformations verify that data meets specified criteria and business rules before proceeding to analysis stages. Unlike cleaning, which corrects identified issues, validation acts as a quality gate that prevents problematic data from entering downstream processes. Comprehensive validation frameworks implement multiple check types: data type verification ensures numeric fields contain numbers rather than text, range validation confirms values fall within expected boundaries, and referential integrity checks verify that foreign keys correspond to valid primary keys in related tables. Business rule validation adds another layer by confirming that data relationships align with organizational logic—ensuring that order dates precede shipping dates, for example. The validation process becomes increasingly critical as data volumes grow and sources multiply. Automated validation rules catch issues that would be impossible to identify manually, while comprehensive logging provides audit trails that support troubleshooting and compliance reporting. ## Enrichment: adding context and value Data enrichment enhances existing datasets by incorporating external information that provides additional context for analysis. This transformation type moves beyond cleaning and organizing existing data to actively augment it with new dimensions that support deeper insights. External data sources commonly used for enrichment include demographic databases, geographic information systems, weather data, economic indicators, and social media feeds. A logistics company might enrich shipment data with weather forecasts to predict delivery delays, while a retail organization could augment customer records with demographic information to improve segmentation accuracy. The enrichment process requires careful consideration of data quality, licensing, and update frequency for external sources. Teams must establish processes for monitoring external data quality and handling situations where enrichment sources become unavailable or unreliable. ## Integration: unifying disparate data sources Data integration represents one of the most complex transformation types, combining information from multiple sources into unified datasets that provide comprehensive views of business entities or processes. This transformation addresses the reality that most organizations store related information across numerous systems, creating analytical challenges when insights require cross-system perspectives. Customer data integration exemplifies this challenge. A typical organization might maintain customer information across CRM systems, e-commerce platforms, marketing automation tools, support ticketing systems, and financial databases. Integration processes combine these disparate sources into unified customer profiles that support 360-degree analysis and personalized experiences. Integration transformations must address schema differences, data type mismatches, identifier conflicts, and temporal alignment issues. Success requires establishing master data management practices that define authoritative sources for key entities and implement conflict resolution rules when sources disagree. ## Advanced transformation patterns in modern data architectures Contemporary data architectures increasingly support sophisticated transformation patterns that go beyond [traditional ETL approaches](https://www.getdbt.com/blog/extract-transform-load). Real-time streaming transformations enable immediate processing of high-velocity data sources, while incremental processing techniques optimize resource utilization by processing only changed data. Parallel processing architectures allow complex transformations to be split across multiple streams, improving performance while maintaining data consistency. These patterns prove particularly valuable when dealing with large datasets or time-sensitive analytical requirements. The emergence of cloud-native data platforms has also enabled new transformation approaches that leverage elastic computing resources. Teams can now implement transformation logic that automatically scales based on data volume and processing requirements, optimizing both performance and cost. ## Choosing the right transformation approach Selecting appropriate transformation types depends on multiple factors including data characteristics, analytical requirements, performance constraints, and organizational capabilities. Teams must balance transformation complexity against maintainability, considering both immediate needs and long-term scalability requirements. [Modern transformation tools like dbt](https://www.getdbt.com/product/what-is-dbt) have revolutionized how organizations approach these decisions by providing frameworks that support multiple transformation types within unified workflows. These platforms enable teams to implement complex transformation logic using familiar SQL syntax while maintaining software engineering best practices like version control, testing, and documentation. The key to successful transformation strategy lies in understanding that different transformation types serve different purposes and often work together to achieve comprehensive data preparation. A single pipeline might implement cleaning to address quality issues, normalization to ensure consistency, aggregation to improve performance, and enrichment to add analytical value. ## Building scalable transformation workflows As organizations mature their data capabilities, transformation workflows must evolve to support increasing complexity and scale. This evolution requires implementing practices that ensure transformation logic remains maintainable, testable, and discoverable as the number of data sources and use cases grows. Modular transformation design enables teams to build reusable components that can be combined in different ways to support various analytical requirements. This approach reduces development time, improves consistency, and simplifies maintenance by centralizing common transformation logic. Comprehensive testing strategies become essential as transformation complexity increases. Teams must implement automated tests that verify transformation logic produces expected results, validate data quality at each stage, and ensure that changes don't introduce regressions in existing workflows. The most successful transformation implementations treat data transformation as a software engineering discipline, applying practices like code review, continuous integration, and deployment automation to ensure reliable, scalable operations. This approach enables organizations to build transformation capabilities that support both current analytical needs and future growth requirements. Understanding and effectively implementing these various transformation types enables data engineering teams to build robust pipelines that convert raw data into trusted analytical assets. The key lies in recognizing that transformation is not a single operation but a comprehensive process that requires careful orchestration of multiple techniques to achieve reliable, scalable results. ## ETL Data transformation FAQs **What are the different types of ETL data transformation?** The main types of ETL data transformation include data cleaning and quality assurance, normalization, aggregation, generalization and discretization, validation, enrichment, and integration. Data cleaning addresses quality issues by correcting inaccuracies and removing duplicates. Normalization standardizes data formats and ranges across sources. Aggregation summarizes granular data for performance and insights. Generalization creates hierarchical structures while discretization converts continuous data into categories. Validation ensures data meets specified criteria before processing. Enrichment adds external context to existing datasets. Integration combines information from multiple disparate sources into unified datasets. **What normalization techniques are described, and when would you use each?** The article focuses on normalization as a process for creating consistency across data sources rather than detailing specific mathematical techniques. It describes normalization as transforming data into standardized ranges and formats to enable meaningful comparisons. The process involves converting currencies to a single standard, standardizing date formats, and aligning measurement units. This type of normalization is particularly valuable for global organizations dealing with data from multiple regions that arrive in various currencies, time zones, and formatting conventions. It works alongside data cleaning to establish consistency and reliability by addressing structural and format variations. **What's the best way to deal with missing values?** Missing values are addressed through data cleaning processes that form the cornerstone of data transformation. The approach involves implementing completeness checks to identify null or empty critical fields, followed by filling missing values as part of comprehensive validation layers. The cleaning process uses multiple validation techniques including format validation, completeness checks, and uniqueness constraints. Teams should implement automated validation rules and comprehensive logging to catch issues that would be impossible to identify manually, while ensuring that missing value treatment aligns with business rules and maintains data integrity throughout the pipeline. --- --- title: "Why moving from stored procedures to dbt drives trust, talent, and AI-readiness" description: "Migrating from stored procedures to dbt boosts trust, talent, and AI-readiness, for faster, more efficient data transformation." url: "https://www.getdbt.com/blog/why-moving-from-stored-procedures-to-dbt-drives-trust-talent-and-ai-readiness" date: "2025-09-18" authors: ["Ryan Bennett"] categories: ["Insights"] --- # Why moving from stored procedures to dbt drives trust, talent, and AI-readiness _This guest blog post is from Ryan Bennett, lead data engineer at [phData](https://www.phdata.io/). _ As a strategic data leader, you focus on driving business value, building productive teams, instilling company-wide trust in data, and setting your organization’s technological direction (just to name a few). You focus on the “what” and the “why”, asking and answering broad questions like: - “What are we trying to achieve? And why does it matter?” - "Is our strategy aligned with the strategic goals of the organization?" - "Are we driving actionable insights, or just generating reports?" - "How do we quantify and measure the ROI of our work?" Maybe your customer transaction data holds untapped signals that could unlock new growth opportunities. Or, perhaps your organization spends large amounts of time verifying untrusted data. The specifics on how to achieve your vision rest within the domain of your team. The foundation they build upon, though, can either accelerate your mission or tether you to the past. One critical, but often overlooked, foundation is data transformation. For decades, many organizations have relied on stored procedures: SQL code embedded in the database or even loose SQL scripts sitting on desktops to manage data transformations. While these approaches may work in the short term, they create hidden risks: siloed logic that only a few people understand, fragile pipelines that break under scale, and limited visibility for leaders who need to trust the data. In this blog, we’ll explore the hidden ways that stored procedures hinder your strategic goals, and how migrating to a modern platform like [dbt](https://www.phdata.io/dbt/) can unlock greater strategic value. **** ## Building organizational trust with modern data engineering Trust is critical in the world of data, and stored procedures make building organizational trust difficult. Often existing as complex database objects, these black boxes are understood only by the few technical staff who build and maintain them. Stored procedures often lack documentation, tests, or a full view of the data lineage. This opacity makes debugging inefficient, and teams may spend days or weeks unravelling an issue. Beyond eroding trust with data consumers, these delays can have significant business costs. Additionally, without a clear understanding of downstream impacts or tests to give confidence, changes to stored procedures can introduce risk. With built-in tools for documentation, testing, and version control, dbt empowers organizations to create trustworthy data products. Understanding the data lineage from a raw source to a curated dataset is transparent in dbt. ## Setting technological direction As a strategic data leader, you vet and choose platforms and tools for your organization’s future, and that future is increasingly linked with AI. Legacy platforms are often slow to adopt new capabilities or neglect them altogether. Stored procedures are not the language of the modern AI ecosystem. Projects powered by dbt, with their rich metadata and structure, are fertile grounds for AI-assisted development. dbt Copilot allows developers to generate code, documentation, and tests accelerating your organization’s delivery of insights. [The dbt Semantic Layer](https://www.getdbt.com/blog/semantic-layer-introduction) can provide context to other AI tools, enabling capabilities like “Talk to Your Data”. Ever-growing training data from public dbt projects, coupled with the aforementioned metadata and structure, gives dbt an advantage in the modern AI landscape. This advantage will only increase as time passes. ## Attracting and retaining talent with modern tools The technologies and platforms you choose directly impact your organization’s ability to attract and retain talent. Data professionals desire modern tools and systems that align with software development best practices. Technology centered around stored procedures can be a red flag for potential candidates and a source of frustration for your existing team. Additionally, the development of stored procedures can resemble the wild west, with differing code styles and a lack of consistency between developers. This can hinder collaboration and slow down the onboarding of new hires. As an opinionated platform, dbt offers consistent styling and structure between projects. Once familiar with dbt, developers can jump into most projects and provide value quickly. With its standardized workflows, dbt reduces onboarding time for new hires by 30%, lowering ramp-up costs. As your team matures, upskilling SQL-savvy analysts to dbt is an easier proposition than training them on stored procedure development. ## Driving business value with dbt Ultimately, every strategic decision needs to drive business value. The inefficiencies of stored procedures are a tax on your organization’s ability to move quickly, generate insights, and build trust. From complex debugging, manual testing, and risky deployments, stored procedures slow your organization’s development cycle. This negatively impacts your ability to make data-driven decisions. dbt provides value by streamlining the development process. With integrated tooling for documentation, tests, and AI-assisted development, dbt increases development efficiency and decreases time to actionable insights. Data products can be built, tested, and deployed with more confidence and speed. Teams have reported spending significantly less time debugging data issues after adopting dbt, thanks to its clear error messages and modular structure. With fewer outages and rework, dbt frees up to 20% more development capacity. This is a competitive advantage in today’s fast-paced world. Choosing between legacy tools like stored procedures and modern platforms like dbt is a choice between the past and the future. ## FAQ ### What are the risks of using stored procedures for data transformation? Stored procedures are difficult to maintain and scale. They are oftentimes black boxes filled with hidden, complex logic. ### How does dbt support AI-driven data projects? dbt empowers teams to build, test, and deploy data products with the same rigor as software engineering. This leads to cleaner, higher-quality data that produces better results for AI applications. --- --- title: "How to reduce your data pipeline tech support burden" description: "Are your data teams turning into tech support teams? Here’s a more scalable approach to running your data pipelines." url: "https://www.getdbt.com/blog/reduce-tech-support-data-pipelines" date: "2025-09-15" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # How to reduce your data pipeline tech support burden Deriving business value from data isn’t a simple task. It requires combining and transforming raw data from multiple sources to create high-quality [data products](https://www.getdbt.com/blog/data-product-examples). Doing this at scale requires more than just writing a one-off data transformation script and setting up a cron job. It requires creating scalable and robust automated [data pipelines](https://www.getdbt.com/blog/data-pipelines) that can continuously validate and publish new data as it arrives. Many data teams start off trying to build this functionality themselves. It doesn’t take long before they find themselves in tech support hell. Pretty soon, it seems all their available time is spent supporting users and fixing pipeline errors instead of investing in next-generation solutions. ## The components of data pipeline automation To prepare data for commercial use, data engineering teams need to ensure that, at a minimum, all data is: - Processed using high-quality data transformation code that’s been thoroughly tested - Packaged and deployed into easily discoverable data products - Continuously monitored for data quality and governance issues To facilitate this, most spin up some form of data pipeline automation, turning a manual, error-prone, and inconsistent process into one that’s standardized, centralized, and well-governed. These pipelines typically consist of several open-source and commercial data tools, which shepherd data through a multi-stage process: - Ingestion via tools such as [Fivetran](https://www.fivetran.com/blog/navigating-complex-data-landscapes-with-data-observability-through-monte-carlo) and [Apache Kafka](https://kafka.apache.org/) - Data transformation, integration, and testing using [dbt](https://getdbt.com/) - Orchestration and deployment via [Apache Airflow](https://airflow.apache.org/) and [Dagster](https://dagster.io/) - Monitoring via tools such as [Prometheus](https://prometheus.io/) and [Grafana](https://grafana.com/) - Infrastructure support via on-demand cloud systems like [AWS](https://aws.amazon.com/) and [Google Cloud](https://cloud.google.com/?hl=en) ## The challenges with a DIY data pipeline To be sure, there are benefits to rolling your own data pipelines. There are a number of fantastic open-source data tools on the market. This enables data teams to mix and match components to meet their exact needs, often with minimal licensing costs. Unfortunately, as most teams quickly realize, the tech support burden of a DIY approach escalates quickly. A few of the problems that rear their ugly heads include: - Local installation is a headache - Git-based workflows are complicated for many users - Maintaining reliable CI/CD infrastructure is resource-intensive - Little to no support for local testing - Lack of self-service tools ### Local installation is a headache Most DIY systems require anyone who wants to contribute to data pipelines to download, install, and configure a vast array of tools - data warehouse connectors and CLIs, Git, dbt, [sqlfluff](https://www.sqlfluff.com/), etc. Less technical users might struggle to install all of these dependencies correctly - especially if one of their dependencies conflicts with another software package. Data engineering team members are the ones on call to troubleshoot and resolve these issues. The more time spent on tools debugging, the less time they have to spend on more fundamental work, such as improving the company’s overall data architecture. ### Git-based workflows are complicated for many users [Git](https://git-scm.com/) is the source code control powerhouse that powers nearly all automation in the software and data engineering worlds. Using Git repositories, teams can easily collaborate on data analytics code, tracking and reviewing all changes. Pull requests to a repository act as a trigger to kick off an automated testing and deployment pipeline. While Git is ridiculously useful, it’s also complicated. [Even engineers struggle with groking it for the first time](https://news.ycombinator.com/item?id=25123014). It’s so easy to paint yourself into a corner with Git that entire websites exist to explain how to get out of them. ([As one popular site put it](https://ohshitgit.com/), “Git is hard: screwing up is easy, and figuring out how to fix your mistakes is f*****g impossible.”) To make matters worse, not everyone using data is a technical user. A mature [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) involves data engineers, analytics engineers, analysts, business stakeholders, and other roles that overlap or fall in between these catch-all job descriptions. Even for those who are SQL experts and may be able to contribute to data pipeline code, working with Git creates a high barrier to entry. And supporting their issues can become a full-time job in and of itself. ### Maintaining a CI/CD infrastructure is resource-intensive Automated data pipelines are the lifeblood of a data-driven organization. If they go down, business grinds to a halt. Keeping them running is a Herculean effort that requires: - Scaling up to handle incoming data spikes - and scaling back down to avoid unnecessary cloud spend - Detecting or fixing issues - invalid or unexpected data from an upstream source, data drift, cloud computing platform issues, etc. - that can result in pipeline stoppage - Continuously monitoring and improving performance as data workloads grow over time Given the effort involved, it’s no wonder many large companies need a small operations team just to keep their data pipelines running smoothly. ### Little to no support for local testing A good data pipeline will run automated tests on any code changes. This verified that the changes are error-free before running the code in production. However, this involves a time-consuming round-trip. Data engineers have to check in a change to a data pipeline. Then, they have to wait for the pipeline to trigger (which may take a while if others are ahead of them in the build queue) and the test suite to run. If a test fails, they have to sift through the logs, find the root cause, make a fix, and check in another change. That starts the whole time-consuming process all over again. It would be much faster if data engineers could thoroughly test and debug changes on their dev boxes before checking in code. Usually, this involves creating a separate dev data warehouse, with isolated environments for each engineer. This is time-consuming to create. It’s even harder to keep synced with the current state of production. It also significantly increases data development costs. ### Lack of self-service tools Data isn’t worth anything unless the people who need it can find it. One of the most frequent drains on engineering teams is answering basic questions about data, such as: - Where do I find data for [x]? - Where did the data come from? - When was it last updated? - Can I trust it? Most homegrown data pipelines don’t provide a way for users to answer such basic questions for themselves. That results in a large queue of support tickets that fall into the data engineering team’s lap. ## Building low-overhead data pipelines with dbt All of this tech support burden means that data engineering teams spend less time building useful new data products. The good news is that it doesn’t have to be this way. dbt is a [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) that teams can use to build, deploy, and monitor data transformation pipelines at a fraction of the time and cost of a DIY solution. What’s more, dbt provides tooling that enables all participants in the data lifecycle to find, understand, and utilize analytics code - not just data engineers. ### Create scalable data pipelines with a few clicks Using dbt, data engineers can set up [full CI/CD promotion pipelines](https://docs.getdbt.com/docs/deploy/continuous-integration) for their analytics workflows with just a few clicks. All changes to analytics code are checked into Git source control, where other team members can review them before approving for deployment. The CI/CD pipeline can then run all associated data tests in a pre-production environment, and deploy changes only if all tests succeed. This automated release process means deploying analytics code changes isn’t an error-prone manual process that pulls the data engineering team away from more critical work. It also enforces a series of gates that improve the quality of all code shipped to production, reducing the time that engineering spends diagnosing and fixing data issues. ### Analytics that are accessible to everyone Not everyone who touches analytics code wants to write their own SQL or memorize Git commands. That’s why dbt offers multiple ways to create or revise analytics code: - [dbt Studio](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/develop-in-the-cloud), a fully-integrated cloud IDE for all personas that simplifies both editing and checking in code - [dbt Canvas](https://docs.getdbt.com/docs/cloud/canvas), a tool for analysts and business stakeholders that provides a visual editing experience for data transformations - [The dbt extension for Visual Studio Code](https://docs.getdbt.com/docs/install-dbt-extension), for seasoned engineers who are most comfortable with traditional programmer IDEs dbt enables producing data products that are easy to discover, understand, and use: - Built-in support for [documentation](https://docs.getdbt.com/docs/build/documentation) and [data lineage](https://www.getdbt.com/blog/what-is-data-lineage) means stakeholders can more easily understand the origin and purpose of data - All data products are propagated to [dbt Catalog](https://www.getdbt.com/product/dbt-catalog), where stakeholders can self-service answers to their most common questions about data - Stakeholders can also verify the health and quality of data using [data health tiles](https://docs.getdbt.com/docs/explore/data-tile) ### Debug and test locally, as you write The dbt Fusion engine is a rewrite of the dbt engine that speeds up development by implementing a full SQL compiler locally. This enables Fusion to parse their entire projects and understand data dependencies, giving contributors immediate feedback on data model errors as they code. The dbt Fusion engine understands and emulates all major data warehouses, so engineers can test code locally - no need to set up dev instances of the warehouse. It’s written in Rust as a single installable binary, making it simple to install locally. You can power your dbt projects with the Fusion engine in the dbt platform. ## Conclusion DIY data pipeline solutions start with the best of intentions. Ultimately, however, the attending tech support burden means they fail to scale with your business. By migrating your data pipelines to dbt, you can ship more analytics code to production in less time and at less cost. Reduced infrastructure costs, self-service features, and AI-powered productivity tools mean your data engineering teams can spend less time on tech support and more time building tomorrow’s solutions. Try it for yourself today by [signing up for a free dbt account](https://www.getdbt.com/signup). --- --- title: "How data engineering drives business transformation" description: "Data engineering fuels business transformation by scaling analytics, enabling AI, and making trusted data accessible to everyone." url: "https://www.getdbt.com/blog/data-engineering-drives-business-transformation" date: "2025-09-15" authors: ["Joey Gault"] categories: ["Pulse"] --- # How data engineering drives business transformation Business transformation fundamentally depends on the ability to make informed decisions quickly and accurately. Data engineering creates this capability by establishing robust pipelines that move data from various sources (databases, APIs, streaming platforms, and external services) into centralized repositories where it can be analyzed and acted upon. This infrastructure work directly impacts business outcomes. When data is fragmented, inconsistent, or slow to access, teams waste valuable time reconciling conflicting information and making decisions based on incomplete data. Strong data engineering eliminates these friction points by ensuring data is well-organized, governed, and readily available when needed. The transformation from intuition-based to data-driven decision making requires more than just collecting information: it demands systems that can handle the scale, variety, and velocity of modern data. Data engineers build these systems with scalability in mind, creating architectures that can grow with the business and adapt to changing requirements without requiring complete rebuilds. ## Enabling self-service analytics and democratizing data One of the most significant ways data engineering contributes to business transformation is by democratizing access to data across the organization. Traditionally, data analysis was confined to specialized teams with technical expertise. Modern data engineering practices break down these silos by creating self-service capabilities that empower business users to find and analyze data independently. This democratization happens through careful design of data models, implementation of semantic layers, and creation of well-documented, business-friendly datasets. When data engineers build transformation pipelines that clean, standardize, and structure data according to business logic, they enable analysts, product managers, and other stakeholders to focus on generating insights rather than wrestling with data quality issues. The shift toward self-service analytics fundamentally changes how organizations operate. Business teams can respond more quickly to market changes, test hypotheses in real-time, and make decisions without waiting for technical teams to prepare custom reports. This agility becomes a competitive advantage, allowing organizations to adapt faster than competitors who rely on traditional, centralized analytics approaches. ## Supporting advanced analytics and AI initiatives As organizations increasingly invest in machine learning and artificial intelligence, data engineering becomes even more critical to business transformation. AI models are only as good as the data they're trained on, and successful AI initiatives require clean, consistent, and well-governed datasets that data engineers provide. The modern data stack, with its emphasis on [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) workflows, is particularly well-suited for AI applications. By loading raw data into cloud data warehouses first and then applying transformations, organizations maintain flexibility in how they prepare data for different use cases. This approach supports both traditional analytics and the more complex data preparation requirements of machine learning workflows. Data engineers also play a crucial role in implementing the [data governance and quality controls that AI initiatives require](https://www.getdbt.com/blog/understanding-data-governance-ai). They build testing frameworks that validate data quality, implement monitoring systems that detect drift or anomalies, and create documentation that helps data scientists understand the provenance and characteristics of their training data. These capabilities are essential for building trustworthy AI systems that can be deployed in production environments. ## Operational efficiency through automation Business transformation often involves automating manual processes and eliminating inefficiencies that slow down operations. Data engineering contributes to this transformation by creating automated data pipelines that reduce the manual effort required to collect, process, and deliver information across the organization. Modern data engineering practices emphasize automation at every level. [Continuous Integration and Continuous Deployment (CI/CD) pipelines](https://docs.getdbt.com/guides/custom-cicd-pipelines?step=1) automatically test and deploy changes to data transformations, reducing the risk of errors and accelerating the pace of development. Orchestration tools manage complex workflows, ensuring that data processing happens reliably and on schedule without manual intervention. This automation extends beyond technical processes to business operations. When data engineers build real-time streaming pipelines, they enable automated decision-making systems that can respond to events as they happen. For example, e-commerce platforms can automatically adjust pricing based on demand signals, or fraud detection systems can flag suspicious transactions in milliseconds rather than hours. ## Scaling data operations with business growth As organizations grow, their data needs become more complex. What works for a startup with a few data sources and hundreds of users may not scale to an enterprise with dozens of systems and thousands of stakeholders. Data engineering provides the architectural foundation that allows data operations to scale alongside business growth. This scalability manifests in several ways. Cloud-native data architectures can handle increasing data volumes without requiring significant infrastructure changes. Modular pipeline designs allow teams to add new data sources and transformations without disrupting existing workflows. Well-designed data models can support growing numbers of users and use cases without performance degradation. The ability to scale data operations efficiently has direct business implications. Organizations that can handle growing data volumes and user demands without proportional increases in cost or complexity maintain their competitive advantage as they expand. Those that struggle with data scalability often find themselves constrained by their own success, unable to capitalize on growth opportunities because their data infrastructure can't keep pace. ## Breaking down organizational silos Traditional organizational structures often create silos between different business functions, with each department maintaining its own data sources and analysis capabilities. Data engineering contributes to business transformation by creating shared data infrastructure that breaks down these silos and enables cross-functional collaboration. When data engineers build centralized data platforms, they create a single source of truth that all departments can rely on. Marketing teams can access the same customer data that product teams use, ensuring consistent metrics and aligned decision-making. Finance can analyze the same transaction data that operations teams monitor, eliminating discrepancies and improving coordination. This integration goes beyond just technical data sharing: it changes how organizations operate. Cross-functional teams can work more effectively when they have access to the same information. Strategic initiatives that span multiple departments can be executed more efficiently when everyone is working from the same data foundation. ## Enabling real-time business operations The pace of modern business often requires real-time or near-real-time responses to changing conditions. Data engineering makes this possible by building streaming data pipelines that process information as it's generated, rather than waiting for batch processing windows. Real-time data capabilities transform how organizations operate across multiple dimensions. Customer service teams can access up-to-the-minute information about customer interactions and account status. Supply chain managers can respond immediately to disruptions or demand changes. Marketing teams can adjust campaigns based on real-time performance metrics. The shift from batch to real-time processing represents a fundamental change in business operations. Instead of making decisions based on yesterday's data, organizations can respond to current conditions. This responsiveness becomes particularly valuable in fast-moving markets where delays in decision-making can result in missed opportunities or competitive disadvantages. ## The evolution toward AI-enhanced data engineering The data engineering field itself is undergoing transformation as AI capabilities mature. Modern AI tools can already automate many routine data engineering tasks, from writing transformation code to debugging pipeline failures. This evolution will likely accelerate the pace of business transformation by making data engineering capabilities more accessible and efficient. AI-enhanced data engineering tools can generate SQL transformations from natural language descriptions, automatically detect and resolve data quality issues, and optimize pipeline performance without manual intervention. These capabilities reduce the technical barriers to implementing sophisticated data operations, allowing organizations to achieve data maturity faster than previously possible. However, this technological evolution doesn't diminish the importance of data engineering: it amplifies it. As AI tools handle more routine tasks, data engineers can focus on higher-value activities like architectural design, stakeholder collaboration, and strategic data initiatives. The result is likely to be faster, more comprehensive business transformation as organizations can implement data-driven capabilities more quickly and effectively. ## Building sustainable data practices Long-term business transformation requires sustainable practices that can evolve with changing needs and technologies. Data engineering contributes to sustainability by implementing governance frameworks, documentation standards, and quality controls that ensure data systems remain reliable and maintainable over time. Sustainable data practices include version control for data transformations, automated testing to catch regressions, and comprehensive documentation that helps teams understand and maintain complex systems. These practices become increasingly important as data systems grow in complexity and as organizations become more dependent on data for critical operations. The investment in sustainable data engineering practices pays dividends over time. Organizations with well-governed, documented, and tested data systems can adapt more quickly to new requirements, onboard new team members more efficiently, and maintain confidence in their data as they scale. Those without these foundations often find themselves constrained by technical debt and reliability issues that slow down transformation initiatives. ## Measuring and demonstrating value Successful business transformation requires the ability to measure progress and demonstrate value to stakeholders. Data engineering enables this measurement by creating the metrics, dashboards, and reporting capabilities that track transformation initiatives and business outcomes. When data engineers build comprehensive data models that capture key business metrics, they provide the foundation for measuring transformation success. These models can track everything from operational efficiency improvements to customer satisfaction changes, giving leaders the visibility they need to guide transformation efforts. The measurement capabilities that data engineering provides also create accountability and alignment across the organization. When everyone can see the same metrics and understand how their work contributes to broader business objectives, transformation initiatives are more likely to succeed. This transparency helps maintain momentum during long-term transformation efforts and ensures that data-driven decision making becomes embedded in organizational culture. Data engineering represents far more than technical infrastructure: it serves as the enabling foundation for comprehensive business transformation. By creating reliable, scalable, and accessible data systems, data engineers empower organizations to make better decisions, operate more efficiently, and adapt more quickly to changing market conditions. As the volume and importance of data continue to grow, the role of data engineering in driving business transformation will only become more critical. Organizations that invest in strong data engineering capabilities position themselves to capitalize on the full potential of their data assets and maintain competitive advantage in an increasingly data-driven world. ## Data engineering FAQs **What is data engineering?** Data engineering is the practice of building robust pipelines and infrastructure that move data from various sources (databases, APIs, streaming platforms, and external services) into centralized repositories where it can be analyzed and acted upon. It involves creating scalable systems that handle the scale, variety, and velocity of modern data, implementing governance frameworks, and establishing automated processes that ensure data is well-organized, consistent, and readily available for business decision-making. **What does a data engineer do? ** Data engineers build and maintain the technical infrastructure that enables data-driven decision making across organizations. They create automated data pipelines, implement data quality controls and governance frameworks, design scalable architectures that can grow with business needs, and build transformation workflows that clean and standardize data according to business logic. They also implement monitoring systems, create documentation, establish testing frameworks, and design data models that enable self-service analytics capabilities for business users. **What is the difference between data engineering and related fields like data science and data analysis?** Data engineering focuses on building the foundational infrastructure and pipelines that collect, process, and organize data, while data science and data analysis focus on extracting insights from that prepared data. Data engineers create the systems that enable analysts and data scientists to access clean, consistent datasets without having to wrestle with data quality issues. While data scientists build machine learning models and analysts generate business insights, data engineers ensure the underlying data infrastructure is reliable, scalable, and well-governed to support these downstream activities. --- --- title: "How to work smarter (not harder) in data" description: "Women in data share real stories and advice for working smarter—not harder—on The View on Data podcast." url: "https://www.getdbt.com/blog/how-to-work-smarter-not-harder-in-data" date: "2025-09-12" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # How to work smarter (not harder) in data In the third episode of _The View on Data_, hosts Faith McKenna, Erica (Ric) Louie, and Grace Goheen dive into a familiar challenge for anyone working in data: trying to do too much. From over-engineering solutions to chasing edge cases that may never happen, they explore what it actually looks like to “work smarter” in a fast-paced, high-stakes field. They also reflect on burnout, perfectionism, and the practical habits that have helped them create better balance, both in and out of work. This episode is for anyone who's ever tried to brute force their way through a project, fix every edge case, or write 100+ lines of regex to solve a problem that could’ve been handled with copy/paste. Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions. 🎧 Listen & subscribe: [Spotify](https://open.spotify.com/show/0Itxg3BoF6zaQdORfVuY6I?si=6133f88257624765) | [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-view-on-data/id1835755886) | [Amazon Music](https://music.amazon.com/podcasts/84c47dbe-adf7-4d33-a371-a1fe55255495/the-view-on-data) | [YouTube](https://youtu.be/iBDkLyqwkOw) ## Why data practitioners tend to over-engineer For Ric, working harder felt like a badge of honor, until it wasn’t. “I thought working a lot meant that I was good at my job… I would spend my evenings going through all these different solutions rather than just tapping on the board and being like, ‘What are we actually trying to solve here?’” Grace shared a similar story from her days as a practitioner, when she wrote a massive regex script for the Codegen package. “It’s so gross and stupid… I probably would’ve been faster just copying and pasting SQL. But I wanted to make it reusable, I wanted it to be perfect and it totally didn’t need to be.” Perfectionism is a powerful force in data work, especially when the job is to make sense of messy systems. But that instinct to fix everything can often get in the way of making progress. ## How to define “good enough” in data work One of the biggest lessons the group shared is the importance of scoping tightly and building toward a minimum viable product (MVP). “If you find yourself tunneling, step back and ask: what are we really trying to do here?” Ric said. “What’s the smallest version of this that actually works?” Having a shared definition of “done” helps, too. Faith talked about the pull request templates and model documentation standards she used as a practitioner: “We had three questions every model had to answer: who built it, who asked for it, and are we planning to upgrade it in the future? It gave us a shared sense of when something was good enough to ship.” ## When automation is helpful, and when it isn’t The team also shared stories about how they’ve learned to automate wisely, focusing on repetitive tasks and known processes. “Automate the boring stuff,” said Ric. “But if it’s a high-risk workflow or something ambiguous, you probably still need a human in the loop.” Grace brought up her work on CI features that compare changes between dev and production environments: “It’s one of those things that felt so repetitive when I was a practitioner—doing manual checks on every PR. Automating that saved a ton of time.” Faith noted that she’s started using voice-to-AI chat to track progress across her different responsibilities: “It’s helped lessen the mental load, because I can say, ‘Hey, remember when I told you about those four projects? I finished two. What’s left?’ And it remembers.” The key theme? Use automation to reduce friction, not add complexity. ## Managing context switching in data jobs The hosts all agreed: constant context switching is part of the job, but it doesn’t have to be painful. Ric described how she sets up her workweek using a stream deck and structured Notion workspace: “Every Monday, I press a button, and all the docs I need pop up. It’s like a daily routine I can follow, instead of having to figure it out every day.” Grace relies on time blocking to stay focused: “I try to do more time-boxed deep work. Like, all of Wednesday is for one thing. And when I switch, it’s a conscious choice, not just reacting to every Slack ping.” Faith shared how she uses a daily work journal to reflect and prioritize: “At the end of each day, I write down three things: what I did, what I’m proud of, and what I need to do tomorrow. It helps me remember where I left off and what matters most.” ## Creating space for life outside of data While working smarter matters inside the office, the team also talked about what it means to _fully_ log off. Grace finds balance through acting: “I do theater, which means I have rehearsal at 6 p.m. And that forces me to stop working because I have to physically go somewhere else.” Faith echoed the need for built-in rituals: “My husband and I go to a weekly yoga class. I also do improv at a theater near my house, which is great because it activates a totally different part of my brain.” For Ric, the lesson came after hitting burnout: “At the end of 2023, I took a three-month leave. I realized that just because I love my job doesn’t mean it should take over my life.” Now, she holds herself to a simple rule when deciding whether to keep working late: “Does this need to be done right now? Does it need to be done by me? What would happen if I waited until tomorrow?” ## Episode takeaways This episode is a practical and personal reminder that doing good work in data isn't about solving every problem. It's about solving the right problems, at the right time, in a way that's sustainable. Here are a few takeaways from the conversation: - Over-engineering often comes from perfectionism, not necessity - MVP thinking creates clarity and momentum - Smart automation frees up focus, but should be intentional - Time-blocking and rituals can ease context switching - Work-life balance requires clear boundaries and regular check-ins --- --- title: "What to expect from sessions at Coalesce 2025" description: "Rewrite your future. This is where the most impactful data conversations happen." url: "https://www.getdbt.com/blog/what-to-expect-from-sessions-at-coalesce-2025" date: "2025-09-10" authors: ["Daniel Poppy"] categories: ["Community"] --- # What to expect from sessions at Coalesce 2025 The dbt community gathers at Coalesce to share the best ways to get data work done. Not in theory, but in practice and in production. The Coalesce 2025 breakout sessions are built for doers. They’re curated for the challenges that data teams face daily and the tools and practices they’re‌ using to solve them. Whether you're scaling globally or just getting started with dbt, there's a track waiting for you at Coalesce 2025. Here’s what to expect: ## Track: **Data modernization** Learn how to go from legacy to leading edge. This track showcases real before-and-after stories of platform migrations and process overhauls with dbt as the linchpin. You’ll hear how teams prioritized their modernization efforts, tackled complex architectures, and achieved tangible results while future-proofing their analytics stack. **This track is perfect for:** Data platform owners, modernization champions, and technical leaders rewriting the next chapter of their organization’s data journey. Featured sessions: - **Bumble** – _Starting from day one: 7 ways we reduced our tech debt and prepared for ML/AI_. Setting up a dbt project the right way from the start is critical to ensuring long-term maintainability, scalability, and efficiency. Learn how Bumble uses dbt to minimize tech debt that had caused ‌operational headaches and performance bottlenecks. - **Riot Games** – _Leveling up data engineering at Riot: How we transformed DevEx with dbt_. Riot Games reduced its Databricks compute spend and accelerated development cycles by migrating from bespoke Databricks notebooks and Spark pipelines to a scalable, testable, and developer-friendly dbt-based architecture. You’ll learn how they designed and operationalized dbt to support Riot’s evolving data needs. - **Ericsson** – _Ascending data Everest: Ericsson's climb to scalable analytics._ Learn how Ericsson’s Enterprise Wireless Solutions team modernized its analytics stack by replacing legacy tools like Alteryx and ThoughtSpot with scalable, modular data models powered by the dbt platform. Learn how they tackled SOX compliance, streamlined operations, and expanded from internal analytics to external analytics to create enterprise-grade analytics performance. - **Airbyte** – [_One platform, two-way value: How data activation drives business impact_](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda?session=e89540ff-ef24-4684-9d69-b6bd35531db4&shareLink=true). A leading life sciences company discovered the hidden cost of disconnected tools when it was locked into an expensive, vertically-focused data integration platform. Find out how the company consolidated its data activation workflow (ELT and reverse ETL) on a unified platform with Airbyte connected to dbt. ## Track: **Embracing AI** AI is only as good as the data behind it. This track explores how companies are identifying real-world AI use cases, building reliable data foundations, and accelerating time to value with dbt as a key part of their AI workflows. **This track is perfect for**: Leaders exploring best practices for AI adoption and for practitioners building the infrastructure to support it. **Featured sessions:** - **Bilt Rewards**: _Bilt's conversational data layer: How we connected data to LLMs with dbt._ Bilt Rewards turned their dbt project into a natural language interface. By connecting their semantic layer and underlying data warehouse to an LLM, business users and data analysts can ask real business questions and get trusted and creative insights. This session shows how Bilt modeled their data for AI, how they kept accuracy intact, and how they increased data-driven conversations across the business. - **JP Morgan**: _Building autonomous AI agents for finance._ How can AI agents make informed financial decisions—without compromising compliance, accuracy, or trust? In this session, learn how a major finanical institution is bridging the gap between experimentation and enterprise-grade financial intelligence. Drawing from real-world use cases, explore how autonomous agents can parse annual reports, extract insights from earnings calls, forecast financial outcomes, and evaluate options strategies. - **Hex** – _[How to measure your data team’s ROI in the age of AI](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda?session=d219ed20-736a-466c-adf8-6b51f771c711&shareLink=true)._ New approaches to teaming, tracking, and tech are helping data teams shift from cost centers to strategic powerhouses. We’ll share how Hex customers are leading this change, how AI is evolving the space, and actionable takeaways for all data roles. - **Omni **– _[700 users, 100 models, $20: Cribl’s blueprint for scalable AI-powered insights](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda?session=c17c6422-0298-4b68-aa7a-c6947c2b6731&shareLink=true)._ Cribl scaled self-service analytics with just 100 dbt models, a $20 auto-doc hack, and smart governance practices. Learn how Git, SDLC workflows, and Omni’s dbt integration helped streamline development, reduce manual effort, and keep AI outputs grounded in clean, trusted data. - **AWS** – _[Building with Kiro + dbt MCP Server and Amazon Bedrock AgentCore](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda?session=e5744bda-03d9-47bb-a9fe-9462e4390ee3&shareLink=true)._ See how AI is shaping developer tools for the next generation. AWS Kiro provides a powerful general purpose IDE great for a wide variety of programming languages and tasks. Pairing Kiro with the dbt MCP Server makes this even more powerful by providing access to dbt specific functionality, context and agents. In addition to this, interact with an external agent deployed in Amazon Bedrock AgentCore. ## Track: **dbt at scale** More teams. More models. More impact. If your data environment is growing faster than your ability to manage it, this track is for you. See how enterprises are using dbt to mature their development processes, enable cross-functional collaboration, and unlock faster delivery of insights—at real organizational scale. **This track is perfect for**: Platform teams, data enablement leads, and enterprise data architects managing growth. Featured sessions: - **Okta** – _From silos to AI-ready: How Okta built governed analytics at scale_. Okta is rebuilding its data platform, moving from disconnected pipelines to a scalable, well-governed foundation with Snowflake and dbt. In this session, you’ll learn how they improved CI/CD, enabled self-service, and are now testing a prototype to make dbt models ready for AI. - **Zendesk** – _How we built a cross-domain, multi-platform data strategy at Zendesk_. Serving hundreds of engineers and analysts across dozens of domains takes more than good tooling. Learn how Zendesk scaled dbt across its distributed Snowflake architecture and multiple development environments. They’ll also dive into topics like advanced CI/CD, hybrid mesh, and workload portability. ## Track: **Empowering self-service** See how dbt customers are enabling business users to explore and use data without compromising on governance, performance, or trust. These sessions focus on what it takes to scale self-service across an organization. **This track is perfect for**: Data teams supporting self-serve environments and business users who want better, faster answers. **Featured sessions:** - **Sweetgreen** – _From chaos to clarity: How Sweetgreen streamlined data with AI and dbt_. In this session, you’ll learn how Sweetgreen is transforming its data ecosystem with dbt to enable self-serve analytics and accelerate development. Sweetgreen will walk through how they built a trusted semantic layer, simplified their Snowflake assets, and began using AI to boost delivery speed. Whether you're focused on reliability, scalability, or empowering more users with data, this session will offer practical takeaways from Sweetgreen’s data journey. - **Virgin Media O2** – _From merge to momentum: How Virgin Media O2 built scalable self-service with dbt across two org_s. Merging two large organizations with different tools, teams, and data practices is never simple. Virgin Media O2, we used dbt to help bring consistency to that complexity, building a hybrid data mesh that supported self-serve analytics across domains. In this session, we’ll share how we gave teams clearer ownership, put governance in place using dbt, and set up secure data sharing through GCP’s Analytics Hub. If you’re working in a federated or fast-changing environment, this session offers practical lessons for making self-serve work at scale. - **Snowflake** – [_How WHOOP unlocks smarter decisions with dbt and Snowflake_](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda?session=79d9dc81-1d49-4f9b-bc45-4a4d0fc32ce8&shareLink=true). Join Snowflake and WHOOP to explore how product integrations with dbt are revolutionizing data transformation. Snowflake will cover real-world business use cases they're helping their clients unlock. WHOOP will share how they use dbt and Snowflake to drive smarter decisions and enterprise-wide collaboration with a custom chat app. ## **Track: Analytics development best practices** Rewrite how how data work gets done. Explore proven approaches data teams have used to work more efficiently and collaboratively, and deliver real value for their organizations. Sessions in this track focus on establish robust analytics processes, improving team workflows, and turning analytics into a strategic advantage. **This track is perfect for**: Data practitioners and data leads looking to level up their analytics development lifecycle with best-in-class practices. **Featured sessions:** - **Daiichi Sankyo** – _How we are building a federated, AI-augmented data platform that balances autonomy and standardization at scale_. **Platform engineers from the global pharmaceutical company share their journey in creating a cloud-native, federated data platform using dbt, Snowflake, and Data Mesh. Learn how they established foundational tools, and standards to develop automation and self-service capabilities. - **Tatango** – _dbt data migration: Best practices for large-scale projects_. In this session, attendees will learn best practices for migrating a large-scale dbt project from one platform to another. Learn how to run dual-execution environments, address syntactic differences and implement data integrity tests to ensure your models are running smoothly. - **AWS** – [_Best practice for leveraging Amazon Analytic Services + dbt_](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda?session=7e944f44-6f77-4053-94c9-a3768512f4be&shareLink=true). As organizations increasingly adopt modern data stacks, the combination of dbt and AWS Analytics services emerged as a powerful pairing for analytics engineering at scale. This session will explore proven strategies and hard-learned lessons for optimizing this technology stack to use dbt-athena, dbt-redshift, and dbt-glue to deliver reliable, performant data transformations. They will cover case studies, best practices, and modern lakehouse scenarios with Apache Iceberg and Amazon S3 Tables. ## Track: **Data quality and trust** Trust isn’t given. It’s built. Learn how leading organizations are maintaining data quality at scale. Learn how they’re operationalizing testing, enabling governed collaboration across teams, and creating environments where data consumers trust what they see—and use it to drive impact. **This track is perfect fo**r: Data stewards, governance pros, and any team fighting the “what does this number even mean?” battle. **Featured sessions**: - **United Services Automobile Association (USAA)** – _Towards a more perfect pipeline: CI/CD in the dbt Platform_. In this session USAA will show how they integrated CI/CD dbt jobs to validate data and run tests on every merge request. Attendees will walk away with a blueprint for implementing CI/CD for dbt, lessons learned from their data journey, and best practices to keep data quality high without slowing down development. - **The Information Lab** – _Achieve governance harmony: Powering Tableau with dbt_. Explore the growing dbt and Tableau partnership. Learn how integrations enhance the modern analytics stack, streamline workflows, and provide accurate, actionable insights with strong data governance. - **Datadog** – [_How trusted data powers growth at Ramp._](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda?session=b32bdd91-8788-4326-9fb0-dbb9d378cf65&shareLink=true) Ramp is one of the fastest-growing software companies of all time, serving over 40,000 businesses with modern spend management. Their growth is fueled by the close partnership between product and data teams, who together turn messy, real-world inputs into trusted insights and data-powered capabilities. In this session, Kevin Hu, former CEO of Metaplane now leading Data Observability at Datadog, and Drew Pinta, Head of Growth Analytics at Ramp, will walk through how trusted data powers critical growth efforts: from ad tech optimization to answer engine optimization. ## Track: **dbt Product** The roadmap to what’s next. Find out what’s new, and what’s coming soon, across the dbt platform and dbt Core. These sessions come straight from the product teams building the future of dbt. Learn how to use new capabilities, give feedback, and stay ahead of the curve. This track is perfect for: Everyone. Power users, architects, and anyone responsible for maximizing the value of dbt in their organization. **Featured sessions:** - _Next-gen data development with dbt Fusion engine_. The next-generation dbt Fusion engine brings a faster, more responsive, and more intuitive development experience to dbt. In this session, you’ll see firsthand what Fusion can do for your team. From real-time, intelligent validation of your code to state-aware runs that skip the rebuilds you don’t need, discover how Fusion helps you move faster, ship with confidence, and stay in flow. - _Shaping the future of self-service with the dbt MCP server and Semantic Layer_. In this session, you’ll learn how the dbt Semantic Layer and dbt MCP server enable safe, AI-powered access to data through natural language interfaces. As business users start exploring data conversationally, the analyst role is evolving from writing queries to curating trusted logic, guiding usage, and enabling scalable self-service. We’ll break down what this shift looks like and how dbt helps analysts stay at the center of decision-making. - _Governed data meets context-aware AI: How analysts build better in dbt._ Learn why dbt is the best place for analysts to build reliable data products quickly, combining structured workflows with context-aware AI. We will explore how dbt Canvas, dbt Studio, and more give analysts the visibility, control, and flexibility they need to move from exploration to production without relying on engineers. You will also see how AI agents in dbt accelerate development by recommending logic, surfacing relevant models, and helping troubleshoot issues—making self-service both faster and more trusted. ## Experience Coalesce your way Breakout sessions at Coalesce aren’t just informative, they’re transformative. You’ll leave with guidebooks to make your data work more scalable, trusted, and impactful. Join us in Las Vegas to catch any session you want, or join online for a selection of breakout sessions livestreamed for virtual attendees. [Explore the full breakout agenda](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda) and [register for Coalesce 2025 now](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/summary). --- --- title: "ETL is over: faster, smarter pipelines with dbt" description: "Learn how Tredence and dbt accelerate migration from legacy ETL to trusted, scalable analytics with speed and lower cost." url: "https://www.getdbt.com/blog/etl-is-over-faster-smarter-pipelines-with-dbt" date: "2025-09-09" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # ETL is over: faster, smarter pipelines with dbt In data analytics, there's always been a trade-off between speed, data trust, adoption, and cost. Today, AI has completely reset expectations around the time it takes to deliver insights. Everyone, but especially executives, now expects insights in minutes, not days. They want to know their decisions are based on trusted business and operational data. They want self-service. And they want to know that asking deeper questions leads to deeper insights, not to data teams scrambling to rebuild pipelines. Even if you’re not embracing AI today, your data teams are under pressure to move faster than ever. We’ll look at how dbt and GenAI-enabled data migration tools from [Tredence](https://www.tredence.com/services/data-modernization#service) can help you meet those demands for speed without sacrificing trust, adoption, or cost efficiency. ## Why legacy ETL is failing Unfortunately, [the old extract-transform-load (ETL)](https://www.getdbt.com/blog/extract-transform-load) paradigm can’t meet the need for speed, trust, adoption, and collaboration in today’s AI-driven data analytics paradigm. Data engineers spend more time fixing broken pipelines than building robust new ones. Analysts, distanced from the ETL process, either doubt the metrics or lose momentum waiting for slow transformations. Even when they trust the data, sluggish delivery limits adoption. Teams may resort to pulling their own extracts, fragmenting insights, and weakening the company’s overall intelligence. When engineers do keep pace, the sheer data volume inflates cloud costs, and pipeline rework triggers unpredictable spikes. And when trust erodes or speed lags, stakeholders stop asking the big questions, allowing your faster‑moving competitors to turn data into insights first. ### ETL pipeline challenges Here are some of the common challenges you may face with a legacy ETL pipeline process: **Performance and scalability bottlenecks:** Typical ETL pipelines are hardware‑ and resource‑intensive, slow to run, and costly to scale, often requiring expensive infrastructure upgrades just to keep pace with demand. **Collaboration gaps:** Without modern version control and [CI/CD integration](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud), teams struggle to share, review, and test changes in real-time. This slows delivery and increases the risk of errors. **Data quality and validation challenges:** Manual checks for every job slow down developers and stall pipeline creation. The lack of [integrated version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) and [automated analytics code deployments](https://docs.getdbt.com/docs/deploy/jobs) further hinders accuracy and rapid deployment. **High total cost of ownership (TCO): **The higher demand for performance and scalability, combined with the need to process data further from the source, results in costly hardware and added licensing fees. That, in turn, drives up infrastructure overhead and TCO. **Cloud cost volatility: **Inefficient pipelines, frequent rework, and poor visibility make monthly cloud spend unpredictable. That translates into sudden spikes that are hard to forecast or control. Legacy ETL forces teams into trade‑offs between speed, trust, adoption, and cost — trade‑offs modern data teams can no longer afford. By contrast, [dbt](https://www.getdbt.com/) eliminates those compromises by embedding software‑engineering best practices like CI/CD, version control, [modular code](https://www.getdbt.com/product/data-modernization), and [automated testing](https://www.getdbt.com/product/test-and-observe) directly into your [modern analytics workflow](https://www.getdbt.com/resources/the-analytics-development-lifecycle), so trust is engineered in from the start. ## How dbt and Tredence overcome legacy ETL challenges By addressing the trust, speed, and cost constraints of legacy ETL, dbt creates the foundation for faster, more collaborative, and more cost‑efficient analytics. Getting there from a legacy stack is where the [Tredence migration framework](https://www.tredence.com/services/data-modernization) and T‑Converter accelerator tool come in. T‑Converter automates the conversion of legacy ETL jobs into dbt‑native models that run directly in your cloud warehouse. This enables you to adopt dbt’s test-driven, version-controlled workflows, minimizing the need for costly hardware overhauls. T-Converter’s AI-powered accelerator parses ETL metadata, generates dbt-native SQL and YAML, and integrates with Git-based CI/CD for automated testing, documentation, and deployment. The result: clean, trusted data ready for analysis in a fraction of the time. In our _“[ETL is Over: Faster, Smarter Pipelines with dbt”](https://www.getdbt.com/resources/webinars/etl-is-over-faster-smarter-pipelines-with-dbt)_[ webinar](https://www.getdbt.com/resources/webinars/etl-is-over-faster-smarter-pipelines-with-dbt), [Devang Pandya](https://www.linkedin.com/in/dpandya-data/), VP of Growth Partnerships at Tredence, showcased the company’s five-step migration methodology and demonstrated the value of the T‑Converter accelerator. This proven approach, successfully deployed across multiple enterprises, speeds migration from legacy ETL platforms, like [Informatica](https://informatica.com/) and [SAP DS](https://www.sap.com/products/data-cloud/data-services.html), to [Snowflake](https://www.snowflake.com/) and dbt. This reduces time to value by unlocking faster data workflows. ### Exploring the Tredence migration journey Tredence breaks migration into five steps, each designed to move you from legacy ETL to a modern, dbt‑powered cloud stack with speed and confidence. **1. Discover and define**: Assess the current data landscape, map dependencies, and set the migration roadmap using the Tredence T‑Analyzer to accelerate discovery. **2. Data migration:** Move historical and incremental data into the target cloud platform (e.g., Snowflake), optimizing architecture along the way. **3. Convert and build: **Use T‑Converter to translate legacy ETL such as Informatica and SAP DS into dbt-native SQL/YAML, enabling test‑driven, version‑controlled workflows. **4. Retrofit and optimize:** Validate outputs, tune performance, and align processes for cost‑efficient operation. **5. Sustain**: Embed governance, monitoring, and continuous improvement to protect long‑term value. In Tredence’s broader modernization framework, two additional stages extend the value. The **Enable AI/ML Ops **step integrates machine learning workflows to support advanced analytics. Finally, **Validate and Optimize** fine‑tunes workloads for performance, scalability, and cost efficiency while ensuring outputs match legacy baselines. ### The Value of the Tredence Approach Using this methodology and its AI-powered accelerators, Tredence has helped customers accelerate the time-to-value of their data migrations, improving accuracy and trust while reducing costs. An analysis by Pandya highlighted the following benefits: - 80% faster data discovery - 30% productivity improvement - 50% faster data migration - 80% savings in platform costs in one year One global travel and hospitality company, with a large Informatica and SAP DS footprint, utilized Tredence to modernize its infrastructure to a Snowflake and dbt stack. Tredence ingested all legacy jobs into its converter, analyzed the complexity of the existing landscape, and built a migration roadmap identifying components to retire, consolidate, redesign, or directly convert. The result was 50% faster delivery, a 90% reduction in errors by eliminating manual code rewrites, and a 30% productivity gain. Tredence’s approach provides customers with a modern, cloud-native foundation that is faster, more accurate, and less costly than their legacy ETL environment. Paired with dbt’s collaboration, testing, and version control, teams can maintain trusted data, adapt quickly, and prepare for the broader advantages dbt offers at scale. ## The dbt platform vision for modern data pipelines dbt helps data teams go from raw data to reliable insights faster, with tools to help you build robust transformations, manage production pipelines, and analyze data securely at scale. Using dbt, you can bring software engineering best practices to analytics with the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle). Once you’ve migrated your data to dbt with a tool like Tredence, you no longer need to face those traditional ETL tradeoffs between speed, trust, adoption, and cost. ### Speed Speed is about more than just writing SQL faster or producing more transformations. It’s about removing blockers so everyone, from engineers to analysts, can work with confidence. dbt acts like a baton that passes context, trust, and automation along the workflow to accelerate the whole team. The team gets shared definitions, integrated workflows, and consistent experiences that reduce rework and increase velocity. Unlike traditional ETL or BI tools built for a single persona, dbt provides a shared context and common language for all roles, preserving both trust and speed across the workflow. ### Trust One of dbt’s founding principles is that data teams should work like software teams. That’s why we built features like CI/CD, version control, modular code, and testing into the framework from the start. Today, these practices are formalized in the ADLC, helping teams reduce rework, collaborate across roles and organizations within a company, and move fast and iterate without fear of breaking things. Trust isn’t an afterthought — it’s engineered into every step. ### Adoption and scale dbt began as the standard for data transformations, turning raw data into analysis‑ready insights. Today, we offer a [unified data control plane](https://www.getdbt.com/resources/whitepaper-the-control-plane-for-data-collaboration-at-scale) that integrates orchestration, observability, cost management, catalog, and semantics into a single interface powered by the [dbt Fusion engine](https://www.getdbt.com/product/fusion). Processing still happens in‑database, and now increasingly across multiple warehouses and engines, thanks to our [data mesh architecture](https://www.getdbt.com/blog/data-mesh-getting-started). The platform supports a wide range of personas, including data engineers, analytics engineers, analysts, and business stakeholders. We’ve recently introduced the [dbt MCP Server](https://www.getdbt.com/blog/build-reliable-ai-agents-with-the-dbt-mcp-server), which provides capabilities for [AI agents](https://www.ibm.com/think/topics/ai-agents) and accelerates the value of dbt. This breadth makes it possible for virtually anyone in the company to participate in the analytics workflow. ### Cost dbt doesn’t just improve developer productivity. It helps your whole organization become leaner by reducing duplication, waste, and confusion across the stack. One of the most significant expenses for modern data teams is platform cost. dbt’s [cost optimization](https://www.getdbt.com/product/cost-optimization) features help drive that down. With observability and automation embedded directly in the transformation layer, teams can spot inefficiencies, prevent unnecessary compute costs, and align data investments with business impact. [ROI tooling](https://www.getdbt.com/resources/study-forrester-tei) makes it easier to measure and demonstrate these savings. ### How leading companies win with dbt Don’t just take our word for it. Organizations across industries are demonstrating what’s possible when speed, trust, scale, and cost efficiency come together in a single transformation framework. With dbt, teams move fast while preserving trust in their data, driving adoption across roles, and keeping control over business logic, pipelines, and cloud spend. **Proven results include:** - **[Enpal](https://www.getdbt.com/case-studies/enpal): **30X faster delivery - **[Pepperstone](https://www.getdbt.com/case-studies/pepperstone): **80% decrease in inconsistent reports - **[Siemens](https://www.getdbt.com/case-studies/siemens): **300,000 employees and 70,000+ data consumers powered by dbt at massive scale - **[Roche](https://www.getdbt.com/blog/roche-unifies-data-enables-ai): **70% cost reduction by consolidating tools like Informatica, Talend, and Microsoft into a single, globally adopted transformation framework Today’s dbt — with [AI copilots](https://docs.getdbt.com/docs/cloud/dbt-copilot), intelligent observability, and conversational interfaces — is the control plane for modern data teams, uniting transformation, governance, and AI to supercharge every step of the data workflow. Overcome your ETL challenges - watch the [webinar](https://www.getdbt.com/resources/webinars/etl-is-over-faster-smarter-pipelines-with-dbt) or [request a demo of dbt today](https://www.getdbt.com/contact). --- --- title: "Inside the dbt Canvas Ambassador Program" description: "The dbt Canvas Ambassador Program helps teams adopt dbt Canvas faster with hands-on practice, feedback, and shared best practices." url: "https://www.getdbt.com/blog/dbt-canvas-ambassador-program" date: "2025-09-09" authors: ["Faith McKenna", "Hrishi Kulkarni"] categories: ["Product"] --- # Inside the dbt Canvas Ambassador Program dbt’s training and customer marketing experts teamed up to pilot a hands-on, champion-led program to help teams adopt dbt Canvas, our visual modeling workspace built for governed self-serve. Instead of a traditional one-to-many session, ambassadors participated in hands-on, one-on-one practice, with opportunities to network and exchange best practices. Ambassadors praised the program for accelerating their onboarding and providing them with a direct feedback loop for continuous improvement. Our customers shared that dbt Canvas makes complex models easier to read, safer to change, easier to both document and explain, and have an overall positive product experience. ## Why did we launch an Ambassador Program? When you introduce a new product line, adoption starts with a single internal champion. That single person rallies their colleagues, proves value on one problem, thus driving both adoption and growth As a lead technical instructor, I have known this intimately from leading data teams around the world through their dbt Platform onboarding. I designed the **dbt** **Canvas Ambassador Program** in partnership with our customer marketing lead to deliver customer participants a personalized one-on-one experience and a space to both build and ship, not just watch. Participants completed an asynchronous course on dbt Canvas, then met live with peers from other companies to share best practices, discuss how dbt Canvas fit into their workflows, network, and complete guided practice together. These customer participants left as Canvas Ambassadors with a repeatable playbook to unblock their teams and a direct feedback loop to our product team. ## The problem dbt Canvas is tackling across engineering and analytics Modern data teams move fast, but two issues potentially slow them down: - **Understanding and safely evolving complex SQL.** New joiners or partner teams struggle to parse someone else’s logic. Refactoring is risky; small changes feel costly. - **Scaling self-serve without sacrificing governance.** Analysts teams need to iterate without waiting in line, without creating the shadow pipelines and parallel logic that often happen in external visual ETL tools, like Alteryx. Meanwhile, leaders need consistency, tests, and reviews. [**dbt Canvas**](https://www.getdbt.com/blog/dbt-canvas-is-ga) was launched in May 2025 to address both problems by adding a governed, visual modeling workspace that sits on top of your governed dbt project. It maps in a left-to-right flow, so teams can see the logic of a model at a glance. They can make targeted changes using drag-and-drop tooling, and open their visual models seamlessly in dbt Studio to make edits directly to the model’s underlying SQL. dbt Canvas supports code commenting on each individual operator to enhance code clarity, and will soon support standard dbt documentation. And unlike legacy visual tools, dbt Canvas compiles to warehouse-native SQL and runs where your data lives. Models stay version-controlled in Git, validated with dbt tests, reviewed via your CI process, and preserved in dbt lineage and docs, so speed never comes at the cost of trust. **** ## What ambassadors said about their dbt Canvas experience > “dbt Canvas is a excellent addition for both our data engineers, data scientists and production support roles. The visual feature enables quicker understanding and troubleshooting of large complex SQLs. It’s intuitive, fast to pick up and the left-to-right data pipeline visual concept flow of inputs, transformations, and outputs paints a clear picture of the logic within the sqls.” -**Ron Barimo, Lead Data Engineer, USAA** > > “I’m fairly SQL proficient, but not a daily user, so dbt Canvas is a great way to make small changes without worrying about writing perfect SQL. The drag-and-drop interface makes it easier and faster to update existing models without breaking anything. It’s especially helpful for less frequent or less advanced SQL users.” -**Tony Mayer, Senior LOB Reporting & Analytics Manager, Fifth Third Bank** > > “dbt Canvas is great for SQL developers too, not just those that prefer drag and drop. It’s like having lineage within a single query which makes it much easier to understand how data flows through complex models, especially ones you didn’t build. It’s also a powerful tool for documentation, helping you clearly see and explain each step of the transformation.” -**Marcelo Bour, Analytics Engineer, Dynamic Data** ## How dbt Canvas helps data engineers dbt Canvas accelerates code comprehension with a clear left-to-right visual flow, so unfamiliar models become approachable. Engineers can spot joins, filters, and aggregations faster and spend less time parsing dense SQL. It also surfaces step-by-step logic, which makes refactoring easier to reason about and validate before opening a PR. The visuals double as living documentation, reducing handoff friction and improving reviews and onboarding. ## How dbt Canvas helps analyst teams dbt Canvas lowers the barrier to making safe changes. The drag-and-drop interface lets teams tweak logic, test hypotheses, and ship quick wins, without memorizing every SQL nuance. Because all the work happens inside their governed dbt project, [analysts move faster](https://www.getdbt.com/product/analyst) without spawning shadow pipelines. The visual steps make stakeholder walkthroughs straightforward, so results are trusted and easier to explain. ## What ambassadors said about the program experience: - Telecom analytics engineer (Enterprise) — 10/10: “Incredible instruction with clear structure that set everyone up for success, while allowing discussions to flow naturally. Discussion and breakouts were especially valuable.” - Fifth Third Bank — 10/10: “Hands‑on training was super helpful, with clear walkthroughs of Canvas and end‑of‑chapter problems to apply concepts in session.” - Sports analytics practitioner — 9/10: “Engaging, helpful course with in‑class walkthroughs that made concepts easy to follow.” - Manufacturing data engineer — 10/10: “Clear explanations, valuable across sections, and directly applicable to active projects.” ## What’s next We’ll continue to run small, hands-on cohorts so champions can co-build with us and bring back repeatable practices. If you’d like to nominate a champion for a future cohort, or you are that champion, please reach out to your dbt representative or connect with us in the dbt Community. **Want to see dbt Canvas in action?** dbt Canvas is available for Enterprise or Enterprise+ users on the dbt platform. [Request demo](https://www.getdbt.com/signup) tailored to your use case or check out our [docs page](https://docs.getdbt.com/docs/cloud/canvas) to get started. Thank you to our first ambassadors and to Ron, Tony, and Marcelo for sharing how dbt Canvas is helping their teams move faster with more clarity. If you’re exploring ways to scale self-serve while keeping governance intact, dbt Canvas gives your engineers and analysts a shared, visual language for building and evolving trusted data. --- --- title: "Agentic coding in analytics engineering" description: "Mikkel Dengsøe, the cofounder of SYNQ, discusses his tests (and tips) with agentic coding tools." url: "https://www.getdbt.com/blog/agentic-coding-in-analytics-engineering" date: "2025-09-08" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Agentic coding in analytics engineering What does agentic coding look like in analytics engineering? Mikkel Dengsøe, co-founder at SYNQ, recently [wrote](https://medium.com/@mikldd/using-ai-for-data-modeling-in-dbt-975838054cb1) a [series](https://medium.com/@mikldd/using-ai-to-build-a-robust-testing-framework-4e034dfd014f) of [posts](https://medium.com/@mikldd/using-omnis-ai-assistant-on-the-semantic-layer-0572f997451d) on his experiences as an analytics engineer with agentic coding tools. In this episode of The Analytics Engineering Podcast, he walks through a hands-on project using Cursor, the [dbt Fusion engine](https://www.getdbt.com/product/fusion), the [dbt MCP server](https://www.getdbt.com/blog/mcp), Omni’s AI assistant, and Snowflake. Tristan and Mikkel cover where agents shine (staging, unit tests, lineage-aware checks), where they’re risky (BI chat for non-experts), and how observability is shifting from dashboards to root-cause explanations delivered to the right person at the right time. Along the way: practical prompts, why “one model at a time” keeps you in control, and a testing philosophy that avoids alert fatigue while catching what matters. [**To see real-world use cases of agentic coding and to learn directly from data and AI leaders, join us at Coalesce 2025 in Las Vegas, Oct. 13-16**](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/summary/?utm_medium=social&utm_source=substack&utm_campaign=q3-2026_coalesce-2025_aw&utm_content=coalesce____&utm_term=all_all__). _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways ### Can you talk a little bit about your background? **Mikkel Dengsøe:** Yeah, so I can start from the beginning. I've been in data for, I think it's coming up to 15 years now, and started my career in data at a Danish shipping company, which was very much zero to one. When I came in, there was no data warehouse, and the only way we could know how many containers were shipped was by an IT guy pulling that out of the system every six months. I then spent two years there building up their data warehouse on SQL Server, which was super fun. After that, I spent five years at Google, which was a very different gear. ### That's a natural transition. Just global shipping company straight to Google. Exactly. And that was very much a hundred-to-end where, in my case, I worked with the ads data and you get a perfectly curated data table that you can work with and everything kind of works. Then after that I joined a company called Monzo. For those who are not familiar, it's a scaling fintech out of the UK and that was very much the one to a hundred. When I joined we were 30 data people, but scaled to a hundred over two years. We had 10,000 dbt models and we built every internal tool under the sun for dbt. Super interesting. And then three and a half years ago I went on to found SYNQ alongside Peter and Steve, which is a data observability platform. ### Tell us a little bit more about SYNQ. We are a data platform that primarily works with companies using tools like dbt already, but have issues going from important data to business-critical data. That might be customer-facing dashboards, machine learning models, or something else. They want better monitoring—we often deploy anomaly monitors—and they also want workflows such as incident management for when things go wrong. We were founded in 2022, so now we're in early stages of working with scale-ups and startups, and now also onboarding enterprises and larger companies. It's been a fun journey. ### In your series of blog posts, you went through the modern data stack and said, “What's the most current version of this tool and how effectively can I AI-ify that?” Whether that's using Cursor to build dbt models or using the agent experience inside of Omni—what made you decide to get into this and write about it? The first part of it is just: it's super fun to tinker with these tools and try them out. It's magic. And we were also building an MCP server at SYNQ, so I had a lot of interest in seeing how it works with others and what we can learn. It was also driven by being able to have conversations with our customers. When they ask about it, being able to speak from the point of view of having actually tried this and seen what works and what doesn't. ### The early days of using Redshift were such a visceral experience relative to what came before. If I hadn't interacted with it directly, I wouldn't have understood how big a state change cloud data was. This feels like another one of those moments: if you don't have hands-on experience, you're not going to really get it. Fair? Spot on. And I think pretty much every data team should be doing this unless they have a very good reason not to. The risk and the stakes can be pretty low if you use it for internal workflows like data modeling and writing tests. You're still in control. I recommend everybody do it. ### What tasks did you try to accomplish? It's three different blog posts: the data modeling part, the testing part, and then exposing it in Omni's AI agent where people can ask questions about the data. There's a fourth post: once the data is live, how can you use the SYNQ MCP to do things like root-cause analysis and planning changes. I started with data modeling. I had raw data from different JSON sources, some XMLs, some profiles—extracted and put into Snowflake—and then did the data model. ### So the data was already loaded into Snowflake? Yeah, exactly. For the data modeling, I started from the sources and then worked through staging, marts, and finally metrics using the semantic layer. Each step looks a little different when you use AI tools because the behavior differs. In terms of tooling, I used Cursor with the dbt-MCP plugged in. If you're not familiar, dbt-MCP lets you, via prompt, interact with dbt tools—execute `dbt build`, get models, or get everything upstream of a given model—so you can chain work without explicitly doing it. ### Cursor + dbt-MCP. What model did you use? I just used the default in Cursor, which I believe is Claude. There's an important distinction: Cursor is really good at writing code, but it can't execute queries on your behalf. If you want to extract raw data and query Snowflake to get rows out, you have to do that in Claude Desktop. That became key. Early on, as I built models, the first thing I did was get a snapshot of sample data from Snowflake—10,000 rows of a source. I fed that into Cursor and said, “These are examples of what this data looks like.” Using that data, Cursor could model in a clever way. For example, a column called `quarter` like “2025 Q1”—Cursor understood to translate it into a datetime and do the transformations. ### I've used the dbt MCP server a decent amount—less in Cursor, more in Claude Desktop. Your stack was Cursor + Claude models + Claude Desktop. And Cursor cannot directly execute queries in Snowflake, but Claude Desktop can. Is that because there’s tool use Claude has that Cursor doesn't? I believe so. In Claude Desktop, if you write queries against dbt-MCP, Claude can visualize a graph, show outputs of a SQL statement, etc. Cursor, as far as I know, couldn't. My middle ground was to take sample data out of Snowflake, put it into a CSV, and feed that back into Cursor so it could look at raw data. ### As part of its own context window? Exactly. That was key for my workflow. Then when I wanted to write unit tests, I could use real data examples from the sample. Or when automatically documenting the data, I asked Cursor to specify examples in the docs based on the most common occurrences within a column. Letting Cursor peek at raw data was a core pillar. ### It's a little hacky, right? Cursor should really be able to interact directly with Snowflake or Databricks to investigate the shape of the data. Agents should be empowered to do that. I would say so. There might be a way I didn’t know about, but I patched the gaps by uploading into the context window. ### So that's the state of the art today. Seems so. To be clear, I think the limitation is IDE differences—Cursor vs. Claude Desktop—rather than dbt-MCP itself. ### Once you had sample data in context, did you have to suggest conversions, or did it naturally do them? It got the defaults pretty right, but I guided it on what I wanted from the source data. I wanted control over everything, so I asked it to do one model at a time rather than auto-generate a whole stack. That way I could review each step and stay in control. ### Your prompt workflow was “Build me a model with this name that stages the data from this table,” basically? Yeah. When it proposed code I didn't like, upstream it was usually simple (regex to parse dates, etc.). Downstream, in marts and metrics, I started describing my ideal data product: user jobs-to-be-done and the final output. That’s when Cursor got creative and invented metrics I hadn’t anticipated—like “apartment price relative to time on market.” I pruned ones I didn’t want, but some were good surprises. ### Which layer did it help most? Testing. Modeling was good—especially staging—but testing accelerated significantly. SQL is a bit like English; for simple datasets you can express intent easily. Testing can be much harder and more verbose. ### Roughly how much more effective did you feel? Modeling: multiples faster. It nailed the tedious parts—regex, casting, pass-throughs—so staging/intermediate layers flew. In marts/semantic metrics, the benefit was brainstorming. It helped me think of metrics I wouldn't have. ### Did the dbt Fusion engine help? Yes. Fusion shows lineage and whether a column is pass-through. For example, if a column is pass-through with no transforms, don't add another `not_null` or `unique` if there's one upstream. I bounced between the IDE to check this and codified it as a testing strategy. That's already top-10% testing hygiene. ### Any MCP feature requests surface? The more context and tools the agent has, the more it can do. In the fourth post, for root cause analysis, we used the SYNQ MCP. We collect all your Git commits and have history, so the agent could correlate recent code changes with incidents. Requests depend on the job at hand. ### Let's move to testing—why was it the most additive? Testing is hard; many teams don't know how to do it and alert fatigue is common. A huge share of tests we see are `not_null`/`unique`, which doesn't reflect real data risks. First thing I did in Cursor for testing was provide our internal testing philosophy as guidelines: test heavily at the source, don't retest pass-through columns, focus on business and metric anomalies in marts. That worked really well. For sources and staging, it generated relevant tests. Then for marts, I asked for unit tests and gave it a thousand sample rows from Snowflake. It wrote very relevant unit tests I’d otherwise spend a lot of time on. ### Examples? Simple ones like: when you pass a string value in the date column, does it transform correctly to datetime and match the expected format? These just worked. Then at the metric level, it looked at raw data and proposed assumptions—like square-meter price should be between X and Y—sometimes segmenting by postcode. Very thoughtful, though I'd replace static thresholds with anomaly monitors so they don't go stale as prices move. ### So at least 5× on testing? At least. Apart from swapping static thresholds for anomaly detection, it nailed testing and did so in a lineage-aware, layer-appropriate way. ### Tell me about the BI layer. Many teams start at the BI layer with a chat interface. I think that's risky because it's used by business users and you only get so many chances before trust drops. I moved into Omni. You create a “topic” (a data model you can join with others) and then specify an AI context: instructions for how the LLM should behave. For example: if a user asks about price, always return square-meter price; never make up fields not present in the mart; if asked about provenance, mention the source. Writing AI context is a new skill for our industry. ### Were you using Omni’s AI assistant to create assets faster, or to let users self-serve? The latter—so users could ask questions instead of going to a dashboard. It could have been any BI tool with similar functionality; we just use Omni internally. ### And how was the experience as a consumer? Amazing when it works, but I'd hesitate to give my VP of Marketing access. It gets things wrong maybe one in five times, and it's not obvious why if you're not a data person. For analysts doing exploratory work, it's great—they can inspect and dig in. I wouldn't replace company-wide dashboards with a chat bot yet. Omni does log freeform queries and feedback, so there's a path to iterate the AI context over time. ### The last thing you did was use AI plus SYNQ to monitor production infrastructure. What does observability look like in the future? Historically it's looked like dashboards—Datadog for data pipelines. Is it just more effective monitors, or fundamentally different? Fundamentally different. We’re heading to a place where observability tools can tell you what's wrong at the right time, with just the right context, delivered to the right person—inside or outside the data team. Done well, there may be few dashboards; instead you get an LLM-summarized root cause delivered from a monitor that might be auto-created. Less “active tool you poke at,” more “proactive explanation.” ### Still technical observability (pipelines/data issues), or business observability? More the former. Teams at the edges—Sales Ops managing Salesforce, engineering teams creating web events—often need to be notified about data issues. Business KPI movements require a different experience for marketers, etc. ### Automated remediation? Gradual. You can imagine an issue occurs without a dedicated test; the system proposes a new test. But 80% of issues come from root systems elsewhere (someone typing in Salesforce), and closing that loop is still hard. In the article’s fourth part, we had a data issue and I asked the SYNQ MCP through Claude Desktop to do root cause analysis. It walked the same steps a data person would: inspect the model, check errors, examine lineage and upstreams, review recent commits, and documented each step to the root cause. That works now. ### At the beginning you said there’s no good reason not to use these tools today. What reasons do you hear for not trying? People are busy. But if you look at a risk curve, lowest risk is modeling and testing—you're in the driver's seat and it boosts productivity. Higher risk is replacing your BI tool with a chat bot; higher still is customer-facing experiences. The first two are hard to argue against. ### Enterprise IT approvals might be one blocker—approved models, data access, etc. True. For example, our MCP can query raw data to detect if an issue happens in a segment, and enterprises might hesitate there. Also, “MCP” as a term can be confusing. But it's actually simple and explainable, not a black box. Setting up dbt-MCP can still feel hacky in enterprises; if it lived natively in cloud environments, it’d be easier to adopt. ### You can set it up locally—no permissions/procurement—and just play. We also shipped the MCP server as a remote MCP in cloud, though that introduces auth/permissions considerations. If I had to pick a persona, it's the analyst. Analysts have had a tough decade: more tools, harder workflows, less time to tinker. MCPs and AI workflows are a turning point. At Monzo, we had a philosophy that you should be able to have an idea on your commute and have it implemented by midday. As we grew to 10,000 dbt models and long CI checks, that faded. I can see a world where this returns. MCPs can help. I'm excited. ### I love that. Analytics engineers think “infrastructure, correctness.” Analysts think “idea to validation fast.” Excel was always the analyst’s best friend because it's fast and flexible. MCPs make it easy to plug tools together and get answers quickly again. One company we work with—Voi, a scooter company out of Sweden—has a strong data leader, Magnus, who is very bought into metrics. Their data team doesn't produce dashboards; they produce metrics. In an AI world with MCPs, flows, and curves, that's a clear decision. ### I believe there's no such thing as the wrong BI tool—different tools have different trade-offs. Probably true for models/IDEs too: Claude Desktop vs. Claude Code vs. Cursor—no single “right answer” as long as the underlying context and metric definitions are shared. Agreed. What really matters across workflows: consistent metric definitions, documentation for columns and fields, and high-quality data. Those foundations matter even more when an LLM is in the loop; you may not have a human sanity-checking every result. ## Chapters - **00:00** — Tristan’s intro - **01:10** — Mikkel’s background: shipping → Google → Monzo → SYNQ - **03:08** — What SYNQ does (data observability for business-critical data) - **04:15** — Running the experiment - **06:23** — Scope: modeling, testing, BI agent, observability - **07:17** — Tooling: Cursor + dbt MCP server + Snowflake + Omni - **09:38** — Sampling real data into the agent’s context - **13:14** — Modeling workflow: one model at a time - **15:14** — Where agents help most: testing > modeling - **18:10** — dbt Fusion engine: lineage-aware checks, fewer redundant tests - **19:50** — Feature requests and root-cause via commit history - **20:57** — Testing philosophy: source-heavy, pass-through aware, metric-level - **22:49** — Unit tests from samples; thresholds vs anomaly monitors - **25:10** — BI agents: great for analysts, risky for broad rollout - **31:54** — The future of observability: explain first, dashboards second - **36:10** — Adoption curve: safe places to start - **40:49** — Analyst superpowers return - **42:04** — Metrics over dashboards --- --- title: "From traditional to AI data engineering: What's different?" description: "Learn how AI is changing how data engineers work—from routine tasks to strategic roles in modern data stacks." url: "https://www.getdbt.com/blog/traditional-to-ai-data-engineering" date: "2025-09-03" authors: ["Joey Gault"] categories: ["Pulse"] --- # From traditional to AI data engineering: What's different? Traditional data engineering has long focused on manual processes for building and maintaining data pipelines. Data engineers spend considerable time writing SQL transformations, creating documentation, building tests, and troubleshooting pipeline failures. These tasks, while essential, are often repetitive and time-intensive, creating bottlenecks in data delivery. AI data engineering leverages generative AI and large language models to automate many of these routine tasks. Rather than replacing data engineers, AI augments their capabilities, allowing them to focus on higher-value strategic work. This represents a shift from manual coding and maintenance to AI-assisted development and orchestration. The distinction goes beyond simple automation. AI data engineering incorporates deep contextual understanding of data relationships, metadata, and lineage to generate more accurate and relevant outputs. This contextual awareness enables AI systems to make intelligent decisions about data transformations, testing strategies, and documentation that generic automation tools cannot achieve. ## Task automation and efficiency gains The most immediate difference lies in how routine tasks are executed. In traditional data engineering, creating new data transformation assets requires manually writing SQL code, often involving extensive research into existing schemas and business logic. Data engineers must remember syntax peculiarities, look up column names, and ensure consistency across models. AI data engineering transforms this process through natural language interfaces. Engineers can describe their requirements in plain English, and AI systems generate the corresponding SQL code, complete with proper formatting and adherence to established style guides. This approach is particularly valuable for complex queries involving multiple tables or intricate business logic. Testing represents another area of significant divergence. Traditional approaches require engineers to manually craft test cases, often leading to inconsistent coverage or tests being deprioritized under time pressure. AI data engineering can automatically generate comprehensive test suites based on data model context, including unit tests, data quality checks, and integration tests. The AI understands the data structure and relationships, enabling it to suggest relevant validation rules and edge cases. [Documentation](https://docs.getdbt.com/docs/build/documentation), historically one of the most neglected aspects of data engineering due to time constraints, becomes effortless with AI assistance. Instead of manually documenting hundreds of tables and fields, AI systems can generate initial documentation based on schema analysis, existing metadata, and similar data assets within the project. This creates a foundation that teams can iteratively improve over time. ## The critical role of frameworks and standards One of the most important distinctions in AI data engineering is the heightened importance of frameworks and standardization. While [traditional data engineering benefits from consistent approaches, AI data engineering makes standardization essential for effectiveness](https://www.getdbt.com/blog/how-ai-will-disrupt-data-engineering). AI systems perform optimally when working with codebases that are concise, consistent, and well-documented. Heterogeneous environments with multiple languages, frameworks, and conventions create challenges for AI systems, just as they do for human developers. However, AI systems are particularly sensitive to these inconsistencies because they rely on pattern recognition and established conventions from their training data. Frameworks like [dbt](https://www.getdbt.com/product/what-is-dbt) become even more valuable in AI-enabled environments because they provide the structure and consistency that AI systems need to generate reliable outputs. When AI tools can leverage well-documented frameworks with extensive community examples, they produce higher-quality code with fewer errors. The standardized patterns and conventions inherent in mature frameworks create an ideal environment for AI assistance. This emphasis on frameworks extends to CI/CD processes, logging, and observability practices. AI systems can more effectively debug issues and suggest optimizations when working within standardized environments with consistent error messages and monitoring approaches. ## Enhanced incident resolution and maintenance Traditional data engineering requires significant manual effort for troubleshooting pipeline failures and performance issues. Engineers must analyze logs, trace data lineage, and identify root causes through time-intensive investigation processes. AI data engineering introduces the possibility of automated incident resolution. By providing complete log outputs and project context to AI systems, engineers can receive detailed diagnoses and proposed solutions within minutes rather than hours. Some implementations can even generate complete pull requests with fixes, ready for review and deployment. This capability extends to proactive maintenance tasks like performance optimization and refactoring. AI systems can analyze entire codebases to identify opportunities for consolidation, suggest performance improvements, and implement large-scale changes across multiple files simultaneously. These multi-file refactoring capabilities represent a significant advancement over traditional manual approaches. ## Stakeholder interaction and self-service capabilities Traditional data engineering involves substantial time spent answering stakeholder questions about data availability, trustworthiness, and usage. These interactions, while valuable, create friction and slow down both data teams and business users. AI data engineering enables more sophisticated self-service capabilities through natural language interfaces. Instead of requiring stakeholders to learn SQL or rely on data engineers for every query, AI systems can translate business questions into appropriate technical queries. However, this capability requires robust metadata management and semantic layer implementations to ensure accuracy. The development of context protocols represents a significant advancement in this area. These protocols allow AI systems to access comprehensive business context about data assets, including their reliability, appropriate use cases, and relationships to other data sources. This contextual awareness enables more accurate responses to stakeholder queries and reduces the need for direct data engineer intervention. ## Evolving skill requirements and role definitions The shift to AI data engineering is changing the skill sets required for success in the field. While technical proficiency remains important, the ability to effectively collaborate with AI systems becomes crucial. This includes understanding how to craft effective prompts, validate AI-generated outputs, and integrate AI tools into existing workflows. Traditional data engineering roles are evolving into three primary directions. Data platform engineers focus on the infrastructure and systems that support AI-enabled workflows, requiring deep technical expertise in performance, governance, and reliability. Automation engineers bridge the gap between data insights and business actions, building systems that automatically respond to data-driven triggers. Domain-focused data engineers work closely with business stakeholders to ensure data products meet specific business requirements. ## Quality assurance and governance considerations AI data engineering introduces new considerations for quality assurance and governance. While AI can generate code more quickly than manual approaches, the outputs require careful validation to ensure accuracy and adherence to business requirements. This creates a need for robust testing frameworks and review processes specifically designed for AI-generated code. The[ governance implications extend beyond code quality to include AI model management](https://www.getdbt.com/blog/understanding-data-governance-ai), prompt engineering standards, and audit trails for AI-assisted decisions. Organizations must establish clear guidelines for when and how AI tools should be used, ensuring that critical business logic receives appropriate human oversight. ## Infrastructure and tooling evolution The infrastructure requirements for AI data engineering differ significantly from traditional approaches. Organizations need access to large language models, either through cloud services or on-premises deployments. Integration with existing development environments becomes crucial, requiring tools that can seamlessly incorporate AI assistance into established workflows. The tooling ecosystem is rapidly evolving to support AI-enabled data engineering. Traditional data engineering tools are incorporating AI capabilities, while new specialized tools emerge to address specific AI data engineering use cases. This creates both opportunities and challenges as organizations evaluate and integrate new technologies. ## Looking ahead The transformation from traditional to AI data engineering represents more than a technological upgrade: it's a fundamental reimagining of how data work gets done. Organizations that successfully navigate this transition will find themselves able to deliver higher-quality data products more quickly and at greater scale. The key to success lies in understanding that AI data engineering is not about replacing human expertise but augmenting it. The most effective implementations combine the speed and consistency of AI with the strategic thinking and domain knowledge of experienced data engineers. As this field continues to evolve, the organizations that invest in both AI capabilities and the frameworks to support them will be best positioned to capitalize on the opportunities ahead. The [differences between traditional and AI data engineering are profound](https://www.getdbt.com/blog/how-ai-will-disrupt-data-engineering), touching every aspect of how data teams operate. From task execution to stakeholder interaction, from infrastructure requirements to skill development, the shift represents a new era in data engineering that promises greater efficiency, higher quality outputs, and more strategic focus for data professionals. ## AI data engineering FAQs **Will AI replace data engineers?** AI will not replace data engineers but will augment their capabilities, allowing them to focus on higher-value strategic work. The shift represents a move from manual coding and maintenance to AI-assisted development and orchestration. Data engineering roles are evolving into three primary directions: data platform engineers who focus on infrastructure and systems, automation engineers who bridge data insights with business actions, and domain-focused data engineers who work closely with business stakeholders to ensure data products meet specific requirements. **What is the single biggest data engineering challenge you anticipate in enabling these autonomous workflows?** The biggest challenge is establishing robust frameworks and standardization. AI systems perform optimally when working with codebases that are concise, consistent, and well-documented. Heterogeneous environments with multiple languages, frameworks, and conventions create significant challenges for AI systems because they rely on pattern recognition and established conventions. Organizations must also implement proper quality assurance and governance processes specifically designed for AI-generated code, including testing frameworks, review processes, and audit trails for AI-assisted decisions. **How can you design and implement end-to-end batch and streaming data pipelines on AWS to support machine learning and analytics use cases?** AI data engineering transforms pipeline design through natural language interfaces and automated code generation. Engineers can describe requirements in plain English, and AI systems generate corresponding SQL code with proper formatting and adherence to style guides. The process leverages AI's contextual understanding of data relationships, metadata, and lineage to make intelligent decisions about data transformations. This approach requires standardized frameworks like dbt to provide the structure and consistency that AI systems need to generate reliable outputs, along with robust CI/CD processes and observability practices for effective debugging and optimization. --- --- title: "Scaling FinOps at Workday: Turning fragmented cloud costs into actionable insights" description: "How Workday cut cloud costs with a custom FinOps platform built on dbt, Lightdash, Trino, and Delta Lake." url: "https://www.getdbt.com/blog/scaling-finops-at-workday" date: "2025-09-03" authors: ["Hrishi Kulkarni", "Chakshu Mehta", "Ernesto Ongaro"] categories: ["Product"] --- # Scaling FinOps at Workday: Turning fragmented cloud costs into actionable insights Workday is a leading cloud-based software enterprise that streamlines HR, finance, and IT management. With thousands of customers worldwide, Workday helps some of the world’s largest organizations plan, execute, and analyze their most critical business operations. Like many enterprises, Workday relies on multiple cloud services—primarily Amazon (AWS) and Google (GCP). But doing [cost optimization](https://www.getdbt.com/product/cost-optimization) across such large, complex environments is no small task. Without clear visibility into cloud spend, it’s nearly impossible to identify inefficiencies or waste. We’ll explore how Workday tackled this challenge by building a custom FinOps platform called Opus. We’ll also look at how dbt played a central role in helping Workday structure its data to generate trustworthy, actionable insights. ## Fragmented and siloed cloud-spend data For Workday, reconciling cloud-cost data from the two providers was incredibly difficult. Each provider reports cost data in a different way, with different schemas, naming conventions, and levels of granularity. Workday’s finance and engineering teams had limited visibility into where money was going, and it was challenging to answer basic questions like: - How much are teams spending? - What are the costs associated with different services and customers? - How much of that spend is efficient or wasteful? - What are the cost forecasts for the next quarter? - Where should the business reallocate budget? On top of that, Workday's previous reporting stack was slow and required a lot of manual work to use. Key business logic was trapped in the reporting tool, making it hard to audit or reuse. Because the reporting tool was closed-source, customizations were limited, leading to bottlenecks and performance issues. “We had little to no governance, which led to silos,” reflects Pattabhi Nanduri, Senior Data Engineer at Workday. “Without a framework for development testing or CI/CD, it was difficult to maintain code and deliver timely, data-driven insights to the business.” ## A FinOps platform, powered by dbt To address these challenges, Workday built Opus, an internal FinOps intelligence platform that sits on top of dbt, Lightdash, and a custom ETL stack running on Trino and Delta Lake. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/8f52ddac60fb131d4cc63dd67244e5723e2b5e21-1600x955.png) Opus brings together cost and usage data from AWS and GCP, standardizes it, and models it into clean, usable formats. Today, it powers dashboards used across finance, engineering, and other teams. Its key capabilities include: - **Unified dashboards in Lightdash.** Now, finance and engineering teams have a single place to monitor cloud costs, model future spend, and identify opportunities to cut waste. - **Deep visibility** into compute usage, instance types, and spending trends. Teams can better understand which workloads drive costs and where to make optimizations. - **Role-based access control (RBAC) and SAR compliance.** Access to sensitive financial data is restricted based on user roles, ensuring regulatory compliance. - **Cost categorization** powered by dbt. Costs are organized into dynamic, multi-level hierarchies—like compute, storage, or credits—using reusable dbt macros. The structure can easily evolve as business needs change. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/d2773287bec35b5d1c591defca2895fcb019b7c5-1600x930.png) “We delivered a new cost-categorization hierarchy quickly and iteratively, thanks to dbt’s templating,” summarizes Nanduri. “Additionally, dbt’s Jinja macro framework made it easy to restructure our models without rewriting large amounts of code.” ## Building structured, scalable data models with dbt To build Opus, dbt gave Workday the technical foundation to support a high-impact, high-visibility FinOps platform. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/fca61bc9eed3216fa31f44ff4228ed93bd3bf9f6-1600x907.png) Under the hood, here’s what that looks like: - **Macro-driven model design.** Workday built an internal dbt macro library to template and enforce source models, staging models, intermediate models, and mart models. As a result, Workday can enforce modeling best practices across the organization. - **Streamlined hierarchies and flexibility.** dbt's templating allowed dynamic modeling of an 8-level product dimension hierarchy. Categories like compute, storage, databases, and credits were easily shifted across levels through filter logic updates—saving significant refactoring time. - **Integrated testing.** With custom macros and dbt-expectations, Workday automated join key validation, freshness checks, and custom filter and logic validation. This streamlined data QA without burdening users with complex testing syntax while driving trust in the data. - **YAML generation and documentation.** Workday’s macros also auto-generated YAML files and documentation from model specs, eliminating manual sync work between SQL and metadata. This ensured documentation was always accurate and discoverable through dbt Catalog. - **Execution optimization with lineage graphs. **Workday used dbt's lineage graphs and model profiling to build dynamic execution plans. These plans fed into Airflow and Cosmos to scale resources up or down, maximizing performance and minimizing cost. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/7e2fc055d85aacf156d602ebf69985e7e727f10e-1600x923.png) “What’s great about dbt is how quickly new users can get up to speed with it,” says Eric Pu, Senior Software Engineer at Workday. “They can dive right in without spending months learning how to use it. Meanwhile, advanced users can go deeper and add customizations as it makes sense.” ## Cleaner data and faster answers Today, thousands of internal users leverage Workday’s cloud-cost data through dbt powered Opus. Now, finance and engineering teams have a single source of truth for cloud costs. Dashboards load quickly, users can find data for specific services or time periods, and spend is properly attributed across departments and environments. Teams are onboarding faster, and they’re experiencing performance improvements across the board. It’s a truly self-serve model: users can access the data they need without waiting for a data analyst to pull a custom report. Crucially, Workday’s FinOps layer is no longer a black box. All model logic lives in version-controlled dbt code. Every transformation is testable, reviewable, and documented via dbt Docs, making audits far more trustworthy. “dbt is a powerful and flexible tool for enterprises,” concludes Pu. “It’s easy to plug in to your existing environment and customize to fit your exact needs.” _Your stakeholders need fast answers they can trust. If you’re thinking about how to build data infrastructure that delivers actionable insights, we’d love to chat. Get in touch with us to [book a live demo](https://www.getdbt.com/contact), or [sign up](https://www.getdbt.com/signup) for free to connect to your data warehouse and start building._ --- --- title: "dbt Labs Is Named to the 2025 Forbes Cloud 100" description: "dbt Labs has been named to the 2025 Forbes Cloud 100, the definitive ranking of the top 100 private cloud companies in the world." url: "https://www.getdbt.com/blog/dbt-labs-named-to-2025-forbes-cloud-100" date: "2025-09-03" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Is Named to the 2025 Forbes Cloud 100 **SAN FRANCISCO** (September 3, 2025) – [dbt Labs](http://getdbt.com), the leader in standards for AI-ready structured data, has been named to the Forbes 2025 Cloud 100, the definitive ranking of the top 100 private cloud companies in the world, published by [Forbes](http://www.forbes.com) in partnership with [Bessemer Venture Partners](https://www.bvp.com/). As AI completely changes the way organizations interact with data, dbt Labs is driving the next phase of innovation in the market. In May, the company unveiled the [dbt Fusion engine](https://www.getdbt.com/blog/dbt-labs-redefines-dbt-with-new-fusion-engine), a complete rewrite of the foundational technology that powers dbt. It introduces robust SQL comprehension into the platform, enables faster analytics delivery, drives down cloud costs, and allows organizations to scale analytics. With this massive upgrade, dbt provides the context around an organization’s structured data that is critical to successful AI initiatives. This focus sparked significant momentum this year, as dbt Labs surged past [$100M ARR](https://www.getdbt.com/blog/dbt-labs-100m-arr-milestone) and [acquired SDF Labs](https://www.getdbt.com/blog/dbt-labs-announces-sdf-labs-acquisition). Globally, more than 80,000 teams rely on dbt to power their data work, including those at Siemens, Roche, Bilt Rewards and Condé Nast. “The future of the data ecosystem is being written today, and the main drivers of that future are open standards and AI. dbt accelerates both, and inclusion on this year’s Forbes Cloud 100 is another signal that the market recognizes our trajectory and potential,” said Tristan Handy, founder and CEO, dbt Labs. For the tenth consecutive year, the Cloud 100 reviews submissions from hundreds of cloud startups and private companies each year. The Cloud 100 evaluation process involved ranking companies across four factors: market leadership (35%), estimated valuation (30%), operating metrics (20%), and people & culture (15%). For market leadership, the Cloud 100 enlists the help of a judging panel of public cloud company CEOs who assist in evaluating and ranking their private company peers. “For the last decade, the Forbes Cloud 100 list has recognized the most innovative private cloud companies in the world, and this year’s standouts are among the most impressive we’ve ever seen,” said Richard Nieva, the Forbes editor of the Cloud 100. “Our honorees highlight the massive sea change that AI has brought to the enterprise, with sky-high growth and valuations.” “As we mark the 10th year of the Cloud 100 with our 2025 rankings, we celebrate a highly competitive cohort of companies that, for the first time, collectively exceed $1 trillion in value,” said Elliott Robinson, partner at Bessemer Venture Partners. “The cloud is in a period of rapid, AI-driven transformation, with this year’s cohort demonstrating how AI is fundamentally reshaping how the best cloud companies grow, scale, and compete.” dbt Labs will host its annual [Coalesce](http://coalesce.getdbt.com) event October 13-16 in Las Vegas and online, bringing together thousands of data leaders and practitioners to discuss the future of data, analytics and AI. Ashley Kramer, Head of Revenue at OpenAI, will deliver a featured keynote as part of this year’s agenda. To learn more and register for the event, visit [coalesce.getdbt.com](http://coalesce.getdbt.com). The Forbes 2025 Cloud 100 and 20 Rising Stars lists are published online at[ www.forbes.com/cloud100](http://www.forbes.com/cloud100). Highlights of the list appear in the August/September 2025 issue of _Forbes_ magazine. **About dbt Labs** Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 80,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at [getdbt.com](http://getdbt.com), and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). **About Bessemer Venture Partners** Bessemer Venture Partners helps entrepreneurs lay strong foundations to build and forge long-standing companies. With more than 150 IPOs and 300 portfolio companies in the enterprise, consumer and healthcare spaces, Bessemer supports founders and CEOs from their earliest days through every stage of growth. Bessemer’s global portfolio has included Pinterest, Shopify, Twilio, Yelp, LinkedIn, PagerDuty, DocuSign, Wix, Fiverr, Toast and ServiceTitan, and has $19 billion of assets under management. Bessemer has investment teams located in San Francisco, Silicon Valley, New York, Boston, London, Bangalore, and Tel Aviv. Born from innovations in steel more than a century ago, Bessemer’s storied history has afforded its partners the opportunity to celebrate and scrutinize its best investment decisions (see [Memos](https://www.bvp.com/memos)) and also learn from its mistakes (see [Anti-Portfolio](https://www.bvp.com/anti-portfolio)). **About Forbes** Forbes champions success by celebrating those who have made it, and those who aspire to make it. Forbes convenes and curates the most influential leaders and entrepreneurs who are driving change, transforming business and making a significant impact on the world. The Forbes brand today reaches more than 140 million people worldwide through its trusted journalism, signature LIVE and Forbes Virtual events, custom marketing programs and 43 licensed local editions in 69 countries. Forbes Media’s brand extensions include real estate, education and financial services license agreements. --- --- title: "What AI really means for data engineering workflows" description: "See how AI assistants are cutting time for data engineers—from code generation to testing—and what still requires human oversight." url: "https://www.getdbt.com/blog/ai-tools-data-engineering-efficiency" date: "2025-08-29" authors: ["Joey Gault"] categories: ["Pulse"] --- # What AI really means for data engineering workflows Recent industry data reveals the rapid pace of AI adoption among data professionals. According to the [2025 State of Analytics Engineering Report](https://www.getdbt.com/resources/state-of-analytics-engineering-2025), 80% of data practitioners now use AI in their day-to-day workflows, up from 30% one year prior. This surge reflects not just curiosity about new technology, but practical recognition of AI's ability to address longstanding efficiency challenges in data engineering. The most common applications center on code development, with 70% of analytics professionals using AI to assist in writing and maintaining data transformation logic. Documentation and metadata development represent the second most frequent use case, addressing one of the persistent pain points in data engineering workflows. These adoption patterns suggest that AI tools are proving most valuable in areas where data engineers traditionally spend significant time on routine, albeit important, tasks. ## Transforming core data engineering activities AI tools are demonstrating measurable impact across the spectrum of data engineering responsibilities. In coding and transformation work, AI assistants can generate SQL statements, create complex regular expressions, and perform bulk edits on existing code bases. This capability proves particularly valuable when working with [dbt](https://www.getdbt.com/product/what-is-dbt), where the structured nature of the transformation framework provides rich context that AI tools can leverage to produce more accurate, relevant code suggestions. In dbt specifically, teams use [dbt Copilot](https://www.getdbt.com/product/dbt-copilot) to generate tests, documentation, metrics, semantic models, and starter SQL. [Fusion’s state-aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-about) reduces unnecessary recompute by building only impacted nodes, accelerating safe iteration without changing engineering standards. The impact extends beyond simple code generation. AI tools excel at multi-file refactoring tasks that traditionally consume substantial engineering time. For example, when a new field is added to a source system, AI can trace that field through the entire data pipeline, updating models and dependencies accordingly. Similarly, when consolidating duplicate logic across different parts of a data transformation graph, AI can identify opportunities for optimization and implement changes across multiple files simultaneously. Testing and quality assurance represent another area where AI tools provide significant efficiency gains. Rather than manually crafting [test cases for each data model](https://www.getdbt.com/blog/data-testing), AI can analyze the structure and logic of transformations to suggest appropriate validation tests. These might include checks for data completeness, referential integrity, or business rule compliance. The AI understands the context of the data model: its dependencies, transformations, and expected outputs to recommend tests that are both comprehensive and relevant. Documentation, long considered a necessary but time-intensive aspect of data engineering, becomes far more manageable with AI assistance. AI tools can generate initial documentation for tables and columns based on their names, the logic that creates them, and patterns observed in similar data assets. While human review and refinement remain essential, this automated first pass eliminates the blank page problem and provides a foundation that engineers can build upon. ## The critical role of frameworks and standards The effectiveness of AI tools in data engineering contexts depends heavily on the underlying structure and consistency of the codebase. Frameworks like dbt provide the standardization that makes AI assistance more reliable and valuable. When data transformations follow consistent patterns, use well-documented conventions, and employ standard testing and deployment practices, AI tools can better understand the context and generate more accurate suggestions. This relationship between frameworks and AI effectiveness creates a virtuous cycle. Teams using structured approaches to data engineering with consistent coding standards, modular design patterns, and comprehensive metadata see greater benefits from AI tools. In turn, the efficiency gains from AI assistance make it easier to maintain these high standards, as the tools can automatically enforce formatting rules, suggest appropriate tests, and generate documentation that follows established patterns. The importance of this standardization becomes clear when considering the alternative. In environments with inconsistent coding practices, multiple languages or frameworks, and ad-hoc development approaches, AI tools struggle to provide reliable assistance. The heterogeneity that makes codebases difficult for humans to maintain also makes them challenging for AI systems to understand and work with effectively. ## Addressing stakeholder interactions and self-service analytics Beyond direct code development, AI tools are beginning to transform how data engineers interact with business stakeholders. Traditionally, data engineers field numerous questions about data availability, quality, and appropriate usage. These interactions, while valuable, can become bottlenecks that slow both the engineering team and the business users who need data insights. AI-powered interfaces that understand an organization's data catalog, lineage, and quality metrics can handle many routine inquiries without human intervention. Business users can ask questions about which datasets to use for specific analyses, understand data freshness and reliability, and even generate basic queries using natural language. This shift doesn't eliminate the need for data engineering expertise, but it allows engineers to focus on more complex, strategic work rather than repetitive support tasks. The development of context protocols (standardized ways for AI systems to access and understand enterprise data metadata) promises to accelerate this trend. When AI assistants can access comprehensive information about data sources, transformations, and quality metrics, they become far more capable of providing accurate, helpful responses to business users. **** ## Measuring efficiency gains and managing expectations While the potential for efficiency improvements appears substantial, organizations are still learning to measure and realize these benefits effectively. Early adopters report significant time savings in specific tasks, particularly in areas like documentation generation, basic test creation, and code formatting. However, the overall impact on data engineering productivity varies considerably based on factors like existing code quality, team practices, and the specific AI tools employed. The most successful implementations combine AI assistance with strong governance and review processes. AI-generated code still requires human oversight to ensure accuracy, adherence to business requirements, and integration with existing systems. Teams that treat AI as a powerful assistant rather than a replacement for human judgment tend to see the best results. There's also an important distinction between efficiency gains in individual tasks and overall productivity improvements. While AI can dramatically reduce the time required to write initial documentation or generate test cases, the broader impact depends on how these time savings translate into faster delivery of data products, improved data quality, or enhanced team capacity for strategic initiatives. ## Challenges and limitations Despite the promising developments, AI tools in data engineering face several important limitations. Generic large language models, while powerful, often lack the specific context needed to generate truly robust data pipelines. They may suggest code that works syntactically but fails to account for business logic, data quality requirements, or performance considerations specific to an organization's environment. The accuracy of AI-generated code remains a concern, particularly for complex transformations or edge cases. While AI tools excel at routine tasks and common patterns, they can struggle with nuanced business requirements or unusual data scenarios. This limitation underscores the continued importance of human expertise in reviewing, testing, and refining AI-generated outputs. Security and compliance considerations also present challenges. In highly regulated industries, the use of AI tools for code generation must be carefully managed to ensure that generated code meets all relevant compliance requirements. Organizations need clear policies about when and how AI assistance can be used, particularly when dealing with sensitive data or critical business processes. ## The evolution of data engineering roles As AI tools become more sophisticated and widely adopted, they're beginning to reshape data engineering roles themselves. Rather than eliminating positions, the technology appears to be pushing data engineers toward more specialized and strategic work. Some engineers are focusing more heavily on data platform architecture and infrastructure, ensuring that the systems supporting AI-assisted development are robust, scalable, and well-governed. Others are moving toward closer collaboration with business teams, using their technical expertise to translate business requirements into data solutions while leveraging AI tools to handle routine implementation tasks. Still others are specializing in automation and orchestration, building systems that can act on data insights rather than simply generating reports. This evolution reflects a broader trend toward higher-value work enabled by AI assistance. As routine coding, documentation, and testing tasks become more automated, data engineers can focus on problems that require human creativity, business understanding, and strategic thinking. ## Looking forward The integration of AI tools into data engineering workflows represents more than a simple productivity enhancement: it's a fundamental shift in how data teams operate. Organizations that successfully navigate this transition will likely see significant competitive advantages in their ability to deliver reliable, high-quality data products quickly and efficiently. The key to success lies in thoughtful implementation that combines AI capabilities with strong engineering practices, appropriate governance, and clear understanding of the technology's limitations. Teams that view AI as a powerful tool to augment human expertise, rather than replace it, are positioning themselves to realize the full potential of this technological shift. As AI tools continue to evolve and improve, their impact on data engineering efficiency will likely grow. However, the fundamental principles of good data engineering (clear requirements, robust testing, comprehensive documentation, and strong governance) remain as important as ever. AI tools make it easier to implement these principles consistently, but they don't eliminate the need for human judgment, creativity, and expertise in building systems that truly serve business needs. The future of data engineering will be characterized by this partnership between human expertise and AI assistance, with both elements essential for delivering the reliable, efficient data systems that modern organizations require. **** ## AI in data engineering FAQs **How should data engineers integrate vector databases, embeddings, and RAG pipelines into their 2025 data stacks?** Data engineers should approach vector databases and RAG pipelines as extensions of their existing data infrastructure, applying the same principles of standardization and governance that make AI tools effective. The key is treating embeddings as another data type that requires proper lineage tracking, quality monitoring, and consistent transformation patterns. Integration works best when vector operations follow established frameworks like dbt for transformations, ensuring that AI assistants can understand and maintain these pipelines effectively. Organizations should also implement clear protocols for managing embedding updates and version control, as the dynamic nature of vector representations requires careful handling to maintain system reliability. **How can standardized frameworks and homogeneous codebases make AI assistants more effective at authoring, testing, and maintaining data pipelines?** Standardized frameworks create a virtuous cycle with AI effectiveness by providing consistent patterns that AI tools can reliably understand and replicate. When data transformations follow established conventions through frameworks like dbt, AI assistants can generate more accurate code suggestions, appropriate test cases, and comprehensive documentation because they have clear context about expected structures and practices. This standardization allows AI tools to perform complex multi-file refactoring tasks, trace dependencies across entire pipelines, and enforce consistent coding standards automatically. Teams with well-structured, homogeneous codebases see dramatically better results from AI assistance compared to those with ad-hoc development approaches, as the consistency makes it easier for AI systems to understand context and generate reliable outputs **Will AI replace data engineers?** No, AI will not replace data engineers, but rather transform their roles toward more strategic and specialized work. While AI tools excel at routine tasks like code generation, documentation, and basic testing, they still require human oversight for accuracy, business logic validation, and complex problem-solving. The technology is pushing data engineers toward higher-value activities such as platform architecture, business collaboration, and automation system design. Data engineers remain essential for translating business requirements into technical solutions, ensuring data quality and governance, and building the robust frameworks that make AI assistance effective in the first place. The fundamental principles of good data engineering (clear requirements, comprehensive testing, and strong governance) still require human expertise, creativity, and judgment that AI cannot replace. --- --- title: "How data cleaning boosts transformation quality and reliability" description: "Discover how to improve quality and maintain trust in your data workflows." url: "https://www.getdbt.com/blog/data-cleaning-transformation-quality" date: "2025-08-27" authors: ["Joey Gault"] categories: ["Pulse"] --- # How data cleaning boosts transformation quality and reliability Data cleaning is the process of identifying and correcting errors, inconsistencies, and inaccuracies in datasets. This includes finding and filling missing values, correcting inaccuracies, removing duplicated records, and standardizing formats. However, data cleaning extends beyond simple error correction—it's fundamentally about preparing data for reliable analysis and ensuring that downstream transformations can operate on a solid foundation. Within the broader transformation workflow, cleaning often occurs in tandem with other processes like normalization and validation. This overlap is intentional and beneficial. When you normalize currency values to a standard format like USD, you're simultaneously cleaning inconsistent representations and ensuring comparability across datasets. Similarly, validation checks that verify data adheres to specified criteria often catch cleaning issues before they propagate through your transformation pipeline. The modern [ELT (Extract, Load, Transform](https://www.getdbt.com/blog/extract-load-transform)) approach has made data cleaning more strategic than ever. Unlike traditional ETL processes where cleaning happened before loading, ELT allows teams to clean data within the warehouse environment, leveraging the full computational power of modern cloud platforms. This shift enables more sophisticated cleaning operations and makes it easier to iterate on cleaning logic as business requirements evolve. ## Core data cleaning techniques that enhance transformation quality Effective data cleaning during transformations requires a systematic approach to common data quality issues. Missing value handling represents one of the most critical cleaning operations. Rather than simply dropping incomplete records, sophisticated cleaning processes evaluate the context of missing data. For customer records, missing phone numbers might be acceptable for email-based campaigns, but missing customer IDs would break downstream join operations. The cleaning logic must understand these business contexts to make appropriate decisions. Duplicate detection and removal requires more nuance than simple exact matching. Customer records might have slight variations in name spelling or address formatting while representing the same entity. Advanced cleaning processes implement fuzzy matching algorithms to identify these near-duplicates and establish canonical representations. This cleaning step is crucial for accurate customer analytics and prevents double-counting in downstream aggregations. Data type standardization ensures that transformation operations can execute reliably. When date fields arrive in multiple formats—some as strings, others as timestamps with different timezone information—cleaning processes must standardize these representations. This standardization prevents transformation failures and ensures that date-based calculations produce consistent results across different data sources. Format consistency extends beyond data types to business-specific standards. Phone numbers might arrive as "(555) 123-4567", "555-123-4567", or "5551234567". Cleaning processes establish canonical formats that downstream transformations can depend on. This consistency is particularly important for data integration operations where multiple sources must be joined on potentially formatted fields. ## The compounding benefits of clean data in transformation pipelines Clean data creates a virtuous cycle throughout the transformation process. When initial cleaning operations remove inconsistencies and errors, subsequent transformation steps can focus on business logic rather than error handling. This separation of concerns makes transformation code more maintainable and reduces the likelihood of bugs that could corrupt final outputs. Aggregation operations particularly benefit from thorough data cleaning. When calculating customer lifetime value, dirty data can skew results dramatically. A single customer record with an incorrectly entered purchase amount of $100,000 instead of $100.00 will distort averages and percentiles across the entire customer base. Cleaning processes that implement reasonable bounds checking and outlier detection prevent these errors from propagating through complex aggregation logic. [Data integration](https://www.getdbt.com/blog/data-integration) becomes significantly more reliable when source data has been properly cleaned. Joining customer data from a CRM system with transaction data from an e-commerce platform requires consistent customer identifiers. If the cleaning process hasn't standardized email addresses—removing extra whitespace, converting to lowercase, handling common typos—the join operation will miss legitimate matches and create incomplete customer profiles. The reliability gains from proper cleaning compound over time. As transformation pipelines grow more complex and interdependent, the cost of data quality issues increases exponentially. A cleaning error that affects a foundational customer dimension table will impact every downstream model that depends on customer data. Conversely, robust cleaning at the foundation level ensures that complex business logic can operate on trustworthy inputs. ## Implementing scalable data cleaning with modern tools Managing data cleaning at enterprise scale requires more than ad hoc scripts and manual processes. Modern [data transformation platforms like dbt](https://www.getdbt.com/product/what-is-dbt) provide structured approaches to implementing and maintaining cleaning logic. By treating cleaning operations as code, teams can version control their cleaning rules, test them systematically, and deploy changes through proper CI/CD processes. dbt's approach to data cleaning emphasizes repeatability and transparency. Cleaning logic written as [dbt models](https://docs.getdbt.com/docs/build/models) can be reviewed, tested, and documented alongside other transformation code. This integration ensures that cleaning operations receive the same engineering rigor as business logic transformations. When cleaning rules need to change—perhaps to handle new data sources or evolving business requirements—the changes can be implemented, tested, and deployed through established workflows. The [testing capabilities built into dbt](https://docs.getdbt.com/docs/build/data-tests) are particularly valuable for data cleaning operations. Teams can write tests that verify cleaning logic produces expected results: ensuring that phone number standardization handles all expected input formats, confirming that duplicate detection doesn't inadvertently merge distinct records, or validating that missing value imputation follows business rules. These tests run automatically as part of the transformation pipeline, catching cleaning failures before they affect downstream processes. [Version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) becomes crucial when cleaning logic evolves. As data sources change or business requirements shift, cleaning rules must adapt. With dbt, these changes are tracked in Git, making it possible to understand why cleaning logic changed and to roll back problematic updates. This historical context is invaluable when debugging data quality issues or explaining analytical results to stakeholders. ## Monitoring and continuous improvement of cleaning processes Data cleaning isn't a set-it-and-forget-it operation. As data sources evolve and business requirements change, cleaning logic must adapt accordingly. Effective monitoring helps teams understand when cleaning processes need attention and provides insights for continuous improvement. Automated monitoring should track both the volume and types of cleaning operations performed. If the percentage of records requiring duplicate removal suddenly increases, it might indicate changes in upstream data collection processes. Similarly, if missing value imputation rates spike for certain fields, it could signal data quality degradation in source systems that requires investigation. Performance monitoring ensures that cleaning operations scale with data volumes. As datasets grow, cleaning logic that worked well on smaller datasets might become bottlenecks. Monitoring query performance and resource utilization helps teams optimize cleaning operations before they impact overall pipeline performance. Modern cloud data warehouses provide detailed performance metrics that make this monitoring straightforward. Quality metrics provide feedback on cleaning effectiveness. Teams should track measures like the percentage of records that pass validation tests after cleaning, the consistency of key business metrics before and after cleaning operations, and the rate of downstream transformation failures that can be attributed to data quality issues. These metrics help quantify the business value of cleaning investments and guide prioritization of improvement efforts. ## Strategic considerations for data engineering leaders For data engineering leaders, data cleaning represents both a technical challenge and a strategic opportunity. Organizations that invest in robust cleaning processes gain competitive advantages through more reliable analytics and faster time-to-insight for new use cases. However, these investments require careful planning and resource allocation. The build-versus-buy decision for cleaning capabilities depends on organizational context. While custom cleaning solutions offer maximum flexibility, they require ongoing maintenance and specialized expertise. Modern transformation platforms like [dbt](https://www.getdbt.com/product/dbt) provide built-in cleaning capabilities that can handle most common scenarios while allowing customization for specific business needs. This approach often provides the best balance of capability and maintainability. Team structure and skills development play crucial roles in cleaning success. Data cleaning requires understanding both technical implementation details and business context. Teams need members who can write efficient SQL for cleaning operations while also understanding the business implications of different cleaning decisions. Cross-training between data engineers and business analysts often produces the best outcomes for cleaning initiatives. [Governance](https://www.getdbt.com/blog/data-governance) and compliance considerations become more complex as cleaning operations scale. Cleaning processes that modify or remove data must comply with regulatory requirements and audit standards. Documentation becomes crucial—not just for technical maintenance, but for demonstrating compliance with data handling regulations. Modern transformation platforms provide audit trails and documentation capabilities that support these governance requirements. ## Conclusion Data cleaning serves as the foundation for reliable data transformations, creating a cascade of quality improvements throughout the analytics pipeline. When implemented systematically using modern tools like dbt, cleaning processes become scalable, maintainable, and auditable components of the data infrastructure. The investment in robust cleaning capabilities pays dividends through improved analytical reliability, reduced debugging overhead, and increased stakeholder confidence in data-driven decisions. The key to successful data cleaning lies in treating it as an integral part of the transformation process rather than a separate preprocessing step. By embedding cleaning logic within transformation workflows, teams can ensure that quality improvements compound throughout their data pipelines. As data volumes continue to grow and analytical requirements become more sophisticated, organizations with strong data cleaning foundations will be better positioned to extract value from their data assets while maintaining the trust and reliability that stakeholders demand. ## Data cleaning FAQs **What is data cleaning?** Data cleaning is the process of identifying and correcting errors, inconsistencies, and inaccuracies in datasets. This includes finding and filling missing values, correcting inaccuracies, removing duplicated records, and standardizing formats. However, data cleaning extends beyond simple error correction—it's fundamentally about preparing data for reliable analysis and ensuring that downstream transformations can operate on a solid foundation. **Why is data cleaning important?** Data cleaning is crucial because it creates a foundation for reliable data transformations and analytics. Clean data prevents errors from propagating through complex transformation pipelines, makes aggregation operations more accurate, and enables reliable data integration. When initial cleaning operations remove inconsistencies and errors, subsequent transformation steps can focus on business logic rather than error handling, making the entire process more maintainable and reducing the likelihood of bugs that could corrupt final outputs. **How does data cleaning differ from data validation, and when is each applied?** Data cleaning focuses on identifying and correcting errors, inconsistencies, and inaccuracies in datasets, while data validation checks that data adheres to specified criteria. These processes often work in tandem within transformation workflows, with validation checks frequently catching cleaning issues before they propagate through the pipeline. In modern ELT approaches, both cleaning and validation occur within the warehouse environment, allowing teams to leverage full computational power and iterate on logic as business requirements evolve. --- --- title: "Overcoming impostor syndrome in tech" description: "Women in data share real stories and advice for managing impostor syndrome in tech on The View on Data podcast." url: "https://www.getdbt.com/blog/overcoming-impostor-syndrome-in-tech" date: "2025-08-26" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Overcoming impostor syndrome in tech **In episode 2 of _The View on Data_, four women in tech share honest stories about self-doubt, resilience, and navigating impostor syndrome in the data industry.** If you've ever wondered whether you're qualified enough for your role, if everyone else is more prepared than you, or if you're the only one struggling, you’re not alone. This episode is a must-listen for anyone in data feeling impostor-y. Please reach out at **podcast@dbtlabs.com** for questions, comments, and guest suggestions. 🎧 **Listen & subscribe:** [Spotify](https://open.spotify.com/show/0Itxg3BoF6zaQdORfVuY6I?si=6133f88257624765) | [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-view-on-data/id1835755886) | [Amazon Music](https://music.amazon.com/podcasts/84c47dbe-adf7-4d33-a371-a1fe55255495/the-view-on-data) | [YouTube](https://youtu.be/3LPk3R1kIyQ) ## Why impostor syndrome is so common in data careers The tech industry, and data roles in particular, attract high achievers. But they also attract uncertainty. Whether you're an analytics engineer, data analyst, or just starting your career in data, you've likely felt that moment: _Do I really belong here?_ Lauren Benezra, Paige Berry, Faith McKenna, and Erica Louie from dbt Labs dive into how impostor syndrome shows up in their daily work and why it hits even the most experienced professionals. > “I’ve been doing data work for over 20 years… and I still feel like an impostor.” – Paige Berry ## What impostor syndrome feels like on a data team This episode highlights the emotional reality of working in tech, especially as a woman or an underrepresented person. The pressure to be perfect. The fear of being "found out." The exhausting emotional labor of proving yourself again and again. It also unpacks how impostor syndrome can creep in even after you’ve landed the job or gotten the promotion—especially when you’re the “first” or “only” in the room. > “When someone says ‘you’re the expert,’ and you don’t feel like one—it’s a mental spiral.” – Faith McKenna ## Tools and tactics to fight impostor syndrome in tech The group shares practical ways to reframe self-doubt and build confidence in your data career: - **Prep with intention:** Spend time researching, planning, and writing things out—but set a boundary. Don’t overdo it. - **Use a work journal or brag folder:** Track wins, compliments, and proud moments. Refer back when self-doubt kicks in. - **Talk to people you trust:** Managers, peers, and mentors can offer a perspective you might not see yourself. - **Use AI tools like ChatGPT:** Draft, review, or reframe your work in a way that builds confidence. ## The importance of psychological safety on data teams One of the most powerful themes in the episode is the role of culture in managing impostor syndrome. When companies prioritize transparency, psychological safety, and feedback over perfectionism, it changes everything. At dbt Labs, the team credits open channels, supportive leadership, and an emphasis on learning with helping them move past self-doubt and grow into their roles. > “If you’re not getting support or feedback, it might not be imposter syndrome—it might be your environment.” – Erica Louie ## Reframing your inner critic Impostor syndrome isn’t just a tech problem. It’s a human problem. But as this episode shows, it _can_ be managed—especially when you shift how you see it. Feelings aren’t facts. And self-doubt doesn’t disqualify you. Here are a few closing tips from the team: - Name your lizard brain. (Seriously. Paige’s is called Karen.) - Collect positive feedback like receipts. - Build your own story—not someone else’s narrative. - Don’t aim to eliminate doubt. Aim to recover faster from it. > “Negative feelings are just signals. They don’t have to be the whole story.” – Lauren Benezra --- --- title: "Building resilience: Observability for modern data teams" description: "Explore how dbt and observability tools like Monte Carlo help data teams improve reliability, performance, and trust at scale." url: "https://www.getdbt.com/blog/data-observability" date: "2025-08-25" authors: ["Joey Gault"] categories: ["Pulse"] --- # Building resilience: Observability for modern data teams Data observability has emerged as a distinct discipline that goes beyond traditional data quality monitoring. While data quality focuses on the accuracy and completeness of data itself, observability encompasses the broader visibility into data systems—understanding not just what happened, but when, why, and how it impacts downstream processes. The need for observability becomes particularly acute during periods of data downtime, when data is partial, erroneous, missing, or inaccurate. These incidents can cascade through an organization, affecting everything from executive dashboards to automated marketing campaigns. The cost of poor data quality extends beyond immediate operational impacts—it erodes trust in data systems and can lead to hesitation in making data-driven decisions. Modern data stacks, while powerful and flexible, have actually increased the surface area for potential issues. The proliferation of data sources, transformation tools, and consumption endpoints creates more points of failure and makes it harder to maintain end-to-end visibility. This fragmentation makes observability not just helpful, but essential for maintaining reliable data operations at scale. ## Learning from SurveyMonkey's approach SurveyMonkey's data engineering team provides a compelling example of how organizations can build resilience through observability. When the team conducted an internal survey to understand data challenges across the organization, they discovered that 53% of respondents cited data quality issues as a primary concern. More surprisingly, 50% identified data processing issues—problems with pipeline performance and reliability—as equally significant challenges. This discovery led to a systematic approach that combined two complementary strategies: reactive monitoring through Monte Carlo and proactive data management through dbt. Rather than treating these as separate initiatives, SurveyMonkey integrated them to create a comprehensive observability framework that addressed both data quality and processing performance. The company's data landscape provides context for the scale of this challenge. SurveyMonkey processes data from nearly 900 sources, maintains 600 dbt models with 2,000 test cases, and runs over 180 workflows daily. This complexity requires sophisticated monitoring and alerting capabilities that can operate reliably regardless of pipeline success or failure. ## Integrating reactive monitoring and proactive testing The integration of Monte Carlo and dbt at SurveyMonkey demonstrates how reactive monitoring and proactive testing can work together synergistically. dbt serves as the foundation for data transformation, providing models that clean raw data from various sources and create high-quality, usable datasets. The dbt testing framework verifies data quality through automated checks, while dbt documentation creates consistent, well-documented models that serve as single sources of truth. Monte Carlo complements this foundation by monitoring dbt models and pipelines in production, detecting anomalies and ensuring ongoing data accuracy. The real power emerges from how these tools work together rather than in isolation. One of the most effective practices SurveyMonkey implemented was converting Monte Carlo anomalies into dbt test cases. When the monitoring system detected an anomaly that indicated a serious data quality issue, the team would create a corresponding dbt test that would prevent the pipeline from proceeding if the same condition occurred again. This approach shifts the responsibility for data quality upstream, enabling business users to address issues at their source rather than waiting for data engineering intervention. This integration also standardized quality checks across all models. Every dbt model became subject to mandatory testing, creating a consistent baseline for data quality expectations. Regular performance reviews ensured that models didn't degrade over time, while automated monitoring and alerting provided proactive notifications when issues arose. ## Measuring the impact of observability The results of SurveyMonkey's observability implementation demonstrate the tangible business value of comprehensive monitoring and testing. The most striking outcome was a 73% reduction in Snowflake credit usage across nearly 10,000 credit jobs. This dramatic improvement came primarily from performance monitoring that identified inefficient queries, unused models, and optimization opportunities. Performance monitoring revealed long-running jobs performing unnecessary cross-joins, enabling the team to simplify SQL statements and merge redundant queries. The systematic identification and removal of unused models and tables further contributed to cost reduction. In one particularly impressive case, obsoleting unused models and revamping inefficient queries resulted in a 94% reduction in pipeline runtime and a 97% reduction in Snowflake credit usage. Beyond cost savings, the observability framework enabled SurveyMonkey to scale while keeping costs stable. Despite adding many more dbt models—including bringing a marketing analytics platform in-house—job execution times remained stable while cost per credit steadily decreased. This demonstrates how effective observability can support growth without proportional increases in operational overhead. Data quality improvements were equally significant. When SurveyMonkey integrated the third-party marketing analytics platform, Monte Carlo initially detected a spike in data anomalies as the team learned to work with the new data sources. However, the systematic conversion of anomalies into dbt test cases led to a corresponding decrease in anomalies over time, creating a self-improving system that became more robust with experience. ## Building observability with dbt artifacts While third-party observability tools provide valuable capabilities, teams can also build significant observability using dbt's native artifacts and metadata. dbt generates detailed artifacts after every run, test, or build command, containing granular information about model execution, test results, and pipeline performance. These artifacts serve as a rich data source for custom observability solutions. The project manifest provides complete configuration information for the dbt project, while run results artifacts contain detailed execution data for models, tests, and other resources. When combined with data warehouse query history, these artifacts enable deep insights into model-level performance that can inform optimization decisions. Teams have successfully built lightweight ELT systems that ingest artifact data into their data warehouses, then use dbt itself to transform this metadata into structured models that power dashboards and alerting systems. This approach leverages existing infrastructure and skills while providing customizable observability tailored to specific organizational needs. The key components of such a system include orchestration that reliably captures artifacts regardless of pipeline success or failure, storage that preserves historical artifact data for trend analysis, modeling that transforms raw artifacts into actionable insights, and alerting that notifies relevant stakeholders when issues arise. ## Implementing effective alerting strategies Effective alerting represents one of the most critical aspects of data observability, yet it's often implemented poorly. The goal is to provide timely, actionable notifications to the right people without creating alert fatigue or overwhelming teams with false positives. SurveyMonkey's approach demonstrates several best practices for data alerting. First, they implemented domain-specific tagging that allows alerts to be routed to appropriate team members based on model ownership. Every model in their dbt deployment includes domain tags like "growth," "finance," or "catalog," which correspond to Slack user groups containing relevant stakeholders. This targeted alerting ensures that model owners receive notifications about their specific models rather than broadcasting alerts to entire teams. The alerts include sufficient context for debugging, including error messages, model names, and timestamps, enabling recipients to quickly understand and address issues. Importantly, the team learned not to introduce anomaly notifications to business users at the beginning of new data integrations. When data engineers themselves don't fully understand new data sources, it's counterproductive to alert business users about anomalies. Taking time to let pipelines stabilize and understand normal data patterns before involving business users prevents unnecessary noise and maintains alert credibility. ## Performance optimization through observability Beyond alerting and quality monitoring, observability data provides valuable insights for performance optimization. By combining dbt artifacts with data warehouse query history, teams can identify models that are candidates for different materialization strategies, clustering improvements, or warehouse sizing adjustments. Performance dashboards can surface models with high execution times, excessive data spillage, or inefficient partition scanning patterns. Time series views of individual models help identify performance degradation over time, while pipeline-level visualizations reveal bottlenecks that affect overall execution times. These insights enable data teams to make informed decisions about optimization priorities. Rather than guessing which models might benefit from incremental materialization or increased warehouse sizes, teams can use concrete performance data to guide their efforts and measure the impact of changes. ## Lessons learned and best practices The experiences of teams implementing data observability reveal several important lessons. First, the combination of proactive testing and reactive monitoring creates more resilient systems than either approach alone. dbt tests catch many issues before they reach production, while observability tools detect the problems that slip through. Second, stopping pipelines on bad data gets everyone's attention and creates accountability for data quality. When important anomalies are detected, creating test cases that prevent pipeline execution until upstream issues are resolved shifts responsibility appropriately and prevents the propagation of known data quality problems. Third, onboarding new users to observability systems from the beginning ensures adoption and effectiveness. Training business users to leverage observability tools for incident routing and data discovery creates a self-service culture that reduces the burden on data engineering teams. Finally, the tools and approaches used for observability should integrate well with existing workflows and infrastructure. Solutions that require significant additional overhead or specialized expertise are less likely to be maintained effectively over time. ## The future of data observability As data systems continue to evolve, observability practices must adapt to new challenges and opportunities. The emergence of data mesh architectures and domain-driven data ownership creates new requirements for observability that spans organizational boundaries while maintaining appropriate access controls and governance. Artificial intelligence and machine learning are beginning to enhance observability capabilities, from automated anomaly detection to intelligent alerting that reduces false positives. However, the fundamental principles of comprehensive monitoring, proactive testing, and effective alerting remain constant. The most successful data teams will be those that treat observability as a core competency rather than an afterthought. This means investing in the tools, processes, and skills necessary to maintain visibility into data systems at scale, while fostering a culture that values data reliability and quality as essential business capabilities. Building resilience through observability requires commitment and investment, but the returns—in terms of cost savings, improved reliability, and increased trust in data—justify the effort. As data becomes increasingly central to business operations, the organizations that master data observability will have a significant competitive advantage in their ability to make reliable, data-driven decisions at scale. ## Data observability FAQs **What is data observability?** Data observability is a distinct discipline that goes beyond traditional data quality monitoring. While data quality focuses on the accuracy and completeness of data itself, observability encompasses broader visibility into data systems—understanding not just what happened, but when, why, and how it impacts downstream processes. It provides comprehensive monitoring and testing capabilities that can operate reliably regardless of pipeline success or failure. **Why is data observability important? ** Data observability becomes particularly crucial during periods of data downtime, when data is partial, erroneous, missing, or inaccurate. These incidents can cascade through an organization, affecting everything from executive dashboards to automated marketing campaigns. The cost of poor data quality extends beyond immediate operational impacts—it erodes trust in data systems and can lead to hesitation in making data-driven decisions. Modern data stacks have increased the surface area for potential issues, making observability essential for maintaining reliable data operations at scale. **What are the five pillars of data observability? ** The core components of effective data observability include orchestration that reliably captures artifacts regardless of pipeline success or failure, storage that preserves historical data for trend analysis, modeling that transforms raw information into actionable insights, alerting that notifies relevant stakeholders when issues arise, and performance monitoring that identifies optimization opportunities. These elements work together to create comprehensive visibility into data systems and enable proactive management of data quality and reliability. --- --- title: "Collaborative analytics engineering: A guide to scalable development" description: "Discover how dbt supports scalable, collaborative analytics engineering with modular code, testing, and cross-team workflows." url: "https://www.getdbt.com/blog/collaborative-analytics-engineering" date: "2025-08-25" authors: ["Joey Gault"] categories: ["Pulse"] --- # Collaborative analytics engineering: A guide to scalable development Successful collaborative analytics engineering starts with establishing shared standards and practices across teams. Unlike traditional software development, analytics engineering must balance technical rigor with business context, making collaboration both more critical and more complex. The most effective analytics engineering teams operate with a clear division of responsibilities while maintaining strong communication channels. Analytics engineers typically own the transformation layer, working in tools like dbt to convert raw data into analysis-ready datasets. They collaborate upstream with data engineers who manage ingestion and infrastructure, and downstream with analysts and business users who consume the transformed data for insights and reporting. This collaboration requires more than just technical coordination. Analytics engineers must understand business logic deeply enough to encode it correctly in their transformations, while also maintaining the technical discipline to ensure code quality, performance, and reliability. The 2025 State of Analytics Engineering Report shows that 57% of analytics professionals spend most of their time maintaining and organizing datasets, highlighting the operational nature of this work and the importance of sustainable development practices. Version control becomes particularly crucial in collaborative analytics engineering. Unlike traditional analysis work that might live in isolated notebooks or ad-hoc queries, analytics engineering code forms the foundation for multiple downstream use cases. Changes to core data models can impact dashboards, reports, and other analytical work across the organization. Teams must implement branching strategies, code review processes, and deployment practices that ensure stability while enabling rapid iteration. Documentation and knowledge sharing take on heightened importance in collaborative environments. Analytics engineering code often encodes complex business logic that may not be immediately obvious to other team members. Comprehensive documentation of data lineage, transformation logic, and business rules becomes essential for team members to understand, maintain, and extend each other's work. ## Building scalable development workflows As analytics engineering teams grow, ad-hoc collaboration approaches quickly become insufficient. Scalable development requires systematic workflows that can accommodate multiple contributors working on interconnected data models without creating conflicts or quality issues. The most successful teams adopt development workflows borrowed from software engineering but adapted for analytics use cases. This typically involves feature branching for new development work, automated testing to catch regressions, and staged deployment environments that allow for safe testing before production releases. However, analytics engineering presents unique challenges that require specialized approaches. Data dependencies create complex coordination requirements. Unlike traditional software where modules can often be developed independently, analytics engineering work frequently involves shared data models and interdependent transformations. Changes to upstream models can break downstream dependencies, requiring careful coordination and communication between team members working on different parts of the data pipeline. Testing strategies must account for both code correctness and data quality. While traditional software testing focuses on logic and functionality, analytics engineering testing must also validate data freshness, completeness, and business rule compliance. Teams need automated testing frameworks that can catch both technical errors and data quality issues before they impact downstream users. The modular design principles that enable scalable software development apply equally to analytics engineering. Teams should structure their data models as reusable components with clear interfaces and well-defined responsibilities. This modular approach reduces code duplication, centralizes business logic, and makes it easier for team members to understand and modify each other's work. Environment management becomes more complex in collaborative analytics engineering. Teams need development, staging, and production environments that can accommodate multiple concurrent development streams while maintaining data consistency and performance. This often requires sophisticated orchestration and resource management capabilities that go beyond traditional software development needs. ## Organizational structures for scale The organizational structure of analytics engineering teams significantly impacts their ability to collaborate effectively and scale their impact. Different organizational models work better for different company sizes, data maturity levels, and business requirements. Many organizations start with a centralized analytics engineering team that serves the entire company. This approach works well for smaller organizations or those in the early stages of analytics engineering adoption. A centralized team can establish consistent standards, build reusable components, and develop deep expertise in the tools and practices of analytics engineering. However, as organizations grow, centralized teams can become bottlenecks, struggling to understand the nuanced requirements of different business areas. The hybrid model, where analytics engineers are distributed across business functions while maintaining connection to a central team, has gained popularity as organizations scale. This approach allows analytics engineers to develop deep domain expertise while still benefiting from shared standards and practices. The 2025 State of Analytics Engineering Report indicates that this hybrid approach is becoming more common, with work distributed by both business area and function. Embedded analytics engineers work directly within business teams, providing dedicated support for specific domains like marketing, finance, or operations. This model maximizes business alignment and domain expertise but can lead to inconsistent practices and duplicated effort across teams. Organizations using this model need strong governance frameworks and regular cross-team collaboration to maintain consistency. Regardless of organizational structure, successful analytics engineering teams maintain regular communication and knowledge sharing practices. This might include regular technical reviews, shared documentation repositories, and cross-team rotation programs that help spread knowledge and maintain consistency across the organization. ## Technology and tooling considerations The technology stack choices for collaborative analytics engineering significantly impact team productivity and scalability. While individual analytics engineers might be productive with basic SQL editors and manual processes, collaborative teams require more sophisticated tooling that supports version control, testing, documentation, and deployment automation. dbt has become central to many collaborative analytics engineering workflows because it brings software engineering practices to data transformation work. Its built-in support for version control, testing, documentation, and modular development makes it well-suited for team environments. The dbt ecosystem also provides specialized tools for collaborative development, including cloud-based development environments and automated deployment capabilities. Cloud data warehouses like Snowflake, BigQuery, and Redshift provide the computational foundation for collaborative analytics engineering. These platforms offer the performance and scalability needed for complex transformations while providing the isolation and resource management capabilities that teams need for safe collaborative development. Orchestration and scheduling tools become more important as teams scale. While individual analytics engineers might manually run their transformations, collaborative teams need automated scheduling that can handle complex dependencies and coordinate work across multiple contributors. Tools like Airflow, Prefect, and cloud-native orchestration services provide the reliability and observability needed for production analytics engineering workflows. Data quality and observability tools help teams maintain reliability as they scale. Automated data quality monitoring can catch issues before they impact downstream users, while lineage tracking helps teams understand the impact of changes across complex data pipelines. These capabilities become essential as the number of contributors and the complexity of data transformations increase. ## Managing complexity and technical debt As analytics engineering teams and codebases grow, managing complexity becomes a critical challenge. Unlike traditional software applications, analytics engineering projects often accumulate technical debt through incremental business requirement changes, evolving data sources, and the natural entropy that occurs in long-running data pipelines. Refactoring strategies for analytics engineering must account for the interconnected nature of data transformations. Changes to core data models can have far-reaching impacts that are difficult to predict and test comprehensively. Teams need systematic approaches to refactoring that include comprehensive impact analysis, staged rollouts, and rollback capabilities. Performance optimization becomes more challenging in collaborative environments where multiple team members are making changes to shared data models. Teams need monitoring and profiling capabilities that can identify performance regressions and attribute them to specific changes. This requires sophisticated observability tools and development practices that include performance testing as part of the standard workflow. Code organization and architecture decisions have long-term impacts on team productivity. Teams should establish clear conventions for naming, file organization, and dependency management that make it easy for team members to understand and navigate the codebase. These conventions become more important as teams grow and new members join the project. Technical debt management requires ongoing attention and dedicated resources. Teams should regularly assess their codebase for opportunities to consolidate duplicated logic, improve performance, and simplify complex transformations. This work often doesn't have immediate business impact but is essential for maintaining long-term productivity and reliability. ## Measuring success and continuous improvement Collaborative analytics engineering teams need metrics and feedback mechanisms that help them understand their impact and identify opportunities for improvement. Traditional software development metrics like code coverage and deployment frequency provide useful insights, but analytics engineering teams also need metrics that capture data quality, business impact, and user satisfaction. Data quality metrics should track both technical correctness and business relevance. This includes traditional measures like data freshness and completeness, but also business-specific measures like metric consistency and stakeholder satisfaction. Teams should establish service level agreements for data quality and track their performance against these commitments. Development velocity metrics help teams understand their productivity and identify bottlenecks in their development process. This might include measures like time from development to production, frequency of deployments, and time to resolve data quality issues. These metrics should be used for continuous improvement rather than individual performance evaluation. Business impact measurement helps teams demonstrate their value and prioritize their work. This might include measures like the number of business users enabled by analytics engineering work, the reduction in time to insight, or the elimination of manual data processing work. These measures help justify continued investment in analytics engineering capabilities and guide strategic decisions about team growth and tooling investments. User feedback and satisfaction surveys provide qualitative insights that complement quantitative metrics. Regular feedback from analysts, business users, and other stakeholders helps teams understand how well their work is meeting business needs and identify opportunities for improvement. The future of collaborative analytics engineering will likely see continued evolution in tools, practices, and organizational approaches. AI-powered development tools are already beginning to impact how analytics engineers work, with 70% of professionals using AI for code development according to recent surveys. However, the fundamental challenges of collaboration, quality, and scale will remain central to successful analytics engineering programs. Organizations that invest in building strong collaborative analytics engineering capabilities will be better positioned to leverage their data for competitive advantage. This requires not just technical tools and practices, but also organizational commitment to the discipline and ongoing investment in team development and capability building. The most successful organizations will be those that treat analytics engineering as a core competency rather than a supporting function, building teams and practices that can scale with their data and business requirements. ## Collaborative Analytics Engineering FAQs **What is an analytics engineer? ** Analytics engineers work in the transformation layer of data pipelines, using tools like dbt to convert raw data into analysis-ready datasets. They bridge the gap between data engineers who manage ingestion and infrastructure, and analysts who consume the transformed data for insights and reporting. Unlike traditional analysts, analytics engineers must balance technical rigor with deep business context, encoding complex business logic into their transformations while maintaining code quality, performance, and reliability. **What technical skills do you need to know? ** Analytics engineers need proficiency in SQL and transformation tools like dbt, which has become central to collaborative workflows due to its built-in support for version control, testing, documentation, and modular development. They should understand cloud data warehouses like Snowflake, BigQuery, and Redshift, along with orchestration tools like Airflow or Prefect for automated scheduling. Additionally, they need skills in version control, automated testing frameworks that validate both code correctness and data quality, and data quality monitoring tools for maintaining reliability at scale. **How does data analytics engineering differ from data engineering and data science? ** Analytics engineering occupies a distinct middle layer between data engineering and data science. While data engineers focus on infrastructure, ingestion, and raw data management, analytics engineers work in the transformation layer, converting that raw data into business-ready datasets. Unlike data scientists who primarily consume data for analysis and modeling, analytics engineers build the foundational data models that enable downstream analytical work. They must understand business logic deeply enough to encode it correctly while maintaining the technical discipline of software engineering practices like version control, testing, and documentation. --- --- title: "Under the hood of Apache Iceberg" description: "Christian Thiel, the cofounder of Lakekeeper, walks through the state of the Iceberg ecosystem." url: "https://www.getdbt.com/blog/under-the-hood-of-apache-iceberg" date: "2025-08-24" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Under the hood of Apache Iceberg _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/under-the-hood-of-apache-iceberg). _ If you're a data practitioner, you likely understand Iceberg as a user, why it's important, and how it's changing the way that we build data systems. But you may not know a lot about what's going on beneath the surface. There are multiple ways to interface with Iceberg catalogs, multiple versions of the Iceberg REST spec. There's several leading catalogs that implement that spec. All this in an ecosystem that includes companies of all sizes, in proprietary and open-source code, and in academic and commercial contexts. In a few years, all this ambiguity will be behind us, but right now it's very much evolving in real-time. To get an update on the status of the Iceberg ecosystem and to walk through all the developments, Tristan talks with Christian Thiel. Christian is one of the lead architects of Lakekeeper, of one of the most widely used Iceberg catalogs. [**To learn more from some of the leaders in the Iceberg ecosystem, join us at Coalesce 2025 in Las Vegas, Oct. 13-16**](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/summary/?utm_medium=social&utm_source=substack&utm_campaign=q3-2026_coalesce-2025_aw&utm_content=coalesce____&utm_term=all_all__). _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways ## Walk us through your background **Christian Thiel:** I started in natural language processing, then moved into machine learning applications in manufacturing. Like many people, I realized that the biggest barrier wasn’t the algorithms but the data—its availability, quality, and accessibility. That led me deeper into data architecture and engineering, eventually to building Lakekeeper. ## What is Lakekeeper, and what are you building now? Lakekeeper is an Iceberg catalog implementation—a technical requirement for building distributed, composable analytic systems based on Apache Iceberg. But our vision goes beyond that. We see the future in data collaboration and reliable sharing of data, supported by clear contracts. ## For listeners new to Iceberg, what makes it so important? Iceberg allows organizations to store data once, in an open format, and then use the compute engine best suited for each workload. It’s a foundation for building modern, composable data platforms while avoiding vendor lock-in. If there’s one thing that should be open, it’s the data at the center of your platform. ## Some folks might say this sounds like Hadoop all over again—lots of open standards that are hard to integrate. Why is this time different? The ecosystem has matured. Even big vendors like Snowflake and Databricks are embracing Iceberg, which shows there’s a strong shift toward openness. Plus, the tooling and infrastructure are much easier to deploy today. A modern Iceberg setup is far less complex than a Hadoop environment used to be. ## Let’s talk about what’s happening under the hood. How does Iceberg work? Iceberg organizes data using a metadata hierarchy. At the top, there’s a JSON file that stores high-level table information: snapshots, schema, and locations. Below that are manifests and other layers that keep track of files. This hierarchy is what makes things like time travel, atomic transactions, and schema evolution possible. ## What about ongoing maintenance? There are two key tasks. First, expiring old snapshots so you don’t accumulate unnecessary files. Second, compaction—combining many small files into larger ones ## Catalogs are another critical piece. What role do they play? Catalogs manage the top layer of metadata and coordinate transactions. They make atomic updates possible, allow multiple writers, and handle governance—things like access control and multi-table transactions. ## How enterprise-ready is Iceberg today? Very ready. A year ago, there were still gaps, but today, performance and feature parity with native tables on platforms like Snowflake and BigQuery are strong. Governance and authorization models are still evolving, and different catalogs implement them differently, but the core functionality is there. ## Speaking of catalogs, how should someone pick between options like Lakekeeper, Polaris, Unity, AWS Glue, or Gravitino? Christian Thiel: It depends on priorities. Lakekeeper focuses on performance, extensibility, and ease of use. Polaris is developer-focused but less user-friendly. Unity is tightly integrated into Databricks. Glue now supports the Iceberg REST spec, which makes it more interoperable than before. Gravitino is another option aimed at enterprise-scale environments. ## Recently, DuckDB announced DuckLake. What’s your take on that? It’s interesting, but there are two concerns. First, it uses a database schema directly for the catalog, which creates interoperability issues—similar to the early JDBC catalog in Iceberg that the community eventually moved away from. Second, it was built without community involvement, and openness without adoption isn’t really openness. That said, for heavy DuckDB users, it could offer optimizations that make queries extremely fast, and if the broader ecosystem adopts it, it could become a viable open format. ## What’s next for Lakekeeper? We’re continuing to invest in table optimization, enterprise features, and data collaboration tools. Our vision is what we call the “unbreakable lakehouse,” where contracts and collaboration guardrails make shared data more reliable. Long-term, we see Lakekeeper as enabling truly collaborative, open data ecosystems. ## Chapters - **00:00 – Introduction**Tristan Handy introduces the episode and the focus on Apache Iceberg. - **01:40 – Christian Thiel’s background**From natural language processing to data engineering. - **04:30 – Introduction to Lakekeeper**What Lakekeeper is and its role in the Iceberg ecosystem. - **06:00 – Why Iceberg matters**How open table formats enable flexibility and reduce vendor lock-in. - **11:40 – How Iceberg works under the hood**Metadata hierarchy, catalogs, and how state is managed. - **21:30 – Maintenance and optimization**Snapshot expiration, compaction, and keeping tables performant. - **24:20 – Catalogs and governance**Access control, multi-table transactions, and security. - **31:40 – Enterprise readiness**How Iceberg is evolving for production use in large organizations. - **42:10 – Choosing the right catalog**Overview of Lakekeeper, Polaris, Unity, Glue, and Gravitt. - **47:20 – DuckLake discussion**Pros, cons, and ecosystem adoption challenges. - **52:00 – The future of Lakekeeper**Data contracts, collaboration, and building the “unbreakable lakehouse.” --- --- title: "Aligning data strategy with business objectives: challenges & solutions" description: "How to bridge the gap between data and business strategy—tackle silos, inconsistent metrics & scale with modern practices." url: "https://www.getdbt.com/blog/align-data-strategy-business-objectives" date: "2025-08-22" authors: ["Joey Gault"] categories: ["Pulse"] --- # Aligning data strategy with business objectives: challenges & solutions The traditional approach to data management, characterized by static, top-down governance processes, is increasingly inadequate for today's business environment. Organizations are dealing with exponentially growing data volumes from diverse sources, while simultaneously facing pressure to support AI initiatives that require high-quality, well-governed datasets. This shift has created new complexities in aligning data strategy with business objectives. The rise of [generative AI](https://www.databricks.com/discover/generative-ai) has particularly intensified these challenges. Unlike traditional analytics use cases, AI applications require data that meets stringent quality standards and can be traced through clear lineage paths. The probabilistic nature of large language models means that poor-quality input data can lead to unpredictable and potentially harmful outputs. This reality has forced many organizations to reconsider their approach to data governance and quality management. Furthermore, the regulatory environment surrounding AI is evolving rapidly, creating additional compliance requirements that data teams must navigate. Organizations need data strategies that are not only technically sound but also responsive to changing regulatory demands. This requires a more dynamic, [continuous approach to data governance](https://www.getdbt.com/blog/understanding-data-governance-ai) that can adapt quickly to new requirements while maintaining the reliability and consistency that business users depend on. ## Common challenges in aligning data strategy with business objectives ### Data silos and fragmented tooling One of the most persistent challenges facing data engineering leaders is the proliferation of data silos across the organization. Different teams often use disparate tools and systems for data storage, transformation, and analysis, leading to inconsistent approaches and duplicated efforts. This fragmentation makes it difficult to establish enterprise-wide standards and creates barriers to collaboration between teams. The problem is compounded when teams implement their own ad hoc solutions for data transformation and analysis. While these solutions may address immediate needs, they often lack the documentation, testing, and version control practices necessary for long-term maintainability. As organizations scale, these fragmented approaches become increasingly difficult to manage and can undermine trust in data across the enterprise. ### Inconsistent data quality and definitions Without standardized approaches to data transformation and quality management, organizations often struggle with inconsistent data definitions across different teams and use cases. The same business metric might be calculated differently by different teams, leading to conflicting reports and undermining confidence in data-driven decision making. This problem becomes particularly acute when supporting AI initiatives, where data quality issues can have far-reaching consequences. The challenge extends beyond technical consistency to include semantic consistency. Teams may use different naming conventions, apply different business rules, or make different assumptions about data relationships. These inconsistencies create confusion for business users and can lead to incorrect conclusions being drawn from data analysis. ### Lack of collaboration and visibility Traditional data management approaches often create barriers between data producers and consumers. Data engineering teams may work in isolation, creating datasets without sufficient input from business users about their actual needs. Conversely, business users may struggle to find and understand available datasets, leading them to request new data products that duplicate existing capabilities. This lack of collaboration is exacerbated by poor visibility into data lineage and dependencies. When business users can't easily understand where data comes from or how it's been transformed, they lose confidence in its reliability. Similarly, data engineers may struggle to understand the downstream impact of changes they make to data models or pipelines. ### Scaling challenges with manual processes Many organizations rely heavily on manual processes for data quality management, documentation, and deployment. While these approaches may work for small teams or limited use cases, they become increasingly unsustainable as data volumes grow and the number of use cases expands. Manual processes are also prone to human error and can create bottlenecks that slow down data delivery. The challenge is particularly acute when [supporting AI initiatives](https://www.getdbt.com/product/ai), which often require rapid iteration and experimentation. Manual deployment processes can significantly slow down the development cycle, making it difficult for organizations to keep pace with business demands or competitive pressures. ## Solutions for better alignment ### Establishing a unified data control plane The foundation for aligning data strategy with business objectives lies in establishing a unified approach to data transformation and management. By adopting a [single data control plane](https://www.getdbt.com/blog/data-control-plane-introduction), organizations can ensure consistency across teams while maintaining the flexibility needed to support diverse use cases. This approach enables all teams to work with the same tools and follow the same best practices, reducing fragmentation and improving collaboration. [dbt](https://www.getdbt.com/product/what-is-dbt) serves as an effective data control plane, providing a SQL-first transformation workflow that allows teams to collaborate on data models while maintaining software engineering best practices. By standardizing on dbt, organizations can ensure that all data transformations follow consistent patterns, include appropriate testing, and are properly documented. This consistency is crucial for[ building trust in data](https://www.getdbt.com/product/build-trust-in-data-and-data-teams) and enabling teams to build on each other's work. The unified approach also facilitates better governance and compliance. When all transformations follow the same patterns and use the same tools, it becomes much easier to implement enterprise-wide policies and ensure that all data products meet required standards. This is particularly important for organizations operating in regulated industries or those [implementing AI initiatives that require strict data governance](https://www.getdbt.com/blog/understanding-data-governance-ai). ### Implementing collaborative data development practices Modern data development should mirror software engineering practices, with emphasis on collaboration, peer review, and continuous integration. By implementing these practices, organizations can improve data quality while fostering better collaboration between data producers and consumers. This includes establishing clear processes for code review, testing, and deployment that ensure only high-quality transformations make it to production. dbt's built-in support for version control and collaborative development makes it easier for teams to work together on data models. Multiple team members can contribute to the same project, with changes tracked and reviewed before being deployed. This collaborative approach helps ensure that data models meet business requirements while maintaining technical quality standards. The collaborative approach extends beyond the data engineering team to include business stakeholders. By making data models more accessible and understandable, organizations can involve business users in the development process, ensuring that data products actually meet their needs. This involvement is crucial for ensuring that data strategy remains aligned with business objectives. ### Enabling self-service data access with governance One of the most effective ways to align data strategy with business objectives is to enable self-service data access while maintaining appropriate governance controls. This approach allows business users to find and use data independently, reducing the burden on data engineering teams while ensuring that data usage follows established policies and standards. [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) provides a centralized location where users can discover available datasets, understand their lineage, and access documentation. This self-service capability reduces the time spent on data discovery and helps ensure that users are working with the most appropriate datasets for their needs. The catalog also provides visibility into data usage patterns, helping data teams understand which datasets are most valuable to the business. Self-service access must be balanced with appropriate governance controls. Role-based access control ensures that sensitive data is only accessible to authorized users, while automated testing and monitoring help maintain data quality. This balance between accessibility and control is essential for enabling business users while maintaining the trust and reliability that enterprise data systems require. ### Standardizing metrics and business logic Inconsistent metric definitions are a major source of confusion and mistrust in data-driven organizations. By centralizing the definition of key business metrics, organizations can ensure that everyone is working from the same understanding of important business concepts. This standardization is crucial for maintaining alignment between data strategy and business objectives. [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) provides a framework for defining metrics using standardized formulas and naming conventions. This approach eliminates the confusion that arises when different teams implement their own versions of the same metric, and it ensures that business reports are consistent regardless of which tool or team produces them. The semantic layer also makes it easier to maintain metrics over time, as changes can be made in a single location and automatically propagated to all downstream uses. Standardized metrics also facilitate better decision-making by ensuring that business leaders are working with consistent, reliable information. When everyone uses the same definitions and calculations, it becomes much easier to have productive discussions about business performance and make data-driven decisions with confidence. ### Implementing continuous integration and automated testing Reliable data products require robust testing and deployment processes. By implementing continuous integration practices, organizations can ensure that data transformations are thoroughly tested before being deployed to production. This approach reduces the risk of data quality issues while enabling faster, more reliable delivery of data products. [dbt's built-in testing framework](https://docs.getdbt.com/docs/build/data-tests) makes it easy to implement comprehensive test suites that verify data quality at multiple levels. Tests can check for data completeness, accuracy, and consistency, helping to catch issues before they impact business users. Automated testing also provides confidence when making changes to existing data models, as teams can quickly verify that their changes don't introduce new problems. The continuous integration process should include peer review of all changes, ensuring that multiple team members examine new code before it's deployed. This review process helps maintain code quality while also facilitating knowledge sharing across the team. When combined with automated testing, peer review provides a robust quality assurance process that helps maintain trust in data products. ### Building data products with clear ownership and SLAs Treating datasets as products with clear ownership, documentation, and service level agreements helps ensure that data strategy remains aligned with business needs. Data products should be designed to solve specific business problems, with clear success metrics and ongoing maintenance responsibilities. The [data product approach](https://www.getdbt.com/blog/guide-to-ai-data-products) encourages data teams to think more strategically about their work, focusing on business value rather than just technical implementation. By defining clear success metrics and SLAs, data teams can better understand whether their work is meeting business needs and make adjustments as necessary. This business-focused approach helps ensure that data strategy remains aligned with organizational objectives. Data products also facilitate better collaboration between data teams and business stakeholders. When datasets are treated as products, it becomes natural to involve business users in the design process and to gather feedback on whether the products are meeting their needs. This ongoing dialogue helps ensure that data strategy evolves in response to changing business requirements. ## Measuring success and maintaining alignment Successful alignment between data strategy and business objectives requires ongoing measurement and adjustment. Organizations should establish clear metrics for data quality, user satisfaction, and business impact, and regularly assess whether their data strategy is delivering the expected value. This measurement should include both technical metrics (such as data quality scores and system reliability) and business metrics (such as user adoption and decision-making speed). Regular feedback loops with business stakeholders are essential for maintaining alignment over time. Data teams should actively seek input from business users about whether data products are meeting their needs and what additional capabilities would be valuable. This feedback should inform prioritization decisions and help guide the evolution of data strategy. The measurement approach should also include assessment of team productivity and collaboration effectiveness. Are teams able to work together effectively? Are data products being delivered on time and meeting quality standards? These operational metrics are important indicators of whether the data strategy is sustainable and scalable. ## Conclusion Aligning data strategy with business objectives requires a comprehensive approach that addresses both technical and organizational challenges. By establishing unified tooling and processes, implementing collaborative development practices, and focusing on business value, data engineering leaders can build data systems that truly serve their organization's needs. The key is to move beyond purely technical solutions and embrace approaches that facilitate collaboration, maintain quality, and remain responsive to changing business requirements. Tools like dbt provide the technical foundation for this alignment, but success ultimately depends on implementing the right processes and maintaining a focus on business value. As the data landscape continues to evolve, particularly with the growth of AI initiatives, the importance of this alignment will only increase. Organizations that successfully align their data strategy with business objectives will be better positioned to capitalize on new opportunities and maintain competitive advantage in an increasingly data-driven world. ## Data strategy FAQs **What is a data strategy?** A data strategy is a comprehensive approach to managing and utilizing an organization's data assets to support business objectives. It encompasses the tools, processes, governance frameworks, and organizational practices needed to ensure data quality, accessibility, and alignment with business goals. An effective data strategy moves beyond purely technical solutions to include collaborative development practices, standardized metrics, and clear data product ownership that enables organizations to make reliable, data-driven decisions. **What are the key components of an effective data strategy?** The key components include establishing a unified data control plane for consistent transformation and management across teams, implementing collaborative data development practices with proper testing and peer review, enabling self-service data access with appropriate governance controls, standardizing metrics and business logic to eliminate confusion, implementing continuous integration and automated testing for reliability, and building data products with clear ownership and service level agreements. These components work together to ensure data quality, foster collaboration between teams, and maintain alignment with business objectives. **How do you measure the effectiveness of your data strategy implementation?** Measuring effectiveness requires tracking both technical and business metrics through ongoing assessment and feedback loops. Technical metrics include data quality scores, system reliability, and team productivity indicators, while business metrics focus on user adoption, decision-making speed, and business impact. Regular feedback from business stakeholders is essential to understand whether data products meet their needs and what additional capabilities would be valuable. Organizations should also assess operational metrics like collaboration effectiveness, on-time delivery of data products, and whether quality standards are being met consistently. --- --- title: "Enterprise data governance strategy: The essential elements" description: "See how dbt helps enterprises scale governance across teams and systems—built for AI, quality, and collaboration." url: "https://www.getdbt.com/blog/enterprise-data-governance-strategy-elements" date: "2025-08-21" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Enterprise data governance strategy: The essential elements Enterprise data governance isn’t optional - it’s essential. A solid enterprise data governance strategy leads to crafting high-quality, reliable, and secure data with minimal errors. This builds confidence in data and leads to better decision-making. The problem is crafting a data governance strategy that can keep pace with the speed of your business. The emergence of [Generative AI (GenAI)](https://cloud.google.com/ai/generative-ai?hl=en) use cases, in particular, has created a skyrocketing need for high-quality data drawn from across the enterprise. Keeping up with the pace of change requires more than “data governance as usual.” It requires adapting new strategies, paired with good tools, that enable all teams in an enterprise to improve and monitor data quality, collaborate on data across teams, and automate data pipelines to reduce downtime and minimize human error. In this article, we’ll look at how enterprise data governance is changing in the age of AI and the strategies you need to implement to meet the moment. We’ll also see how dbt supports turning these strategies into reality by enabling teams to collaborate on data governance easily and effectively. ## How enterprise data governance strategy is changing Traditionally, enterprise data governance has been a largely static, top-down process. A central authority laid out standards and policies for all teams. Data stewards worked with teams to implement and enforce policies on a local level. This was largely a slow, established, and manual approach to data governance. And it worked…for a while. These traditional approaches to enterprise data governance, which were already strained, don’t scale in the age of AI. There’s too much data, from too many different teams, for a manual approach to governance. In addition, AI technology - and the regulatory environment that governs it - is changing daily. This means AI requires an approach to enterprise data governance that is: - Dynamic, continuous, and multidisciplinary - Automated and dynamic - Proactive and responsive to a fast-changing regulatory environment ### The additional governance concerns raised by AI Additionally, AI raises enterprise data governance concerns that either didn’t previously exist or existed on a smaller scale. [Large language models (LLMs)](https://www.cloudflare.com/learning/ai/what-is-large-language-model/), which serve as the heart of most AI applications, are trained on vast amounts of data to generate a probabilistic model. The responses are stochastic - i.e., not directly correlated to the inputs. This makes LLM output hard to predict. It also raises concerns around the LLMs themselves, which operate as black boxes. This leads to issues such as: - Bias in the underlying data, leading to biased responses that harm certain individuals or groups; - A lack of transparency or auditability in how a given LLM arrived at a certain output; and - An inability to explain or control why the model made the decisions it did. In addition, [AI is susceptible to a number of unique threats](https://www.sentinelone.com/cybersecurity-101/data-and-ai/ai-security-risks/). These include data poisoning, prompt injection, model inversion, and private leakage attacks, among others. ### The essential elements of a modern data governance strategy The good news is that these challenges aren’t insurmountable. Tackling them, however, requires that your enterprise approach managing data differently. In most enterprises, data is siloed and fragmented. Data engineering teams all use different data storage systems and different tools to transform, move, and manage data. The result is inconsistency, little cross-team collaboration, and a lack of overall trust in data. dbt acts as a [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) for data across your enterprise. Using dbt, your enterprise can manage data in a flexible and cross-platform way that avoids vendor lock-in, encourages collaboration, and produces trustworthy outputs. With dbt, you can implement a flexible and proactive data governance strategy centered around three core principles: - Enable all teams to create and publish high-quality datasets - Emphasize data collaboration - Define a continuous release process centered on quality Let’s look at each one of these in detail and at how dbt helps you turn each of these strategies into business reality. ### Enable all teams to create and publish high-quality datasets In one way, AI has changed the game. But in a fundamental sense, nothing’s changed. Producing high-quality data consistently is still the name of the game. The key is “consistently.” In most companies, data quality varies considerably from team to team. Even enterprises with data quality standards in place may struggle to determine which teams are adhering to them. Adopting a single data control plane like dbt means that every team across the enterprise uses a single toolset to [model](https://docs.getdbt.com/docs/build/models) and transform data. dbt supports a built-in [testing framework](https://docs.getdbt.com/docs/build/data-tests), enabling data engineers to build out test suites that verify all changes before release. This ensures that data transformations generate the correct outputs before data is made available to downstream consumers. In addition, the new [dbt Fusion engine](https://www.getdbt.com/product/fusion) accelerates the development of high-quality data transformation code. dbt Fusion’s deep SQL comprehension emulates the SQL syntax of all popular data warehouses. That means that data engineers can see errors in their code as they type, and run tests locally on their development machines - no code check-in or remote data warehouse required. **** ### Emphasize data collaboration Do-it-yourself approaches to data quality make it hard, if not impossible, for teams to work together or share their results. There are two key reasons for this. First, everyone is using different tools for data transformation. One team might use dbt for their data pipelines, while another is running Python scripts, and a third is editing stored procedures directly in a data warehouse. This makes it impossible for teams to work together on common problems or share data transformation code. Second, there isn’t a single place to find data once it’s ready. Without an easily accessible single source of truth, data consumers may spend hours, days, or weeks hunting across various [Snowflake](https://snowflake.com) instances for the data they need. Using dbt as your data control plane, anyone who knows SQL can understand and contribute to data models. Data engineering teams can share common transformation code across projects, reducing the need to start from scratch with each new dataset. Engineering teams can further increase data governance for AI by packaging and deploying their datasets as [data products](https://www.getdbt.com/blog/data-product-examples). Data products are structured and curated assets designed to solve a specific business problem. They’re a great fit for AI workloads, as they make it easier to verify the origin and quality of data that comprises a given AI solution. Once ready for production, data models are published and discoverable via [dbt Catalog](https://www.getdbt.com/product/dbt-catalog). Anyone with the appropriate permissions can find the data, read its accompanying documentation, see its [lineage](https://www.getdbt.com/blog/what-is-data-lineage), and leverage it for their own work - whether that’s a quarterly report, a machine learning model, or an AI application. You can even use dbt to centralize the definition of critical business metrics. With [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl), you can define metrics - such as quarterly revenue - using a single formula with standardized naming. This eliminates data quality issues caused by teams implementing their own definitions based on potentially unreliable sources. ### Define a continuous release process centered on quality Writing data transformation code is only part of the battle. The hard part is ensuring that only high-quality code makes its way to production. Error-riddled code results in data errors and privacy leaks that can undermine both decision-making and customer trust. Using dbt, you can easily create [a continuous integration (CI) release process](https://docs.getdbt.com/docs/deploy/continuous-integration) to deploy data transformation code changes. After an engineer checks in a change, it goes to another team member for peer review. This ensures no change goes live without at least two sets of eyes on it. Once deployed, the CI process can automatically run all data tests against a non-production database, verifying that the new code produces correct outputs before it touches production data. Once in production, these tests can be continuously run against incoming data and [monitored via data health signals in dbt](https://docs.getdbt.com/docs/explore/data-health-signals). When data is distributed across multiple systems, it can be difficult to control who has access to production systems. By defining a CI process in dbt, you can use [role-based access control](https://docs.getdbt.com/docs/mesh/govern/model-access) to define who’s authorized to make changes to specific models. ## dbt as your data control plane for better enterprise data governance Data sprawl, data siloes, and inconsistent tooling make a comprehensive approach to enterprise data governance impossible. Using dbt as your data control plane, you can provide a consistent, governed approach to managing your enterprise data, no matter where it lives. To learn more about how dbt can improve your enterprise data governance, [schedule a demo today](https://www.getdbt.com/contact). --- --- title: "Build reliable AI agents with the dbt MCP Server" description: "Learn how to build reliable AI agents with dbt MCP Server by connecting them to governed context from your dbt projects." url: "https://www.getdbt.com/blog/build-reliable-ai-agents-with-the-dbt-mcp-server" date: "2025-08-21" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Build reliable AI agents with the dbt MCP Server 61% of attendees at a recent Gartner conference stated they’d made investments in agentic AI. However, Gartner itself predicts that over [40% of those projects will be canceled](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027) by 2027. They cite escalating costs, unclear business value, inadequate risk controls, and agents’ lack of maturity and true autonomous agency as the primary reasons for failure. In addition, we have observed that, as AI agents become more capable, their ability to interact with data pipelines and analytics workflows is hitting a wall. They lack context. Without knowing which models are trusted, how metrics are defined, or what transformations were applied, agents risk misleading or incomplete outputs. You can take steps to ensure that your agentic AI doesn’t end up in the dustbin of abandoned AI projects. This article explains how you can build reliable AI agents with structured context from your dbt projects, using the new dbt MCP Server to expose your trusted data to AI. ## The importance of structured metadata and context for agentic AI By now, it’s common knowledge that [large language models (LLMs) can hallucinate](https://www.redhat.com/en/blog/when-llms-day-dream-hallucinations-how-prevent-them). They can make mistakes, fabricate data, and even introduce security risks. That’s because LLMs are ultimately probability models. They generate responses based on patterns in their training data, predicting what sounds right, not necessarily what is correct. They aren’t grounded in your business logic and don’t understand your data’s specific context. This is where structured context becomes essential. **Structured context is the governance layer** that defines how your data is modeled, tested, and interpreted. It includes [lineage](https://www.getdbt.com/blog/what-is-data-lineage) (how datasets relate to each other), versioning, testing frameworks, metric definitions, SQL logic, personal identification tagging, and other reusable components. Together, this context tells an AI system not just what data exists, but how it’s connected, what it means, and how it should be used. It’s the difference between an AI that’s guessing and one that delivers accurate, trustworthy answers. By providing your AI agents with structured metadata and context, you're giving them a reliable playbook. Context unlocks five key components that transform AI agents from predictive guessers into trusted collaborators: - **Understanding your data and business logic. **With access to your models, joins, and metrics, agentic outputs are accurate and aligned to how your business functions (and there’s no more guessing what “revenue” or “customer ID” means). - **Keeping multi-step workflows on track. **By understanding how data is connected, agents can follow the correct path without causing downstream issues. - **Improving memory. **When agents learn from clean, governed examples, they build on experience instead of fabricating responses. - **Reusing verified SQL logic.** Instead of starting from scratch, agents can draw from version-controlled models, macros, and definitions that are already built into [dbt](https://docs.getdbt.com/). - **Providing an end-to-end audit trail.** Using lineage and [version control](https://www.getdbt.com/blog/building-a-mature-analytics-workflow) enables you to trace how decisions were made, validate results, and resolve issues quickly—whether it's a broken SQL query, a failed test, or a number that suddenly appears incorrect. Your dbt projects already provide the structured foundation AI agents require, including models, relationships, and metrics defined in version-controlled, documented, and tested code. With the launch of [dbt MCP Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server), we can now make that structure accessible to AI in a way that’s usable, safe, and available in real time, laying the groundwork for trustworthy agentic AI workflows. ## dbt MCP Server: Enabling a playbook for AI agents The **dbt MCP Server** is an open-source implementation of the [Model Context Protocol (MCP)](https://www.anthropic.com/news/model-context-protocol) that enables AI systems to dynamically access structured data. MCP Server exposes everything already built into your dbt project—structured metadata, the dbt compiler, the MetricFlow query engine, and compute integrations—and serves it to AI agents using a standard, open protocol. This enables agents to: - Query the semantic layer. - Execute transformations at scale. - Generate SQL or documentation based on existing, tested logic. - Reuse logic safely across teams without duplication or drift. - Maintain governance guardrails, such as freshness, ownership, and version history. - Collaborate and interoperate reliably. Because agents answer business questions using the metrics and definitions from the [dbt Semantic Layer](https://docs.getdbt.com/docs/build/semantic-layer/overview), the answers they return are correct and reliable. By centralizing logic in one layer, you decrease compute, reduce redundant queries, and achieve faster results. And, as your AI stack evolves with new models, tools, and interfaces, your dbt-structured foundation remains consistent. The dbt MCP server is the key to unlocking the value of your structured data with AI applications. It delivers: - Fewer hallucinations - Faster development - Stronger governance and security - Trusted outputs that scale with your data Currently, the MCP Server enables [three key use cases](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) that bridge the gap between your dbt project and realizing the value of your structured data through agentic AI: - **Data discovery: **Understanding what data assets exist in your dbt project and how they relate to each other. - **Data querying: **Accessing trusted metrics and running queries against your data models. Uses the dbt Layer as a single source of truth for metrics reporting, enabling you to execute SQL queries for data exploration and development. - **Project execution:** Running dbt commands and managing your project through conversational interactions with AI agents. In a recent webinar, dbt and Indicium demonstrated how the dbt MCP Server enables AI agents to [execute a data migration](https://www.getdbt.com/resources/webinars/build-reliable-ai-agents-with-the-dbt-mcp-server), converting legacy systems, such as PySpark notebooks, ETL scripts, and other artifacts, into a modern dbt project. ## MCP Server in action: dbt migration Migrating legacy systems to dbt is often a complex, manual process that demands extensive engineering hours, including multiple interviews and handoffs. With the dbt MCP Server, AI agents can take the lead, streamlining execution from mapping to validation. AI agents need context, which means having a solid understanding of the target architecture and dbt project structure before engaging AI. With that foundation in place, the dbt MCP Server can power an agentic AI workflow that automates the migration process. ### Data mapping and assessment The migration workflow begins with mapping the legacy system and translating it into a knowledge graph. Knowledge graphs are powerful because they capture relationships and legacy data lineage. This enriched context is especially valuable for LLMs and AI agents because it’s expressed in language they can understand and use to drive the migration. This stage uses an **Assessment Agent** to: - Profile the legacy system’s schemas, code, and dependencies. - Translate those artifacts into a knowledge graph that captures tables, columns, joins, and data-flow relationships. ### Planning and project breakdown Once your data has been mapped, so the AI agent has the necessary context, you can create a **Planning** **Agent** to break the migration into manageable tasks. The Planning Agent: - Breaks the migration into waves based on complexity and dependencies. - Generates specific tasks such as creating staging models, intermediate models, or merging tables to handle differences in business logic. These tasks are stored in memory for execution. ### Execution In the final stage, an **Executor Agent** takes over to execute the migration tasks. Specific Execute tasks include: - **Analyze**: The Executor Agent ingests file schemas via MCP and writes a memory file that lists all migration subtasks. - **Implement**: The Executor Agent auto-generates source files and staging models in SQL, following dbt best practices, then runs dbt compile and dbt run through MCP to catch errors early. - **Documentation and testing**: The Executor Agent produces doc blocks and schema tests for each new model, then verifies test results via dbt test. - **Audit**: The Executor Agent uses the [dbt audit-helper package](https://github.com/dbt-labs/dbt-audit-helper) to build models that compare legacy tables to new tables, both at the table and column levels, and aggregate match rates over time. - **Final build: **The Executor Agent calls dbt build and dbt test one more time, then surfaces a 100 % match as the migration’s definition of done. The final result is a production-ready target data system for both business and technical users. Thanks to the context provided by the knowledge graph and MCP, documentation is no longer an afterthought. It’s generated early, with column descriptions and model metadata captured automatically. ## Real-world impact: Migration use case Recently, dbt Partner [Indicium](https://indicium.ai/) helped [Aura Minerals](https://indicium.ai/knowledge-hub/blog/ai-in-data-migration-aura-minerals/) harness the power of the MCP Server. Aura wanted to adopt a future-ready framework, powered by dbt, governance, and automation. Indicium created a proprietary AI Migration Agent that followed the above workflow to map, convert, validate, and migrate the company’s PySpark estate to dbt. The results were stunning: - **Scale**: 400+ PySpark notebooks—spanning bronze, silver, and gold layers—and 130 complex workflows migrated to dbt. - **Speed**: 87% improvement in pipeline—from 45 hours to six hours. - **Quality**: ~99% code conformity to organizational best practices. - **Collaboration Overhead**: 66% reduction in team dependency by shifting context into a knowledge graph, minimizing stakeholder interviews and back-and-forth, and enabling agents to handle translation. By combining context-grounded AI with dbt MCP, Aura now has a governed dbt environment with models that both business and data teams can trust and understand. ## Building reliable AI agents with dbt and dbt MCP Server Agentic AI workflows promise speed and automation, but without structured context, they deliver hallucinations and best guesses. The MCP Server is the bridge between your dbt project and any MCP-enabled client. Together, dbt and the dbt MCP Server help organizations build reliable AI agents by: ### Providing metrics as a single source of truth The dbt [Semantic Layer](https://docs.getdbt.com/docs/build/semantic-layer/overview) provides version-controlled models, macros, documented code, and metrics. When AI agents query business metrics directly in the Semantic Layer, they ground their outputs in real business context, ensuring consistent logic and reducing the likelihood of hallucinations. ### Exposing rich project metadata for discovery The MCP Server exposes knowledge about your data assets to LLMs and AI agents, enabling powerful discovery capabilities. Agents can automatically discover and understand the available data models, their relationships, and their structures without human intervention. This allows agents to navigate complex data environments and produce accurate insights autonomously. ### Accelerating discovery and reuse dbt’s modular project structure is built on consistent naming conventions and layered model organization. dbt [lineage](https://www.getdbt.com/blog/what-is-data-lineage) lets you visualize dependencies as data flows through models, sources, and transformations. When you expose this rich context via the MCP Server, AI agents can automatically trace relationships, query metadata, and execute dbt commands, accelerating discovery, reuse, and development workflows without needing direct access to production systems. ### Strengthening governance and trust dbt helps you build a robust [governance framework](https://www.getdbt.com/lp/data-governance) with software best practices like version control, testing, CI/CD, and auto-documentation. Key metadata, such as freshness, ownership, and sensitivity, provide clear guardrails to help AI agents make decisions based on your governance policies. The [_audit_helper_ package](https://github.com/dbt-labs/dbt-audit-helper) adds another layer of control by enabling easy comparisons between old and new models, helping you validate changes and prevent unintended impacts. ### Standardizing AI-data integration via open protocol. MCP replaces fragmented AI integrations with a [universal standard](https://www.anthropic.com/news/model-context-protocol), enabling AI agents to connect to dbt projects with reliability and scale. Using the MCP Server helps you accelerate future AI initiatives by allowing the creation and reuse of agents that operate on trusted, well-documented data models. ## Take the next steps with the dbt MCP Server The dbt MCP Server will fundamentally change how AI interacts with your data. By providing the [missing glue](https://www.getdbt.com/blog/mcp) between your dbt projects and AI agents, MCP Server ensures that your AI will deeply understand your data so you can trust the outcome of your AI projects. The dbt MCP Server promises safe, reliable access to your structured data. Its built-in security features and access controls help ensure that AI agents operate within trusted guardrails. When dbt becomes the [central control plane](https://www.getdbt.com/resources/whitepaper-the-control-plane-for-data-collaboration-at-scale) for your data, it provides a solid foundation for AI-driven insights, grounding every output in organizational truth. The MCP Server is now available on GitHub for prototyping AI agents that will benefit from a deep understanding of how your data is structured and utilized. To start building your reliable AI agents with the MCP Server, watch our webinar to [view a demo](https://www.getdbt.com/resources/webinars/build-reliable-ai-agents-with-the-dbt-mcp-server) of the MCP Server agentic AI workflow described above, and [talk to a dbt expert](https://www.getdbt.com/contact) to explore MCP pilot programs. --- --- title: "dbt Labs Launches Reimagined Global Partner Ecosystem Program to Accelerate Strategic Growth" description: "New partner tiers, go-to-market alignment, and enablement investments accelerate momentum across the partner ecosystem" url: "https://www.getdbt.com/blog/dbt-labs-launches-reimagined-global-partner-ecosystem-program" date: "2025-08-20" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Launches Reimagined Global Partner Ecosystem Program to Accelerate Strategic Growth **PHILADELPHIA, August 20, 2025** – dbt Labs, the leader in standards for AI-ready structured data, announced the launch of its newly architected global partner ecosystem program. It introduces structured partner tiers, unified go-to-market models, and increased investment in enablement, laying the foundation for predictable, scalable growth across the dbt Labs partner ecosystem. This strategic initiative is purpose-built to more quickly deliver impact to enterprise customers through expanded partner capability and collaboration. The new partner program launches as dbt Labs and its partner ecosystem experience significant momentum. After surpassing the [$100M ARR](https://www.getdbt.com/blog/dbt-labs-100m-arr-milestone) mark earlier this year, the company’s growth is now coupled with increased demand for frictionless access to dbt across cloud providers; Marketplace transactions have grown more than 190% year-over-year. In addition, partner portal applications have recently increased more than 300% over historical weekly averages. This energy around joining the program, sparked by [the recent dbt Fusion engine launch](https://www.getdbt.com/blog/dbt-labs-redefines-dbt-with-new-fusion-engine) and product opportunities, underscores the accelerating interest in dbt. dbt Labs’ reimagined program is built to support and foster this momentum while demonstrating the strategic importance of the dbt Labs partner ecosystem to customers. “This program is architected for long-term scale and global consistency, which are critical to delivering the best possible outcomes for joint customers,” said Shawn Toldo, Vice President, Worldwide Partner Ecosystem. “By delivering a clear, consistent framework for engagement, we are empowering partners to invest more confidently, collaborate more effectively, and deliver greater impact to our joint customers,” added Raymond Wong, Sr. Director of Partner Programs and Strategy. dbt Labs’ new partner program delivers a consistent, global structure that enables partners to engage more effectively and accelerate go-to-market success. The program introduces three formal tiers – Visionary, Advanced, and Registered – applicable to Consulting & Services and Technology partners, while aligning investments across all four strategic partner types: Data & AI Platform, Consulting & Services, Technology, and Reseller. The program is anchored around three key pillars: - **Global consistency and clarity **– A unified framework standardizes partner engagement and benefits across regions, with structured tiers for Consulting & Services and Technology partners to create predictability and clarity on how partners engage with dbt Labs. - **Increased investment** – Visionary and Advanced partners gain access to expanded co-marketing and co-sell funding, along with structured go-to-market collaboration opportunities. - **Customer-centricity** – Enhanced resources including streamlined onboarding, opportunity registration tools, and expanded technical enablement help partners reduce time-to-value and deliver outcomes faster for joint customers. The new partner program is [now live](https://www.getdbt.com/partners), with transition plans and onboarding resources available for existing and new partners. dbt Labs continues to prioritize investment in the broader partner ecosystem and its strategic relationships, illustrated by recent announcements alongside AWS and Snowflake. **AWS Collaboration Continues to Accelerate** As part of its evolving partner strategy, dbt Labs has signed a [strategic collaboration agreement](https://www.getdbt.com/blog/dbt-labs-expands-signs-strategic-collaboration-agreement-with-aws) (SCA) with Amazon Web Services (AWS) to deepen product integrations and accelerate joint go-to-market efforts. This agreement expands dbt’s reach across AWS services including Amazon Redshift, Amazon Athena, and Amazon SageMaker Lakehouse, while expanding [AWS Marketplace accessibility](https://aws.amazon.com/marketplace/pp/prodview-tjpcf42nbnhko) for enterprise customers. **Snowflake Embeds the dbt Fusion Engine** At [Snowflake Summit 2025](https://www.snowflake.com/en/summit/), Snowflake, the AI Data Cloud company, introduced [Snowflake dbt Projects](https://www.snowflake.com/en/news/press-releases/snowflake-openflow-unlocks-full-data-interoperability-accelerating-data-movement-for-ai-innovation/), soon to be powered by the dbt Fusion engine **– **dbt’s next-generation, Rust-powered transformation engine. This integration will allow users to run dbt-native workflows directly within Snowflake, deepening technical alignment between two of the most widely used technologies in the modern data stack. “We are looking forward to offering our joint customers the power of Fusion embedded in Snowflake for a powerful way to transform data natively,” said Chris Child, VP of Product, Data Engineering at Snowflake. “We are excited about strengthening our relationship with dbt through this work to accelerate data-driven decision making for the enterprise.” Ryan Segar, dbt Labs Chief Product Officer, adds: “Snowflake’s decision to offer the dbt Fusion engine in its platform reflects what we’re hearing across the ecosystem. Fusion represents the next chapter of performant, cost-effective data transformation, and we’re thrilled to see Snowflake leading the charge.” This announcement builds on a longstanding collaboration between the two companies. Snowflake recently named dbt Labs its [2025 Monetization Data Cloud Product Partner of the Year](https://www.getdbt.com/blog/dbt-labs-named-snowflake-monetization-data-cloud-product-partner-of-the-year), marking dbt Labs’ third consecutive annual Snowflake partner award. For more information about the dbt Labs global partner program, visit [https://www.getdbt.com/partners](https://www.getdbt.com/partners). **About dbt Labs** Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 60,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at [getdbt.com](http://getdbt.com), and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "Fusion and the dbt VS Code extension are now in Preview for local development" description: "Get lightning-fast, next-gen development for your local environment" url: "https://www.getdbt.com/blog/fusion-and-dbt-vs-code-extension-preview-launch" date: "2025-08-19" authors: ["Elias DeFaria", "Alexis Jones"] categories: ["Product"] --- # Fusion and the dbt VS Code extension are now in Preview for local development Big news—the [dbt Fusion engine](https://www.getdbt.com/product/fusion) is now in Preview for local development on Snowflake, Databricks, BigQuery, and Redshift, available in both the dbt VS Code extension and the CLI. Fusion marks a step-function improvement in what it means to build on dbt, with 30x faster parse times than dbt Core, local query validation straight from the command line, and a new compile cache to drastically speed up the developer loop (we heard you!). And if you’re tapping into Fusion from the [dbt VS Code extension](https://docs.getdbt.com/guides/fusion) (also in Preview today), you’ll enjoy sleek development features like rapid, in-line SQL validation, code autocompletion, automated refactoring, CTE preview, and more. Developing dbt code in Fusion is designed to be dramatically faster, more reliable, and more delightful than ever before. **** _Note that a Preview release of Fusion for the cloud-hosted [**dbt platform**](https://www.getdbt.com/product/dbt) will follow soon ([currently in beta](https://docs.getdbt.com/docs/fusion/supported-features))._ ## So, what’s new in Preview? Since the Fusion [beta launch in late May](https://www.getdbt.com/blog/get-to-know-the-new-dbt-fusion-engine-and-vs-code-extension), our teams have been shipping daily improvements to both the [engine](https://github.com/dbt-labs/dbt-fusion) and the [dbt VS Code extension](https://docs.getdbt.com/docs/about-dbt-extension). We’ve added support for the four primary adapters (Snowflake, Databricks, BigQuery, and Redshift), we’ve incorporated feedback from early adopters in an intense beta program, and have already seen hundreds of teams move to adopt Fusion. A snapshot of the progress made since May: - 450+ weekly average projects running on Fusion - 5,500+ weekly average users of the dbt VS Code extension - Launched support for the four most used adapters in the dbt ecosystem: Snowflake, Databricks, BigQuery, and Redshift - 3-5 new features shipped every week (including exposures, the clone command, `--empty` and `--sample`, Cloud CLI credentials, onboarding improvements, and more) - Revamped logging to make it easier to debug and audit runs - Added first class support for Windows operating systems, both for the Fusion binary and the dbt VS Code extension - 700+ bugs addressed and stability improvements shipped - Rolled out a new language server (LSP) cache to dramatically improve subsequent compile times in the VS Code extension. This is particularly relevant for large scale dbt projects with significant introspection query usage - Published [8 editions of the Fusion Diaries](https://github.com/dbt-labs/dbt-fusion/discussions/289), the weekly changelog on our path towards Fusion general availability **** We know that [replatforming an entire engine](https://www.getdbt.com/blog/where-we-re-headed-with-the-dbt-fusion-engine) that serves a passionate community of 60,000+ teams is…both _hard work_ and a _process_! The team has been putting in the very hard work to deliver the results in the stats above, and today’s Preview release is testament to that. As for the process, this release moves us one step closer on our [path to GA](https://docs.getdbt.com/blog/dbt-fusion-engine-path-to-ga), making Fusion compatible with more warehouses, projects, and dbt features than ever before. > > > — Sonja Strempel, DPG Media But we know we’re not done yet. While Fusion doesn’t support every existing dbt project or feature yet, each week we get incrementally better. In the meantime, [auto-fix scripts](https://github.com/dbt-labs/dbt-autofix) make it easier to resolve deprecation warnings and to get prepared to run on Fusion. If you use dbt Core and haven’t checked out Fusion since May, we’re confident that now is a great time to give the new engine a whirl. In the next sections we’ll walk through our approach to conformance with dbt Core and which projects are a great fit for Fusion today. ## Current progress on conformance with dbt Core As we ship improvements to Fusion, we've been measuring a metric we call “Fusion conformance.” Fusion conformance aims to prove Fusion will perform exactly the same as dbt Core for a given project. Conformance breaks down into several stages: first parse, then compile, then build (like `dbt parse`, `dbt compile`, and `dbt build`). - Parse conformance measures the projects with valid dbt code that Fusion can successfully understand. - Compile conformance measures whether Fusion can understand your dbt project, comprehend its underlying SQL, and is ready to run your models against the warehouse and power all the new features in the VS Code extension’s LSP. - Build conformance tests everything, ensuring that Fusion runs, tests, seeds, compiles, and parses models in the same way dbt Core does. Our data now shows that after running our [auto-fix script](https://github.com/dbt-labs/dbt-autofix), a sufficient percentage of user projects are passing compile conformance tests. Along with an active community of users using Fusion on their projects for development and in production, this metric gives us confidence to designate the engine as Preview today. ## The best candidates for Fusion today To gauge whether your project will easily run on Fusion, there are two key factors to keep in mind: project complexity and blocking features. We continue to tackle regular improvements to both. - Model count is the best proxy for project complexity. Through testing, we’ve found that projects with fewer than 500 models typically perform well with Fusion. As project size grows, performance may vary, and projects with more than 3,000 models are likely to experience the most challenges today, particularly if they rely heavily on [introspection queries](https://docs.getdbt.com/faqs/Warehouse/db-connection-dbt-compile#introspective-queries). - Blocking features are certain dbt capabilities not yet supported in Fusion. Projects that depend on any of these features will not be able to run in Fusion today. You can review the current list of [blocking features here](https://docs.getdbt.com/docs/fusion/supported-features). Even if Fusion isn't currently recommended for your project, you can still explore its capabilities with the Jaffle Shop sample project by following our [Fusion Quickstart Guide](https://docs.getdbt.com/guides/fusion?step=1). A critical next step is investing _even more_ into build conformance, so that when we roll Fusion out more broadly to our customer base of dbt platform users, we’ll be able to communicate proactively and precisely the projects and customers that we know will be Fusion-compatible. This is a journey we’re thrilled to be on, and one that’s redefining the developer experience; let’s dive into all the features and capabilities that are shipping with Fusion and the dbt VS Code extension today 👇. ## Build faster, locally with the dbt Fusion engine Fusion is the next-generation engine powering the future of dbt. It’s designed for speed, intelligence, and developer delight. And now, you can bring all of that power to your local development environment. When you develop with Fusion, you get: - **Lightning-fast performance**: Fusion is built in Rust, a systems-level language known for speed and memory safety, and is optimized to parse even the largest dbt projects 30x faster than dbt Core. This allows developers to iterate rapidly, stay in flow longer, and ultimately, ship high-quality data products faster. - **Built-in SQL comprehension**: With Fusion, dbt can emulate your cloud data warehouse locally and reason about your code with the intelligence of a modern compiler. It’s a major shift that enables powerful new capabilities that profoundly up-level the development experience. For the best possible local development experience with Fusion, including real-time error validation, consider using the [dbt VS Code extension](https://docs.getdbt.com/docs/about-dbt-extension). See the next section for more information. - **Shift left for better data quality**: Fusion lets you spot issues in your code as you’re writing it, and surface errors locally on your machine—without hitting your cloud data platform or impacting production pipelines. This includes rapid impact analysis that instantly detects whether your changes break any downstream processes as you make them. > > > — Matt Karan, Obie Insurance Local development with Fusion helps you move fast, catch issues early, and stay fully in control, all while working with the tools and development environments you’re already most familiar with. This workflow may be a good fit for teams that prefer to run dbt on their own infrastructure, work in compliance-heavy environments, or simply want the performance and experience of Fusion without relying on commercial software. ## Best-in-class developer experience with the dbt VS Code extension We’re bringing the power of Fusion right to your favorite editor. The** [dbt VS Code extension](https://docs.getdbt.com/guides/fusion) is now in Preview**, delivering a native dbt Fusion engine experience directly in VS Code, Cursor, or Windsurf. It’s the fastest, smartest, and most context-aware way to work with dbt locally, and the only way to unlock the _full_ potential of Fusion in your development environment. The extension supercharges the dbt development experience with features that let you: - Autocomplete SQL functions, model names, columns, macros, and more. - Instantly refactor across your project: Rename models or columns and see references update project-wide. - Jump to the definition of any `ref`, macro, model, or column with a single click. Particularly useful in large projects with many models and macros. - Preview column types and schema information on hover - Preview a CTE’s output directly from inside your dbt model for faster validation and debugging. - See lineage at the column or table level as you develop—no context switching or breaking flow. - View detailed logs to make it easier to spot and troubleshoot issues and audit performance - … and more! > > > — Bruno Souza de Lima, pHData To use the extension, your project must run on the dbt Fusion engine. It is not compatible with dbt Core. ## Get started with Fusion: the new standard for analytics development Today’s release marks a major milestone in [our path to GA](https://docs.getdbt.com/blog/dbt-fusion-engine-path-to-ga) as Fusion becomes the go-to foundation for development in dbt. And this is just the beginning. We're actively working on expanding Fusion support to more environments and platforms in the coming months. Are you ready to start building with Fusion? You can: - [Install Fusion locally via the CLI](https://docs.getdbt.com/docs/fusion/install-fusion) (now in Preview) - [Install the dbt VS Code extension](https://marketplace.visualstudio.com/items?itemName=dbtLabsInc.dbt) (now in Preview) - [Explore the Fusion Quickstart](https://docs.getdbt.com/guides/fusion) project to try Fusion in a sandboxed environment - [Read the docs to learn more](https://docs.getdbt.com/docs/fusion/about-fusion) We’re excited to bring the Fusion experience to our users everywhere, and we can’t wait to hear what you build next. --- --- title: "How to start a career in data" description: "Learn how to start and grow your data career with stories and advice from five women in data on The View on Data podcast." url: "https://www.getdbt.com/blog/how-to-start-a-career-in-data" date: "2025-08-15" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # How to start a career in data The debut episode of _View on Data_ brings together its five co-hosts, Grace Goheen, Lauren Benezra, Faith Lierheimer, Paige Berry, and Erica Louie, for an honest, often hilarious conversation about building a career in data. This monthly roundtable series is hosted by women in data and tech, and it’s all about sharing candid career stories, lessons learned, and the unfiltered reality of working in this industry. Each episode is part storytelling, part advice, and part “you had to be there” banter. In this first conversation, the hosts swap tales about unconventional career pivots, surprising opportunities, and the skills that really make a difference. Please reach out at **podcast@dbtlabs.com** for questions, comments, and guest suggestions. 🎧 **Listen & subscribe:** [Spotify](https://open.spotify.com/show/0Itxg3BoF6zaQdORfVuY6I?si=6133f88257624765) | [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-view-on-data/id1835755886) | [Amazon Music](https://music.amazon.com/podcasts/84c47dbe-adf7-4d33-a371-a1fe55255495/the-view-on-data) | [YouTube](https://www.youtube.com/playlist?list=PL0QYlrC86xQk2WktL3vbdB1FcWlpay7VQ) ## Many paths lead to data careers One of the biggest takeaways from the episode is that there’s no single “right” way into data. Everyone’s journey looked different: - **Grace Goheen** studied industrial engineering and dabbled in theater before a senior-year Python and R course sparked her interest. After joining a consulting company during the pandemic, she completed an internal SQL and dbt bootcamp and started working on internal data projects. She kept a “venting journal” of product frustrations and ideas, which ended up helping her move into product management. - **Lauren Benezra** began as a “snobby pure math person,” moved into applied math, learned to code, and worked with lab robots before discovering the joy of cleaning messy data. “It’s the thing that makes my brain tingle,” she said. - **Faith Lierheimer** worked in animal behavior research, including catching crickets in Hawaii with a headlamp, and in science education before pivoting during the pandemic through a bootcamp. “I was basically Fivetran-ing my crickets into Excel,” she joked. - **Paige Berry** majored in art, worked in coffee shops, and discovered Excel in an office job, which led to a 21-year career in analytics. She’s also the creator of custom Slack emojis used across the company. “Every time I see one pop up, I get a little dopamine hit.” - **Erica Louie** studied political science and Italian, learned VLOOKUP on YouTube, and cold-emailed dbt Labs’ founders. “They said ‘wear whatever you want’ to the interview, so I showed up in cut-off shorts and a zip-up hoodie.” ## Excel is often the starting point A recurring theme was how many careers started with Excel. “Excel is a gateway drug,” Erica said. “You think it’s just a spreadsheet, and then next thing you know, you’re building dashboards.” For many, Excel was the first time they saw they could manipulate data, find insights, and solve problems, and that spark eventually led to SQL, BI tools, and dbt. ## Soft skills are power skills While technical skills matter, the group agreed that communication, adaptability, and collaboration are just as important. “Knowing how to be a decent human at work is a skill, and it’s actually harder to find in data than you’d think,” Lauren said. These skills can make or break a project, and they often determine whether you’re trusted to lead initiatives or interface with stakeholders. ## How to stand out in interviews The panel shared practical ways to make an impression when applying for data roles: - **Do your homework.** “If you’re [applying to dbt](https://www.getdbt.com/about-us/careers) and you bring up the roadmap post in your interview,” Grace said, “I’m impressed. It means you cared enough to look.” - **Follow up.** Paige told interviewers, “I don’t know, but I can learn”, and then followed up later with researched answers. - **Show adaptability.** Employers value people who can step into unfamiliar territory and figure things out. - **Highlight collaboration.** Data is rarely a solo activity. Demonstrating teamwork is a big plus. ## Mentorship accelerates growth Mentorship came up repeatedly as a career accelerator. Sometimes it was a formal program; other times, it was simply shadowing someone on a project. “Getting to be a fly on the wall was huge for me,” Grace said about shadowing a product manager before officially moving into the role. Erica compared approaching a potential mentor to dating: “You don’t just say, ‘be my mentor.’ You start with a conversation.” ## Advice for breaking into data in 2025 The group acknowledged that entering the field now is more competitive, but far from impossible. Their advice: - Build a strong network, especially in communities like the [dbt Community](https://www.getdbt.com/community) Slack. - Keep your goals short-term. Focus on your next step, not your five-year plan. - Say yes to unusual or unexpected projects; they can open unexpected doors. - Keep learning and share your work publicly to build credibility. “Don’t think five years ahead,” Erica said. “Think one year ahead. You change too much in this industry to plan farther than that.” ## Why there’s no one path into a data career There’s no single formula for breaking into data. Whether your background is in math, art, science, or something else entirely, curiosity, persistence, and a willingness to keep learning can take you far. Sharing stories like these helps break down the gatekeeping that can make data roles feel inaccessible and reminds us that there’s space for many different kinds of practitioners in the field. --- --- title: "Do you need a data integration platform?" description: "How data integration platforms support ingestion and how dbt enhances reliability with testing, documentation, and versioning." url: "https://www.getdbt.com/blog/data-integration-platform" date: "2025-08-14" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Do you need a data integration platform? Data pipelines often fail not because of data volume, but due to frequent changes in upstream systems and data sources. Upstream schema changes, such as a table adding a new field or an API changing its response format, can easily break a data pipeline. Simply connecting to these data sources isn’t enough. Resilient pipelines require a strategy built on modular, testable, and version-controlled workflows. **This is where a modern data integration platform proves its value.** A data integration platform detects schema drift, adds new columns, and handles API updates with minimal manual effort. Pre-built connectors and self-healing routines minimize manual upkeep, ensuring the pipeline runs smoothly. This guide explains what data integration platforms are and how they enhance pipeline stability. It also highlights how they work alongside [dbt](https://www.getdbt.com/product/what-is-dbt) as part of a modern analytics stack. ## What is a data integration platform? Data integration is one stage of the modern [data engineering](https://www.getdbt.com/blog/what-is-data-engineering) lifecycle, preceding transformation and serving as a foundation for subsequent layers. A data integration platform moves data from multiple sources to a centralized repository, usually a data warehouse or data lake. It handles the Extract and Load phases of an [ELT pipeline](https://www.getdbt.com/blog/etl-pipeline-best-practices) (Extract, Load, Transform), automating ingestion so downstream jobs can transform the data later. The platform focuses on data ingestion, bringing raw data into your analytics environment. Controlling connectivity and ingestion processes eliminates data stitching and custom-coded extraction scripts. However, not every team needs a full platform. Custom-built pipelines can be sufficient in cases where the number of data sources is low or the architecture is simpler. The choice depends on scale, complexity, and available resources. ### Why does this matter? Modern data ecosystems are inherently heterogeneous, with each source presenting its schema, update patterns, and quirks. Manually combining these systems is slow and difficult to scale as requirements grow. For example, a single analytics use case might retrieve customer data from [Salesforce](https://salesforce.com) and product data from PostgreSQL, each delivered in a different format. A [data integration platform](http://www.getdbt.com/blog/data-integration-vs-data-transformation) abstracts those differences and provides a managed, repeatable ingestion workflow. It offers: - SaaS tools, APIs, databases, and file connectors that are prebuilt or highly configurable. - Controlled ingestion processes that are scheduled or real-time. - Support of high-volume loads, Change Data Capture (CDC), or streaming ingestion. - Schema-drift detection keeps pipelines running when a source adds columns or changes types. ### Who benefits most? Integration platforms are helpful for teams with limited engineering capacity. Instead of spending hours maintaining fragile ETL scripts, engineers can rapidly ingest and integrate data to create data models, analytics, or applications. Integration platforms are also useful for organizations that rely on new, accurate dashboards. One of the frequent causes of stakeholders losing trust in analytics is the failure of the data pipeline, resulting in missing, outdated, or inconsistent data. Integration platforms can help mitigate this by offering a robust ingestion layer that can flex to upstream changes without breaking. ### Key components of a data integration platform Data integration platforms are not just linkages to data sources. They automate ingestion, monitor pipelines to detect failures, and manage schema changes, making them resilient by default. Typical components include: #### Prebuilt connectors These connectors simplify data extraction from various sources, including SaaS applications, databases, and file systems. They handle authentication, API pagination, and schema translation. The platform provider maintains and updates connectors to handle most API and schema changes. This reduces the overhead of maintaining dozens of brittle, one-off pipelines. #### Orchestration and scheduling Pipelines can run on fixed intervals or trigger in response to events such as new files arriving or webhook notifications. Orchestration tools can manage dependencies, resource allocation, and scaling workers to handle workloads. Failure monitoring and built-in retry logic guarantee seamless data ingestion without external tools or custom code. #### Monitoring and alerting Reliable pipelines require visibility. Most platforms include dashboards that track sync status, latency, row counts, and error rates. The system alerts teams via email or chat when issues arise, enabling quick action to maintain SLAs. #### Change detection and incremental loads Modern platforms enable incremental ingestion, rather than reloading entire datasets every time. They only detect and match new or modified records with methods such as change data capture (CDC) or timestamp filtering. #### Schema mapping Data integration platforms convert raw data into query-ready tables. They map the source fields to target schemas, enforce data types, and establish relationships such as foreign keys or joins to maintain consistency. This can render the data analytics-ready without additional manual modeling. ## Architectural conditions for using a data integration platform In addition to business requirements, the design and complexity of your data architecture will determine whether a data integration platform is necessary. Some of the technical conditions that might inform this decision include: ### Minimal ingestion infrastructure Choosing a data integration platform depends on business needs and data landscape complexity. While basic setups may be enough at first, growing demands often require a stronger, managed solution. - **Basic ingestion logic**: Many pipelines begin as stateless batch jobs. A script queries a source table, dumps the rows to object storage or an ingest buffer, and exits. The data sets are small, and the table structures rarely change. This eliminates the need for cursor-based pagination, rate-limit processing, and dynamic schema evolution. - **Established orchestration layer: **Purpose-built systems such as [Airflow ](https://airflow.apache.org/)or [Dagster](https://dagster.io/) coordinate tasks, manage failures, and ensure timely data flow. They substitute weak scripts with robust, visible processes. - **Homogeneous data sources**: Most data originates internally instead of third-party SaaS APIs. Owning source systems ensures predictable schema changes and manageable rate limits. - **Tolerant latency requirements: **Teams ingest at a low frequency, tolerating delays without affecting downstream processes. This flexibility lowers the operational strain and eases pipeline management. When these conditions hold, internal pipelines can deliver value without introducing another platform layer. ### High-ingestion complexity environments Many custom pipelines struggle with continual change, multiple sources, and tight latency requirements. The following conditions highlight when a data integration platform becomes a practical and strategic choice. - **Diverse or external data sources:** Data originating from various systems, including SaaS platforms, APIs, cloud services, and operational databases, introduces significant variability. Every source can possess its schema, update behavior, and error modes. - **Advanced sync mechanisms**: Modern pipelines often require support for change data capture (CDC), incremental syncs, or polling patterns to prevent full reloads and minimize latency. Manual implementation of these features requires sophisticated state tracking and deduplication logic. - **Operational overhead**: Teams with their ingestion pipelines may experience frequent breakages due to API changes, schema changes, rate limits, or timeouts. Frequent breakages due to API issues, schema changes, rate limits, or timeouts result in data lag and constant firefighting. - **Non-reusable codebase:** With the accumulation of pipelines, ad hoc development frequently results in a set of custom scripts. The lack of a modular or reusable framework makes it difficult to maintain consistency with the sources. A platform forces a standardized approach, so adding a new source is more plug-and-play than reinventing the wheel each time. - **Missing delivery pipeline**: Ingestion code is not part of formal CI/CD pipelines. Tweaks go directly to production without version-control hooks or automated testing. Any untested change can corrupt tables and cause failures to cascade to downstream jobs. A controlled data integration platform bundles connectors, monitoring, and deployment hooks, moving the team beyond pipeline firefighting to stable data operations. ### When you might not need a data integration platform With the right conditions, internal tools and strong pipelines can provide sufficient reliability and scalability, without requiring new platforms. In some cases, existing tools and pipelines provide sufficient reliability and scalability. These scenarios include: - Most data resides in a few internal databases, and ingestion may frequently be addressed through direct access or lightweight scripts. - Stable batch jobs that run regularly and have a low failure rate minimize the necessity of a data integration platform. - Teams want to use custom ingestion logic, particularly where they have practice in testing, version control, and monitoring. - Custom pipelines can still be effective when the volume of data is small and the number of sources is manageable. A data integration platform is most valuable when ingestion is a bottleneck. Meanwhile, teams with stable pipelines, few sources, or strong engineering resources may choose to optimize current processes before implementing new tools. ## How dbt complements your data integration platform Most ingestion-focused platforms stop after landing raw or lightly structured data in an object store or warehouse [staging schema](https://docs.getdbt.com/best-practices/how-we-structure/2-staging). Although it provides stable delivery, connectivity, and schema management, it does not prepare the data for analysis and decision-making. Let’s have a look at what a data integration platform doesn’t offer: - Ingestion tools land tables but rarely let you chain multi-step SQL logic with explicit dependencies and rerun guarantees. - A connector might log “2000 rows synced”, but it will never verify whether the customer_id is unique or revenue suddenly dropped to zero. - The platform doesn’t run unit tests or data validations on pull requests. Faulty logic can reach production without review or safeguards if there is no formal deployment pipeline in place. - SQL code frequently repeats throughout pipelines because it cannot support macros or importable logic modules. This generates maintenance overhead and inconsistencies. - Deployments are frequently done ad hoc without version control, change history, or a rollback strategy. This compromises reliability and traceability within production settings. **This is where [dbt](http://www.getdbt.com/blog/what-exactly-is-dbt) comes in, helping teams do data differently by replacing brittle scripts with modular, testable models.** [dbt](https://www.getdbt.com/product/dbt) builds trust in data pipelines with automated [testing](https://docs.getdbt.com/docs/build/data-tests), clear [data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage), and version-controlled deployments. dbt streamlines data transformation workflows by: - **Modeling business logic in SQL: **Define each transformation step as a modular [SQL query](https://docs.getdbt.com/docs/build/sql-models) that runs directly in the data warehouse. Transformations are categorized into distinct layers (staging, intermediate, marts) that capture the way the business thinks about data. - **Defining reusable, testable data products: **Teams can specify [data expectations](https://docs.getdbt.com/docs/build/data-tests). For example, uniqueness, non-null values, or valid relationships in configuration files. These tests run automatically during the development and deployment phase, guaranteeing the quality of data in all models. - **Building maintainable, documented DAGs**: dbt auto-generates interactive [documentation](https://docs.getdbt.com/docs/build/documentation) that shows how models connect, where data flows, turning projects into a live data map. As the code contains the structure, modifications to a model will modify the entire graph, with no need to manually track or draw diagrams. - **Implement analytics best practices of software engineering**: Every dbt project is in [Git](https://git-scm.com/), and branching and pull requests provide a history of all changes. dbt introduces stability and discipline to analytics workflows with [version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics), testing, and environment-specific settings. The future is built on trusted data. Find out why dbt is the leader in building trusted data workflows - [sign up for a free dbt account now](https://www.getdbt.com/signup) or [book a demo with us today](https://www.getdbt.com/contact). --- --- title: "dbt Labs offers 45-day Enterprise free trial via Startup Perks program from Google Cloud" description: "Get a 45-day free dbt Enterprise trial for Google Cloud Scale Tier startups to accelerate analytics and data governance." url: "https://www.getdbt.com/blog/enterprise-free-trial-startup-perks-program-google-cloud" date: "2025-08-13" authors: ["Natasha Loeffler-Little"] categories: ["Partnerships"] --- # dbt Labs offers 45-day Enterprise free trial via Startup Perks program from Google Cloud If your startup runs on Google Cloud, here’s great news: dbt Labs now offers a 45-day Enterprise trial for qualified startups. This exclusive [Startup Perk ](https://cloud.google.com/startup/perks?hl=en)unlocks access to dbt’s Enterprise license so you can build, test, and scale trusted analytics pipelines faster. ## Who qualifies for the 45-day free dbt Enterprise trial for Google Cloud startups The 45-day Enterprise license trial is exclusive to Scale Tier members of the[ Google for Startups Cloud Program](https://cloud.google.com/startup?hl=en). If you are interested in dbt, but are not in the Google for Startups Cloud Program or a Scale Tier member, you can sign up for a [dbt 14-day free trial](https://www.getdbt.com/signup). ## Why startups should choose dbt early to scale analytics and data governance Startups face a pressure cooker of responsibility to manage resources wisely and create a return on investment for investors. As a data team leader, you play a pivotal role in implementing the central nervous system of your business, your data systems, while also meeting innovation demands by: - Moving quickly - Having a strong governance process - Ensuring you have trusted data - Being wise about using investor resources Which means you also have to: - Minimize tech debt and avoid making stack changes frequently or too late - Attract top talent that is capable of doing all the above and doing it fast Making the decision early about what goes into your tech stack will directly impact your ability to succeed. For those unfamiliar, dbt Labs is data’s great organizer. We pioneered the data analytics framework to transform bits of information into actionable data points for practitioners, teams, and businesses who rely on data they can trust. We operate as the control plane for the analytics stack, continuously innovating the data engineering toolbox and the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) that defines the way data gets done for everything from dimensional models to novel AI. Today, nearly 6,000 customers and a [community](https://www.getdbt.com/community) of more than 100k members create less data work and more data that works. That translates to a lot of great knowledge and enthusiasm you can tap into for professional development and elevating your data playbook. We are thrilled to be a partner in the Startup Perks program from Google Cloud and to partner with data teams at startups in their journey to success. With dbt’s rich feature set, innovative approach, and strong partnership with Google, thousands of organizations have adopted dbt as the essential playbook for their data teams. ## How dbt ensures data governance and trusted analytics pipelines for startups Innovation depends on trustworthy, well-governed data. You could fail if you attempt to be innovative without first setting governance and a method to trust the outputs of your data. Between managing investor resources and compressed timelines, it’s important not to skip the step of a good foundation. dbt changed the approach to create trustworthy, well-governed data and built technology that enables teams to work with a modern approach. dbt takes a code-first approach, meaning all data transformations are written and managed as modular SQL code within a development environment designed for software best practices. This is different from many legacy or GUI-based tools, where transformation logic is built visually, often making it harder to track, test, and collaborate on changes. When you start your 45-day Enterprise license trial, you’ll explore features such as automatic validation of data quality at every step. See how individual developers and analytics engineers can collaborate and iterate quickly to deploy models, all while organization-wide standards, security controls, and documentation are enforced. The real difference, though, is the platform approach and specifically a system that understands metadata so you can get “reliable, trustworthy and nuanced answers” as dbt Lab’s CEO, Tristan Handy, [points out](https://www.getdbt.com/blog/lets-talk-about-ai). ## How dbt helps startups build scalable, reliable analytics pipelines The tools and approaches you implement as a data leader will determine the success of your business, and your data systems are the nervous system. If you’re like me, you want a central nervous system that remains calm and not overtaxed. That’s exactly the metaphor for what dbt does to elevate the work of every data professional. Without dbt, you’re just pulling the next ticket from a queue. With dbt, you’re building organization-wide knowledge systems, allowing your data practitioners to do their best and most important work orders of magnitude faster without tech debt weighing them down. It’s not just the experience of using dbt. It’s one system, deployed smoothly and integrated with your Google suite. Many data leaders at startups experience tech debt from manual scripts and piecemeal workflows that got the initial POC done, but are now bursting at the seams, both in their capacity to handle bigger volumes of data as well as ballooning costs. You know that’s painful and also that you have to meet the needs of the [modern data engineer](https://www.getdbt.com/blog/data-engineering). We want amazing outcomes for startups. Dealing with fewer remedial tasks means you and your team get more time on the interesting, impactful work at the heart of your business. ## Why dbt helps attract and retain top data talent at startups We’re addressing this last, but the number one most important piece of the driving success on data teams is the humans. Tristan Handy said it best in the [Analytics Development Lifecycle blog](https://www.getdbt.com/resources/the-analytics-development-lifecycle) when he said, “We believe that teams are capable of practicing analytics in a way that hits a far higher mark on these dimensions simultaneously. This requires both better analytical systems and more mature analytical workflows.” It’s cumulative of the technology, the approach, and the people who bring it to life. It’s important that there is a human-first approach in making your data org thrive under the immense startup business pressures. Having a large pool of trusted experts is vital for rapidly evolving organizations like yours. There are over 100,000 members in the [dbt Community](https://www.getdbt.com/community). This is a place where data professionals of any skill can ask questions, debate ideas, and push each other forward. There are also hundreds of Google engineers and product leaders active in the community as experts and advocates for your data transformation questions. dbt is the standard in data transformation, and your startup benefits greatly from established standards in education and learning. When you’re looking to build a high-performing team, you want talent that can ramp quickly. Employees with [dbt certifications](https://www.getdbt.com/dbt-certification) can jump into your stack and have an impact quickly. Additionally, dbt is well respected by users, [according to G2](https://www.g2.com/products/dbt/reviews), and regularly, our own SI partners share with us that their consultants request to work on dbt projects because they love the product so much. ## dbt Labs and Google Cloud partnership: Accelerating startup data success Google and dbt share a passion for customer success, and it’s our rich collaboration that will continue to deliver innovation for you to develop your startup. Most recently, we recapped our [releases announced](https://www.youtube.com/watch?v=kHdNkfXHWas) at Google Next 25, launched support of [Apache Iceberg tables in BigQuery](https://www.getdbt.com/blog/dbt-supports-apache-iceberg-tables-bigquery), and [outlined](https://docs.getdbt.com/blog/train-linear-dbt-bigframes) how to train a linear regression model with dbt and BigFrames. dbt is [available on Google Cloud](https://www.getdbt.com/blog/dbt-cloud-google-cloud) and can be purchased on the [Google Cloud Marketplace](https://console.cloud.google.com/marketplace/product/dbt-marketplace-public/dbt-cloud?hl=en&inv=1&invt=Ab5P6g&project=dbt-marketplace-public), so you have the combined benefit of easy deployment and procurement. ## How to claim your 45-day free trial If you’re not yet a Google for Startups Cloud Program member, start [here](https://cloud.google.com/startup?hl=en). For Google Startup Scale Tier organizations, visit the [Startup Perks](https://cloud.google.com/startup/perks?hl=en) site and register to receive the unique URL for dbt’s 45-day Enterprise free trial. Our team looks forward to connecting with you to help you unlock your team’s expertise. --- --- title: "How to ensure data product SLAs and SLOs" description: "How to implement SLAs & SLOs for data products—covering quality dimensions, monitoring, governance and continuous improvement." url: "https://www.getdbt.com/blog/data-product-slas-and-slos" date: "2025-08-13" authors: ["Joey Gault"] categories: ["Pulse"] --- # How to ensure data product SLAs and SLOs Before establishing meaningful SLAs and SLOs, data engineering leaders must first understand the various [dimensions that define data quality](https://www.getdbt.com/blog/data-quality-dimensions). These dimensions provide the framework for measuring and monitoring the reliability of your data products. The key dimensions include usefulness, accuracy, completeness, consistency, uniqueness, validity, and freshness. Each dimension represents a different aspect of data quality that can impact your ability to meet service commitments. Usefulness measures whether data generates actual business value, helping you identify and address dark data that consumes resources without providing returns. Accuracy ensures that data reflects reality, while completeness guarantees you have all required records and fields needed to answer business questions. Consistency maintains data integrity across upstream and downstream sources as information flows through your data lifecycle. This dimension becomes particularly critical when dealing with patient medical systems or financial data, where inconsistencies can have severe consequences. Uniqueness prevents the chaos caused by duplicate records with conflicting information, while validity ensures data values are correct for their column types and within acceptable ranges. Freshness, also known as timeliness, measures whether data has been updated within target timeframes. This dimension directly translates to SLA requirements, as different use cases demand different freshness guarantees. Weekly sales reports might require only daily updates, while real-time operational dashboards need hourly or even minute-level freshness. ## Implementing comprehensive monitoring and testing frameworks Establishing reliable SLAs and SLOs requires robust monitoring systems that can detect issues before they impact end users. This proactive approach involves[ implementing automated testing](https://docs.getdbt.com/docs/build/data-tests) at multiple levels of your data pipeline, from source data validation to final output verification. [dbt](https://www.getdbt.com/product/what-is-dbt) provides powerful capabilities for implementing data tests that verify quality across all dimensions. Generic tests can be reused across projects to check for null values, uniqueness constraints, accepted value ranges, and referential integrity. These tests form the backbone of your quality assurance process, catching issues during transformation rather than after data reaches consumers. Source freshness monitoring represents a critical component of SLA management. [dbt's source freshness functionality](https://docs.getdbt.com/docs/deploy/source-freshness) allows you to define acceptable update intervals for each data source and automatically monitor compliance. The frequency of these checks should align with your SLA requirements: if you have a one-hour SLA on a dataset, monitoring freshness every 30 minutes provides adequate coverage to detect violations promptly. Beyond automated testing, implementing comprehensive data lineage tracking enables rapid issue identification and resolution. When problems occur, understanding the flow of data from source to destination allows teams to quickly isolate the root cause and implement fixes without extensive investigation time. ## Establishing realistic and measurable SLAs [Creating effective SLAs requires balancing business requirements with technical realities.](https://www.getdbt.com/blog/data-slas-best-practices) The most common mistake data engineering leaders make is committing to SLAs that sound impressive but cannot be consistently achieved given current infrastructure and processes. Start by conducting a thorough assessment of your current data pipeline performance. Measure actual refresh times, error rates, and recovery periods over several months to establish baseline performance metrics. This historical data provides the foundation for setting achievable targets while identifying areas requiring improvement. Different data products warrant different SLA commitments based on their business criticality and usage patterns. Customer-facing dashboards displaying real-time metrics require more stringent SLAs than internal reporting used for monthly planning sessions. Segment your data products by business impact and establish tiered SLA structures that reflect these priorities. Consider both availability and performance metrics in your SLAs. Availability measures the percentage of time data products are accessible and functioning correctly, while performance metrics cover data freshness, processing times, and error rates. ## Building automated alerting and response systems Meeting SLAs consistently requires automated systems that can detect violations and trigger appropriate responses without human intervention. Manual monitoring simply cannot provide the coverage and response times necessary for modern data operations. Implement multi-layered alerting that escalates based on severity and duration of issues. Initial alerts might notify the on-call data engineer, while prolonged outages trigger escalation to management and potentially business stakeholders. Configure different alert thresholds for different types of violations: a minor delay in non-critical data might warrant a low-priority notification, while a complete failure of customer-facing analytics requires immediate attention. During an incident which blocks fresh data from arriving, a naively-orchestrated project will still execute as normal, wasting compute resources to recalculate unchanged tables. Pairing freshness checks with [dbt State](https://www.getdbt.com/product/dbt-state) means that a table and its tests will not be executed if nothing has changed since the last invocation. Automated remediation capabilities can resolve many common issues without human intervention. Simple problems like temporary network connectivity issues or resource constraints often resolve themselves through retry mechanisms and auto-scaling infrastructure. More complex issues require human intervention, but automated systems can still perform initial diagnostics and gather relevant information to accelerate resolution. Documentation and runbooks play crucial roles in maintaining consistent response quality. When alerts fire, responders need immediate access to troubleshooting procedures, escalation contacts, and historical context about similar issues. Well-maintained runbooks reduce mean time to resolution and ensure consistent handling regardless of which team member responds. ## Governance frameworks for SLA management Effective SLA management extends beyond technical implementation to encompass organizational processes and governance structures. Clear ownership models, change management procedures, and communication protocols ensure that SLAs remain relevant and achievable as business requirements evolve. Establish clear ownership for each data product and its associated SLAs. Product owners should understand both the technical constraints and business requirements, enabling them to make informed decisions about trade-offs between features, performance, and reliability. This ownership model prevents the common scenario where SLAs are defined without adequate consideration of implementation complexity or resource requirements. Change management processes must account for SLA impacts when evaluating modifications to data pipelines or infrastructure. Seemingly minor changes can have cascading effects on performance and reliability, potentially causing SLA violations if not properly assessed. Implement review procedures that evaluate proposed changes against existing SLA commitments and require explicit approval for modifications that might impact service levels. Regular SLA reviews ensure that commitments remain aligned with business needs and technical capabilities. Quarterly reviews should examine actual performance against targets, assess whether SLAs remain appropriate for current business requirements, and identify opportunities for improvement. These reviews also provide opportunities to celebrate successes and learn from failures. ## Measuring and reporting on SLA performance Transparent reporting on SLA performance builds trust with business stakeholders and provides data for continuous improvement efforts. Effective reporting balances technical detail with business-relevant metrics, ensuring that different audiences receive appropriate information. Create dashboards that provide real-time visibility into SLA compliance across all data products. These dashboards should highlight current status, recent trends, and any active issues requiring attention. Different views serve different audiences: technical teams need detailed metrics about individual pipeline performance, while executives require high-level summaries of overall service reliability. Monthly SLA reports should provide comprehensive analysis of performance trends, root cause analysis for any violations, and plans for addressing systemic issues. These reports serve as historical records and help identify patterns that might not be apparent from real-time monitoring alone. Consider implementing SLA credits or other accountability mechanisms for critical data products. While not always appropriate, formal consequences for SLA violations can provide additional motivation for maintaining high service levels and demonstrate commitment to reliability. ## Continuous improvement and optimization [SLA management](https://blog.happyfox.com/ultimate-guide-sla-management/) is not a one-time implementation but an ongoing process of measurement, analysis, and improvement. Regular assessment of both technical performance and business requirements ensures that your data products continue to meet evolving needs. Analyze patterns in SLA violations to identify systemic issues requiring architectural changes. Frequent violations of freshness SLAs might indicate the need for more powerful processing infrastructure or pipeline optimization. Recurring accuracy issues could signal problems with source data quality or transformation logic that require broader remediation efforts. Invest in infrastructure and tooling improvements that enhance your ability to meet SLAs consistently. Modern data platforms offer features like automatic scaling, improved monitoring capabilities, and more efficient processing engines that can significantly improve reliability and performance. Optimize your data delivery by matching each table's execution cadence to its SLAs, not to a generic "every four hours" command. After you have defined SLAs in conjunction with your stakeholders, you may realize that reprocessing a table multiple times a day for a dashboard with a 24-hour SLA is unnecessary. Instead, use dbt State to define a [lag tolerance](https://docs.getdbt.com/reference/resource-configs/lag-tolerance) for models in your project, enabling you to throttle highly volatile sources while reducing manual orchestration overhead. Foster a culture of reliability within your data engineering organization. This involves training team members on SLA management principles, establishing clear expectations for service quality, and recognizing achievements in maintaining high service levels. When reliability becomes a core value rather than just a technical requirement, teams naturally make decisions that support SLA compliance. The path to reliable data product SLAs and SLOs requires commitment across technical, process, and cultural dimensions. By implementing [comprehensive quality frameworks](https://www.getdbt.com/blog/data-quality-framework-choosing), robust monitoring systems, and clear governance structures, data engineering leaders can build the foundation for consistently meeting service commitments. Success in this area not only reduces the financial impact of poor data quality but also builds the trust necessary for data-driven decision making across the organization. The investment in proper SLA management pays dividends through reduced firefighting, improved stakeholder confidence, and the ability to take on more strategic initiatives. As data becomes increasingly central to business operations, the organizations that master reliable data product delivery will gain significant competitive advantages in their markets. ## Data SLAs and SLOs FAQs **What is a data SLA?** A data SLA (Service Level Agreement) is a formal commitment that defines the expected performance and reliability standards for data products. It establishes measurable targets for key dimensions like data freshness, accuracy, completeness, and availability, providing clear expectations between data teams and their consumers about service quality and response times. **What are the most common SLA metrics for data quality?** The most common SLA metrics for data quality include availability (percentage of time data products are accessible and functioning correctly), freshness (whether data has been updated within target timeframes), error rates (typically maintained below 0.1%), processing times, and data refresh intervals. These metrics are often combined into comprehensive agreements, such as guaranteeing 99.5% uptime with data refreshed within four hours of source updates. **What happens when an SLA is breached?** When an SLA is breached, automated alerting systems trigger multi-layered responses that escalate based on severity and duration. Initial alerts notify on-call data engineers, while prolonged outages escalate to management and business stakeholders. The response includes immediate diagnostics, following established runbooks for troubleshooting, implementing automated remediation where possible, and conducting root cause analysis. For critical data products, some organizations implement SLA credits or formal accountability mechanisms as consequences for violations. --- --- title: "Best of both worlds: dbt + Databricks Lakeflow Jobs" description: "Orchestrate dbt jobs in Databricks Lakeflow to unify data, AI, and ML workflows for faster, trusted pipeline delivery." url: "https://www.getdbt.com/blog/dbt-databricks-jobs" date: "2025-08-13" authors: ["Nina Anderson"] categories: ["Partnerships"] --- # Best of both worlds: dbt + Databricks Lakeflow Jobs Picture this: your data engineering team builds ingestion pipelines in Databricks, your analytics team transforms data with dbt, and your ML engineers train models—all in completely separate workflows. Sound familiar? Those days are officially numbered. The combination of dbt and Databricks has already proven transformative for [countless data teams globally](https://www.databricks.com/blog/how-home-trust-modernized-batch-processing-databricks-data-intelligence-platform-and-dbt-cloud), streamlining transformations while leveraging the lakehouse's power and scale. But what if you could take that winning combination even further? What if you could seamlessly weave dbt jobs directly into your Databricks workflows, creating one unified pipeline that speaks every team's language? That's exactly what we’ve built. ## What's new: dbt Platform Task Type in Databricks Lakeflow The dbt Platform Task Type is now available in Private Preview, allowing you to orchestrate dbt Platform jobs directly within Databricks Lakeflow Jobs.⁠ This represents a big step forward in how data teams collaborate. With Public Preview expected in time for [Coalesce](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/register-1),⁠ this integration represents the first time you can truly unify your ingestion, transformation, and AI pipeline orchestration in one place. ## Why it matters ### End-to-end pipeline orchestration made simple Constantly switching between interfaces can disrupt ‘flow state’, and in so doing, detract from developer productivity. It’s one of our convictions that good developer tools should promote flow state. Gone are the days of managing separate orchestrators for different parts of your data pipeline, and no more custom API calls to get Databricks and dbt orchestration working together. Now, you can manage raw data ingestion, streaming, batch transformation, and AI/BI in a single Databricks Lakeflow Jobs interface ### Meeting your teams where they are dbt offers three [developer interfaces](https://www.getdbt.com/product/develop) – Studio, Canvas, and CLI – and now your teams can productionize their dbt pipelines within Databricks Lakeflow Jobs without losing any of that rich development experience.⁠ ### Unlock the power of dbt Mesh Here's where it gets really exciting: cross-project and cross-platform orchestration capabilities are only available through dbt Platform jobs.⁠ This means you can respect team-specific orchestration (per domain project) while integrating those into global pipelines, delivering true [dbt Mesh](https://docs.getdbt.com/best-practices/how-we-mesh/mesh-1-intro) architecture at scale. ### Centralized control without compromise Organizations standardizing on Databricks and dbt get the centralized scheduling, retries, alerting, and audit trails they need,⁠ while development teams retain the flexibility and power of dbt features like enhanced IDE integration and cost optimization.⁠ ## The technical sweet spot Because dbt deeply integrates with Unity Catalog, you also get full metadata syncing and governance,⁠ while enhanced developer experience through dbt Platform's Jobs integration with Databricks⁠ means your teams get the best of both worlds: enterprise-grade orchestration with developer-friendly transformation tools. Plus, with improved data lineage, documentation, and performance metrics flowing seamlessly between platforms, you finally get that single source of truth. ## What's next? If you're using Databricks with dbt, this integration isn't just a nice-to-have; it's the missing piece that transforms how your entire data organization collaborates. Private Preview customers are already testing this,⁠ with Public Preview launching by [Coalesce](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/summary) in October.⁠ Ready to see what unified data and ML orchestration looks like in practice? The future of collaborative data development is here, and it's running on the platform you already know and love. _Want to learn more about the dbt Platform to Databricks Lakeflow Jobs integration? Contact your dbt Labs or Databricks representative to discuss joining the Public Preview program launching at Coalesce._ --- --- title: "Why you should care about data transformation" description: "Here’s how transformation makes your analytics reliable, scalable, and trusted across the business." url: "https://www.getdbt.com/blog/why-data-transformation-matters" date: "2025-08-12" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Why you should care about data transformation Businesses increasingly strive to be data-driven. And yet, [57% of companies](https://sloanreview.mit.edu/article/building-a-data-driven-culture-four-key-elements/) still struggle to achieve it. The challenge is that raw data is unusable. It flows in from multiple systems, each with its own formats and errors. Without data transformation, this causes unreliable insights, delayed decisions, and wasted time. Teams often rely on ad hoc SQL queries when standardized data transformation practices are absent. However, varying data definitions can lead to different teams making divergent assessments of the same question. Multiple versions of key metrics create uncertainty about which data to trust. This results in everyone working in isolation, duplicating efforts, and increasing inconsistencies. The good news is, it doesn’t have to be this way. This article shows how dbt helps your team replace fragile processes with workflows that deliver consistent, trustworthy data. ## What is data transformation? Data transformation is the process of converting unstructured data into structured formats for reliable analysis and interpretation. It involves: 1. Correcting errors or inconsistencies in the data, like removing duplicates or filling in missing values. 2. Standardizing data formats and structures to ensure consistency across datasets. 3. Summarizing data to provide meaningful insights, such as calculating averages or totals. 4. Enhancing data by incorporating additional information from external sources to provide deeper insights. 5. Reorganizing data structures to facilitate analysis, such as pivoting tables or merging datasets. It’s important to distinguish data transformation from extraction and loading within the data pipeline. Transformation specifically focuses on modifying data to meet the intended use. In a modern data transformation process, [such as an Extract, Load, and Transform (ELT) process](https://www.getdbt.com/blog/extract-load-transform), this can occur multiple times as the data is reshaped for different use cases. For example, type casting, removing duplicates, and creating metrics are all part of data transformation. These tasks ensure data is accurate and aligned with business objectives for analysis and reporting. ## Why does data transformation matter? Data transformation converts raw inputs into usable assets that enable more precise analysis and improved business outcomes. The key benefits are: **1. Faster and more accurate decision-making. **Well-organized data accelerates analysis and boosts confidence in dashboards and [key performance indicators (KPIs)](https://www.getdbt.com/blog/streamlining-kpi-dashboards-dbt-semantic-layer). When data is structured and error-free, analysts spend less time cleaning it. They focus more on interpreting, leading to quicker and reliable insights. **2. Compliance & auditability. **For regulated industries, a robust data management strategy is essential for properly handling sensitive information. Data-centric processes with clear rules for masking sensitive fields and tracking [data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) help meet compliance requirements. This ensures organizations maintain reliable audit trails, so that they can always demonstrate to themselves and regulators that data is accurate and not compromised. **3. Operational integrity. **Data integrity issues, such as inconsistencies, inaccuracies, and duplications, can cascade into critical systems. They skew machine learning outputs, disrupt billing processes, and compromise automated decisions. Implementing data integrity measures, such as validation and consistency checks, before data enters systems, prevents costly errors. ### When to prioritize data transformation Prioritize data transformation when your organization faces inconsistent metrics, scaling, and governance challenges. Common scenarios include: - **Inconsistent metrics across teams. **When different departments report conflicting metrics, it creates confusion and erodes trust in data. This signals a need for standardized definitions and transformations. - **Strict compliance needs. **Regulated industries require robust transformation processes. These include data validation and automated lineage to establish audit trails and demonstrate regulatory compliance. - **Scaling cloud analytics.** If growing data volumes create a complex cloud data warehouse environment, transformation becomes essential. A transformation layer standardizes business logic in one place reducing manual data preparation. This enables teams to deliver world-class data products at scale. ## Common use cases where data transformation creates value Data transformations unlock value across analytics, operations, and AI workflows. ### Reporting and Business Intelligence (BI) Transformation aligns raw data with shared definitions and feeds BI tools with consistent metrics through a semantic layer architecture. A [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) acts as a single source of truth for standardized business metrics. Metrics are delivered to downstream tools on demand, consistently reflecting accurate values. These values can also flow to AI to provide it a single, accurate source for key business metrics. ### Real-time and near-real-time analytics Transformation is applied to high-volume streaming data to extract timely insights. It involves using windowing functions and aggregations, such as session windows, to process raw events. This process enables real-time decision-making and updates to dashboards and operational systems. ### ML and AI feature engineering Transformation cleans and aggregates raw data into [reproducible features](https://www.ibm.com/think/topics/feature-engineering) that help ML pipelines to scale effectively. It uses modular logic and version control to build features quickly and reliably. The consistent application of this process ensures accurate model training and efficient deployment. ### Regulatory reporting & governance Data transformation ensures that datasets are auditable, tested, and traceable, maintaining compliance. Clear data lineage helps produce compliant, validated reports that meet regulatory standards. ## Three ways to optimize your data transformation workflow Efficient data transformation depends on smart compute management, clear ownership, and streamlined pipeline design. Each of these impacts costs, performance, and reliability. Here are three ways to optimize your workflows: - **Warehouse compute costs vs. speed.** Transforming data within the data warehouse consumes compute resources. It is essential to optimize [model materialization](https://docs.getdbt.com/docs/build/materializations) and [pipeline scheduling](https://docs.getdbt.com/docs/deploy/job-scheduler) to balance compute costs with performance. - **Ownership and roles.** Assign clear ownership. Analytics engineers should own the modeled datasets and the associated data quality tests. This ensures that business users receive reliable and consistent answers when using BI tools. - **Incremental models. **Design your models to process only new or updated records. Using [incremental materializations](https://docs.getdbt.com/docs/build/incremental-models) enhances pipeline efficiency by limiting how much data is transformed. ## How dbt empowers data transformation dbt addresses common data transformation challenges by bringing software engineering best practices to analytics. It acts as your [data control plane](https://www.getdbt.com/resources/whitepaper-the-control-plane-for-data-collaboration-at-scale), providing a flexible, cross-vendor, enterprise-wide solution for collaborating on data no matter where it lives. dbt streamlines transformation workflows by supporting: **1. Modular and reusable SQL models.** dbt enables the creation of [modular](https://www.getdbt.com/blog/modular-data-modeling-techniques) and reusable SQL logic. This approach ensures consistency and reduces redundancy across various data models, preventing duplicated work and misaligned metrics. **2. Git-native CI/CD workflows.** dbt ensures transformations are reliable and changes are tracked through [integrated testing](https://docs.getdbt.com/docs/build/data-tests) and [version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics). Version control supports collaboration by providing a single source of truth for all analytics code changes and reconciling conflicts among contributors. Meanwhile, tests run as part of a [Continuous Integration/Continuous Delivery (CI/CD)](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud) deployment process ensure that your data test suite is run with every push to production, ensuring accuracy and consistency of data before a developer’s changes go live. **3. Context-aware development. **[The dbt Fusion engine](https://docs.getdbt.com/docs/fusion/about-fusion) compiles and checks data warehouse code before it** **ever hits your data warehouse.** **Other techniques, such as [incremental materialization](https://docs.getdbt.com/docs/build/incremental-models), ensure that dbt jobs only run on new or updated data. This enhances pipeline efficiency by limiting data volume and churn, and managing compute costs. **4. Automated documentation and lineage.** dbt enables [embedding documentation as part of your data models](https://docs.getdbt.com/docs/build/documentation) and generates docs automatically with every push to prod. It adds transparency through detailed [directed acyclic graphs (DAG)](https://www.getdbt.com/blog/dag-use-cases-and-best-practices), ensuring datasets are well-documented and traceable. dbt supports analytics engineering by integrating with modern data warehouses and complements ingestion and BI tools. It runs transformations within your data warehouse, using systems you already use, such as [Snowflake](https://snowflake.com) or BigQuery, to implement scalable [Extract, Load, Transform (ELT)](https://www.getdbt.com/blog/extract-load-transform) pipelines. This provides a single, authoritative source of truth, ensuring everyone uses the same data definitions. dbt improves data consistency, helping teams build trust in their analytics. **5. Semantic layer. **Your teams can’t collaborate effectively if they have different definitions of basic concepts like “revenue.” The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) centralizes key metrics in a single location using standard business terminology, not technical jargon. This ensures that, when it comes to data, everyone at your company is speaking the same language. ## Conclusion With proper transformation practices, teams spend less time fixing data issues and more time analyzing. Consistent logic, testing, and modular workflows reduce downstream errors and accelerate reliable insights across teams. A modular framework helps adapt to increasing data volume. dbt’s modular, version-controlled models efficiently manage large datasets while maintaining reliability in analytics environments. Adopting these modern practices ensures your workflows remain efficient and responsive. By investing in the right tools and culture, you minimize risks associated with data errors and non-compliance. dbt’s automated documentation, testing, and governance features enhance transparency and improve organizational agility. Start transforming data reliably today—[sign up for dbt for free](https://www.getdbt.com/signup) and walk through [one of our quickstarts](https://docs.getdbt.com/docs/get-started-dbt). --- --- title: "Federated data governance: What makes it different?" description: "Explore how federated data governance works—and how dbt turns principles into scalable, self-serve practices across domains." url: "https://www.getdbt.com/blog/federated-data-governance" date: "2025-08-11" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Federated data governance: What makes it different? As organizations grow, data governance can support data agility and trust or become a barrier to both. Conventional **centralized data governance** offers robust centralized control and uniformity. On the minus side, it can create a data monarchy that often slows decision-making and limits flexibility. In contrast, completely **decentralized governance** empowers individual teams, accelerates decision-making, and enhances localization. However, this can result in [data silos](https://www.techtarget.com/searchdatamanagement/definition/data-silo), policy inconsistencies, and redundant efforts - i.e., data anarchy. **Federated data governance** strikes a balance. It combines a central guiding structure with distributed, domain-level execution to achieve trust and compliance without compromising speed and scalability. In this article, we’ll explore the concept of federated data governance and how it differs from centralized and decentralized models. We’ll look at the central pillars of a federated model, including distributed control and automated policy enforcement, which make it a scalable but controlled solution. ## Understanding federated data governance In a federated model, a central governing unit sets company-wide data policies, standards, and best practices. Individual domain teams implement and enforce these policies within their respective data products and pipelines. Each domain owns its data and customizes governance as long as it meets the minimum standards at the center. This approach provides a consistent level of uniformity and conformance throughout the organization while allowing for localized optimization. Federation works well in large organizations where data is varied, rapidly growing, and used by different departments with varying requirements. A one-size-fits-all strategy is insufficient in such environments. Federated governance flourishes with strong self-service platforms. Data catalogs, metadata systems, and platforms like [dbt](https://www.getdbt.com/product/dbt) let domain teams manage data quality, access, and lineage without centralized control. Governance rules can be coded as reusable elements, which guarantees uniform enforcement across domains. The same documentation and metadata standards create a shared language, facilitating easy cross-team collaboration. ![Circular diagram titled “Understanding Federated Data Governance,” showing interactions between a Central Governance Hub and three Domain Teams (A, B, and C). The hub provides policies, standards, and a catalog. Arrows indicate flows such as “Governance-as-Code,” “Lineage,” “Model Contracts,” and “Self-Serve Data Projects” between the hub and domain teams, as well as peer-to-peer sharing of lineage between teams. The graphic emphasizes decentralization with central coordination for scalable, governed data operations.](https://cdn.sanity.io/images/wl0ndo6t/main/b74f8b6a88d85fec99b699bf5b875343bf80542e-1024x1024.jpg) ### How federated models differ from centralized and decentralized models The table below illustrates key differences among centralized, decentralized, and federated governance models across core characteristics: ## Core pillars of federated data governance Federated data governance is not only a structural change, but also a cultural shift. It combines automation, ownership, transparency, and collaboration into a model that can scale with contemporary data organizations. When done right, it turns governance into a driver of trust, innovation, and efficiency through data products. ### Distributed control Federated governance gives data ownership to the teams most familiar with it: the data domain experts. These groups have the authority to establish and control their data products, develop local quality and access rules, and align them to the requirements of their functional area. A finance team may implement more rigid reconciliation rules, and a marketing team may customize customer segmentation models. However, both teams operate under standard enterprise policies, such as data privacy requirements, naming conventions, and access control rules. This non-central ownership can help avoid siloing, as domain teams continually maintain and enhance their assets. Federated models are inherently more agile than centralized models, which tend to slow down innovation and require top-down decisions - often by people who aren’t as familiar with that domain’s particular rules and assumptions. ### Automation and scalability Manual processes and periodic compliance audits don’t scale for modern governance. In federated governance, policies are enforced as executable code and embedded in data pipelines, transformation logic, and deployment processes. This approach, known as [policy as code](https://aws.amazon.com/blogs/security/governance-at-scale-enforce-permissions-and-compliance-by-using-policy-as-code/), automates and integrates governance into the daily workflow. Some common examples include common [data quality checks](https://www.getdbt.com/blog/data-quality-checks), such as null checks, data uniqueness, validation, access restrictions, and data retention timelines. These are often built into tests and scripts within production environments to ensure everything runs smoothly. The rules are then versioned with Git and deployed through CI/CD pipelines, similar to software code. This makes it easier to track changes, automate testing, and revert if changes cause issues. Automation enables organizations to scale governance in response to the increasing volume and complexity of data. As more areas generate data, they receive centralized logic and templates that ensure compliance without adding unnecessary overhead or requiring engineers to reinvent the wheel. Automation also improves time-to-insight, since governance is built into the development pipeline instead of being added as an audit layer after deployment. Embedding policies early in the workflow enables teams to identify issues before they impact downstream systems. ### Transparency and accountability End-to-end visibility is crucial in a federated governance model where responsibilities are shared among teams. Accountability and coordination within domains are ensured by clear tracking of data ownership, policy application, and change history: - Governance logic, rules, transformations, and model definitions must be maintained in both human and machine-readable formats to allow automation and facilitate cross-team processes. - Artifacts such as YAML configurations, schema files, policy specifications, and lineage graphs serve as living documentation and binding agreements between producers and consumers, making updates transparent, testable, and enforceable across different environments. - Version control systems also enhance accountability, as all policy updates, rule changes, and schema modifications are documented and traceable. This creates a reliable audit trail that is beneficial for compliance, internal learning, and team coordination. Ultimately, transparency and accountability will turn governance into a vibrant and adaptive system that responds to the needs of both business and regulation. ### Community-driven practices A key pillar of federated governance is that it functions not as an authoritative enforcement system, but as a cooperative ecosystem. Domain teams, platform engineers, and central policy stewards co-create governance decisions. This bottom-up contribution ensures that governance policies are not only compliant but also practical and context-sensitive. An effective federated model requires open channels of communication, regular forums, documentation hubs, shared [Slack](https://slack.com/) spaces, or community-owned repositories. These serve as the town square where teams can agree on standards, share reusable components (e.g., macros to mask PII), and help evolve governance practices. This community-based model resembles internal open-source models, in which governance artifacts are shared, reviewed, and extended between teams. When a domain team develops a data minimization method, the central team must convert it into a generic template that is reusable across the board. They must also provide any additional tools and advice applicable to all domain teams in general. This is an expression of contemporary platform thinking, in which the core capability empowers domain teams by standardizing best practices without constraining local adaptability. ## Benefits of federated data governance A federated solution to data governance changes the way organizations handle control, collaboration, and accountability. The following benefits highlight its effectiveness in contemporary data settings. - **Democratized analytics:** Federation enables domain teams to curate their data products, allowing for self-service analytics and reducing dependency on central IT teams. This leads to a data-driven culture through improved data literacy and the faster generation of insights across the organization. - **Effective use of dark data:** Federated governance can help organizations identify and classify [dark data](https://www.gartner.com/en/information-technology/glossary/dark-data), including logs and documents, to convert it into usable intelligence. This increases the value of current assets and AI/ML initiatives that depend on diverse, governed data sources. - **Cross-system interoperability:** Federated models promote interoperability of hybrid and multi-cloud environments by aligning domain-level implementations with central metadata standards. This facilitates the smooth flow of data and its combination with the other domains. ## Challenges of federated data governance Although federated data governance has many benefits, it comes with new challenges. Tackling these issues is critical to addressing for long-term success and data product growth. - **Operational complexity:** Managing numerous autonomous teams increases both architectural and operational complexity. This requires platforms and trained personnel to incorporate metadata, apply policies, and coordinate governance effectively. - **Performance and latency overheads:** Distributed systems may experience network latency, slow responses, or inconsistent performance when executing federated queries. These problems can impact real-time analytics unless addressed through caching or query optimization techniques. - **Security fragmentation: **Decentralization creates the risk of uneven security enforcement across domains. In the absence of rigorous central auditing and cohesive security policies, the environment of each domain may become a soft entry point. ## How dbt enables federated data governance To effectively execute federated data governance, companies require more than a strategy; they need the appropriate tools to translate principles into practice. This is where [dbt](https://www.getdbt.com/product/governance) can be particularly useful. dbt allows autonomy and control by incorporating governance into the daily data workflow. Federated governance is scalable and enforceable because it allows domain teams to govern their data and stay aligned with centralized policies. ### Domain ownership self-serve data projects With dbt, domain teams can own and control their data projects, creating, testing, and deploying models themselves. This fosters a self-serve culture, where [domains](https://docs.getdbt.com/guides/core-to-cloud-1?step=1) govern themselves in line with central policies. This aligns with the [data mesh](https://www.getdbt.com/blog/data-mesh-getting-started) principle of federated computational governance, where teams operate independently within a shared governance framework. ### Version-controlled governance Governance rules, such as data quality tests, model contracts, and schema validation, are coded [as reusable packages](https://docs.getdbt.com/docs/build/packages) in dbt and applied when the pipeline is executed. These rules are versioned through [Git ](https://docs.getdbt.com/docs/cloud/git/git-configuration-in-dbt-cloud)and delivered through [CI/CD](https://docs.getdbt.com/guides/set-up-ci?step=1), making them traceable, allowing rollbacks, and ensuring audit readiness. ### Cross-project lineage and visibility using dbt Mesh The cross-project references feature of [dbt Mesh](https://docs.getdbt.com/best-practices/how-we-mesh/mesh-1-intro) enables domains to share and reference public models across projects, establishing modular [dependencies](https://docs.getdbt.com/docs/mesh/govern/project-dependencies). End-to-end lineage between domains is visualized and controlled via the dbt Catalog UI, where users can navigate project-level and account-level lineage graphs. ### Reliability and collaboration model contracts [Model contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) in dbt provide a framework for specifying and enforcing schema-level expectations directly within your data pipelines. This ensures reliability and collaboration between teams. You can define constraints like not allowing null values or duplicates by writing contracts in YAML in your project or packages. dbt will test these rules at build time to identify any violations before the materialization happens in your warehouse. If a [contract](https://docs.getdbt.com/reference/resource-configs/contract) is violated, dbt will raise an evident error on dbt run. This prevents breaking changes and ensures the responsible domain owner addresses data-quality issues promptly before they impact production. ### Self-serve platform central registration & monitoring The self-serve platform capabilities of dbt facilitate the central registration and monitoring of domain projects. [dbt Catalog](https://docs.getdbt.com/docs/explore/access-from-dbt-cloud) and the [Discovery API](https://docs.getdbt.com/docs/dbt-cloud-apis/discovery-api) enable a single, searchable list of all data assets and metadata across all environments. The Catalog uses metadata created with every dbt run. Combined with your warehouse-external metadata, this allows users to find tables, views, models, tests, and lineage visualizations in a single location. The dbt API can programmatically sync metadata into external governance or catalog platforms. dbt makes governance a self‑serve and integrated practice by adding these capabilities to the seamless navigation between the Catalog and other platform capabilities. ## Conclusion Federated data governance offers a scalable, adaptable solution for organizations balancing autonomy with oversight. By combining centralized standards with domain-level execution, federated models ensure data quality, compliance, and trust—without sacrificing speed or flexibility. dbt helps operationalize this model by embedding governance directly into development workflows. Domain teams can create and own production-grade data products, while central teams maintain visibility through lineage, testing, documentation, and contracts—all version-controlled and CI/CD-enabled. Tools like the [dbt VS Code extension](https://docs.getdbt.com/docs/install-dbt-extension) enhance this experience further. They give practitioners real-time linting, model suggestions, and inline error detection—all within their preferred development environment. This makes it easier for teams to build governed data products confidently and efficiently. As data ecosystems grow, federated governance with dbt ensures that agility scales with accountability. --- --- title: "Here’s why you (and your team) should attend Coalesce 2025" description: "Coalesce 2025 is the ultimate offsite for your team." url: "https://www.getdbt.com/blog/here-s-why-you-and-your-team-should-attend-coalesce-2025" date: "2025-08-05" authors: ["Daniel Poppy"] categories: ["Community"] --- # Here’s why you (and your team) should attend Coalesce 2025 [**Coalesce 2025 **is approaching quickly: Oct. 13-16 in Las Vegas](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/summary)**.** This is the premier event for data teams to come together to rewrite the future of data work, but if you haven’t been before, you might not know what to expect. Join your peers to share the ideas, learn about the latest innovations changing how data gets done, and learn how dbt can help you advance your career. Coalesce is the highlight of the year for the dbt community. Here are a few reasons to attend Coalesce 2025 in Las Vegas: ## Reason #1: Connect with the data community This is your chance to be in person with the leading minds doing data work. Connect in-person with the people you only know from Slack, or get back in touch with old friends you haven't seen in a year. Coalesce is where everyone in the dbt community gets together to swap ideas, celebrate wins, and talk data & AI. > “There’s a real electricity here.” — Coalesce 2024 attendee. The data world is growing fast. Here’s your chance to make the human connection with the community that you’ve chosen to be a part of. These are your people. ![Coalesce 2024 attendees having lunch](https://cdn.sanity.io/images/wl0ndo6t/main/83257d5c08b18f59434f9be7561397b74dd97438-4032x3024.jpg) ## Reason #2: Solve problems together Year after year, we hear feedback that the [**Coalesce peer exchanges**](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda) are some of the most powerful and important components of the week. These small-group forums bring the data community together to tackle some of the most pressing challenges in analytics work today. Wherever you are in your data journey, there’s a good chance someone at Coalesce has been there and is ready to swap notes. > “Coalesce is about solving data problems together. The problems you try to solve involve people and culture and community. It’s important to put ‌humanity into data work.” — Coalesce 2024 attendee. Here are some of the 2025 peer exchanges (so far): - Analysts-on-dbt: A collaborative forum for analytics practitioners - Iceberg, the truth and the hype of multi-engine data stacks - Onboarding junior analytics engineers: Lessons, pitfalls, and practical strategies - Test smarter, not harder: A peer exchange ## Reason #3: Discover the tools and learn the practices that are rewriting data work Breakout sessions, training, hands-on labs, mingle, and repeat. Coalesce is unlike any other data gathering in its opportunities to learn, get hands-on practice, and leave with skills you can apply immediately. Join any of the [**70+ of breakout sessions**](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/agenda) on analytics development best practices, embracing AI, fostering data quality and trust, empowering self-service analytics, building dbt at scale, and bringing data modernization to your organization. This is an unparalleled chance to hear directly from data experts who are rewriting how teams collaborate at enterprise scale, who gets to build with data, and how trust is earned in data. > “It’s really helpful to see how others are using features and sharing ideas for ways to overcome the common challenges we all face.” — Coalesce 2024 attendee. [**Sign up for training**](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/training-certification) with dbt experts, and leave with skills that you and your team can apply immediately: - Getting started with dbt - Becoming a dbt Architect - Developing with dbt Canvas - Cross-platform Mesh with Iceberg Tables - Managing costs with dbt - Building a data quality framework with dbt - Upgrading to Fusion Join hands-on labs with dbt resident architects, solutions architects and technical instructors. We provide the sandbox, and you get hands-on experience with the new dbt Fusion Engine, dbt Copilot, dbt Semantic Layer, dbt Canvas, and analytics workflows in dbt. And if you’re ready, get your [**dbt Analytics Engineering Certification**](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/training-certification) or your **[dbt Architect Certification](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/training-certification).** ## Reason #4: Coalesce is the ultimate data team offsite Data teams that come to Coalesce together leave as better data teams. [And if you register 3+ people in one order, you automatically** save 30% on Coalesce registration**](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/regProcessStep1?rp=1d16b637-27b7-4986-b153-7c3e5f6d667f). No awkward trust falls. Just you and your team learning from others' successes and failures and gaining skills that directly benefit your team and organization. This is your chance to make your team feel connected with each other and to the larger data community that they're a part of. The conference is built for teams to divide and conquer during breakout sessions, share learnings at breaks, and regroup energized for the epic parties and events. Coalesce is fun. And it’s designed to be fun for all sorts. It’s Las Vegas after all. You're probably the dbt champion at your organization. You can also be the champion of your data team by bringing them to the best place for a data team to be. ![Coalesce 2024 attendees at the afterparty ](https://cdn.sanity.io/images/wl0ndo6t/main/3fa84a64ebdc3416c5cef21284ff252336954495-4000x2667.jpg) ## Reason #5: Experience a community that wants you to win > “The dbt community is my second home. Everyone is kind, everyone is helpful, everyone wants you to succeed. And you want them to succeed too.” — Coalesce 2024 attendee. There’s nothing else like the dbt community in the data space. You may have joined a dbt meetup. You may be active in dbt Slack. But the place where the dbt community really shows up is at Coalesce. And how the community shows up is by helping each other grow. Join the Women in Data session to connect and share notes with folks on their journeys through tech. Grab a seat at a product roundtable and build with your peers. Celebrate the wins of the dbt community this year, and come together to point where we head next. This is your community. This is your Coalesce. Rewrite how data work gets done. ## [**Register for Coalesce 2025 now.**](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/regProcessStep1?rp=1d16b637-27b7-4986-b153-7c3e5f6d667f) ![Community wall at Coalesce 2024 that reads, "How has dbt transformed your career and your life?"](https://cdn.sanity.io/images/wl0ndo6t/main/7b569dce1e79d43433498e891d25acc1e0adcf58-2854x3461.jpg) --- --- title: "From stored procedures to dbt: A modern migration playbook" description: "Stored procedures are holding back your data teams. Here’s how to migrate to dbt for better scalability, reliability, and quality." url: "https://www.getdbt.com/blog/stored-procedures-dbt-migration-playbook" date: "2025-08-04" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # From stored procedures to dbt: A modern migration playbook For decades, stored procedures served as the backbone of enterprise data transformation. They were fast, efficient, and kept everything contained within a single database. However, as data teams embrace cloud-native architectures and modern analytics practices, these once-reliable workhorses are showing their age. That’s why more companies are migrating out of stored procedures to more transparent, cross-platform data transformation systems. The shift from legacy stored procedures to modern transformation tools represents more than just a technology upgrade. It's a fundamental change in how organizations approach data transformation, collaboration, and governance. Companies are discovering that migrating from stored procedures to dbt not only solves immediate technical challenges but unlocks new levels of productivity, transparency, and data quality. We’ll look at why stored procedures fall short in the modern age, how dbt provides a new model for managing data transformations consistently across your enterprise, and how to conduct a successful migration from the old world to the new. **** ## Why stored procedures fall short in modern data stacks Stored procedures were designed for a different era of data management. However, before cloud data platforms and modern [Extract, Load, and Transform (ELT) pipelines](https://www.getdbt.com/blog/extract-load-transform), they delivered a number of benefits by: - Condensing complex transformations into repeatable, callable units of logic - Reducing data movement and improving performance - Centralizing logic within a single database platform, such as SQL Server, Oracle, or Teradata, which simplifies data governance However, as data stacks evolved toward cloud-native architectures, cracks started to appear in the stored procedure story. ### The black box problem Perhaps the most significant challenge with stored procedures is their opacity. These transformation processes operate as closed systems where the logic remains hidden from most team members. When business stakeholders ask where a particular metric comes from or how it's calculated, teams often struggle to provide clear explanations because the logic is buried deep within procedural code. This lack of transparency creates bottlenecks when changes are needed. If the original developer leaves the organization, institutional knowledge walks out the door with them. New team members face steep learning curves trying to understand undocumented transformation logic, often leading to costly delays and potential errors. ### Missing modern development practices Stored procedures predate many of the software engineering practices that modern data teams consider essential. Version control, continuous integration, automated testing, and collaborative code review—all standard practices in software development—are difficult or impossible to implement with traditional stored procedures. Without Git-based version control, teams can't easily track changes, revert problematic updates, or collaborate effectively on transformation logic. The absence of automated testing means data quality issues often go undetected until they've already impacted downstream systems or business decisions. ### Vendor lock-in and portability challenges Stored procedures are typically written in database-specific languages like [T-SQL for SQL Server](https://learn.microsoft.com/en-us/sql/t-sql/language-reference?view=sql-server-ver17) or [PL/SQL for Oracle](https://www.oracle.com/database/technologies/appdev/plsql.html). This tight coupling to particular platforms creates significant migration challenges when organizations want to modernize their data infrastructure or switch to cloud-based solutions. Teams that have invested heavily in stored procedure development often find themselves trapped by this vendor lock-in. They can’t leverage more cost-effective or feature-rich cloud data platforms without rewriting their transformation logic extensively. ### Security and governance gaps Traditional stored procedures often lack the granular permissioning and governance capabilities that modern data teams require. [Overprivileged access](https://www.oasis.security/glossary/overprivileged) becomes common because fine-grained control is difficult to implement and maintain. This creates compliance risks and makes it challenging to implement proper data governance practices. ## dbt: Managing data at scale All of these factors make it hard - if not impossible - to scale data projects. And that leaves teams wrestling with common problems, such as figuring out why key metrics are different across tools. Modern data teams require a modern solution that doesn’t necessitate starting from scratch with every project, provides improved data quality and trust, and enables cost optimization. dbt manages this complexity in a way that’s modular, scalable, repeatable, and governed—all from directly inside your data platform. Data teams can use dbt to transform data into clean and prepared datasets that are ready to power every downstream use case. With dbt, you can be confident that your data is accurate, consistent, and shipped to downstream teams with agility. dbt enables this by supporting numerous features that stored procedures lack: ### Transparency and documentation Unlike stored procedures, dbt transformations are written as [data models](https://docs.getdbt.com/docs/build/models) in SQL or Python. These models are version-controlled, documented, and easily understood by both technical and business teams. Every transformation includes metadata about its purpose and the columns it creates. This enables tracing [data lineage](https://www.getdbt.com/blog/what-is-data-lineage), the journey that data takes as it travels throughout your company. This transparency extends to the automatically generated [documentation](https://docs.getdbt.com/docs/build/documentation) that dbt creates, providing a searchable catalog of all data assets, their relationships, and their business context. Teams can finally answer questions like "where does this metric come from?" with confidence and clarity. ### Modern development workflow Delivering accurate analytics is more about writing code. Analytics code must be developed, tested, deployed, and monitored like any other piece of software within your company. dbt enables the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle), bringing proven software engineering practices to data transformation. Teams can implement Git-based version control, [automated testing](https://docs.getdbt.com/docs/build/data-tests), [automated deployment](https://docs.getdbt.com/docs/deploy/continuous-integration), and collaborative code review as standard parts of their workflow. This dramatically reduces the risk of introducing errors into production systems while enabling faster iteration and more reliable deployments of data transformation code. **** ### Platform flexibility Because dbt compiles to standard SQL, organizations aren't locked into specific database platforms. The same transformation logic can run on [Snowflake](https://www.snowflake.com/), [BigQuery](https://cloud.google.com/bigquery), [Databricks](https://www.databricks.com/), or other modern cloud data warehouses, providing flexibility as business needs evolve. ## dbt vs. stored procedures in action How, exactly, does dbt code differ from stored procedure code? For the nitty-gritty details, read our article that shows step-by-step [how to implement a standard stored procedure within dbt](https://docs.getdbt.com/blog/migrating-from-stored-procs). We also give an extensive, hands-on example in our recent webinar on stored procedure migration. You can [register and watch it for free here.](https://www.getdbt.com/resources/webinars/from-stored-procedures-to-dbt-your-playbook-for-successful-migrations) ## A practical migration playbook You can begin using dbt to implement new workloads, no matter your current data architecture. The trickier part is migrating your existing code from stored procedures into dbt. Odds are you don’t have only a handful of stored procedures running. Most companies have code that’s been running across multiple data warehouses for years, if not decades. That’s why migration requires a structured approach that balances speed of delivery with long-term maintainability. Based on real-world migration experiences, successful teams follow a five-phase methodology. ### Discovery and assessment The first step involves cataloging existing stored procedures and understanding their scope, complexity, and business criticality. Teams should categorize procedures by business domain, frequency of use, dependencies, and technical complexity to create a manageable framework for migration planning. This phase also includes identifying procedures that may no longer be needed. Legacy systems often accumulate transformations that are no longer used or have been superseded by other processes. Eliminating unnecessary procedures from the migration scope can significantly reduce project complexity. ### Prioritization strategy Rather than attempting to migrate everything at once, successful migrations focus on high-impact, low-complexity procedures first. This approach delivers quick wins that demonstrate value while giving teams time to develop expertise with dbt. Procedures that change frequently are often ideal early candidates because they cause ongoing maintenance overhead in the legacy system. Similarly, transformations that support critical business reports or have high stakeholder visibility can provide compelling early wins. ### Migration design choices Teams must decide between rewriting transformation logic versus refactoring existing procedures. Complete rewrites make sense when business requirements have changed, when legacy logic is incompatible with new platforms, or when the existing code is poorly structured or documented. Refactoring works well when the existing logic is sound and the output meets current business needs. In these cases, teams can extract the core business logic into dbt-friendly SELECT statements while leveraging dbt's built-in capabilities for testing, documentation, and dependency management. The key principle is avoiding direct copy-and-paste approaches. While it might seem efficient to replicate stored procedure logic exactly, this approach brings forward technical debt and fails to take advantage of dbt's modern capabilities. ### Building with best practices Successful migrations establish strong foundations from the beginning rather than trying to retrofit best practices later. This includes implementing consistent naming conventions, [folder structures](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview), testing strategies, and documentation standards. Data engineers should leverage dbt's built-in testing capabilities rather than building custom validation logic. The platform includes [common data quality checks](https://www.getdbt.com/blog/data-quality-checks) for null values, uniqueness, referential integrity, and accepted values. [Community packages](https://hub.getdbt.com/) extend these capabilities even further, often eliminating the need for custom test development. Modularizing your code into [packages](https://docs.getdbt.com/docs/build/packages) and small units of execution is crucial for maintainability. Breaking monolithic stored procedures into discrete, focused dbt models makes the transformation logic easier to understand, test, and modify over time. ### Test, validate, and iterate Migration success depends on thoroughly validating that new dbt models produce the same results as legacy stored procedures. Often, this results in more testing than the original stored procedures received in the first place! With dbt, testing begins during development. Using [dbt Fusion](https://www.getdbt.com/product/fusion), data engineers can run tests on their dbt models without even wiring up a connection to a data warehouse. The data engineering team can set up a [continuous integration process](https://docs.getdbt.com/docs/deploy/ci-jobs) that ensures all code changes are reviewed before they’re approved for production. They can also create a promotion pipeline with [multiple environments](https://docs.getdbt.com/docs/deploy/deploy-environments), so that all tests are run against an isolated environment (e.g., a staging database) before promoting changes to production. ## Real-world success: B2B SaaS transformation A large B2B SaaS company recently demonstrated the power of this migration approach when facing unprecedented growth that stressed its legacy infrastructure to the breaking point. Their stored procedure-based transformation system couldn't handle increased data volumes and began crashing frequently. This led to performance issues that affected customer-facing systems and degraded business decision-making capabilities. That led to an overall loss of trust in data. #### The modernization solution The company asked data and AI services vendor [phData](https://www.phdata.io/) to help them create a comprehensive modernization strategy centered on dbt as the transformation layer. phData helped them migrate their data warehouse to Snowflake for scalable compute and advanced analytics capabilities, while implementing dbt as the heartbeat of their transformation pipeline. This approach enabled several key improvements: **Scalability and reliability**: Using systems like the dbt platform and Snowflake, the customer could more easily architect, deploy, and monitor new data products. No one was waking up at midnight any longer worried that everything had failed - the new systems just ran. **Separation of environments:** Development, staging, and production environments could be properly isolated, reducing the risk of development work affecting production systems. **Version control and collaboration:** phData and the customer moved all transformation logic into Git, enabling proper code review, change tracking, and collaborative development practices. **Automated testing and deployment:** Comprehensive testing ensured data quality while automated deployment processes reduced manual errors and deployment time. **Improved troubleshooting:** When issues occurred, teams could quickly identify root causes and implement fixes using dbt's lineage and testing capabilities. ### Measurable outcomes The migration delivered significant measurable benefits within 12 months: **Unified data ingestion:** Multiple disparate data sources were consolidated into a coherent, well-documented data pipeline that provided richer customer and prospect insights. **Enhanced team capabilities:** The modern tooling enabled better onboarding processes and allowed team members to work within their domains of expertise while following consistent best practices. **Improved decision-making:** The customer and phData were able to build features such as lead scoring and prioritization that rebuilt customer trust. **Cost optimization:** The new architecture lowered the total cost of ownership (TCO) while providing better performance and reliability. **Reduced operational overhead:** Teams no longer woke up wondering if pipelines had run successfully—the new system just worked reliably. ## Getting started with dbt migration Moving off stored procedures isn’t a matter of if - it’s a matter of when. The old methods of data transformation don’t scale to handling data in the age of AI. dbt provides a centralized, vendor-neutral platform for data modeling that follows software engineering best practices. This enables data teams to ship new and updated data products quickly and with high quality. To get started, create a free dbt account today and step through [our tutorial on migrating from DDL, DML, and stored procedures](https://docs.getdbt.com/guides/migrate-from-stored-procedures?step=1) to get a feel for how everything works. --- --- title: "The dbt Fusion engine public beta is now available on Redshift" description: "The dbt Fusion engine is now in public beta for Redshift—experience faster workflows, smarter orchestration, and cost savings." url: "https://www.getdbt.com/blog/dbt-fusion-engine-public-beta-redshift" date: "2025-08-04" authors: ["Azzam Aijazi"] categories: ["Product"] --- # The dbt Fusion engine public beta is now available on Redshift We’re excited to share that teams using Amazon Redshift can now participate in the public beta of the [dbt Fusion engine](https://www.getdbt.com/product/fusion). This brings the number of [supported data platforms](https://docs.getdbt.com/docs/fusion/supported-features) to four: BigQuery, Databricks, Snowflake, and now Redshift. This marks another [key step toward general availability](https://docs.getdbt.com/blog/dbt-fusion-engine-path-to-ga) of the dbt Fusion engine. Redshift has long been one of the most widely used data warehouses across the dbt community, and now teams on Redshift can experience the power, speed, and intelligence of Fusion. [Watch video](https://youtu.be/NiNkdThkKAI?si=Cgyjd0lKsoYt9YvU) ## The dbt Fusion engine now supports Redshift The dbt Fusion engine isn’t just a small iteration. [It’s a complete re-architecture and re-imagining of dbt](https://www.getdbt.com/blog/where-we-re-headed-with-the-dbt-fusion-engine). Built in Rust (a language known for performance) and capable of deep SQL comprehension, Fusion enables blazing-fast development workflows, real-time intelligence, and smarter orchestration of your data transformations. What does that mean for Redshift users? It means: - **Significant performance gains**. Parse performance is up to 30x faster than with dbt Core, meaning developers can iterate more rapidly, stay in flow, and reduce time to insight. - **A vastly improved development experience.** Fusion doesn’t just pass your SQL to your warehouse: it also _understands_ it. With native support for Redshift SQL, developers get real-time feedback, error checking, and intelligent suggestions as they write code. All of this happens without needing to query your warehouse. Excitingly, our powerful new [dbt VS Code extension](https://docs.getdbt.com/docs/install-dbt-extension) brings the full power of Fusion to your local development experience as well. > **“The Fusion engine with VS Code Extension is how folks will want to develop with dbt moving forward. This is awesome. It’s the experience we’ve all been waiting for.” —** Bruno Souza de Lima, Lead Data Engineer at pHData - **Built-in cost optimization**. Fusion introduces state-aware orchestration, which allows dbt to understand what’s _changed_ in your project and intelligently run only the necessary models. That means fewer redundant workloads, unlocking a reduction of ~10% in cloud compute spend. ![Fusion enhanced awareness](https://cdn.sanity.io/images/wl0ndo6t/main/274fbc3f87b1e3f8bcb6f5dd1eaa3a662cbab5b4-1630x878.png) - **Enhanced governance (coming soon)**. In the coming months, Fusion will also get support for sensitive data classification and tracking, allowing for even more robust data governance with dbt. This will help teams manage risk, enforce policies, and lay the groundwork for safe AI adoption and regulatory compliance. ## Who benefits from using Fusion on Redshift? Data developers and data analysts using Redshift with dbt will notice a dramatic boost. Fusion helps teams iterate faster, boost data quality, and scale confidently. Meanwhile, organizations prioritizing cost optimization or data governance will benefit from Fusion’s deep metadata awareness and more intelligent orchestration. > **“We anticipate the dbt Fusion engine will mark a new chapter for our data team — one where speed and efficiency are baked into every part of the analytics lifecycle.” —** Matt Karan, Senior Data Engineer at Obie Insurance ## Getting started with Fusion on Redshift Fusion is currently in public beta, which means: - We’re actively collecting feedback to help us improve it - You can expect the experience to continue to evolve and get more refined - You can enable or disable it at any time ([if your project(s) are eligible](https://docs.getdbt.com/docs/fusion/supported-features)) Note that, at this time, username + password is the only supported authentication method on Redshift. Expect more additions here soon. Ready to see it in action? If you’re a managed dbt customer and your Redshift project is eligible, you can start [upgrading to Fusion](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-fusion) now. If you’re not, you can also use Fusion locally through the [dbt VS Code Extension](https://docs.getdbt.com/docs/install-dbt-extension) or by downloading the permissively licensed Fusion source code and binary. ## See what’s next The best place to experience Fusion in action, and hear directly from customers and dbt Labs leadership, is [**Coalesce 2025**](https://coalesce.getdbt.com/) (Coming Oct 13-16). Join us to explore how the dbt Fusion engine is transforming the dbt developer experience and reshaping modern data workflows. Lock in your spot today. --- --- title: "The pragmatic guide to AI agents in the enterprise" description: "Demystifying AI agents with Sean Falconer, Confluent's senior director of AI strategy." url: "https://www.getdbt.com/blog/pragmatic-guide-to-ai-agents-enterprise" date: "2025-08-03" authors: ["Daniel Poppy"] categories: ["Insights"] --- # The pragmatic guide to AI agents in the enterprise _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/the-pragmatic-guide-to-ai-agents). _ What does it mean to be agentic? Is there a spectrum of agency? In this episode of The Analytics Engineering Podcast, Tristan Handy talks to Sean Falconer, senior director of AI strategy at Confluent, about AI agents. They discuss what truly makes software "agentic," where agents are successfully being deployed, and how to conceptualize and build agents within enterprise infrastructure. Sean shares practical ideas about the changing trends in AI, the role of basic models, and why agents may be better for businesses than for consumers. This episode will give you a clear, practical idea of how AI agents can change businesses, instead of being a vague marketing buzzword. _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways ### **Sean, can you give us the TLDR on your career and what you're working on today?** **Sean Falconer:** I've always worked at the intersection of data, engineering, and AI. From academia studying computer science, into industry as a founder, then to Google, I worked on conversational systems and privacy/security in AI. Currently, at Confluent, I'm leading our AI product strategy, balancing both technical and go-to-market roles. ### **You moved from being deeply technical into marketing and sales. What drove that transition?** I was forced into it as a founder. Initially uncomfortable, but it taught me huge respect for marketing and sales. I had to learn by making many mistakes, eventually building out entire marketing and sales functions. I realized how challenging and critical these roles are. ### **You were at Google before ChatGPT launched. Did you foresee the transformative nature of these technologies?** Honestly, no. Having seen earlier disappointments in conversational AI (like Microsoft's Alice), I was skeptical initially, even as ChatGPT emerged. It wasn’t obvious we'd soon experience this revolution. ### **You’ve written about three waves of AI. Can you describe these?** Yes. Wave one was predictive AI, traditional ML models trained for specific tasks like fraud or spam detection—effective but rigid. Wave two introduced generative AI, or foundation models, trained on vast general datasets, flexible but lacking specific business context. The third wave, agentic AI, involves AI systems that can reason, dynamically choose tasks, gather information, and perform actions as a more complete software system. ### **Do foundation models replace traditional ML methods?** Sometimes they can, but it doesn’t always make sense. An LLM might do sentiment analysis well enough, but a traditional model may be more efficient and cheaper. Think of using an LLM as cutting steak with a chainsaw—possible, but unnecessary. ### **Let's clarify "agents." What makes software truly agentic?** It’s software that can dynamically decide its own control flow: choosing tasks, workflows, and gathering context as needed. Realistically, current enterprise agents have limited agency to ensure reliability. They're mostly workflow automations rather than fully autonomous systems. ### **You mentioned a spectrum of agency. Is this similar to autonomy in self-driving cars?** Exactly. Highly autonomous agents are appealing but not practical yet. Most enterprise success stories involve structured workflows with clearly defined boundaries. ### **Why have agents taken off more in enterprises than consumer apps?** Enterprises have many well-defined, high-value tasks perfect for automation. Consumer scenarios demanding high agency—like planning complex trips—are still too unreliable. Enterprises can benefit significantly even from limited agentic capability. ### **Is an agent just a microservice?** In many ways, yes. An agent functions like a microservice with extra capabilities (using LLMs for decisions). Deployment considerations like state management and long-running tasks differ slightly, but fundamentally it’s similar. ### **What tools and frameworks help build effective agents?** Start with frontier models like GPT-4 or Claude. Frameworks include LangChain, Microsoft Autogen, and CrewAI. But for real-world deployment, treat it as rigorous software engineering with observability, scalability, and robustness in mind. ### **Are organizational barriers bigger than technical challenges?** Yes. AI efforts are often mistakenly tasked to data science teams rather than cross-functional software teams. Successful companies create dedicated teams blending software engineering skills and data expertise to build reliable agentic systems. ### **What pitfalls should teams avoid?** Avoid building monolithic agents. Break systems into smaller, well-defined units in a multi-agent architecture. Use event-driven frameworks to avoid rigid, hard-to-maintain dependencies. ## Chapters - [00:00] Introduction: What's all the hype about agents? - [01:10] Meet Sean Falconer: A journey from engineer to AI strategist - [04:10] Learning marketing as an engineer-founder - [05:50] Inside Google's AI efforts before ChatGPT - [09:00] What does it mean to run AI strategy? - [10:45] Three waves of AI: Predictive, Generative, and Agentic - [16:30] Will foundation models replace traditional ML? - [18:30] Defining agents clearly: Beyond the buzzword - [22:00] The spectrum of agency: From controlled workflows to open-ended tasks - [25:30] Why agents fit better in enterprises than consumer apps - [28:00] Agents as microservices: A practical view - [35:00] What tech stack is needed to build effective agents? - [37:50] Organizational challenges in adopting agents - [39:30] Models that are favorites for developers - [43:30] Why software engineers are best placed to build agents - [46:00] The technical stumbling blocks in building agents - [48:00] Concluding thoughts: Beyond POCs to production agents --- --- title: "Who should own the data transformation layer?" description: "Compare data team ownership, analytics ownership and hybrid models for modern data systems." url: "https://www.getdbt.com/blog/data-transformation-layer-ownership" date: "2025-08-01" authors: ["Joey Gault"] categories: ["Pulse"] --- # Who should own the data transformation layer? The data transformation layer serves as the bridge between raw data storage and business intelligence. It's where data cleaning, validation, modeling, and business logic implementation occur. This layer ensures that downstream consumers—whether analysts, data scientists, or business users—have access to consistent, trustworthy datasets. In modern [ELT architectures](https://www.getdbt.com/blog/extract-load-transform), transformation happens after data has been loaded into the warehouse, leveraging the computational power of cloud platforms like Snowflake, BigQuery, or Databricks. This approach has democratized data transformation, making it more accessible to teams beyond traditional data engineering. However, this accessibility has also created new questions about ownership and governance. The transformation layer encompasses several critical functions: data quality assurance through testing and validation, business logic implementation that reflects organizational definitions and rules, performance optimization to ensure efficient query execution, and documentation that makes data assets discoverable and understandable. Each of these functions requires different skill sets and organizational perspectives, which complicates the ownership question. ## The case for data engineering ownership [Data engineering teams](https://www.getdbt.com/blog/analytics-engineer-vs-data-analyst-vs-data-engineer) have traditionally owned the transformation layer, and there are compelling reasons for this arrangement. Data engineers possess deep technical expertise in building scalable, reliable data pipelines. They understand the intricacies of data warehouse optimization, can implement complex performance tuning strategies, and have experience managing production systems at scale. From an operational perspective, data engineers are well-positioned to ensure transformation pipelines run reliably. They can implement proper monitoring, alerting, and error handling mechanisms. They understand how to design systems that can handle increasing data volumes and complexity without degrading performance. This operational expertise becomes crucial as organizations scale their data operations. [Data engineers also bring software engineering best practices to the transformation layer.](https://www.getdbt.com/resources/the-analytics-development-lifecycle) They can implement version control, automated testing, and CI/CD workflows that ensure code quality and deployment reliability. These practices become increasingly important as transformation logic grows in complexity and business criticality. However, data engineering ownership can create bottlenecks. When business users need new metrics or modifications to existing logic, they must communicate requirements to data engineers, who then implement and deploy the changes. This process can slow down analytics iterations and reduce the agility that modern businesses require. ## The case for analytics engineering ownership The [emergence of analytics engineering](https://www.getdbt.com/blog/what-is-analytics-engineering) as a discipline has created a compelling alternative ownership model. Analytics engineers combine technical skills with deep business understanding, making them natural owners of the transformation layer. They understand both the technical requirements of building reliable data pipelines and the business context that drives transformation logic. Analytics engineers typically work more closely with business stakeholders than traditional data engineers. They understand the nuances of business metrics, the context behind data requirements, and the trade-offs involved in different modeling approaches. This business context is crucial for building transformation logic that truly serves organizational needs. Tools like [dbt](https://www.getdbt.com/product/what-is-dbt) have made analytics engineering more accessible and powerful. Analytics engineers can build sophisticated transformation pipelines using SQL, implement comprehensive testing strategies, and maintain detailed documentation—all while working within frameworks that enforce software engineering best practices. This combination of business understanding and technical capability makes analytics engineers strong candidates for transformation layer ownership. The analytics engineering model also promotes faster iteration cycles. When business requirements change, analytics engineers can quickly modify transformation logic without the communication overhead that often exists between business users and data engineering teams. This agility can be a significant competitive advantage in fast-moving business environments. **** ## The distributed ownership model Some organizations adopt a [distributed ownership model](https://www.getdbt.com/blog/data-governance-ownership) where different teams own different aspects of the transformation layer. In this approach, data engineers might own the foundational infrastructure and core data pipelines, while analytics engineers or domain experts own business-specific transformation logic. This model can work well in larger organizations with mature data teams. Data engineers focus on building reliable, scalable infrastructure that supports transformation workloads. They implement the underlying systems, monitoring, and operational processes that keep the transformation layer running smoothly. Meanwhile, analytics engineers or domain experts build the business logic that sits on top of this infrastructure. The distributed model requires strong governance frameworks to succeed. Organizations need clear standards for code quality, testing, documentation, and deployment processes. They need well-defined interfaces between different layers of the transformation stack. Without these governance mechanisms, distributed ownership can lead to inconsistency, technical debt, and operational challenges. Domain-specific ownership can be particularly effective in organizations implementing data mesh architectures. Different business domains can own their transformation logic while adhering to organizational standards for data quality, security, and governance. This approach can scale well as organizations grow, but it requires significant investment in governance and platform capabilities. ## Factors influencing ownership decisions Several organizational factors should influence transformation layer ownership decisions. Team size and structure play a crucial role. Smaller organizations might not have the luxury of specialized analytics engineering teams, making data engineering ownership more practical. Larger organizations might benefit from distributed models that leverage domain expertise. Technical maturity is another critical factor. Organizations with mature data engineering practices and robust operational capabilities might be better positioned for data engineering ownership. Organizations with strong analytics cultures and business-savvy technical teams might benefit from analytics engineering ownership. The complexity and criticality of transformation logic also matter. Simple transformations might be suitable for broader ownership, while complex business logic might require specialized expertise. Mission-critical transformations that directly impact business operations might need the operational rigor that data engineering teams provide. Business requirements and agility needs should also influence ownership decisions. Organizations that need rapid iteration on analytics might benefit from analytics engineering ownership. Organizations with more stable requirements and higher emphasis on operational reliability might prefer data engineering ownership. ## The role of tooling in ownership decisions Modern transformation tools like dbt have significantly influenced ownership discussions by making transformation development more accessible while maintaining engineering rigor. dbt enables teams to build transformation pipelines using familiar SQL syntax while incorporating software engineering best practices like version control, testing, and documentation. The [accessibility of dbt](https://www.getdbt.com/product/dbt) has enabled analytics engineers and even business analysts to take ownership of transformation logic. The tool's emphasis on modularity and reusability makes it easier to maintain complex transformation projects across different teams. Its built-in testing and documentation capabilities help ensure quality regardless of who owns the code. However, tooling alone doesn't solve ownership challenges. Organizations still need clear governance frameworks, operational processes, and skill development programs. The choice of tooling should support the chosen ownership model rather than dictate it. ## Governance and collaboration framework Regardless of the ownership model chosen, successful transformation layer management requires strong [governance and collaboration frameworks](https://www.getdbt.com/blog/data-governance-framework). These frameworks should define standards for code quality, testing, documentation, and deployment processes. They should establish clear interfaces between different teams and systems. Governance frameworks should also address data quality standards, security requirements, and compliance obligations. They should define processes for handling schema changes, managing dependencies, and coordinating deployments across different teams and systems. Collaboration frameworks are equally important. Teams need clear communication channels, shared understanding of business requirements, and aligned incentives. Regular reviews of transformation logic, performance metrics, and business outcomes help ensure that the transformation layer continues to serve organizational needs effectively. ## Making the ownership decision The question of who should own the data transformation layer doesn't have a universal answer. The right choice depends on organizational context, team capabilities, business requirements, and strategic priorities. However, several principles can guide the decision-making process. First, consider the skills and capabilities of available teams. The transformation layer requires both technical expertise and business understanding. The owning team should have or be able to develop both capabilities. Second, think about operational requirements. The transformation layer is often business-critical infrastructure that requires reliable operation, monitoring, and maintenance. Third, consider agility requirements. How quickly does the organization need to iterate on transformation logic? How important is self-service capability for business users? These requirements might favor ownership models that reduce communication overhead and enable faster iteration. Finally, think about long-term scalability. As the organization grows, will the chosen ownership model continue to work effectively? Can it handle increasing data volumes, complexity, and team sizes? The ownership model should be sustainable as the organization evolves. The most successful organizations often start with clear ownership assignments but maintain flexibility to evolve their approach as they learn and grow. They invest in governance frameworks, tooling, and skill development that support their chosen model while enabling future adaptations. Ultimately, the goal isn't to find the perfect ownership model but to establish clear accountability, enable effective collaboration, and ensure that the transformation layer reliably serves business needs. With the right combination of people, processes, and tools, various ownership models can succeed in delivering high-quality, trustworthy data products that drive business value. ## Data transformation layer FAQs **What is a data transformation layer?** The data transformation layer serves as the bridge between raw data storage and business intelligence. It's where data cleaning, validation, modeling, and business logic implementation occur to ensure that downstream consumers—whether analysts, data scientists, or business users—have access to consistent, trustworthy datasets. This layer encompasses several critical functions including data quality assurance through testing and validation, business logic implementation that reflects organizational definitions and rules, performance optimization to ensure efficient query execution, and documentation that makes data assets discoverable and understandable. **What factors should guide the choice between real-time and batch processing in the processing layer?** Several organizational factors should influence processing decisions, including team size and structure, technical maturity, and business requirements. The complexity and criticality of transformation logic also matter: simple transformations might be suitable for broader processing approaches, while complex business logic might require specialized handling. Mission-critical transformations that directly impact business operations might need more operational rigor, while organizations that need rapid iteration on analytics might benefit from more agile processing approaches. Business requirements and agility needs should also guide decisions, with organizations having more stable requirements potentially preferring different approaches than those needing frequent changes. **What defines a really strong data transformation layer in 2025?** A strong data transformation layer in 2025 requires robust governance and collaboration frameworks that define standards for code quality, testing, documentation, and deployment processes. It should establish clear interfaces between different teams and systems while addressing data quality standards, security requirements, and compliance obligations. The layer should combine both technical expertise and business understanding, with reliable operation, monitoring, and maintenance capabilities. Modern transformation tools enable teams to build pipelines using familiar syntax while incorporating software engineering best practices like version control, testing, and documentation. Most importantly, it should be scalable and flexible enough to handle increasing data volumes, complexity, and team sizes while maintaining the ability to iterate quickly on business requirements. --- --- title: "MCP Servers: How to prepare your data" description: "A guide to how MCP Servers work and how to curate data ready for real-time data access via agentic AI." url: "https://www.getdbt.com/blog/mcp-servers" date: "2025-07-31" authors: ["Daniel Poppy"] categories: ["Learn"] --- # MCP Servers: How to prepare your data [Alexandr Wang](https://fortune.com/2024/05/21/scale-ai-funding-valuation-ceo-alexandr-wang-profitability/), founder of Scale AI, notes that even the most advanced AI systems perform only as well as the quality of data they process. This makes data preparation a crucial step for the success of agentic AI. Without clean, accessible, and contextual information, even the most sophisticated AI agents deliver unreliable outcomes. Model Context Protocol (MCP) servers address this problem by providing a secure, standardized interface that connects AI agents to trusted data models. Moreover, they offer a consistent and reliable access layer providing clean, accessible, and contextual information essential for agentic AI. In this article, we cover the key data preparation principles for MCP servers and show how dbt can ensure data is ready for this new autonomous AI era. ## Why agentic AI demands new data standards Agentic AI is a set of intelligent systems powered by autonomous agents that perceive, reason, and reach goals with minimal human help. These agents make decisions, run workflows, and adapt in real time based on changing environments and inputs. Unlike traditional AI, these systems reshape enterprise operations through autonomous decision-making, reasoning, and goal-driven action. That goes beyond traditional AI applications, which primarily analyze but leave actions up to users. AI agents actively run multi-step workflows, call APIs, and dynamically adapt to evolving inputs with minimal human intervention. For example, in manufacturing, an AI agent can identify a fault, order necessary parts, and adjust production, all while handling the entire response with minimal human intervention. This change requires redesigning data strategies beyond conventional pipelines to accommodate agentic systems. ### Beyond traditional pipelines The core distinction between agentic AI and earlier AI systems, such as rule-based bots, lies in their real-time, contextual access requirements. Unlike static models or rule-based systems, agentic AI operates in dynamic environments where it must make decisions within milliseconds. Legacy data standards fall short because agentic AI operates independently, adapts in real time, and learns from experience. This adaptability and continuous learning require immediate data ingestion and processing, a capability that is largely absent in legacy data infrastructures. For example, an AI agent managing fleet logistics needs instant traffic and weather data to reroute deliveries. Traditional pipelines can’t deliver such data in real time. The absence of real-time, contextual information flow creates blind spots for agentic AI. This hinders agents’ ability to react effectively to changing conditions and make insightful decisions. To overcome this, agents require a secure and consistent access layer to trusted data models. Here, the limitations of legacy systems become particularly evident in complex operational settings. For example, industrial environments often struggle with integrating various systems such as [Supervisory Control and Data Acquisition (SCADA) platforms](https://csrc.nist.gov/glossary/term/supervisory_control_and_data_acquisition), each with unique data formats, access protocols, and latency characteristics. This fragmentation creates fragile pipelines that can’t support the seamless, real-time access that agentic AI demands. Organizations without real-time and scalable data architectures risk being outpaced by their competitors. ## Five essential data pillars Agentic AI depends on five essential data pillars: ### Governance _Audit trails ensure compliance and build trust._ Robust governance guarantees agents receive monitored, auditable, and secure access while meeting evolving regulatory demands. Without this framework, autonomous decisions lack accountability and risk compliance failures. ### Structure _Raw data is fundamentally inadequate._ Agents require modeled, tested, and contextualized datasets transformed into analytics-ready assets. This structured foundation enables precise reasoning and prevents costly errors from unvalidated inputs. ### Semantics _Ambiguity sabotages AI decision-making._ Business-critical terms, such as revenue or churn, must be precisely defined and consistently applied across all systems. Semantic unity ensures agents interpret metrics identically, eliminating conflicting outputs. ### Governed access _Security and accessibility must coexist._ Strict permission controls prevent data leaks while also enabling the agents to securely retrieve contextual information via APIs. ### Observability _Real-time monitoring is non-negotiable._ Teams require immediate visibility into agent actions, their reasoning, and the resulting outcomes. Automated alerts and self-healing mechanisms enable rapid intervention when anomalies arise. ## Understanding MCP servers in the agentic AI ecosystem Organizations require infrastructure that provides secure, structured, and real-time access to various systems. [MCP servers](https://docs.getdbt.com/docs/dbt-cloud-apis/mcp) are designed to meet these demands for agentic AI. They act as standardized interfaces exposing models, metadata, lineage, and system capabilities to AI agents through a unified protocol. ![Diagram showing the architecture of an MCP (Modular Communication Protocol) system. On the left, an MCP Host contains an MCP Client, which communicates over a transport layer with an MCP Server. The MCP Server then connects to various external systems—a Database, API, and Gmail—via Web API calls. Each component is represented in its own colored box, highlighting the modular and interoperable nature of the system.](https://cdn.sanity.io/images/wl0ndo6t/main/c8fc977a805a114d026d754cd369dfc4d00f0a6b-1600x635.jpg) MCP servers abstract the complexity of backend integrations, providing agents with seamless access to both structured and unstructured data. This access enables consistent reasoning and real-time action across disparate systems. MCP servers offer a framework to build and support these capabilities: - Standardized access to models, APIs, and data sources. - Real-time operations with low-latency run times. - Real-time access to contextual data for dynamic, millisecond-level decision-making. - Secure interfaces with built-in authentication and encryption. - Interoperability across tools, services, and platforms. Once authenticated, agents use lightweight, protocol-driven requests to: - Initiate multi-step workflows. - Call external APIs. - Query logs and system telemetry. - Retrieve semantic models and business rules. MCP servers apply business rules to agent requests and return structured responses like JSON or XML that enable agents to decide and act in real time. ### Benefits of the MCP architecture MCP servers help agents manage tasks across systems. Their architecture offers clear benefits for real-world use: 1. **Reduces development overhead:** Reusable, plug-and-play interfaces remove the need for hardcoded integrations. 2. **Democratizes access:** Business users can use agentic systems via intuitive and natural language interfaces. 3. **Supports scalability:** MCP servers are containerized and load-balanced, with support for modular expansion. 4. **Enhances security and auditability:** Every transaction is logged and governed, aligning with enterprise compliance needs. ## Preparing your data for MCP integration For MCP servers to effectively empower agentic AI, organizations must proactively refine their data infrastructure to meet the unique demands of autonomous systems. Here's how to build a solid foundation: ### Step 1: Map data dependencies Agents must understand how changes in one system affect others. Begin by documenting the complete lineage of your data, from its origin in source systems through transformations to the final business reports. For example, if a supplier changes their ID in your procurement system, agents must recognize how this impacts production forecasts and compliance reports. Without this visibility, agents can’t anticipate ripple effects; however, with full lineage mapping, they transform from simple task runners to proactive problem solvers. ### Step 2: Embed quality checks Autonomous systems increase the risk of bad data, so teams must build [data quality validation checks](https://www.getdbt.com/blog/data-quality-checks) directly into their pipelines. - Set freshness alerts to flag delayed data streams before they impact real-time pricing agents. - Implement uniqueness checks to prevent duplicate customer records during automated onboarding. - Create null constraints to block incomplete transactions in payment processing. - Enforce referential integrity to ensure warehouse inventory counts match shipping manifests. These validations act as safety nets, catching errors before agents act on flawed information. ### Step 3: Standardize metadata Agents can’t work together if the data means different things in different places. Establish clear, consistent language across your data ecosystem: - Clearly define business terms. For example, specify an “active customer” as a user who has logged in within the past 30 days. - Tag sensitive information, such as patient health records, with clear ownership labels. - Make data lineage accessible through protocol-driven requests, enabling agents to verify context before making decisions. - Store policies directly in metadata, allowing agents to automatically comply with regional regulations. This eliminates ambiguity, ensuring that a sales agent in Europe and a logistics agent in Asia interpret "order fulfillment rate" identically. ### Step 4: Design secure interfaces Agent systems create new vulnerabilities if not properly contained. Implement strict access protocols: - Restrict agents to the minimal necessary permissions, like allowing customer service agents to access only the support ticket history. - Require authentication through [single sign-on](https://www.onelogin.com/learn/how-single-sign-on-works) for all data requests. - Mask sensitive details in responses, like partially redacting credit card numbers. - Log every agent action to maintain audit trails and meet compliance standards. These controls let agents operate freely while keeping sensitive data protected. ### Step 5: Optimize for low-latency Delayed data breaks real-time agents, as even a few seconds of stock update lag can cost millions. In fast-moving environments, latency isn’t a metric; it’s a liability. Design your systems for speed: - Use buffering techniques to temporarily hold incoming sensor data and prevent overload during sudden surges. - Build failover systems to instantly switch to backup servers during cloud outages. - Implement priority routing so that critical operations, such as emergency shutdowns, bypass queues. - Continuously monitor performance at peak traffic moments. If your pipelines can't deliver data within 500 milliseconds, agentic workflows will fail when they're needed most. ## How dbt prepares your data for MCP servers dbt transforms data infrastructure into an agent-ready foundation by embedding governance, testing, and semantic consistency directly into transformation workflows. In addition, dbt supports its own [MCP server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) to serve data to any MCP-ready LLM for powering real-world, operational scenarios. ### Governance by design With dbt, organizations can build trust in their data by automating governance. Every [model change](https://docs.getdbt.com/docs/build/models) is [version-controlled](https://docs.getdbt.com/docs/cloud/git/version-control-basics) and [documented](https://docs.getdbt.com/docs/build/documentation), with sensitive fields protected using column-level security for traceability and compliance. When regulations change, teams can update definitions centrally, allowing those updates to automatically flow across all dependent models. This ensures agents only access authorized and up-to-date data. ### Transformation engine dbt turns raw and fragmented data into [tested](https://docs.getdbt.com/reference/resource-properties/data-tests) and documented analytics-ready models. Built-in testing features like uniqueness checks, not-null constraints, and [freshness](https://docs.getdbt.com/reference/resource-properties/freshness) alerts catch errors early in the pipeline. These automated safeguards make sure agents interact with accurate and current data. dbt also scales with modern cloud platforms like Snowflake and BigQuery, so performance holds steady even during peak activity. ### Semantic clarity The [dbt Semantic Layer](https://www.getdbt.com/blog/semantic-layer-introduction) enforces a consistent definition of business metrics across the organization. Centralized models ensure that terms like "churn" or "revenue" have the same meaning in every tool. These definitions are versioned, documented, and available to downstream systems. ### MCP-ready metadata dbt exposes metadata such as model documentation, column-level security, and data lineage to enable seamless integration with MCP servers. Agents query this metadata via standardized interfaces to understand data sources, transformation logic, and freshness timestamps. dbt inherits access policies from enterprise identity providers like Okta or Azure AD, reducing redundant permission configuration for MCP interactions. ### Observability and feedback loops dbt enables continuous data quality monitoring by running automated tests on every change to catch issues before agents access the data. Logging and alerts provide insights into freshness metrics, test outcomes, and pattern deviations. This visibility helps teams refine data quality and ensures agents operate with trustworthy inputs. ### Enterprise impact dbt addresses key operational needs for MCP servers: - Unified access to governed data through documented models - Consistent metric definitions across all systems - Built-in compliance via audit trails and access controls - Scalability with cloud platform automation These capabilities ensure autonomous systems receive structured and reliable data without compromising on speed or governance. ## Conclusion MCP servers create a secure, standardized way for agents to access data. However, agents are only as effective as their upstream data. dbt turns raw sources into tested, governed models by catching errors early through built-in checks like freshness, uniqueness, and null constraints. Data is modeled once and reused everywhere, so agents see consistent definitions and context no matter where they run. Through [dbt Cloud's MCP Server](https://docs.getdbt.com/docs/dbt-cloud-apis/mcphttps://docs.getdbt.com/docs/dbt-cloud-apis/mcp), teams can securely expose: - Model-level lineage for dependency mapping - Semantic Layer metrics for unified business logic - Critical metadata for contextual decision-making With dbt, you can both prepare high-quality data for MCP and serve it directly to LLMs, all from a single platform. To get started, [sign up for a free dbt account today](https://www.getdbt.com/signup). --- --- title: "How to solve data SLA challenges in modern pipelines" description: "Data SLAs define expectations for freshness and quality—this guide helps you meet them in modern, complex data environments." url: "https://www.getdbt.com/blog/data-sla-challenges-guide" date: "2025-07-30" authors: ["Joey Gault"] categories: ["Pulse"] --- # How to solve data SLA challenges in modern pipelines [Data SLAs](https://www.ibm.com/think/topics/data-sla) define the expected performance standards for data delivery, including metrics like freshness, availability, and quality. Unlike traditional IT SLAs that focus primarily on uptime, data SLAs must account for the complex interdependencies inherent in modern data pipelines. The most common SLA metrics include data freshness (how recently data was updated), data latency (how long it takes for upstream changes to propagate downstream), and data quality measures across multiple dimensions. Each metric requires different monitoring approaches and remediation strategies, making SLA management a multifaceted challenge. Consider a sales reporting pipeline that requires daily updates by 9 AM. This seemingly simple SLA actually encompasses multiple requirements: source systems must complete their overnight processing, extraction jobs must run successfully, transformations must execute without errors, and quality checks must pass. A failure at any stage can cascade through the entire pipeline, making SLA achievement dependent on the reliability of every component. ## The freshness challenge [Data freshness ](https://www.getdbt.com/blog/data-quality-checks#freshness-and-recency)represents one of the most visible SLA challenges. When reports show stale data, business users immediately notice, and their confidence in the data system diminishes. The challenge intensifies as organizations adopt more complex data architectures with multiple sources, transformation layers, and consumption endpoints. Modern data stacks have made freshness monitoring more sophisticated. [dbt](https://www.getdbt.com/product/what-is-dbt) provides built-in capabilities for tracking source freshness, allowing teams to define acceptable timeframes for each data source and automatically monitor compliance. When configuring freshness monitoring, teams should run checks with at least double the frequency of their lowest SLA. For example, if you have a 1-hour SLA on a particular dataset, freshness checks should run every 30 minutes to ensure adequate monitoring coverage. The key to managing freshness SLAs lies in understanding the entire data lineage. A single delayed upstream source can impact dozens of downstream models and reports. By implementing comprehensive freshness monitoring across all critical data sources, teams can identify bottlenecks before they violate SLAs and take proactive remediation steps. Freshness challenges often stem from unrealistic expectations about data availability. Business stakeholders may request hourly updates for data that originates from systems that only refresh daily. Data engineering leaders must work closely with business teams to establish achievable SLAs that balance business needs with technical constraints. ## Quality as an SLA component [Data quality](https://www.getdbt.com/blog/data-quality-metrics) issues represent another significant challenge in meeting SLAs. Poor quality data can render timely delivery meaningless if the data cannot be trusted for decision-making. Quality-related SLA failures often prove more damaging than freshness issues because they can lead to incorrect business decisions. Effective quality SLAs require clear definitions across multiple dimensions. Accuracy ensures that data values reflect reality. Completeness verifies that all required records and fields are present. Consistency maintains data integrity across different systems and time periods. Validity ensures that data values conform to expected formats and ranges. Each dimension requires specific monitoring and testing approaches. [dbt's testing framework provides a foundation for quality SLA management. ](https://www.getdbt.com/product/build-trust-in-data-and-data-teams)Teams can implement automated tests that verify data quality after each transformation, catching issues before they propagate to downstream consumers. Generic tests for common quality checks can be reused across projects, making quality monitoring more efficient and consistent. The challenge with quality SLAs lies in balancing comprehensiveness with performance. Extensive quality checks can slow pipeline execution, potentially causing freshness SLA violations. Teams must carefully design their quality testing strategy to provide adequate coverage without creating new bottlenecks. ## Pipeline reliability and dependency management Modern data pipelines involve complex dependencies between multiple systems, tools, and processes. A failure in any component can cascade through the entire pipeline, making SLA achievement dependent on the reliability of every element in the chain. This interdependency creates significant challenges for SLA management. Dependency-related SLA failures often occur when upstream systems experience unexpected delays or failures. A source system maintenance window that runs longer than expected can delay all downstream processing. Network issues can prevent data extraction jobs from completing on schedule. Resource constraints in cloud data warehouses can slow transformation processing beyond acceptable limits. The [modern data stack's modular architecture](https://podcasts.apple.com/us/podcast/ai-data-engineering-and-the-modern-data-stack/id1740178076?i=1000713838763), while providing flexibility and scalability, also introduces additional points of potential failure. Data ingestion tools, transformation engines, orchestration platforms, and storage systems must all function correctly for SLAs to be met. Each tool in the stack requires monitoring, maintenance, and optimization to ensure reliable performance. Effective dependency management requires comprehensive observability across the entire data pipeline. Teams need visibility into the performance and health of each component, along with automated alerting when issues arise. This observability must extend beyond individual tools to encompass the relationships and dependencies between different pipeline stages. ## Monitoring and alerting strategies Proactive monitoring represents the foundation of successful SLA management. Teams cannot address issues they cannot see, making comprehensive observability essential for maintaining SLA compliance. However, many organizations struggle with monitoring strategies that provide adequate coverage without overwhelming teams with false alerts. Effective SLA monitoring requires a layered approach that covers different aspects of pipeline performance. Infrastructure monitoring tracks the health of underlying systems and resources. Application monitoring focuses on the performance of specific tools and processes. Data monitoring examines the quality and freshness of the data itself. Each layer provides different insights and requires different monitoring tools and techniques. The challenge lies in correlating signals across these different monitoring layers to provide actionable insights. A data freshness alert might indicate a problem, but determining whether the root cause lies in source system delays, extraction job failures, or transformation bottlenecks requires additional investigation. Effective monitoring systems help teams quickly identify the source of problems and take appropriate remediation actions. Alerting strategies must balance responsiveness with practicality. Teams need to know about SLA violations quickly enough to take corrective action, but excessive alerting can lead to alert fatigue and reduced responsiveness. Successful alerting strategies focus on actionable alerts that clearly indicate when human intervention is required. ## Building resilient data architectures Architectural decisions significantly impact SLA achievement. Systems designed with resilience in mind can better handle the inevitable failures and delays that occur in complex data environments. However, building resilient architectures requires careful consideration of trade-offs between reliability, performance, and cost. Redundancy represents one key aspect of resilient architecture design. Having backup data sources, alternative processing paths, and failover capabilities can help maintain SLA compliance when primary systems experience issues. However, redundancy adds complexity and cost, requiring careful evaluation of which components justify the additional investment. Error handling and recovery mechanisms play crucial roles in SLA achievement. Systems that can automatically retry failed operations, skip problematic records, or fall back to alternative processing methods can often maintain SLA compliance despite encountering issues. The key lies in designing these mechanisms to handle expected failure modes while providing appropriate visibility into what occurred. Resource management also impacts SLA reliability. Cloud data warehouses that automatically scale to handle varying workloads can better maintain consistent performance. However, auto-scaling must be configured appropriately to balance performance with cost considerations. Teams must understand their workload patterns and configure resources to handle peak demands without excessive over-provisioning. ## Organizational and process considerations Technical solutions alone cannot solve SLA challenges. Organizational factors, including team structure, communication processes, and stakeholder management, significantly impact SLA success. Many SLA failures stem from organizational issues rather than technical problems. Clear communication channels between data teams and business stakeholders help prevent unrealistic SLA commitments and ensure appropriate prioritization when issues arise. Business users need to understand the technical constraints that impact data delivery, while data teams need to understand the business impact of SLA violations. Regular communication helps maintain alignment and realistic expectations. Incident response processes become critical when SLA violations occur. Teams need clear procedures for identifying, escalating, and resolving issues that threaten SLA compliance. These processes should include communication protocols for keeping stakeholders informed about the status of ongoing issues and expected resolution timeframes. Change management processes also impact SLA reliability. Uncoordinated changes to source systems, data models, or pipeline configurations can introduce unexpected failures. Effective change management ensures that modifications are properly tested and coordinated to minimize the risk of SLA violations. ## Practical implementation approaches Successfully addressing data SLA challenges requires a systematic approach that combines technical improvements with organizational changes. Teams should begin by establishing baseline measurements of current performance across all relevant SLA metrics. This baseline provides the foundation for identifying improvement opportunities and tracking progress over time. Prioritization becomes essential when addressing multiple SLA challenges simultaneously. Teams should focus first on the issues that have the greatest business impact or the highest likelihood of success. Quick wins can build momentum and demonstrate value while longer-term improvements are implemented. Incremental improvement often proves more effective than attempting comprehensive overhauls. Teams can implement monitoring for critical data sources, establish automated testing for key quality dimensions, and improve error handling for common failure scenarios. Each improvement builds on previous work and contributes to overall SLA reliability. Documentation and knowledge sharing help ensure that SLA improvements are sustainable over time. Teams should document their monitoring approaches, alerting configurations, and incident response procedures. This documentation helps new team members understand the system and provides reference material for troubleshooting issues. ## Measuring success and continuous improvement Effective SLA management requires ongoing measurement and refinement. Teams should regularly review their SLA performance, identify trends and patterns in violations, and adjust their approaches based on lessons learned. This continuous improvement mindset helps organizations adapt to changing requirements and evolving technical landscapes. Key performance indicators should extend beyond simple SLA compliance rates to include metrics like mean time to detection, mean time to resolution, and the business impact of violations. These additional metrics provide insights into the effectiveness of monitoring and response processes and help identify areas for improvement. Regular retrospectives on significant SLA violations can provide valuable learning opportunities. Teams should examine not just the technical root causes but also the organizational factors that contributed to the issue. These retrospectives often reveal process improvements that can prevent similar issues in the future. The modern data landscape continues to evolve rapidly, with new tools, techniques, and requirements emerging regularly. Successful SLA management requires staying current with industry best practices and continuously evaluating new approaches that might improve reliability and performance. Data SLA challenges are complex and multifaceted, requiring both technical expertise and organizational coordination to address effectively. However, teams that invest in comprehensive monitoring, resilient architectures, and effective processes can achieve reliable data delivery that supports their organization's analytical needs. The key lies in taking a systematic approach that addresses both immediate issues and long-term sustainability, ensuring that data systems can meet their commitments consistently over time. ## Data SLA FAQs **What is a data SLA?** A data SLA (Service Level Agreement) defines the expected performance standards for data delivery, including metrics like freshness, availability, and quality. Unlike traditional IT SLAs that focus primarily on uptime, data SLAs must account for the complex interdependencies inherent in modern data pipelines and encompass multiple requirements such as successful source system processing, extraction jobs, transformations, and quality checks. **What metrics are commonly included in data SLAs?** The most common data SLA metrics include data freshness (how recently data was updated), data latency (how long it takes for upstream changes to propagate downstream), and data quality measures across multiple dimensions. Quality metrics encompass accuracy (ensuring data values reflect reality), completeness (verifying all required records and fields are present), consistency (maintaining data integrity across systems and time periods), and validity (ensuring data values conform to expected formats and ranges). **What happens when you miss your data SLA?** When data SLA violations occur, teams need clear incident response procedures for identifying, escalating, and resolving issues that threaten SLA compliance. These processes should include communication protocols for keeping stakeholders informed about the status of ongoing issues and expected resolution timeframes. SLA failures often have cascading effects; for example, a single delayed upstream source can impact dozens of downstream models and reports, and quality-related SLA failures can prove more damaging than freshness issues because they can lead to incorrect business decisions. --- --- title: "dbt Labs Signs Strategic Collaboration Agreement with AWS" description: "The dbt Labs and AWS strategic collaboration agreement will deepen technical integrations and expand AWS Marketplace availability." url: "https://www.getdbt.com/blog/dbt-labs-expands-signs-strategic-collaboration-agreement-with-aws" date: "2025-07-29" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Signs Strategic Collaboration Agreement with AWS **PHILADELPHIA – July 29, 2025 – **[dbt Labs](http://getdbt.com), the leader in standards for AI-ready structured data, today announced it has signed a strategic collaboration agreement (SCA) with Amazon Web Services (AWS). The agreement will expand technical integrations across services including Amazon Redshift, Amazon Athena, and Amazon SageMaker Lakehouse, while also expanding [AWS Marketplace availability](https://aws.amazon.com/marketplace/pp/prodview-tjpcf42nbnhko) to simplify access and procurement for enterprise customers. “By deepening our collaboration with AWS, we’re giving organizations even more flexibility to choose dbt, the standard for AI-ready structured data, as the transformation engine behind their AI initiatives,” said Shawn Toldo, VP of Worldwide Partner Ecosystem at dbt Labs. “Together, we’re meeting customers where they are – enabling faster time-to-value and seamless cloud-native deployments.” Shared dbt Labs and AWS customers including Moderna leverage dbt's tight integration with AWS to break down silos, streamline analytics, and enable AI-driven insights. During AWS re:Invent, [Moderna and dbt Labs](https://youtu.be/Vb6qSMN2AoU?feature=shared) took the stage together to discuss how [Moderna entrusts dbt and AWS](https://www.getdbt.com/blog/aws-reinvent-2024-recap) with ensuring vaccine supply chain resilience so that it can provide critical and timely support to its network of customers. In addition to Moderna, [Culture Amp](https://www.cultureamp.com/) relies on dbt and AWS to scale its data operations and accelerate decision-making across the business. "At Culture Amp, we’re focused on helping organizations build a better world of work. We combine organizational psychology and data science within our products to help organizations deeply understand the experience of their employees, develop talent, drive performance and remove the guesswork around building thriving workplace cultures. To do this, we need a modern data stack we can trust,” said Doug English, Chief Technology Officer & Co-Founder, Culture Amp. “With dbt and AWS, our team can deliver timely, reliable insights at scale, enabling faster experimentation, better decision-making, and ultimately, more impact for our customers. We’re excited to see dbt Labs and AWS investing further in their collaboration to support data teams like ours.” This agreement underscores dbt Labs and AWS’ focus on providing flexibility and unlocking greater business value for customers across industries. “This expanded collaboration with dbt Labs empowers organizations to build faster, more reliable analytics workflows in the cloud," said Allison Johnson, Senior Manager of Americas Technology Partners at AWS. "Together, we're enabling data teams to accelerate time to insight and leverage emerging AI capabilities, all while optimizing for performance, scale, and cost efficiency. This collaboration brings together modern cloud data infrastructure and trusted transformation practices for customers who need to move with both speed and confidence." Simultaneously, dbt Labs continues to innovate its platform to revolutionize developer and [data analyst experiences](https://www.getdbt.com/blog/dbt-labs-launches-ai-powered-features-to-onboard-data-analysts-into-dbt) in the age of AI. The [recent launch](https://www.getdbt.com/blog/dbt-labs-redefines-dbt-with-new-fusion-engine) of the dbt Fusion engine enables faster analytics delivery, lowers cloud costs, and ensures teams can build trusted data pipelines at scale. To learn more about the latest dbt features powering the platform, including the dbt Fusion engine, visit [https://www.getdbt.com/product/fusion](https://www.getdbt.com/product/fusion). For more information on dbt and AWS integrations, please visit [https://www.getdbt.com/data-platforms/redshift](https://www.getdbt.com/data-platforms/redshift). **About dbt Labs** Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 80,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at [getdbt.com](http://getdbt.com), and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "Modernize self-service analytics with dbt" description: "Self-service analytics has, so far, failed to deliver. Here’s how dbt addresses the gaps and delivers collaboration without chaos." url: "https://www.getdbt.com/blog/modernize-self-service-analytics-dbt" date: "2025-07-25" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Modernize self-service analytics with dbt For decades, organizations have chased the promise of [self-service analytics](https://roundup.getdbt.com/p/5-principles-of-self-service). They’ve sought to empower business users to answer their own data questions without constantly relying on engineering teams. Yet despite countless tools and platforms claiming to solve this challenge, most companies still struggle with the same fundamental problems: [data silos](https://www.techtarget.com/searchdatamanagement/definition/data-silo), ungoverned data and code, and misaligned incentives between technical and business teams. The persistent failure of self-service analytics stems from a fundamental disconnect between how these systems are designed and how teams actually work. While technology has advanced dramatically, the organizational and workflow challenges that prevent effective collaboration remain largely unsolved. As the leader in data transformation technologies, dbt has spent years working with dozens of data-driven companies across all industries to solve exactly this challenge. Our latest platform features take aim at the issues that have hobbled self-service analytics, enabling what we call **collaboration without chaos**. ## The persistent pitfalls of self-service analytics Data and analytics consulting firms working across industries consistently observe the same pattern. Despite significant investments in modern data infrastructure, most organizations still operate in what can only be described as "the wild west" of data analysis. [Data analysts](https://www.getdbt.com/blog/data-analyst-closer-to-data) often find themselves forced to do whatever they can to obtain answers, jumping between tools—sometimes visualizing data in one platform, writing queries in another, and creating data in a third. They copy logic from old queries and set up custom tables because everything remains fundamentally siloed. This chaos stems from an organizational structure where data engineers work in one group and analysts work in another, each with different priorities and needs. Rarely are they able to build together. That means engineers and analysts, more often than not, end up talking **at** each other rather than **with** each other. ### An organizational issue The misalignment runs deeper than communication. Data engineering teams are measured on building scalable, well-governed core assets that serve the entire organization. Their sprint cycles often extend six to nine months to ensure proper quality and governance. Meanwhile, business analysts face entirely different pressures. They need answers in days, not months, to support critical business decisions. This creates an impossible situation. A data analyst might need a specific table with a particular data model setup, joins, and columns. The engineering team understands the requirement and agrees it's valuable, but can’t schedule the work for another six months. When the analyst's project deadline is one month away, they're forced to create workarounds—Excel spreadsheets, local databases, or custom SQL queries that bypass established governance processes. The consequences compound quickly. Organizations end up with random SQL queries running everywhere, Excel files being uploaded to create reports, and different definitions for the same business concepts across teams. This means that the way one analyst defines a metric in Excel differs from how their coworker defines it. Both end up building dashboards using those conflicting definitions. This type of divergence erodes trust in data. This situation isn’t anyone’s fault. It’s a natural consequence of how most companies are structured. ## Who should care? Everyone, but that's complicated Ideally, a self-service framework for data would serve **everyone** who works with data. A secure and well-governed self-service data platform needs to accommodate three distinct user types: **Data engineers** represent the most technical end, skilled in SQL and scripting languages, focused on building scalable, efficient, and secure data pipelines that serve the entire company. **Power analysts and analytics engineers** occupy the middle ground. They're skilled at transforming large datasets and comfortable with SQL, but they value rapid iteration and user-friendly interfaces over writing complex code from scratch. **Business users** lean toward the non-technical side while still working directly with data. They understand their business questions deeply but prefer drag-and-drop interfaces and natural language queries over SQL development. These user types are more of a spectrum than rigid categories. The same job titles can mean completely different things at different companies. A "business analyst" at one organization might have advanced SQL skills and build complex data models, while someone with the same title elsewhere might work exclusively with pre-built dashboards and simple filters. Ultimately, **successful self-service analytics needs to be built for humans**: - People who know what questions they want to ask but don't always know where to start - Users wanting fast answers without waiting on engineering teams - Teams needing to move quickly while maintaining governance standards - Analysts uncomfortable writing complex SQL but requiring sophisticated analysis. **** ## dbt’s solution: Empowerment without compromise The fundamental issue is **the universal tension between governance and empowerment**. Organizations typically swing hard in one direction or the other. Some lock down data access so tightly that business teams can't get work done. Others open everything up so broadly that quality and governance disappear. dbt is a framework for delivering trusted data faster, better, and more collaboratively. Our platform **redefines how companies think about trusted self-service**. We do this by adhering to three core principles: **Speaking the same language in a governed environment** by supporting four key capabilities: - Move faster by building on trusted, reusable SQL code; - Structured collaboration to enable reusing logic across teams without duplication or drift; - Automated dependency tracking and [version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics), making it easy to deploy changes safely and scale with confidence; and - Built-in metadata and model relationships to give AI the context it needs to generate accurate, business-aligned outputs. **Trust in data** through comprehensive [data lineage](https://www.getdbt.com/blog/what-is-data-lineage), governance, and transparency. Users can trust data when they know where it comes from, how it was created, and when it was last updated. This represents a dramatic improvement over receiving an Excel file from a colleague who says, "trust me, it's what you want." [Clear documentation](https://docs.getdbt.com/docs/build/documentation) and governance structures eliminate guesswork and reduce human error. **Empowerment without compromise** so that everyone can access and work with data within their skill level, technical capabilities, and available time. Business users shouldn't need to complete SQL training courses to get their work done. Data engineers shouldn't need to learn new systems to maintain governance standards. dbt accommodates different working styles while ensuring everyone can collaborate effectively. ## Speaking the same language in a governed framework To enable these capabilities, dbt implements three key features so that everyone in the organization can easily discover, use, and refine data: - **dbt Catalog** provides global search across dbt and Snowflake assets so that **users know what to ask and where to start** - **dbt Insights** helps analysts and business users generate, validate, and refine queries to **gain insights fast without relying on engineering** - **dbt Canvas** supports low code and no code development of data models so that analysts can **build fast without breaking governance** Let’s take a look at each of these features in detail. ### dbt Catalog: Universal data discovery and trust [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) solves the fundamental problem of data discovery in complex organizations. Many companies inadvertently create multiple versions of similar tables—such as dim_customer, dim_customers, and dim_customer_v2—because analysts are unaware of the existing data or which version represents the authoritative source. dbt Catalog provides global search across all dbt projects and [Snowflake](https://snowflake.com/) data warehouse assets, pointing users to existing high-quality data rather than encouraging them to rebuild from scratch. This prevents duplication of effort and reduces the proliferation of slightly different data definitions. Trust signals embedded throughout the catalog help users evaluate data quality before building reports. These include the status of [data tests](https://docs.getdbt.com/docs/build/data-tests), lineage information, [source freshness](https://docs.getdbt.com/docs/deploy/source-freshness), ownership details, and quality validation through upstream dependency status. [Column-level lineage](https://www.getdbt.com/blog/what-is-data-lineage) provides a clear view of how each field is derived from source systems, enabling users to understand not only what the data represents but also how it was created and transformed. This transparency fosters confidence in data accuracy and appropriateness for specific use cases. For organizations using [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl), dbt Catalog exposes standardized business metrics with complete transparency. Users can see both the high-level metric definition and the compiled SQL that generates results, eliminating guesswork about how key business measures are calculated. ### dbt Insights: Embedded analytics with AI assistance [dbt Insights](https://docs.getdbt.com/docs/explore/dbt-insights) provides embedded analytics capabilities that eliminate the need to switch between discovery, analysis, and development tools. The platform supports the complete [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) from initial data exploration through final insight generation. The experience begins with immediate data validation through pre-populated queries that allow quick assessment of data quality and structure. This addresses common concerns about the origins and trustworthiness of data. Within dbt Insights, [dbt Copilot](https://www.getdbt.com/blog/introducing-dbt-copilot) can transform natural language requests into executable SQL queries so that no one has to become a SQL expert to find their data. Users can ask for "top 10 best-selling parts for the last six months" and receive complete, runnable SQL statements. The system supports iterative analysis through follow-up questions, enabling users to drill down into specific findings, like monthly sales trends for individual products. dbt Insights also includes integrated visualization capabilities with customizable charts and titles, allowing users to create compelling presentations of findings without additional tools. All analysis can be bookmarked and shared across teams via links that include both queries and visualizations, enabling others to build on previous work rather than starting from scratch. ### dbt Canvas: Low-code transformation for visual learners [dbt Canvas](https://docs.getdbt.com/docs/cloud/canvas) addresses the needs of users who prefer visual, drag-and-drop interfaces for building data transformations. Consider a common use case: a data analyst tasked with cohort analysis needs to combine customer information with historical order data from Snowflake to understand purchase patterns and advertising effectiveness. The visual interface allows users to input data sources and immediately see previews of results without running full queries. This real-time feedback loop supports efficient development by catching errors early and ensuring transformations produce expected outputs. The platform enables sophisticated data transformations through intuitive interfaces. Users can create complex logic like case-when statements to clean status fields—for example, converting return_pending to return for consistency—with immediate visual confirmation of results. AI integration through [dbt Copilot](https://www.getdbt.com/blog/introducing-dbt-copilot) enables natural language requests. Users can write something like "Combine stage orders and stage customers and add a column calculating total orders over time." The system will then generate complete workflows with inputs, joins, aggregates, and outputs. The AI makes intelligent suggestions for join keys based on column analysis, though users retain full control over accepting or modifying these recommendations. Technical governance remains intact throughout the visual development process. Users can view the underlying SQL code at any time, seeing exactly what queries their visual transformations generate. All work gets stored as SQL within the project. That means it can be versioned in Git, audited, tested, and even reverted if required. The collaboration benefits extend to code reuse. Users can import existing SQL models created by colleagues into Canvas for visualization, making complex transformations more understandable and modifiable even for those who didn't write the original code. Query history provides a foundation for building on previous work, while dbt Semantic Layer support ensures access to standardized business metrics with transparent calculation logic. ## Making analytics a team sport The evolution of self-service analytics isn't about choosing between governance and agility—it's about creating platforms that enable both simultaneously. Our unified approach addresses the root causes of previous failures through centralized definitions with transparent lineage for data quality, visual tools and AI assistance for improved data literacy, and collaborative workflows that clarify data ownership and responsibilities. By ensuring everything resolves to SQL within a governed framework, organizations can scale self-service analytics without sacrificing the quality standards that enterprise data requires. This transforms analytics from a siloed, ungoverned process into a collaborative discipline where technical and business teams work together effectively. Making analytics a team sport requires addressing organizational dynamics as well as technical capabilities. dbt Canvas, dbt Insights, and our expanded dbt Catalog provide the foundation for this transformation, enabling democratized data access while maintaining and enhancing enterprise governance standards. --- --- title: "What’s new in dbt - July 2025" description: "Get the scoop on all the latest features that just landed in dbt." url: "https://www.getdbt.com/blog/whats-new-in-dbt-july-2025" date: "2025-07-25" authors: ["Sara Gawlinski"] categories: ["Product"] --- # What’s new in dbt - July 2025 Summer is sizzling 🌞 and things are heating up at dbt! In May, we hosted our [dbt Launch Showcase](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase), where we unveiled powerful new features to help your data teams work smarter, faster, and more securely. But we’re not stopping there. We’re continuing to iterate and deliver even more enhancements and capabilities to empower you and your teams. Grab your cold brew and let’s dive in! #### Fusion 🏎️ **The dbt Fusion engine public beta is now available for Snowflake, BigQuery, and Databricks**, marking a significant milestone toward getting Fusion in the hands of more users. [Fusion](https://www.getdbt.com/product/fusion) is more than an incremental improvement to dbt; it represents a fundamental shift in speed, intelligence, and developer experience for anyone building with dbt. Fusion offers: - Instant feedback and SQL comprehension: Fusion understands Snowflake, BigQuery, and Databricks SQL and provides real-time error checking, code suggestions, and smarter development experiences. - Blazing-fast performance: Expect parsing and compilation speeds up to 30x faster than dbt Core. - Cost efficiency: When standardized on Fusion in the dbt platform (and not just running Fusion locally), users can take advantage of state-aware orchestration to automatically only runs models that have new data, reducing cloud compute costs of dbt workloads running in data platforms by 10% or more. - Detailed governance (coming soon): Precise column-level lineage and richer metadata for enhanced data governance. [Read the launch blog](https://www.getdbt.com/blog/get-to-know-the-new-dbt-fusion-engine-and-vs-code-extension) and [docs](https://docs.getdbt.com/docs/fusion/about-fusion) to learn more. 🛠️ **The official dbt** **VS Code Extension is in beta**: dbt’s official VS Code Extension, powered by the dbt Fusion engine, is now in public beta. Unlock powerful capabilities like IntelliSense autocompletion, automatic refactoring, inline CTE previews, and go-to-definition, providing a hyper-responsive and delightful development experience. [Learn more](https://docs.getdbt.com/docs/about-dbt-extension) in the docs. #### Cross-team collaboration 🎨 **dbt Canvas is GA.** This AI-powered visual editing experience empowers users to create and edit dbt models using a drag-and-drop interface, context-aware AI assistance, and intuitive Git-based version control. Users can commit and open PRs directly from Canvas, streamlining collaboration and enhancing productivity. [Check out the launch blog](https://www.getdbt.com/blog/dbt-canvas-is-ga) to learn more. 📈**dbt Insights is in Preview.** Rapidly validate ideas, explore data trends, and build visualizations in a unified, governed environment. Combine SQL, natural language querying, and built-in AI assistance to explore and query data. Learn more about Insights in the [docs](https://docs.getdbt.com/docs/explore/dbt-insights). ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/1ed5cf109e20ac6c94b77413552592528b88de9f-2940x1594.png) 🔎 **Global navigation in dbt Catalog is in Preview.** The dbt Catalog (formerly dbt Explorer) now offers [global navigation](https://docs.getdbt.com/docs/explore/explore-projects#catalog-overview) across all your dbt projects and expands beyond dbt models, empowering you to search and explore Snowflake assets, such as tables and views, directly within dbt. This unified, seamless experience helps teams move faster, answer critical questions, and make more informed decisions - all without switching tools or duplicating efforts. **📈 Power BI integration for the Semantic Layer is in Preview.** At long last! You can now connect Microsoft Power BI directly to the dbt Semantic Layer. Centralize your business logic and metric definitions in dbt and build your Power BI dashboards from consistent, trusted, and low-latency data—in both Power BI Desktop and Power BI Service. [Check out the docs](https://docs.getdbt.com/docs/cloud-integrations/semantic-layer/power-bi) to learn more. **🔗 Trino and Postgres join the Semantic Layer.** You can now query metrics defined in the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) directly from your Trino and Postgres setups. This expansion empowers teams using these databases to unlock seamless cross-team collaboration, unified analytics definitions, and faster insights without ever leaving dbt. Plus, enjoy quicker query responses when leveraging the Semantic Layer Gateway (SLG) for these adapters. #### AI 🤖 **dbt MCP server is in public beta.** The dbt MCP server provides structured, governed context for AI systems. Integrate trusted dbt models, metrics, tests, and lineage directly into your AI workflows, powering smarter, safer, and more trustworthy AI-driven decisions. Check out this [blog post](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) for more information. [Watch video](https://www.youtube.com/watch?v=X3nk_KNrQhU) #### Platform 💰**Cost management dashboard is in Preview. **Quickly track warehouse spend and identify optimization opportunities directly within dbt. Drill down by account, project, and model, clearly tie actions to savings, and reduce unnecessary costs with Fusion’s state-aware orchestration. The cost management dashboard is available for Snowflake customers. Learn more about [cost management in dbt](https://docs.getdbt.com/docs/cloud/cost-management). ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/85d2e9356c4676ad27b24fd5edcd0598aafcfa8d-2880x1800.webp) **🔑 SCIM is officially here.** Say goodbye to manual user management. With SCIM (System for Cross-Domain Identity Management), dbt admins can automate user provisioning, de-provisioning, group management, and license assignments directly from their identity providers like Okta and Entra ID. Available now for Enterprise and Enterprise+ customers (note: not currently available for legacy Enterprise plans). dbt’s integration is fully SCIM 2.0 compliant, supporting any compatible identity provider. Dive into the [docs](https://docs.getdbt.com/docs/cloud/manage-access/scim) to see how SCIM can simplify your life. **🧊 Iceberg catalog integrations**. Iceberg catalog integrations. Beat the summer heat with our coolest catalog integrations yet for Snowflake and BigQuery. dbt can now read and write metadata in built-in catalogs like Snowflake Horizon and Biglake Metastore. You can also effortlessly create Iceberg tables in Snowflake and BigQuery without sacrificing your lineage, observability, or your peace of mind. This exciting step forward lays the groundwork for cross-platform data mesh and opens doors to further enhancements, because optionality is always cooler. Check out the [docs](https://docs.getdbt.com/docs/mesh/iceberg/apache-iceberg-support) for Apache Iceberg support in dbt to learn more. ## What’s next We're excited to hear your feedback on these new features! In the meantime, join us July 30th & 31st for regional sessions on building reliable AI agents with the dbt MCP server where you'll learn how to power intelligent, governed workflows at scale. Save your spot [here](https://www.getdbt.com/resources/webinars/build-reliable-ai-agents-with-the-dbt-mcp-server). Stay cool and stay tuned for more exciting dbt updates coming soon! --- --- title: "Building trust through automated data testing" description: "Learn how automated testing improves data quality, enables faster incident response, and builds trust with business users." url: "https://www.getdbt.com/blog/build-trust-through-data-testing" date: "2025-07-23" authors: ["Joey Gault"] categories: ["Pulse"] --- # Building trust through automated data testing Trust in data systems emerges from consistent reliability over time. Business users develop confidence when they can depend on data being accurate, complete, and available when needed. This reliability stems from systematic validation of data quality dimensions including accuracy, validity, completeness, freshness, and consistency. Automated data testing addresses these dimensions through two primary approaches: correctness validation and freshness monitoring. Correctness testing ensures that key columns maintain uniqueness and non-null constraints, that column values align with expectations, and that transformations produce expected results. Freshness testing monitors data update cadences and alerts teams when source data fails to arrive within expected timeframes. These testing approaches work together to create a comprehensive quality assurance framework. When business users consistently receive accurate, timely data, they develop confidence in both the data itself and the team responsible for delivering it. This confidence translates into increased data adoption, more sophisticated analytical use cases, and ultimately, better business outcomes. ## Strategic testing implementation Effective automated testing requires strategic implementation across three critical phases of the data lifecycle. During development, testing validates both raw source data and newly created transformations. This early validation catches issues before they propagate through downstream systems, reducing the cost and complexity of fixes. Development-phase testing focuses on fundamental data integrity checks. For raw source data, this includes validating primary key uniqueness and non-nullness, ensuring column values meet basic assumptions, and identifying duplicate rows. As data undergoes transformation through cleaning, aggregation, and business logic implementation, additional tests verify that primary keys remain unique, row counts align with expectations, and relationships between upstream and downstream dependencies function correctly. The second critical phase occurs during code integration, where pull request workflows ensure that new data models and transformation logic meet established quality standards before entering production. This peer review process, combined with automated test execution, prevents problematic code from reaching production environments. dbt's continuous integration capabilities enable teams to run comprehensive test suites automatically when code changes are proposed, providing immediate feedback on potential issues. Production testing represents the third essential phase, where automated tests run on scheduled intervals to monitor ongoing data quality. Production environments face constant change as source systems evolve, new features are deployed, and business requirements shift. Automated tests serve as early warning systems, alerting data teams to issues before they impact business users. ## Building comprehensive test coverage Comprehensive test coverage requires a layered approach that addresses different aspects of data quality and transformation logic. Unit tests validate individual model logic by testing specific functions or calculations with known inputs and expected outputs. These tests run quickly and provide immediate feedback during development, making them ideal for iterative development workflows. Integration tests verify that multiple components work together correctly, ensuring that data flows properly between models and that complex transformations produce expected results. These tests are particularly valuable for validating business logic that spans multiple data sources or requires complex calculations. Data tests focus on the actual data content, validating that the information meets business requirements and quality standards. These tests check for expected value ranges, proper formatting, and adherence to business rules. They also monitor for anomalies that might indicate upstream system issues or data corruption. The combination of these testing approaches creates a robust quality assurance framework. Unit tests catch logic errors early in development, integration tests ensure components work together properly, and data tests validate that the final output meets business requirements. This multi-layered approach provides comprehensive coverage while maintaining development velocity. ## Operational excellence through automation Automated testing enables operational excellence by creating predictable, repeatable quality assurance processes. When tests run automatically as part of scheduled workflows, data teams can focus on strategic initiatives rather than manual quality checks. This automation also ensures consistency in testing approaches across different team members and projects. The operational benefits extend beyond efficiency gains. Automated testing creates detailed audit trails that document data quality over time. These records prove invaluable during compliance audits, troubleshooting sessions, and root cause analyses. They also provide objective metrics for measuring data quality improvements and identifying areas that require additional attention. Automated testing also enables faster incident response. When issues occur, comprehensive test coverage helps pinpoint the source of problems quickly. Detailed test results provide context about what changed and when, reducing the time required to identify and resolve issues. This rapid response capability minimizes the impact of data quality problems on business operations. ## Scaling testing practices As data operations grow in complexity and scope, testing practices must scale accordingly. This scaling involves both technical and organizational considerations. From a technical perspective, test execution must remain performant even as data volumes and model complexity increase. Efficient test design, strategic sampling, and parallel execution help maintain reasonable test run times. Organizational scaling requires establishing clear testing standards and practices across the team. This includes defining what should be tested, how tests should be structured, and when they should run. Documentation and training ensure that all team members understand and follow established testing practices. dbt facilitates this scaling through its built-in testing framework and community-driven test packages. Teams can leverage pre-built tests for common scenarios while developing custom tests for specific business requirements. The framework's integration with version control systems ensures that testing practices evolve alongside code changes. ## Measuring testing effectiveness Effective testing programs require ongoing measurement and optimization. Key metrics include test coverage percentages, test execution times, and failure rates. These metrics help teams identify gaps in their testing approach and optimize test performance. More importantly, teams should measure the business impact of their testing efforts. This includes tracking the reduction in data quality incidents, decreased time to resolution for issues that do occur, and increased confidence levels among business users. These business-focused metrics demonstrate the value of testing investments and guide future improvements. Regular review of testing practices ensures they remain aligned with business needs and technical capabilities. As data systems evolve and business requirements change, testing approaches must adapt accordingly. This continuous improvement mindset helps maintain the effectiveness of testing programs over time. ## The trust dividend Organizations that implement comprehensive automated testing realize significant returns on their investment. The most obvious benefit is improved data quality, but the broader impact extends to organizational trust and data adoption. When business users consistently receive reliable data, they become more willing to base critical decisions on analytical insights. This increased trust creates a positive feedback loop. Higher data adoption leads to more sophisticated use cases, which generate additional value for the organization. Business users become advocates for data-driven decision making, creating organizational momentum for further data investments. The trust dividend also manifests in reduced operational overhead. When data quality issues become rare, data teams spend less time on firefighting and more time on strategic initiatives. This shift enables teams to focus on delivering new capabilities and insights rather than maintaining existing systems. Automated data testing represents more than a technical best practice; it's a strategic investment in organizational trust and capability. By implementing comprehensive testing throughout the data lifecycle, data engineering leaders can build the reliable, scalable data operations that modern businesses require. The result is not just better data quality, but stronger relationships with business stakeholders and increased organizational confidence in data-driven decision making. ## Data testing FAQs **What is data-driven testing (DDT)?** Data-driven testing is an approach that validates data quality through systematic testing of key data dimensions including accuracy, validity, completeness, freshness, and consistency. It involves automated testing processes that run throughout the data lifecycle: during development, code integration, and production to ensure data meets business requirements and quality standards before reaching end users. **What are the benefits of data-driven testing?** Data-driven testing provides multiple benefits including improved data quality, increased business user confidence, and reduced operational overhead. It creates predictable quality assurance processes, enables faster incident response through detailed audit trails, and allows data teams to focus on strategic initiatives rather than manual quality checks. Most importantly, it builds organizational trust that leads to higher data adoption and more sophisticated analytical use cases. **How do you manage test automation data and dependencies?** Test automation data and dependencies are managed through a layered approach combining unit tests, integration tests, and data tests. Unit tests validate individual model logic with known inputs, integration tests verify that multiple components work together correctly across data flows, and data tests focus on actual data content validation. This multi-layered framework is supported by automated workflows that run during scheduled intervals and code integration processes, with comprehensive documentation and audit trails to track data quality over time. --- --- title: "Building reliable data pipelines" description: "A foundational guide to reliable, scalable data pipelines using dbt, modular design, and end-to-end quality best practices." url: "https://www.getdbt.com/blog/building-reliable-data-pipelines" date: "2025-07-23" authors: ["Joey Gault"] categories: ["Pulse"] --- # Building reliable data pipelines A data pipeline is fundamentally a series of steps that move data from one or more sources to a destination system for storage, analysis, or operational use. The three core elements remain consistent: [sources, processing, and destination](https://www.databricks.com/glossary/data-pipelines#:~:text=There%20are%20usually%20three%20key,and%20destination%20being%20the%20same.). However, the way these elements interact has evolved significantly with cloud-native architectures. Traditional batch pipelines, which process data on fixed schedules, have largely given way to real-time streaming and cloud-native pipelines. This shift occurred because cloud platforms decouple storage and compute, enabling more flexible and cost-effective data processing at scale. The result is pipelines that can handle huge volumes of data at speed while maintaining reliability and cost efficiency. The architectural shift from [ETL (Extract, Transform, Load)](https://www.getdbt.com/blog/extract-transform-load) to [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) represents a fundamental change in how organizations approach data processing. Traditional ETL architectures transform data before loading it into a warehouse, which works well for on-premises systems with limited compute resources and structured data only. However, this approach becomes slower as data size increases and makes reprocessing or adding data later difficult. ELT pipelines load raw data into the warehouse first, then transform it using the processing power of cloud data warehouses. This approach enables loading and transformation to happen in parallel, provides flexibility for iterative data reprocessing, and scales efficiently with cloud data platforms. Tools like dbt are purpose-built for ELT workflows, allowing transformation to occur after data is already loaded into the warehouse for processing. ## Core components of reliable pipelines Modern data pipelines consist of several interconnected components that work together to ensure reliable data flow from source to destination. Understanding these components and their relationships is essential for building foundational reliability into your data infrastructure. **Ingestion** involves selecting and pulling raw data from source systems into a target system for further processing. Data engineers evaluate data variety, volume, and velocity (either manually or using automation) to ensure that only valuable data enters the pipeline. This evaluation process is crucial for maintaining pipeline efficiency and preventing downstream issues. **Loading** represents the step where raw data lands in cloud data warehouse or lakehouse platforms such as [Snowflake](https://www.snowflake.com/), [BigQuery](https://cloud.google.com/bigquery), [Redshift](https://aws.amazon.com/redshift/), or [Databricks](https://www.databricks.com/). This separation from ingestion emphasizes the "L" in ELT that allows data to be transformed within the data repository, taking advantage of the warehouse's processing capabilities. **Transformation** is where raw data is cleaned, modeled, and tested inside the warehouse. This includes filtering out irrelevant data, normalizing data to standard formats, and aggregating data for broader insights. With dbt, these transformations become modular, version-controlled code, making data workflows more scalable, testable, and collaborative. The transformation layer is often where the most business logic resides and where data quality is established. **Orchestration** schedules and manages the execution of pipeline steps, ensuring transformations run in the right order at the right time. Automated orchestration eliminates manual intervention and reduces the likelihood of human error while providing visibility into pipeline execution and dependencies. **Observability and testing** includes data quality checks, lineage tracking, and freshness monitoring. These components are critical for building trust, ensuring governance, and catching issues before they impact downstream analytics. Without proper observability, teams operate blindly, unable to detect anomalies or trace root causes when problems occur. **Storage** involves storing transformed data within a centralized data repository, typically a data warehouse or data lake, where it can be retrieved for analysis, business intelligence, and reporting. The storage layer must be designed for both performance and cost efficiency. **Analysis** ensures that stored data is ready for consumption by documenting final models, aligning them with a semantic layer, and making them easily discoverable. The semantic layer translates complex data structures into familiar business terms, making it easier for analysts and data scientists to explore data using SQL, machine learning, and BI tools. ## Common challenges in pipeline reliability Building reliable data pipelines requires addressing several persistent challenges that can undermine data quality and organizational trust. These challenges often compound each other, making it essential to address them systematically rather than in isolation. **Lack of observability across the data estate** represents one of the most significant challenges in modern data infrastructure. Today's data stacks are complex and fragmented, making it difficult for businesses to gain visibility across their entire data landscape. Without observability, data engineers cannot detect anomalies, trace root causes, assess schema changes, or ensure reliable data for analytics and AI applications. **Ingestion of low-quality data** creates problems that propagate throughout the entire pipeline. When data enters with missing values, schema drift, or inconsistent formats, it can silently corrupt insights and models downstream. Without clear understanding of data sources (including their reliability, quality, and governance policies) teams risk producing inaccurate outputs and drawing flawed conclusions that impact business decisions. **Pipeline scalability** becomes increasingly challenging as data volumes and sources multiply. Large, monolithic SQL scripts are difficult to debug and maintain, while pipeline bottlenecks slow down data processing and delay insights. The continued reliance on manual rather than automated processes further impacts scalability and speed, creating operational overhead that doesn't scale with business growth. **Untested transformations** represent a critical vulnerability in data pipelines. When transformations are deployed without automated testing, even minor logic errors or schema mismatches can propagate through the pipeline, resulting in broken downstream models, inaccurate dashboards and reports, or flawed machine learning outputs. The cost of these errors compounds as they move further from their source. **Lack of trust in data outputs** undermines the entire purpose of data infrastructure. In complex pipelines, trust breaks down when consumers struggle to identify reliable models or metrics. Data trust depends not only on accuracy, completeness, and relevance but also on transparency. Without clear lineage, semantic definitions, and documented assumptions, consumers may question a metric's validity or ignore it altogether. ## Foundational best practices for reliability Addressing these challenges requires implementing foundational best practices that create systematic reliability rather than ad-hoc solutions. These practices work together to create a comprehensive approach to pipeline reliability that scales with organizational growth. **Adopting a data product mindset** transforms how organizations think about data pipelines. Rather than simply moving data, this approach focuses on creating trustworthy, usable outputs that drive business value. ELT pipelines accelerate this transformation, but organizations must go further by treating data as products with clear ownership, documentation, and quality standards. This mindset reinforces data trust and empowers self-service analytics by ensuring that data outputs meet consumer needs. [dbt](https://www.getdbt.com/product/what-is-dbt) supports this approach through several key capabilities. Column-level lineage in [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) helps consumers understand the journey of individual columns from raw input to final analytical models, building trust by allowing users to trace data origins and transformations. The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) centralizes metric definitions, ensuring consistency across all pipelines and datasets while preventing metric drift and accelerating the creation of reliable, reusable data products. [Workflow governance](https://www.getdbt.com/product/governance) enables teams to standardize on a single platform, ensuring version control, lineage tracking, and access management that makes data transformations auditable and reliable. **Ensuring end-to-end data quality** prevents poor-quality data from creating silent errors, broken dashboards, and flawed insights that erode trust in analytics. Early data quality checks prevent errors from spreading through the pipeline, while automated validation, anomaly detection, and schema enforcement ensure reliable downstream analytics and applications. dbt addresses data quality through integration with industry-leading data quality, observability, and governance tools, ensuring that data entering pipelines won't cause downstream errors. [Built-in tests](https://docs.getdbt.com/docs/build/data-tests) validate uniqueness, non-null values, and referential integrity, while treating data like code with automated tests and reviews prevents failures from reaching production. [Continuous integration capabilities](https://docs.getdbt.com/guides/set-up-ci?step=1) allow teams to test code before deployment and monitor pipeline health, with dbt tracking production environment state and enabling CI jobs to validate modified models and their dependencies before merging. **Optimizing for scalability** ensures that pipelines can adapt to changing data structures, optimize query performance for fast analytics, and support real-time ingestion and integration while balancing scalability with cost. Linear scalability, where teams grow proportionally with pipelines, is unsustainable and must be avoided through architectural and tooling choices. dbt enables scalability through modular, [version-controlled transformations](https://docs.getdbt.com/docs/cloud/git/version-control-basics) that make pipelines easier to debug and maintain. Each model is self-contained, making it easier to isolate and fix errors without affecting the entire pipeline. Git-based version control tracks data changes as code, enabling teams to collaborate, audit, revert updates, and maintain a single source of truth for scalable pipeline management. Incremental model processing transforms only new or updated data, reducing costs, minimizing reprocessing, and improving efficiency while enhancing query performance and lowering warehouse load. Parallel microbatch execution processes data in smaller, concurrent batches, reducing processing time and improving efficiency compared to sequential runs. **Automating orchestration for scalable pipelines** ensures that data pipelines run efficiently by synchronizing processes from ingestion to analysis. Event-driven triggers and CI/CD automation detect failures early, reducing downtime and improving reliability. By automating execution, teams can scale pipelines without manual intervention, ensuring consistent and optimized workflows. dbt supports automated orchestration through [state-aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-about) that optimizes workflows by running models only when upstream data changes, reducing redundant executions and improving efficiency. Hooks automate operational tasks like managing permissions, optimizing tables, and executing cleanup operations, while macros bundle logic into reusable functions that enable parameterized workflows and event-driven execution. Integration with CI/CD workflows automatically tests modified models and their dependencies before merging to production, ensuring that changes don't break existing functionality. ## Implementation considerations Successfully implementing these foundational practices requires careful consideration of organizational context, technical constraints, and long-term goals. The most effective approaches balance immediate needs with sustainable, scalable solutions that grow with the organization. **Workflow processes** provide the foundation for reliable pipeline development and maintenance. Many data pipelines operate on an ad-hoc basis without established SLAs, version control, or systematic response procedures for pipeline issues. This leads to data that diverges from business requirements, untested changes pushed to production, broken pipelines that impair trust in data, and slow decision-making velocity. The [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) addresses these deficiencies by bringing software engineering best practices (version control, deployment pipelines, testing, and documentation) into the data world. Like the Software Development Lifecycle (SDLC), the ADLC breaks down barriers between data team members, ensuring that new pipelines and improvements align with core business objectives and have measurable key performance indicators. **Environmental separation** ensures that development work doesn't impact production systems while providing safe spaces for testing and validation. At minimum, pipeline development should include separate development environments for local testing, staging environments where CI/CD pipelines run comprehensive tests before production deployment, and production environments where only fully tested and approved changes run. **Quality culture** must be embedded throughout the development process rather than treated as an afterthought. According to industry research, poor data quality impacts a significant portion of company revenue, highlighting the need for quality controls at every stage of the ADLC. This includes creating tests for all new pipelines and changes, implementing version control-based processes for quality checkpoints, requiring manual code reviews through pull requests, and using CI/CD processes to test and deploy changes automatically to production. **Documentation and monitoring** enable pipeline maintainers to understand how and why certain decisions were made while making it easier for data consumers to have confidence in resulting datasets. Most pipelines go undocumented, making it harder for others to understand how to modify them, use the resulting data, or understand data derivation and calculation methods. Comprehensive documentation increases overall data trust and enables effective collaboration. Production monitoring ensures that pipelines continue operating smoothly after deployment by testing in production, reporting errors, and recovering gracefully from failures. Monitoring also ensures that code meets defined performance metrics and responds to user requests in a timely fashion, providing the observability necessary for maintaining reliable operations. ## Conclusion Building reliable data pipelines requires a foundational approach that addresses architectural choices, component design, common challenges, and implementation best practices systematically rather than piecemeal. The shift from ETL to ELT architectures, enabled by cloud-native data platforms, provides the technical foundation for scalable, reliable pipelines that can handle modern data volumes and complexity. However, technical architecture alone is insufficient. Reliable pipelines require adopting a data product mindset that focuses on creating trustworthy, usable outputs; ensuring end-to-end data quality through automated testing and validation; optimizing for scalability through modular design and automated orchestration; and implementing comprehensive workflow processes that embed quality and reliability throughout the development lifecycle. The challenges of observability, data quality, scalability, testing, and trust are interconnected and must be addressed holistically. Organizations that implement these foundational practices create data infrastructure that scales with business growth while maintaining the reliability and trustworthiness that stakeholders require for confident decision-making. [dbt](https://www.getdbt.com/product/dbt) provides a comprehensive platform for implementing these foundational practices, offering the tools and capabilities necessary to build, deploy, and maintain reliable data pipelines at scale. By treating data transformation as code, providing built-in testing and documentation capabilities, and enabling collaborative development workflows, dbt helps organizations create the reliable, scalable data infrastructure that modern business requires. The investment in foundational reliability pays dividends as organizations grow and data complexity increases. Rather than constantly fighting fires and rebuilding fragile systems, teams with solid foundations can focus on delivering business value through innovative data products and insights that drive competitive advantage. ## Data pipeline FAQs **What is a data pipeline?** A data pipeline is fundamentally a series of steps that move data from one or more sources to a destination system for storage, analysis, or operational use. The three core elements are sources, processing, and destination. Modern data pipelines have evolved significantly with cloud-native architectures, shifting from traditional batch processing on fixed schedules to real-time streaming and cloud-native pipelines that can handle huge volumes of data at speed while maintaining reliability and cost efficiency. **What is the difference between a data pipeline and ETL?** ETL (Extract, Transform, Load) is a traditional approach within data pipelines where data is transformed before loading it into a warehouse. This works well for on-premises systems with limited compute resources but becomes slower as data size increases. Modern data pipelines increasingly use ELT (Extract, Load, Transform), which loads raw data into the warehouse first, then transforms it using the processing power of cloud data warehouses. ELT enables loading and transformation to happen in parallel, provides flexibility for iterative data reprocessing, and scales efficiently with cloud data platforms. **What are the types of data pipelines?** Data pipelines have evolved from traditional batch pipelines that process data on fixed schedules to more modern approaches. The main types include batch pipelines for scheduled processing, real-time streaming pipelines for continuous data flow, and cloud-native pipelines that decouple storage and compute for flexible and cost-effective processing. Modern architectures also support event-driven pipelines and ELT workflows that leverage cloud data warehouse processing capabilities for transformation after data loading. --- --- title: "How Amazon S3 works" description: "Go under the hood of Amazon S3 with AWS engineering leader Andy Warfield—from virtualization to Iceberg." url: "https://www.getdbt.com/blog/how-amazon-s3-works" date: "2025-07-21" authors: ["Daniel Poppy"] categories: ["Learn"] --- # How Amazon S3 works _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/how-amazon-s3-works-w-andy-warfield). _ In this season of the Analytics Engineering podcast, Tristan is deep into the world of developer tools and databases. If you're following us here, you've almost definitely used Amazon S3 it and its Blob Storage siblings at Microsoft and Google. They form the foundation for nearly all data work in the cloud. In many ways, it was the innovations that happened inside of S3 that have unlocked all of the progress in cloud data over the last decade. In this episode, Tristan talks with Andy Warfield, VP and senior principal engineer at AWS, where he focuses primarily on storage. They go deep on S3, how it works, and what it unlocks. They close out talking about Iceberg, S3 table buckets, and what this all suggests about the outlines of the S3 product roadmap moving forward. _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways ### Operating systems, garage sales, and Xen **Tristan Handy: You’ve done a lot over the last 20 years. Before we get into specifics, can you just share a little about your journey as a software engineer?** **Andy Warfield:** I just like playing with computers. I studied computer science in Ontario for undergrad, then moved to Vancouver for grad school, then to the UK for a PhD. I worked on operating systems, low-level stuff. I got to work on a hypervisor called Xen, which ended up being used by a lot of cloud providers, including Amazon. After that, I did a couple of startups, one around Xen. Then I became a professor at UBC, teaching operating systems, networking, and security. Later, I did another startup in storage, and eventually I joined Amazon. Now I have this highfalutin role—VP and engineer—working across S3, other storage services, and now a bunch of analytics services too. I get to cause trouble in lots of different parts of the cloud. **VP slash distinguished engineer—does that mean you just get to march around telling people how to improve their stuff?** People love that! I’d say about half the time I’m causing trouble—starting things and encouraging new ideas—and the other half I’m helping teams dig out from those ideas. Sometimes I take over a team if we’re doing something especially interesting or innovative, just so I can be closer to the action. **That sounds like a pretty good gig if you can get it.** It’s amazing. I’ve been here nearly eight years, and I still love this job. ### The rise of virtualization and the origin of Xen **I want to talk about Xen. You said you were always interested in operating systems, which is kind of a niche fascination. What drew you in?** When I was a kid, we didn’t have much money, so I built computers from garage sale parts in Ottawa. In high school, I found this federal government warehouse that sold off old equipment. I started a little business buying pallets of hardware for cheap, fixing them up, and reselling. It was chaotic—but I learned a lot. I dealt with machines like IBM DisplayWriters with 8-inch floppy disks and massive dot-matrix printers. Getting them working meant diving into their software and systems. Eventually I played with Linux, hacked on the kernel, and that all led me into OS research and development. **Tristan: So what is a hypervisor, and why did virtualization become so important in the 2000s?** **Andy:** There were two big drivers: server utilization and isolation. Companies had racks full of 1U servers, most of which sat idle most of the time. But they couldn’t share workloads because apps weren’t isolated well—config conflicts, shared resources, etc. Virtualization allowed multiple operating systems to run on the same hardware, with isolation. It also let you consolidate servers, which had big cost and efficiency benefits. There was also a technical challenge: x86 processors weren’t designed to be virtualized. That made it a really interesting research problem. We wanted to see if it could even be done—and done efficiently. **Tristan: And Intel eventually started building virtualization support into the hardware?** **Andy:** Exactly. Our work on Xen and similar projects showed it was possible. That pushed Intel and AMD to add features like VT-x, which made it easier and more performant to run hypervisors. **Tristan: How did AWS end up using Xen?** **Andy:** I wasn’t part of those internal conversations, but the story goes that a small startup in Cape Town, South Africa, was building a control plane for Xen. That team got picked up by AWS and became the basis for EC2. ### Understanding Amazon S3 **Tristan: Let’s switch to S3. I think a common mental model is that S3 is just a big pool of SSDs. But that’s clearly not the whole story. How do you explain what S3 actually is?** **Andy:** That’s one of my favorite questions. Early on, S3 was like a storage locker. You’d rent space to stash things you didn’t need right away—backups, static files, CDN origins. Latency wasn’t great, but durability and availability were. Things really changed when the Hadoop community built S3A—an adapter to let Hadoop use S3 instead of HDFS. Suddenly, we had people doing real analytics on S3. The system had enough drives to support massive parallel reads. Today, workloads are way more demanding. Performance, consistency, and latency matter. We’ve been evolving the system constantly to meet those needs. **Tristan: Are we talking about billions of hard drives?** **Andy:** I can’t share exact numbers, but yes—it's a lot of hard drives. Some of our largest customers have data spread across _millions_ of drives. And most drives are shared across multiple customers. **Tristan: And these aren’t SSDs?** **Andy:** Mostly spinning disks, actually. Hard drives are terrible at latency, but they’re cheap and good for bursty workloads. Spreading your data across many disks lets you take advantage of parallelism. ### S3’s durability, performance, and scale **Tristan: Let’s talk about S3’s durability promise: 11 nines. How do you achieve that?** **Andy:** We use erasure coding—a form of RAID-like redundancy that lets you split data into parts and parity blocks. Then we store those shards across different availability zones. We constantly monitor for failures. Disks die all the time, so we have fleets of processes repairing and maintaining durability. It’s not static. It’s a living system. **Tristan: You must have incredibly precise failure models.** **Andy:** We do. We track failure rates, temperature sensitivity, vendor behavior—everything. That allows us to be proactive and surgical in how we manage risk. ### From Parquet to Iceberg to S3 table buckets **Tristan: I want to talk about table formats. Parquet is everywhere now. And then we got Hive Metastore, then Iceberg. Why did S3 launch table buckets?** Parquet is great, but it’s just files. Customers kept asking for more structured semantics: schema evolution, upserts, ACID transactions. We saw Iceberg adoption grow rapidly—especially among our largest analytics customers. But they were struggling with operational complexity: too many small files, custom compactors, brittle catalogs. So we launched S3 table buckets to bring native Iceberg support to S3. That includes: - Automatic compaction - A REST catalog - High-performance access We wanted to make it easier to treat Iceberg as a storage primitive, not just an analytics backend. **So this is a shift in philosophy—S3 isn’t just object storage, it’s now table-aware?** Exactly. Historically, S3 was just where you stored objects. Now, we’re thinking more about what those objects _mean_. We also launched S3 object metadata tables—a way to semantically describe and query your object store, especially useful for AI workloads using retrieval-augmented generation (RAG). ### The future of open data and S3 **What does the future of S3 look like? Where’s this going?** We’re headed toward more structure, more semantics, and more performance. Inference workloads are scaling fast. AI models are hitting S3 hundreds of thousands of times per second to do vector lookups. That’s changing how we think about indexing, metadata, and latency. We want to make S3 the best place to do open, flexible, high-scale data work—from tables to training data to retrieval. ## Chapters **[01:42] Meet Andy Warfield** Andy shares his background, including startups, professorship, and his current role as VP & Senior Principal Engineer at AWS. **[05:10] From garage sales to hypervisors** Andy describes his early passion for hardware, OS development, and the origin story behind the Xen hypervisor. **[08:50] Why virtualization took off in the 2000s** Exploring why isolation, utilization, and technical curiosity fueled the rise of hypervisors. **[14:30] Xen vs. VMware and the road to AWS** How Xen became the default for EC2 and the technical differences between virtualization approaches. **[17:35] The origin of EC2 and S3** How a team from Cape Town helped launch AWS compute—and the early days of cloud services. **[20:00] What is S3, really?** Andy breaks down the mental model behind S3: not just object storage, but a scalable data platform. **[22:49] How many drives? More than you think** Why S3 storage spans millions of drives—and how AWS uses scale to deliver performance. **[28:10] The 11 nines durability model** Inside S3’s approach to reliability, failure tolerance, and background repairs using erasure coding. **[32:00] Tail latency and engineering for bursty workloads** Why slow requests matter, and how S3 teams optimize for streaming, AI, and analytics use cases. **[35:20] Iceberg, metadata, and table buckets** The emergence of Apache Iceberg as a table format—and AWS’s new structured storage approach. **[38:00] Why S3 added a REST catalog and compaction** How AWS is simplifying the operational burden of working with Iceberg at scale. **[40:00] A new mental model for object storage** S3 is no longer just about storing files—it’s about managing semantics, lineage, and trust. **[44:00] Looking ahead: S3, RAG, and semantic metadata** How S3 is preparing for the next wave of AI, inference, and context-aware applications. **[47:20] Is Iceberg ready for enterprise?** Andy shares thoughts on enterprise readiness, performance tradeoffs, and real-world adoption of table formats. **[49:05] Wrap-up and reflections** Tristan and Andy reflect on the conversation and where data infrastructure is headed next. --- --- title: "Cost-cutting strategies for Amazon Redshift" description: "Learn how to reduce Redshift costs with smarter pricing, efficient transformations, and dbt Fusion’s cost-aware features." url: "https://www.getdbt.com/blog/amazon-redshift-cost-cutting" date: "2025-07-16" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Cost-cutting strategies for Amazon Redshift [Amazon Redshift](https://aws.amazon.com/redshift/) is a fully managed cloud data warehouse built for fast, cost-efficient analysis of large datasets. It handles structured and semi-structured data with optimized SQL execution and seamless integration with BI tools. While designed for performance and [low total cost of ownership](https://aws.amazon.com/redshift/pricing/), Redshift costs can vary depending on configuration, region, query patterns, and storage requirements. This article examines the factors that frequently lead to high usage costs and offers cost-cutting strategies to help you optimize the value of your Redshift warehouse. ## Redshift pricing unpacked [AWS](https://aws.amazon.com/redshift/pricing/) offers several pricing options for Redshift, based on a pay-as-you-go model. Which pricing option you choose depends on how consistent your workloads are, how much flexibility you need, and how quickly your data is growing. In addition to information available on AWS’s pricing page, [CloudZero](https://www.cloudzero.com/blog/redshift-pricing/) provides a fuller description of the following simplified pricing options: - Provisioned (on-demand vs reserved) - Serverless - Concurrency Scaling and Redshift Spectrum - DC2 vs RA3 nodes and Redshift Managed Storage ### Provisioned (on-demand vs reserved instances) **On-demand pricing** - With AWS on-demand pricing, you pay hourly per node, with flexibility to modify or pause clusters as workloads shift. Pausing suspends compute charges but retains backup storage costs. Although this is the [most flexible pricing model](https://www.cloudzero.com/blog/redshift-pricing/), it can be the most costly. **Reserved instances** - Reserved instances offer up to 75% savings for one- or three-year commitments. This option requires accurate forecasting to avoid overpaying for unused capacity or experiencing performance issues due to overcapacity. ### Serverless **Redshift Serverless** - Redshift Serverless bills per second (with a 60-second minimum) based on active Redshift Processing Units (RPUs). It incurs no charges during idle time, making it ideal for dynamic workloads. This option is great for spiky demands or unpredictable ETL. However, it can lead to runaway RPU spend if workloads aren’t adequately monitored. Redshift Serverless includes Concurrency Scaling and Redshift Spectrum. ### Concurrency Scaling and Redshift Spectrum [**Concurrency Scaling**](https://docs.aws.amazon.com/redshift/latest/dg/concurrency-scaling.html) - Automatically adjusts cluster capacity to handle spikes in concurrent queries. You get one free hour every 24 hours (and can bank up to 30 hours); beyond that, billing is per second. Amazon reports that the free credits cover [97% of customer needs.](https://aws.amazon.com/redshift/features/concurrency-scaling/) [**Redshift Spectrum**](https://docs.aws.amazon.com/redshift/latest/dg/c-using-spectrum.html) - Query S3 data directly without importing it into Redshift. Pricing is based on the number of bytes scanned, with a minimum of 10 MB per query. Querying uncompressed or unpartitioned data sets can quickly spike query costs. ### DC2 vs RA3 nodes with Redshift Managed Storage **DC2 nodes **- Bundle storage and compute on SSDs, making them fast and cost-effective for datasets under 1 TB. However, DC2 nodes require simultaneous scaling of both storage and compute, which can become out of sync and costly. **RA3 nodes** - For larger datasets, RA3 nodes separate compute from storage via [Redshift Managed Storage (RMS)](https://hevodata.com/learn/comprehensive-guide-to-redshift-data-storage/), which automatically tiers frequently accessed data to SSDs. You’re charged hourly for storage—not for the number of nodes—offering more flexibility and predictable costs. RMS also powers Redshift Serverless. Finding the right pricing strategy is just the beginning. The following section discusses several additional Redshift cost escalation challenges. ## Cost escalation factors in Redshift Redshift’s pricing flexibility is valuable, but without strategic planning, monitoring, and refinement, costs can escalate quickly. To manage spend effectively, teams must watch for these common risk factors: - Idle or over-provisioned clusters - Inefficient data loads and queries - Uncontrolled scaling features - Inadequate workload oversight and cost attribution ### Idle or over-provisioned clusters **Inactive clusters drive up costs** – On-demand clusters left running during off-hours continue accruing compute charges, even when idle. **Poor planning leads to waste** – Inaccurate forecasting and weak workload monitoring result in underused reservations and excess capacity across both on-demand and reserved clusters. **Node selection impacts flexibility** – DC2 forces joint scaling of compute and storage, while RA3 decouples them—but unmanaged storage growth can still generate unexpected fees. ### Inefficient data loads and queries **Poor query and table design wastes resources** – Inefficient SQL and suboptimal schema choices drive up compute by overusing disk and memory. **Full-refresh transformations inflate costs** – Rebuilding entire tables instead of updating new rows burns excess compute and lengthens runtime. **Unoptimized Spectrum queries scan excessive data – **Running against unpartitioned or uncompressed S3 files triggers unnecessary $5-per-TB scan charges. ### Uncontrolled scaling features **Concurrency Scaling can get expensive fast** – After the one free hour, per-second fees escalate quickly for high-concurrency workloads if usage isn’t tightly controlled. **Serverless auto-scaling drives unpredictable spend** – Rapid scaling during peak demand can balloon per-second RPU charges unless limits are set and usage is monitored. **Elastic Resize and autoscaling need guardrails** – These powerful features improve flexibility, but without caps, they risk runaway costs during surges in demand. ### Inadequate workload oversight and cost attribution **Cost attribution is complex and manual** – Redshift doesn’t offer native model-level reporting, so tracking costs by user or team requires [custom SQL and tagging strategies](https://aws.amazon.com/blogs/big-data/how-to-attribute-amazon-redshift-costs-to-your-end-users/). **Lack of oversight drives hidden spend** – Without strong policies and workload monitoring, inefficient queries, idle resources, and redundant jobs can quietly inflate costs. **CI/CD pipelines can bloat budgets** – Rebuilding unchanged models or rerunning unneeded tests in automated pipelines adds unnecessary compute charges over time. These challenges underscore the need for proactive cost management and architectural awareness. ## Cost cutting strategies for Redshift Redshift cost management requires aligning pricing models, cluster design, data practices, and transformation workflows with real-world usage patterns to optimize your Redshift usage. The following best practices will help you reduce waste, improve performance, and lower costs: - Optimize pricing, clusters, and node types - Streamline data management and query design - Use incremental models to reduce compute spend - Monitor usage and enforce cost controls ### Optimize pricing, clusters, and node types **Choose the right pricing model and cluster configuration** – Align pricing and configuration with actual usage patterns. Pause idle on-demand clusters to avoid charges, and use reserved instances only if you can accurately forecast and predict your workloads. **Right-size clusters and node types** – Avoid paying for excess capacity by selecting the right node type. For growing datasets, choose RA3 nodes to scale storage independently from compute. If your workloads are bursty or unpredictable, configure Redshift Serverless or Concurrency Scaling with strict usage caps to prevent runaway RPU or per-second fees. **Deploy multiple Redshift instances** – Consider adopting [multiple Redshift warehouses](https://aws.amazon.com/blogs/big-data/improve-your-etl-performance-using-multiple-redshift-warehouses-for-writes/) tailored to specific workloads, assigning dedicated clusters for ETL, development, production, testing, or departmental use, etc. This improves performance, prevents resource contention, and enables granular cost tracking by team or environment. ### Streamline data management and query design **Store your data appropriately** – Keep high-use, operational data in Redshift for faster performance and offload infrequently used or historical data to S3, where you can use Spectrum for lower-cost querying. Minimize Spectrum’s $5-per-terabyte scan charges by partitioning and compressing external files and aggregating raw data where possible. **Enable data compression** – Compression reduces disk I/O and improves efficiency, especially for large or frequently updated tables. Run [VACUUM and ANALYZE](https://docs.aws.amazon.com/prescriptive-guidance/latest/query-best-practices-redshift/query-performance-factors.html) to reclaim space and optimize query execution. Use [dbt](https://www.getdbt.com/product/dbt) to apply [performance-aware model configurations](https://docs.getdbt.com/reference/model-configs), such as sort and distribution keys, directly in model files. **Use dbt to manage transformations** – Once semi-structured data is staged in Redshift or made accessible via Spectrum, use dbt to [structure and transform](https://aws.amazon.com/blogs/big-data/manage-data-transformations-with-dbt-in-amazon-redshift/) it. dbt supports incremental updates, modular SQL, and [lineage tracking](https://www.getdbt.com/blog/getting-started-with-data-lineage)—reducing scan volume, cutting ETL overhead, and improving transformation consistency across environments. While dbt doesn’t ingest raw data, it excels at managing transformations once the data is queryable in Redshift. ### Use incremental models to reduce compute spend **Replace full table refreshes with incremental materialization** – Full-table rebuilds are expensive, especially in CI/CD pipelines that reprocess unchanged data. Use [incremental materialization in dbt](https://docs.getdbt.com/docs/build/incremental-models) to transform only new or changed rows, reducing query time and compute usage. Combined with Redshift’s [materialized view](https://docs.aws.amazon.com/redshift/latest/dg/materialized-view-overview.html) refresh, this approach streamlines deployments and cuts overhead. **Use deferred builds in dbt** – [Deferred builds](https://docs.getdbt.com/docs/cloud/about-cloud-develop-defer) enable you to test models in isolation using production artifacts, thereby skipping full rebuilds. This lowers compute costs, shortens development cycles, and avoids redundant storage, especially in large Directed Acyclic Graphs (DAGs) or CI pipelines where full rebuilds are expensive and time-consuming. ### Monitor usage and enforce cost controls **Track Redshift performance and usage **-** **Use [Amazon Redshift Advisor](https://docs.aws.amazon.com/redshift/latest/dg/advisor.html) to identify optimization opportunities and regularly review cluster sizing to eliminate idle capacity. Tag resources and monitor spend with tools like [AWS Cost Explorer](https://aws.amazon.com/aws-cost-management/aws-cost-explorer/), [Budgets](https://aws.amazon.com/aws-cost-management/aws-budgets/), and [Trusted Advisor](https://aws.amazon.com/premiumsupport/technology/trusted-advisor/) for clear cost attribution across teams. **Set usage caps** - For unpredictable workloads, configure Concurrency Scaling and Redshift Serverless with strict usage caps to prevent runaway charges. Monitor query performance and storage growth to ensure your architecture remains aligned with business needs. Redshift cost optimization isn’t a one-time fix. By continuously aligning your pricing models, cluster configurations, data management strategies, and transformation workflows with actual workload patterns, you can effectively manage costs while maintaining optimal performance, maximizing the value of Redshift’s speed and scalability. ## How dbt Fusion helps optimize Redshift costs Best practices lay the groundwork for Redshift cost optimization. [dbt builds ](https://www.getdbt.com/data-platforms/redshift)on these foundations, simplifying transformations and driving more strategic use of Redshift resources: - Smarter configuration and [streamlined execution](https://www.getdbt.com/blog/redshift-dbt). - Aligning sort and distribution keys with query patterns. - Choosing efficient materializations and using incremental models to avoid redundant compute. - [Modular transformation](https://www.getdbt.com/blog/modular-data-modeling-techniques) for standardized testing and clear documentation. [dbt Fusion](https://www.getdbt.com/product/fusion) adds advanced orchestration features that make cost optimization even more consistent across your environment: - Develop and test locally without spinning up Redshift - Optimize Redshift costs with dbt Fusion’s intelligent runtime - Accelerate performance with dbt Fusion’s Rust engine ### Develop and test locally without spinning up Redshift Fusion enables you to build and validate models locally, referencing production artifacts without requiring a remote Redshift instance. This speeds up iteration, reduces development overhead, and avoids unnecessary cloud compute costs, especially in CI workflows or large DAGs. Its Rust-based engine powers fast, accurate SQL [compilation and comprehension](https://docs.getdbt.com/blog/dbt-fusion-engine). **** ### Optimize Redshift costs with Fusion's intelligent runtime Fusion’s [state-aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-setup) runs only models impacted by upstream changes, cutting redundant builds and reducing RPU usage. Teams can [expect ~10% savings](https://www.getdbt.com/blog/where-we-re-headed-with-the-dbt-fusion-engine) on data platform costs. Fusion also supports ahead-of-time compilation and static SQL analysis, helping your teams catch inefficiencies before they hit the warehouse. ### Accelerate performance with Fusion's Rust engine Fusion’s orchestration engine, written in Rust, delivers [30 times faster parsing](https://www.getdbt.com/blog/where-we-re-headed-with-the-dbt-fusion-engine) and compiles twice as fast as dbt Core, improving CI throughput and shortening development cycles. While not a direct Redshift cost reducer, this speed enables more efficient orchestration and earlier error detection, minimizing full-model rebuilds and excess compute. ## Conclusion Cutting costs in Redshift depends on consistently executing best practices and adopting tools built for efficiency. dbt Fusion help teams streamline workflows, reduce compute overhead, and manage resources more strategically. [Request a demo](https://www.getdbt.com/contact) to see how dbt Fusion can help you improve performance and cut Redshift costs. --- --- title: "Who should own the semantic layer?" description: "Semantic layer ownership depends on your org’s size, needs, and tools—but one team shouldn’t own everything. Here’s why." url: "https://www.getdbt.com/blog/semantic-layer-ownership" date: "2025-07-15" authors: ["Joey Gault"] categories: ["Pulse"] --- # Who should own the semantic layer? As organizations scale, so does the complexity of their data. Different teams need different insights, but they all need to agree on the numbers. That’s where the [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) comes in: a centralized translation layer that defines business metrics once and makes them accessible across tools. But without clear ownership, semantic layers can drift into inconsistency, technical debt, or worse—loss of trust. This article explores who should own the semantic layer and how to strike the right balance between technical stewardship and business accountability. ## The case for data team ownership Traditionally, data teams have been the natural custodians of semantic layers, and there are compelling reasons why this arrangement makes sense. Data engineers and [analytics engineers](https://www.getdbt.com/blog/what-does-an-analytics-engineer-do) possess the technical expertise required to navigate the complexities of metric definition, query optimization, and system integration. They understand the underlying data models, the nuances of SQL generation, and the performance implications of different semantic layer configurations. The technical complexity of modern semantic layers cannot be understated. As [Transform's founders ](https://www.getdbt.com/blog/dbt-labs-signs-definitive-agreement-to-acquire-transform-accelerating-development-of-the-dbt-semantic-layer)discovered while building [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow), even seemingly simple metrics can generate hundreds of lines of optimized SQL when accounting for proper joins, time series analysis, and dimensional breakdowns. When you want to analyze entities across multiple tables or compare metrics with different calculation logic, the resulting queries become increasingly sophisticated. This level of technical complexity naturally aligns with the skill sets that data teams have developed. Furthermore, data teams are already responsible for the foundational infrastructure that semantic layers depend upon. They build and maintain the models that serve as the basis for semantic models, manage data warehouse performance, and understand the broader data architecture. Having them own the semantic layer creates a cohesive ownership model where the same team responsible for data transformation also handles metric definition and semantic modeling. From an operational perspective, data teams are equipped to handle the administrative complexities that come with semantic layer deployment. The process of configuring credentials, managing service tokens, setting up integrations with downstream tools, and ensuring proper access controls requires deep technical knowledge of both the data platform and the semantic layer technology itself. Data teams already have established processes for managing these types of infrastructure components. ## The business case for domain ownership However, there's a compelling argument that business teams should own semantic layer definitions, even if they don't manage the underlying infrastructure. The fundamental purpose of a semantic layer is to create precision and consistency around business concepts: revenue, customer count, churn rate, and other metrics that drive organizational decision-making. The people who best understand these concepts are typically not data engineers, but rather the business stakeholders who use these metrics daily. Business teams have the deepest knowledge of how metrics should be calculated, what edge cases need to be considered, and how definitions should evolve as the business changes. They understand the nuances that distinguish a customer from a user, or the specific business logic that should be applied when calculating monthly recurring revenue. This domain expertise is crucial for creating semantic layers that truly serve the organization's analytical needs. The vision articulated by Transform's founders suggests that semantic layers should enable business people to define their own metrics and concepts, focusing purely on business complexity rather than technical implementation details. This approach would free business teams from dependence on data teams for metric definition while allowing them to maintain control over the business logic that drives their analysis. When business teams own metric definitions, they can respond more quickly to changing business requirements. Rather than submitting requests to data teams and waiting for implementation, they can directly modify metric logic as business processes evolve. This agility becomes increasingly important as organizations seek to become more data-driven and responsive to market changes. ## They hybrid approach In practice, the most effective ownership model likely involves a hybrid approach that leverages the strengths of both data and business teams. The technical infrastructure of the semantic layer (the deployment, credential management, integration setup, and performance optimization) naturally falls to data teams. They have the expertise to ensure that the semantic layer operates reliably and efficiently within the broader data architecture. Meanwhile, the definition and governance of metrics and business concepts can be owned by domain experts within business teams. This requires semantic layer tools to provide business-friendly interfaces for metric definition, moving beyond [YAML](https://yaml.org/) file editing toward more intuitive user experiences. The goal is to abstract away technical complexity while preserving the precision required for accurate metric calculation. This hybrid model aligns with the principle that ownership should reside with those most incentivized to ensure correctness. Data teams are motivated to maintain reliable, performant infrastructure, while business teams are motivated to ensure that metrics accurately reflect business reality. By dividing responsibilities along these lines, organizations can leverage the expertise of both groups. The success of this approach depends heavily on the user experience provided by semantic layer tools. As the space matures, we can expect to see more sophisticated interfaces that allow business users to define complex metrics without needing to understand the underlying technical implementation. These tools must strike a balance between simplicity and power, enabling business users to express complex business logic while generating optimized SQL behind the scenes. ## Organizational considerations The ownership decision also depends on organizational maturity and structure. In smaller organizations or those with highly technical business teams, direct business ownership of semantic layer definitions may be feasible. These environments often have fewer stakeholders and simpler governance requirements, making it easier for business teams to manage metric definitions directly. Larger organizations with complex governance requirements may need more structured approaches. They might implement approval workflows where business teams propose metric definitions that are reviewed and implemented by data teams. Alternatively, they might establish centers of excellence that include both business and technical stakeholders, ensuring that metric definitions reflect both business requirements and technical best practices. The choice of semantic layer technology also influences ownership patterns. Some tools are designed with business users in mind, providing graphical interfaces and simplified configuration options. Others are more technically oriented, requiring deeper understanding of data modeling concepts and SQL generation. Organizations should consider their intended ownership model when evaluating semantic layer solutions. ## The path forwards As semantic layers become more prevalent, we're likely to see continued evolution in ownership models. The [acquisition of Transform by dbt Labs](https://www.getdbt.com/blog/dbt-labs-signs-definitive-agreement-to-acquire-transform-accelerating-development-of-the-dbt-semantic-layer) represents a significant step toward making semantic layers more accessible and powerful, potentially enabling new ownership patterns as the technology matures. The integration of [MetricFlow's capabilities](https://docs.getdbt.com/docs/build/about-metricflow) into [dbt's semantic layer](https://www.getdbt.com/product/semantic-layer) will likely influence how organizations think about ownership. Since [dbt](https://www.getdbt.com/product/what-is-dbt) has successfully enabled analytics engineers to own data transformation processes, the enhanced semantic layer may similarly enable business-technical hybrid roles to emerge around metric definition and governance. Ultimately, the question of who should own the semantic layer doesn't have a universal answer. The optimal approach depends on organizational structure, technical capabilities, governance requirements, and the specific tools being used. What matters most is that organizations thoughtfully consider this question and establish clear ownership models that align with their broader data strategy. The semantic layer represents a critical opportunity to bridge the gap between technical data infrastructure and business decision-making. By carefully considering ownership models and implementing appropriate governance structures, organizations can ensure that their semantic layers truly serve their intended purpose: creating consistent, reliable access to business metrics across all analytical and operational systems. As the technology continues to evolve and mature, we can expect to see new patterns emerge that further refine how organizations approach semantic layer ownership. The key is to remain flexible and responsive to both technological capabilities and organizational needs while maintaining focus on the ultimate goal of enabling better, more consistent business decision-making through data. ## Semantic layer ownership FAQs **Who is considered the semantic model owner, and how can ownership be taken over or transferred in the semantic model settings?** The semantic model owner is typically determined by organizational structure and technical capabilities rather than specific system settings. In traditional models, data teams serve as natural custodians due to their technical expertise in metric definition, query optimization, and system integration. However, ownership can shift toward business teams who possess deeper domain knowledge of how metrics should be calculated and what business logic should be applied. The transfer of ownership often involves establishing clear governance structures and approval workflows that align with the organization's data strategy. **What configuration tasks are exclusive to the semantic model owner, such as scheduled refresh, credentials, and automatic aggregations?** Technical infrastructure tasks including deployment, credential management, integration setup, and performance optimization typically remain with data teams regardless of who owns metric definitions. These tasks require deep technical knowledge of data platforms and semantic layer technology. Data teams handle the administrative complexities of configuring credentials, managing service tokens, setting up integrations with downstream tools, and ensuring proper access controls. They also manage the foundational infrastructure that semantic layers depend upon, including dbt models and data warehouse performance. **What is the role of data ownership in a semantic layer?** Data ownership in a semantic layer serves to create precision and consistency around business concepts that drive organizational decision-making. The role involves defining metrics like revenue, customer count, and churn rate, understanding calculation nuances and edge cases, and ensuring definitions evolve appropriately as business requirements change. Effective ownership leverages domain expertise to maintain control over business logic while ensuring metrics accurately reflect business reality. The goal is to bridge the gap between technical data infrastructure and business decision-making, enabling consistent and reliable access to business metrics across analytical and operational systems. --- --- title: "How dbt and Tableau bring better governance to analytics" description: "Self-service data can create a “data Wild West.” Here’s how to tame it using dbt and Tableau." url: "https://www.getdbt.com/blog/dbt-tableau-trust-analytics" date: "2025-07-11" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # How dbt and Tableau bring better governance to analytics With great self-service comes great responsibility. While easier access to data has broken down barriers to information, it’s also led to a “data wild west” that undermines data trust. dbt and [Tableau](https://www.tableau.com/) have partnered to address three critical challenges facing modern analytics teams. Organizations that use them in concert have improved data governance, increased data trust, and accelerated productivity. Deputy, an HR management software company, is a prime example. The company faced numerous challenges with its existing data architecture that eroded trust in data. In this article, we’ll look at these challenges, how the dbt and Tableau integration improves data governance, and the remarkable results that Deputy enjoyed after moving to dbt as their data control plane. **** ## The modern data trust crisis Only 57% of data analytics leaders say they have complete confidence in their data's accuracy. This trust deficit stems from what’s become known as the "wild west" of self-service analytics. Tools like Tableau have democratized data access by enabling self-service access to data for BI. Unfortunately, this also introduces new challenges. Organizations often have hundreds of dashboards running, some built using a myriad of tools, some built on custom SQL queries that connect directly to data warehouses. This freedom to connect anywhere creates environments with no standardization, massive duplication of effort, and unclear data sources. Many analytics teams struggle with data existing across multiple platforms—[Google Sheets](https://sheets.google.com/), [BigQuery](https://cloud.google.com/bigquery?hl=en), [Snowflake](https://www.snowflake.com/)—and often duplicated across systems. Building data models in Tableau becomes unnecessarily complicated when it's unclear which source contains production-ready data. The lack of consistent documentation can make tracing data back to its source nearly impossible. ### Deputy's breaking point Deputy's situation exemplified these challenges at scale. As a fast-growing HR management software company, their already complicated architecture became increasingly convoluted. Simple requests took far too long to complete. Developers created complicated workarounds in dashboards because the architecture wasn't designed with BI in mind. The breaking point came when even the data team—the experts—could no longer efficiently navigate their own system. Deputy's data director, Huss Azfal, admitted that "leadership had nothing but horror stories about working with the data team." This wasn't a people problem. It was an architecture problem that made everyone's jobs harder than necessary. ## The three core challenges The problems faced by Deputy illuminate three core problems with data development: **No consistent approach to building analytics pipelines**. There’s no one, single location to find data. For data producers, this leads to duplicated effort and redundant code. For data consumers, it makes it hard or impossible to find high-quality, polished datasets. **Siloed data leads to inconsistencies**. This lack of visibility into what other teams are doing means everyone ends up doing the same thing, but differently. We’ve seen different teams define the same metric in different ways in order to suit the needs of their particular use case or dashboard. That leads to mass duplication of effort, which further contributes to data trust issues. **Poor visibility into data quality**. We’ve all looked at a dashboard and wondered if the numbers were accurate. Even if data consumers can find data, they lack the tools to verify the data’s quality, determine its freshness, or trace it back to its source. ## How dbt transforms data pipeline development dbt addresses these challenges by allowing analytics teams to prepare data in a modular, centralized, and governed way. It does this in three ways: - Bringing engineering practices to data analytics through the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle), a process that accelerates data velocity while improving data maturity - Centralizing all data transformations in a singular [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction), enabling data discovery and collaboration. - Automated testing and version control to increase data quality, collaboration, and traceability ### From monolithic to modular This transformation is immediately visible in how analytics code is structured. Let’s look at an example. Legacy SQL scripts feature table references scattered throughout, embedded Common Table Expressions (CTEs), and nested subqueries, making them difficult to read and debug. With dbt, these monolithic scripts break down into smaller, manageable "models" that reference each other. You could, for example, structure an order items model into different stages derived from multiple upstream models: two [staging models](https://docs.getdbt.com/best-practices/how-we-structure/2-staging) and one [intermediate](https://docs.getdbt.com/best-practices/how-we-structure/3-intermediate) model. The modular approach creates more readable code with better white space and clearer logic flow. It's also more reusable—models can be referenced in multiple downstream models, eliminating code duplication. When a staging orders model needs updating, developers only need to change it once, rather than hunting down every instance across multiple scripts. This approach saves analysts time, eliminates redundancy, and accelerates the development lifecycle. **** ### Built-in quality and collaboration dbt's GitHub integration enables parallel development, allowing multiple developers to work simultaneously on different models. More importantly, dbt enables specific tests on individual models. Developers can easily test, for example, whether a column contains unique values or has nulls, thereby identifying errors early and preventing them from cascading downstream. The automated lineage is another game-changer. dbt [automatically generates data lineage](https://www.getdbt.com/blog/what-is-data-lineage), showing how metrics like gross profit trace back through revenue and supply cost calculations to source data. This easily produced lineage provides complete transparency without manual effort. ## How dbt and Tableau work together dbt and Tableau have worked closely together to enable access to key features of dbt from Tableau reports. Using dbt’s Semantic Layer and data health tiles, Tableau users can access standardized metrics and track overall data health. This results in less redundancy and greater data trust. ### Semantic Layer: Centralizing business logic The [Semantic Layer](https://www.getdbt.com/blog/semantic-layer-introduction) is one of dbt's most powerful features for ensuring consistency across an organization. It serves as an abstraction layer that translates raw data into business terms that companies use daily. When a company typically defines a metric such as "typical revenue,” this definition lives in multiple places, such as individual dashboards or analyst queries, often with different definitions. The semantic layer defines entities (primary and foreign keys), dimensions, measures, and complex metrics, such as gross profit (revenue minus cost). In dbt, total revenue would live in the Semantic Layer, providing a consistent, discoverable, and accessible result to anyone who requires it. Unlike pre-aggregated views, Semantic Layer metrics are calculated dynamically at query time. This approach enables non-additive metrics—percentages that can't simply be summed to get accurate totals at different aggregation levels. For example, if you have a percentage in a view or a table, you can't sum or average that percentage to get it at a different aggregate level. You have to recalculate it. A Semantic Layer definition of a metric that’s abstracted from the dimension itself means you can do calculations on the fly more efficiently. **** #### Semantic Layer integration with Tableau The integration between dbt's Semantic Layer and Tableau ensures consistency across all reports. After [downloading the dbt Semantic Layer connector from Tableau Exchange](https://exchange.tableau.com/products/1020), analysts connect to predefined metrics rather than raw data. Aggregations are pushed down from Tableau to dbt whenever possible, ensuring consistent metric definitions across all business reports. However, this delegation introduces some UI limitations. For example, when building a chart showing revenue per customer, the aggregation will show as SUM, even though it might be defined differently in dbt. Changing the aggregation in Tableau (e.g., from sum to average) will not change the metric value, because Tableau delegates the aggregation to the dbt Semantic layer. Table calculations, however, remain fully supported. This approach promotes consistency across teams by ensuring everyone is working from the same trusted definitions. ### Data health tiles: building stakeholder confidence [Data health tiles](https://docs.getdbt.com/docs/explore/data-tile) provide real-time visibility into data quality through embedded tiles in Tableau dashboards. These tiles show two key checks: data freshness (stale vs. fresh) and quality assessment (passed/failed tests). The implementation uses [dbt exposures](https://docs.getdbt.com/docs/build/exposures) defined in YAML files. Developers specify which models a dashboard depends on; dbt automatically runs freshness checks and displays test results. The health tile appears as an embedded URL in the Tableau dashboard, requiring minimal setup for maximum impact. Data consumers frequently question whether the data they're viewing is accurate. The health tile provides immediate confidence that data is both fresh and has passed quality tests. Developers also benefit from this during production pushes, as they can verify that new data has surfaced correctly to prod. The real value, however, comes from giving data stakeholders peace of mind about data reliability. ## Deputy: From horror stories to skyrocketing productivity To reap these benefits for themselves, Deputy rebuilt their data architecture with dbt, focusing particularly on modularization and reducing code and metrics redundancy. The numbers demonstrate the significant improvement Deputy's implementation of dbt Cloud with Tableau achieved across multiple dimensions. Dashboard development time plummeted from months to just one week. Computing costs decreased by 20%. Most impressively, their data team earned a +100 [Net Promoter Score (NPS)](https://www.medallia.com/net-promoter-score/) rating from internal stakeholders, meaning people would actively recommend working with the data team. The architectural benefits went beyond metrics. Deputy unified their data sources into a single source of truth. They improved governance through standardized development practices. Data freshness checks, previously nonexistent, became automatic. The modularization reduced redundancy throughout their architecture. Productivity was now "skyrocketing," Deputy’s data director said, rather than generating horror stories from data stakeholders. Gaining a 360-degree view of its data earned Deputy a 180-degree turnaround in how both its data producers and data consumers viewed working with data. ## Getting started with dbt and Tableau The partnership between dbt and Tableau addresses fundamental trust issues that plague modern analytics. By applying software engineering practices to analytics development, organizations create reliable and scalable data products that stakeholders can actually trust. Deputy's success proves that transformation is achievable with concrete, measurable benefits. Their journey from "horror stories" to "skyrocketing productivity" shows what's possible when data teams have the right tools and practices. For teams inspired by Deputy's transformation, several practical resources exist. The [dbt Learn](https://learn.getdbt.com/catalog) platform provides free courses, including certification paths for both dbt Developer and dbt Architect exams. [dbt Quickstarts](https://docs.getdbt.com/docs/get-started-dbt) offer hands-on guides for connecting to Snowflake, Databricks, and other platforms. To discuss your dbt and Tableau integrations needs in more detail, [ask us today for a demo](https://www.getdbt.com/contact). --- --- title: "How Roche unified global data, enabled AI at scale, and improved operational efficiency" description: "Roche unifies global data with dbt, enabling AI at scale and cutting costs by 70% across 80+ countries." url: "https://www.getdbt.com/blog/roche-unifies-data-enables-ai" date: "2025-07-10" authors: ["Hrishi Kulkarni", "Ernesto Ongaro"] categories: ["Product"] --- # How Roche unified global data, enabled AI at scale, and improved operational efficiency Roche is one of the largest biotechnology and pharmaceutical companies in the world. From pioneering cancer treatments to advancing personalized medicine, Roche’s vision is clear: improve health outcomes and build healthier futures for people worldwide. It’s a complex undertaking that’s powered by vast amounts of clinical and commercial data. But over the years, Roche’s data had sprawled into a disconnected ecosystem. In order to scale, Roche needed to unify its data on a global level. Roche decided to standardize and modernize its commercial analytics platform with dbt Labs. The five-year initiative ultimately covered over 80 countries, streamlined thousands of users, and achieved a cost savings of 70%. Here’s how Roche did it. ## Slow decision-making from a fragmented data ecosystem One of Roche’s primary data challenges is understanding its customers: the healthcare practitioners (HCPs) who prescribe its medications. To do so requires synthesizing data from internal systems (like CRM, marketing platforms, and event logs) with external data (such as scientific publications, clinical trials, and market share). All of this data was siloed across more than 80 countries. Adding to the complexity, every global business unit (known as affiliates) managed its own data pipelines, vendors, and business logic. Maintaining infrastructure, training, and contracts across Informatica, Hadoop, Talend, Oracle, and Microsoft was inefficient, unaligned, and expensive. “Each country had its own version of truth,” reflects João Antunes, Lead Engineer at Roche. “We were all trying to answer the same questions but kept reinventing the wheel with different technologies.” ![Roche map](https://cdn.sanity.io/images/wl0ndo6t/main/20a80eb958a9268bb1ecc858a529bd562991ece9-1580x850.png) For example, while one team might use Oracle databases, another would use Informatica pipelines. Some teams relied on drag-and-drop tools, while others built strictly with Hadoop. The myriad of tooling led to duplicate effort and inconsistent business insights. Because every affiliate bought and ingested data separately, with different definitions and pipelines, even basic questions became nearly impossible to answer. ## Implementing global data standards and a modern stack Solving this problem required a global data transformation strategy. To implement one, Roche focused on three pillars: people, process, and technology. For its people, Roche introduced a matrix structure. Teams were organized by both engineering capabilities (like analytics and machine learning) and product workstreams (either global or affiliate-specific). Capability teams defined standards and best practices, while product workstreams focused on delivering business value. This structure allowed Roche to deliver insights with both speed and quality. ![Roche team topology](https://cdn.sanity.io/images/wl0ndo6t/main/ad124a8ff475a57dbefab3fd6ce5baffb4eb1f4e-1582x854.png) To establish consistent processes, Roche made DevOps a cornerstone. By managing everything as code, Roche could enforce global standards and automate deployments. As a result, Roche enabled robust CI/CD pipelines with predictable releases. Today, every team operates in a two-week sprint cycle. “Now we have complete visibility into what’s happening across the entire organization at any given moment,” says Antunes. “If someone starts building something that already exists in another region, we can catch it early and avoid duplicating work.” Finally, Roche implemented a modern, scalable stack with a fully native cloud platform on AWS. Key components include: - **Ingestion**: Amazon Appflow, AWS Lambda, AWS Glue, AWS Transfer Family - **Staging**: Amazon S3, Lake Formation, AWS Glue Data Catalog - **Transformation**: Amazon Redshift and dbt - **Consumption**: Kubeflow for AI/ML; ThoughtSpot and Tableau for BI “dbt is flexible enough to support both technical and non-technical users,” says Antunes. “We can quickly and easily onboard new affiliates, even if their teams aren’t as technical. dbt is a key part of our ability to scale.” ## Enabling AI use cases The result of Roche’s investment has been profound. Today, Roche’s data platform supports operations across 80+ countries. It serves more than 1,000 users and refreshes over 3,000 datasets daily. By decommissioning 4 platforms, Roche achieved approximately **70% cost savings** while establishing a foundation for innovation. ![Roche outcomes](https://cdn.sanity.io/images/wl0ndo6t/main/8049fb7abb75626a089278e8838576165ceadcc7-1600x787.png) What’s more, Roche can now connect external and internal data at scale. For example, it can combine CRM activity with clinical trial participation, publication history, and social media engagement from their HCPs. As a result, sales teams can more easily identify which physicians are emerging thought leaders and tailor outreach accordingly. Now that its data is standardized and centralized, Roche is exploring new AI use cases. To cite just one example, sales reps now receive AI-powered recommendations, right in the CRM system, for content to share with a given physician based on prior interactions. Another standout use case: using Redshift UDFs powered by Amazon Bedrock to classify product complaints and adverse event reports. These are critical tasks for regulatory compliance, and Roche can now meet its requirements faster and more effectively. SQL is dbt’s native language, so Roche embeds the UDF directly into its incremental models and runs it at scale every day. ![Roche AI powered applications](https://cdn.sanity.io/images/wl0ndo6t/main/6b8de6fa7c695b20180eaee46593e1704983474a-1578x914.png) “AI augments what we already do, and it unlocks value that we hadn’t even imagined,” says Antunes. “But none of that would be possible without the foundation of a modern data stack.” ## Bringing data architecture across the organization The work, of course, is just beginning. Next, Roche plans to expand its data architecture upstream into early-stage research and product development. It’s a major step toward breaking down silos across the pharmaceutical value chain. As Roche scales its data platform even further, it plans to expand dbt’s role: - **Use dbt Labs’ Connections API to consolidate projects.** This will reduce the number of data projects Roche manages across affiliates. - **Move more workloads to dbt. **By moving global core projects into dbt, Roche aims to enable a data mesh for the organization. - **Leverage metadata in dbt.** Every day, Roche runs thousands of models. Metadata will help Roche better understand which of these data assets to prioritize and monitor. In the coming years, AI’s impact on health will be transformative. With dbt Labs as a key partner, Roche has built the data foundation to innovate quickly and power the next era of medical breakthroughs. What’s your data strategy look like this year? Whether it’s building a strong foundation for AI, breaking down silos, or empowering teams to make faster decisions, we’re excited to help. Reach out to [book a demo](https://www.getdbt.com/contact), or [sign up now](https://www.getdbt.com/signup) to connect your data warehouse and start building. --- --- title: "Virgin Media O2 rebuilds its data stack and speeds up the customer experience" description: "Virgin Media O2 modernizes its data stack with dbt, slashing call wait times and boosting NPS by 26%." url: "https://www.getdbt.com/blog/virgin-media-o2-modernizes-data-stack" date: "2025-07-10" authors: ["Hrishi Kulkarni", "Ernesto Ongaro"] categories: ["Product"] --- # Virgin Media O2 rebuilds its data stack and speeds up the customer experience When was the last time you upgraded your phone? Like many people now, you probably don’t rush to buy the latest model each year. That’s because today’s devices and network speeds are already fast enough to meet most everyday needs. Virgin Media O2, a telecommunications company, noticed this shift in customer behavior as well. In order to grow, Virgin Media O2 realized it needed a new playbook: one that prioritizes a standout customer experience (and not just selling new phones). ![The phenomenon of good enough](https://cdn.sanity.io/images/wl0ndo6t/main/a8662578b39be282dc440520c08c17e5b51f2321-1420x880.png) To stay competitive, the telco giant needed to anticipate customer needs, personalize interactions, and deliver value quickly with every interaction. But to do that, Virgin Media O2 had to face another big problem: twenty years of technical debt. ## Cutting through two decades of technical debt Because of the Virgin Media and O2 merger, the company’s data lived in two on-prem systems. Both systems were siloed, entrenched in technical debt, and required inefficient workarounds to support daily operations. Nowhere was this more painful than in the call center. Customer support agents had to switch between 8-10 screens across two different systems, just to answer a single customer question. For customers, this led to long wait times and, predictably, poor Net Promoter Scores (NPS). ![Data fragmentation](https://cdn.sanity.io/images/wl0ndo6t/main/f208fba44a43ce41f3341d7a7d247d5c2d6cf30a-1410x872.png) To deliver a better customer experience, Virgin Media O2 had to address its technical debt and modernize. But rather than try to lift and shift these legacy systems to the cloud, its newly formed digital data team made a big bet: take the [Greenfield Approach](https://easy-software.com/en/glossary/greenfield-vs-brownfield-approach/#:~:text=system%20at%20all.-,the%20greenfield%20approach,-With%20this%20strategy) and start over from scratch. It was a bold move, but it paid off. In just a year and with four engineers, the team implemented entirely new standards for data infrastructure inspired by the [Toyota Production System](https://global.toyota/en/company/vision-and-philosophy/production-system/): - **Stability first:** products are designed for consistency and expected to continually improve. - **Standardization:** every workflow, process, and data pipeline follows the same format. - **Stop at the point of failure:** if something breaks early in the data pipeline, stop processing the data immediately to identify the root cause. Resolve the issue and then continue with clean data. - **Just-in-time data delivery:** data is delivered with consistent latency, at the right moment for a specific purpose. - **Customer-centric innovation:** team members are encouraged to innovate and build thoughtfully. ## Powering continuous data delivery with dbt Implementing these standards was made possible with dbt Labs. By making the dbt platform a critical part of its stack, Virgin Media O2 has transformed how its data is modeled, tested, orchestrated, and delivered: - **Modular, incremental data models: **the team rebuilt their pipelines as modular models that update incrementally. This significantly reduced compute load and increased pipeline efficiency. - **Reduced waste with delta-tracking models:** with dbt’s support for incremental logic, delta-tracking, and self-referencing models, the data team cut processing time from 47 minutes to less than 1 minute. It’s a dramatic improvement that aligns with call-center SLAs. ![Virgin Meda O2 results](https://cdn.sanity.io/images/wl0ndo6t/main/abe68399169ab54e1c8bb114e3d1ec335380318c-1600x945.png) - **Continual data flows:** the team used dbt triggers to create a pull-based system. Once the initial step of a model succeeds, the next step begins. Now the team manages 500 runs per day, compared to a single batch job before. - **Automated quality checks:** to ensure data integrity, the team uses dbt’s built-in testing framework to run over 600 automated data quality checks per day. - **GDPR compliance: **data governance is embedded right into the architecture’s foundation. If the entire stack is wiped, the team can rebuild it in under eight hours with one click—honoring GDPR compliance requirements. Today, the team can deploy to production within 24-48 hours. It’s the pace required to innovate and win customer love, and it’s a major competitive advantage.t ## Happier customers, higher NPS, and cost savings With a modern data stack in place, the data team swiftly built a unified call-center dashboard. Now when a customer call comes in, support agents can pull the customer’s information in less than 60 seconds. They can verify the customer’s identity, review account details, and provide upsells based on ML-driven recommendations, without switching between systems. ![Virgin Media O2 business impact](https://cdn.sanity.io/images/wl0ndo6t/main/2796dca72028b580f020ef5c26c68ad99a40c0d8-1350x780.png) Already, the dashboard has had a powerful impact. Customers are spending less time on the phone, and agents are helping more customers than before. The results are evident in Virgin Media’s NPS, too, which** improved by 26%**. Overall customer satisfaction is **up by 3%** and call center efficiency **increased by 3%**. That may seem small, but it’s a significant change that has generated strong cost savings. ## A customer-first transformation that scales In the race to become customer-first, Virgin Media O2 is setting the pace. By integrating best practices from manufacturing and software engineering, Virgin Media O2 has redefined data standards for telco. With dbt, it’s ready to scale with AI and continuously raise the bar for the customer experience. If you're a data leader at an enterprise and you’re thinking about how to modernize your data stack, we can help. [Talk to our team](https://www.getdbt.com/contact) to build your data strategy, or [sign up](https://www.getdbt.com/signup) to connect your warehouse and start building. --- --- title: "Modern deployment strategies for analytics workflows" description: "Learn how to modernize your analytics deployment strategy using CI/CD, testing, and dbt to deliver trusted data." url: "https://www.getdbt.com/blog/modern-deployment-strategies-for-analytics-workflows" date: "2025-07-10" authors: ["Joey Gault"] categories: ["Pulse"] --- # Modern deployment strategies for analytics workflows ### The foundation: CI/CD for analytics Continuous Integration and Continuous Deployment (CI/CD) represents the backbone of modern software delivery, and these principles translate directly to analytics workflows. Continuous Integration ensures that code changes are automatically built and tested as soon as they're committed to version control. When developers complete their work and merge changes from their feature branch into the main codebase, this triggers a series of automated tests and quality checks. Continuous Deployment extends this process by automatically pushing validated changes through pre-production environments and ultimately to production. This automated pipeline includes running tests against realistic data sets, executing any necessary migration procedures, and monitoring system performance to ensure everything operates within expected parameters. The benefits of CI/CD extend beyond automation. This approach builds multiple safeguards into the deployment process: all changes must pass code review by team members, modifications are tested across multiple environments before reaching production, and systems can automatically roll back problematic deployments. Perhaps most importantly, CI/CD encourages teams to scope changes to smaller, more manageable units, which limits the potential impact of any single deployment. These software engineering practices map naturally to analytics workflows built with tools like dbt. Since dbt captures all transformation logic as SQL and Python code stored in version control, teams can apply the same validation, testing, and automated deployment processes that have proven successful in application development. ## Core principles of modern analytics deployment Successful analytics deployment strategies share several fundamental characteristics that distinguish them from legacy approaches. These principles work together to create a deployment process that prioritizes reliability, maintainability, and speed. ### Multi-environment architecture The most critical principle involves deploying changes through multiple isolated environments before they reach production. At minimum, this means establishing a dedicated staging environment that mirrors production as closely as possible. However, mature organizations often implement additional layers, including development environments for individual contributors and integration environments for testing interactions between different components. Each environment serves a specific purpose in the validation process. Development environments allow individual data engineers to experiment and iterate without affecting others' work. Staging environments provide a final testing ground where changes can be validated against realistic data volumes and usage patterns. Only after changes successfully pass through these preliminary stages do they advance to production. This multi-environment approach requires upfront investment in infrastructure and data management. Teams must establish processes for maintaining realistic test data sets and ensuring that non-production environments remain synchronized with production schemas and configurations. However, this investment pays significant dividends by catching errors early in the development lifecycle, when they're far less expensive to resolve. ### Controlled change scope Modern deployment strategies emphasize releasing smaller, more frequent changes rather than large, infrequent updates. This principle runs counter to traditional data warehouse practices, where teams often accumulated weeks or months of changes before deploying them together. While batching changes might seem more efficient, it dramatically increases the risk and complexity of each deployment. By limiting individual deployments to a few table modifications or model updates, teams create change sets that are easier to test, review, and troubleshoot. When issues do arise, the smaller scope makes it much easier to identify root causes and implement fixes. This approach also enables teams to deliver value to stakeholders more frequently, rather than making them wait for large milestone releases. The key to successful small-batch deployments lies in maintaining discipline around change management. Teams must resist the temptation to bundle "just one more" modification into an existing deployment, even when it seems trivial. Maintaining strict boundaries around change scope requires cultural shifts, but it ultimately leads to more predictable and reliable deployments. ### Automated quality gates Automation plays a crucial role in ensuring consistent deployment quality while reducing the manual effort required from data engineers. Modern deployment pipelines incorporate multiple automated checkpoints that validate different aspects of proposed changes. These quality gates begin with automated testing of transformation logic against known data sets. Tests verify that new models produce expected outputs, that data quality constraints are satisfied, and that performance remains within acceptable bounds. Additional checks might validate documentation completeness, ensure naming conventions are followed, or confirm that all dependencies are properly declared. The automation extends beyond testing to include the deployment process itself. Once changes pass all quality gates, automated systems handle the mechanics of promoting code through environments, running migration scripts, and updating production systems. This eliminates the variability and potential errors introduced by manual deployment procedures. ### Zero-downtime deployments A hallmark of mature deployment processes is their ability to update production systems without disrupting ongoing operations. This requires careful coordination of schema changes, data migrations, and application updates to ensure that downstream consumers continue to function throughout the deployment process. Achieving zero-downtime deployments often involves techniques like blue-green deployments, where new versions of data assets are built alongside existing ones before traffic is switched over. Alternatively, teams might use rolling updates that gradually migrate individual tables or models while maintaining backward compatibility. The specific approach depends on the architecture of the data platform and the nature of the changes being deployed. However, the principle remains consistent: production deployments should be invisible to end users, with no interruption to reports, dashboards, or other data-driven applications. ## Implementation strategies Translating these principles into practice requires careful consideration of tooling, process design, and organizational factors. Successful implementations typically follow a phased approach that gradually introduces more sophisticated deployment capabilities as teams build confidence and expertise. ### Source control integration The foundation of any modern deployment strategy is comprehensive source control that captures not just transformation code, but also configuration files, documentation, and deployment scripts. This creates a single source of truth for all changes and enables teams to track the evolution of their analytics infrastructure over time. Effective source control strategies use branching models that support parallel development while maintaining clear pathways for promoting changes to production. Feature branches allow individual developers to work in isolation, while pull requests provide structured opportunities for code review and discussion before changes are merged. The integration between source control and deployment systems should be seamless. When developers merge approved changes into the main branch, this action should automatically trigger the deployment pipeline without requiring additional manual steps. This tight coupling ensures that the deployment process begins immediately and reduces the opportunity for human error. ### Environment management Creating and maintaining multiple deployment environments presents both technical and operational challenges. Each environment must have access to appropriate data sets, maintain consistent configuration with production, and provide sufficient isolation to prevent interference between different development activities. Modern cloud data platforms simplify many aspects of environment management by providing APIs and infrastructure-as-code capabilities. Teams can define environment configurations declaratively and use automated provisioning to create new environments on demand. This approach ensures consistency across environments while reducing the manual effort required to maintain them. Data management across environments requires particular attention. While production data provides the most realistic testing scenarios, privacy and security concerns often prevent its use in non-production environments. Teams must develop strategies for creating synthetic data sets that preserve the statistical properties and edge cases of production data while protecting sensitive information. ### Testing frameworks Comprehensive testing forms the backbone of reliable deployments. Modern analytics testing goes beyond simple data validation to include performance testing, integration testing, and regression testing that ensures new changes don't break existing functionality. dbt's built-in testing capabilities provide a solid foundation for data quality validation. Teams can define tests that check for null values, ensure referential integrity, and validate business logic. These tests run automatically as part of the deployment pipeline and prevent changes from advancing if any tests fail. More sophisticated testing strategies might include performance benchmarks that ensure new models don't degrade query performance, or integration tests that validate interactions with downstream systems. The key is building a comprehensive test suite that provides confidence in the quality of deployed changes while remaining fast enough to provide rapid feedback to developers. ### Monitoring and rollback procedures Even with comprehensive testing, production issues can still occur. Effective deployment strategies include robust monitoring that quickly detects problems and automated rollback procedures that can restore service while teams investigate root causes. Monitoring should focus on both technical metrics (query performance, error rates, data freshness) and business metrics (data quality, completeness, accuracy). Automated alerting ensures that teams are notified immediately when issues arise, rather than waiting for end users to report problems. Rollback procedures must be tested regularly to ensure they work correctly under pressure. The ability to quickly revert to a previous known-good state provides teams with confidence to deploy changes more frequently, knowing that they have a reliable escape hatch if problems arise. ## Advanced deployment patterns As teams mature in their deployment practices, they often adopt more sophisticated patterns that provide additional safety and flexibility. These advanced approaches require more complex tooling and processes but offer significant benefits for organizations with demanding reliability requirements. Blue-green deployments represent one such advanced pattern, where teams maintain two complete copies of their production environment. New changes are deployed to the inactive environment, thoroughly tested, and then traffic is switched over instantaneously. This approach provides zero-downtime deployments and instant rollback capabilities, though it requires significant infrastructure investment. Canary deployments offer another sophisticated approach, where new changes are gradually rolled out to a subset of users or use cases before being applied broadly. This allows teams to validate changes against real production workloads while limiting the blast radius of potential issues. Feature flags provide additional deployment flexibility by allowing teams to deploy code changes without immediately activating new functionality. This separation between deployment and activation enables more frequent deployments while maintaining precise control over when new features become available to users. ## The path forward Modern deployment strategies represent a fundamental shift in how data teams approach production changes. By adopting principles from software engineering and adapting them to the unique requirements of analytics workflows, organizations can dramatically improve the reliability and speed of their data operations. The transition to modern deployment practices requires investment in tooling, training, and process development. However, the benefits (reduced errors, faster time-to-value, and improved stakeholder confidence) far outweigh the costs. As the demand for reliable, timely data continues to grow, organizations that embrace these modern deployment strategies will find themselves with a significant competitive advantage. The key to success lies in starting with the fundamentals: version control, automated testing, and multi-environment deployment, and gradually building more sophisticated capabilities over time. Teams that take this measured approach will develop the expertise and confidence needed to fully realize the benefits of modern analytics deployment strategies. ## Modern analytic workflow FAQs **How should analytics teams design and manage dev, staging, and production environments with realistic test data to validate changes before release?** Analytics teams should implement a multi-environment architecture with at minimum a dedicated staging environment that mirrors production as closely as possible. Development environments allow individual data engineers to experiment without affecting others' work, while staging environments provide final testing against realistic data volumes and usage patterns. Teams must establish processes for maintaining realistic test data sets and ensuring non-production environments remain synchronized with production schemas and configurations. Since privacy concerns often prevent using production data directly, teams should develop synthetic data sets that preserve statistical properties and edge cases while protecting sensitive information. **What CI/CD workflow should trigger on pull request merges to automate testing, migrations, and deployment of dbt models?** When developers merge approved changes from feature branches into the main branch, this should automatically trigger the deployment pipeline without requiring additional manual steps. The workflow begins with Continuous Integration that automatically builds and tests code changes through automated quality gates, including testing transformation logic against known data sets, validating data quality constraints, and checking performance bounds. Once changes pass all quality gates, Continuous Deployment automatically handles promoting code through environments, running migration scripts, and updating production systems while maintaining zero-downtime deployments. **How can we implement automated rollbacks and production smoke tests to achieve downtime-free analytics deployments?** Implement robust monitoring focused on both technical metrics like query performance and error rates, and business metrics such as data quality and accuracy. Automated alerting ensures immediate notification when issues arise. Rollback procedures must be tested regularly and should enable quick reversion to previous known-good states. Advanced patterns include blue-green deployments where new changes are deployed to inactive environments before instantaneous traffic switching, or canary deployments that gradually roll out changes to subsets of users. Feature flags can also separate code deployment from feature activation, providing additional control over when new functionality becomes available. --- --- title: "Iceberg Ahead: dbt now supports Apache Iceberg tables in BigQuery" description: "Materialize Apache Iceberg tables in BigQuery with dbt—unlock open formats, cross-engine access, and AI-ready data lakes." url: "https://www.getdbt.com/blog/dbt-supports-apache-iceberg-tables-bigquery" date: "2025-07-09" authors: ["Stephen Robb"] categories: ["Product"] --- # Iceberg Ahead: dbt now supports Apache Iceberg tables in BigQuery We’ve got cool news—literally. With the latest release of the dbt-bigquery adapter, you can now materialize Iceberg tables directly in BigQuery. That’s right: your dbt models can now land in an open table format designed for massive scale, multi-engine interoperability, and metadata-rich lakehouse architectures. This unlocks new doors for organizations that embrace open formats, AI governance, and cross-platform compatibility—all while staying within the comfort of dbt. Let’s unpack what this means. ## First, what even _is_ a catalog? These days, everyone in data is talking about catalogs. But depending on who you ask, a “catalog” might be a UI, a compliance tool, a metadata layer, or a magic box that solves data governance. So let’s keep it simple. A data catalog is a centralized metadata layer that helps humans and machines understand what data exists, where it lives, and how to interact with it. In dbt’s world, that means knowing what assets already exist, where to materialize new ones, and how different engines (like BigQuery or Spark) can discover those models. ## What makes Iceberg special? Apache Iceberg is an open table format purpose-built for analytics at scale. Think of it as a smarter way to manage massive datasets—complete with time travel, schema evolution, and lightning-fast metadata operations. An Iceberg catalog serves as the logical registry of your tables. It tells compute engines how to find and interact with the underlying files. In BigQuery, that catalog is now fully supported via dbt. Google provides documentation around their implementation [here](https://cloud.google.com/bigquery/docs/iceberg-tables). ## Why should dbt users care? For dbt developers, supporting Iceberg in BigQuery isn’t just a nice-to-have—it’s a powerful unlock. Here’s why: - **Discover faster: **With better metadata awareness dbt can register models in a catalog so that other tools (and dbt itself) can discover them. - **No extra copies:** Cross-engine compatibility so a table built in BigQuery via dbt can now be queried by Spark, Trino, or other Iceberg-aware engines. - **Greater flexibility: **Dynamic execution means that dbt compiles code based on what already exists. Iceberg catalog support makes that compilation smarter and more flexible in modern lakehouse environments. ![With vs. without Iceberg Catalog](https://cdn.sanity.io/images/wl0ndo6t/main/672832dddb399f43f673a3ea6c31c2aedf894f3f-1238x374.png) ## How it works in dbt Using the new Iceberg support is simple and consistent with other dbt workflows: 1. Create a catalogs.yml file in the top level of your dbt project: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f22cef087d2de37158453e48b159f0c64859f737-984x400.png) 2. You’ll then declare Iceberg-specific configs like this in your model: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/43143d5e64a751134ea322c480d74af13b285c12-1046x462.png) 3. Execute the dbt model with a dbt run -s iceberg_model The full starter guide can be found [here](https://docs.getdbt.com/docs/mesh/iceberg/bigquery-iceberg-support). That’s it. Behind the scenes, the dbt-bigquery adapter will: - Create the table as an Iceberg table in BigQuery - Register it in the configured catalog - Allow you to manage and query it just like any other dbt model 🔐 Note: You’ll need the BigQuery Connection Admin role to make this work, as the adapter must register the table with metadata. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f0d6d489eb06aa5bdffd9efa36fb7ad7f34cea95-1600x1049.png) ## What's next? This is just the beginning. As dbt’s semantic layer, Mesh, and multi-platform story continue to evolve, open formats like Iceberg will play a significant role in building AI-ready data systems. We're actively exploring: - Auto-generating storage_uri from your project + model structure - Configurable default catalogs per environment - More granular access control for metadata management ## dbt + BigQuery + Iceberg = A lakehouse dream team This update represents a significant step forward for modern data teams, offering flexible open formats, native BigQuery support, and a dbt-native approach to manage it all. Whether you're building data products, governed metrics, or AI-ready pipelines, Iceberg support in dbt + BigQuery brings speed, scale, and structure to your lakehouse strategy. Ready to try it? Update to the latest version of dbt, claim your permissions, and let the dbt run begin. --- --- title: "Life after Talend and Informatica: Migrating to the future of data" description: "Legacy ETL systems are holding back your AI initiatives. How to migrate to a modern solution in months, not years." url: "https://www.getdbt.com/blog/talend-informatica-migration" date: "2025-07-07" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Life after Talend and Informatica: Migrating to the future of data The AI revolution is transforming how organizations approach data. However, while over 50% of organizations globally plan to deploy AI in 2025, 90% of enterprise data still sits trapped in on-premise legacy systems. If you're one of the countless data leaders wrestling with Informatica, Talend, or other legacy ETL tools, you're not alone in feeling the squeeze between AI ambitions and infrastructure limitations. Your legacy systems weren't designed for today's AI-driven world. They struggle with the unstructured and multi-modal data that modern AI workflows demand. Meanwhile, your teams are spending most of their time managing data pipelines instead of creating strategic value. The cost? Not just the skyrocketing license fees—which can reach millions annually—but the opportunity cost of being left behind in the AI era. The good news? There's a proven path forward to a more modern solution. Many organizations have already blazed the trail, achieving dramatic improvements in performance, cost reduction, and AI readiness. ## The hidden costs of legacy ETL systems Your current pain points aren't just operational inconveniences—they're strategic barriers to your organization's future. Legacy systems like Talend and Informatica present numerous operational challenges in today’s AI-first environment. ### Lack of support for AI Legacy ETL systems were architected for a structured data world. Today's AI applications increasingly rely on unstructured data like video, images, audio, and text. Your Informatica or Talend infrastructure simply wasn't built to handle this shift. The scalability limitations are equally problematic. On-premises systems can’t handle the volume of data modern enterprises generate. That explains why [Infinite Lambda](https://infinitelambda.com/) found that 90% of on-premises data never gets used for analytics. ### Operational burden Perhaps most frustrating is the operational burden. Your data teams shouldn't be getting calls in the middle of the night because pipelines have failed. Sadly, that’s life at many organizations. Data teams told Infinite Lambda that they spend **80% of their time** managing data, frequent schema changes, troubleshooting, etc. The lack of standardization and automation means teams are constantly in firefighting mode, leaving only 20% of their time to create new, unique value. ### A large operational expense Talend and Informatica require enormous licensing fees that, as data volumes and usage grow, make budgeting unpredictable. Even worse, the expensive infrastructure required to run them is a huge capital investment. Additionally, the lack of automation in legacy ETL systems means most problems must be solved through manual intervention. This inflates operational expenses and distracts teams from strategic, value-adding activities. Talend and Informatica recognize they need to get customers off-premises and onto the cloud. They’ve done this largely by raising prices and forcing them into their cloud-based solutions. These offerings are three times the cost of running on-premises and have fewer features. Even worse, they retain the same legacy processes that make on-premises ETL systems so unwieldy. ## Moving to a modern data framework As a result, many Talend and Informatica users remain stuck. Only 25% of today’s data is moving to the cloud. The remaining 75% remains on-premises due to technical debt, compliance fears, and the lack of a clear migration strategy. If you’re moving to the cloud anyway, it’s a good time to consider a more modern alternative - one built to handle unstructured data that integrates with today’s modern AI ecosystem. That may lead you to think about the processes you need and the technologies that could support them. At dbt Labs, we've spent a lot of time thinking about this. The result is the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle)—an eight-phase framework that brings the best practices of software engineering into data and analytics. The ADLC is a mental model for how mature organizations approach data work. Built on the [Software Development Lifecycle (SDLC)](https://aws.amazon.com/what-is/sdlc/) in software engineering, it encompasses development, testing, deployment, and monitoring in a continuous cycle that ensures your data products are reliable, scalable, and maintainable. More importantly, it provides a roadmap for achieving the kind of AI-ready infrastructure that will serve your organization for years to come. ![ADLC](https://cdn.sanity.io/images/wl0ndo6t/main/034af006313d8a5c5c1149a039fe49a270e1a863-1198x672.png) In the context of enterprise AI, the ADLC becomes even more critical. Modern AI workflows require structured data that's accessible, documented, and governed. You need automatic trust frameworks to ensure data integrity, comprehensive documentation to provide context for AI systems, and robust governance to manage how AI applications consume your data. These are capabilities that legacy ETL systems simply can’t provide. The ADLC framework helps you evaluate both your current state and your desired future state. It's a guide for what tooling you need and how different components should work together harmoniously. Most importantly, it recognizes that data work never stops—you need processes that support continuous improvement and adaptation. **** ## dbt: The modern data transformation solution for AI This is where dbt comes in. dbt encapsulates everything you need to achieve the ADLC vision, sitting on top of your data warehouse to make software engineering best practices accessible across your entire organization. It supports all the key components of a data solution - [data transformation](https://www.getdbt.com/blog/data-transformation), development, observability, cataloging, and semantics. ![Data control plane](https://cdn.sanity.io/images/wl0ndo6t/main/767c319018af66ea4fd55275e61e325fa6b1feaa-2073x1000.png) ### An AI-first data transformation solution We’re used to the idea that [Large Language Models (LLMs)](https://aws.amazon.com/what-is/large-language-model/) consume unstructured data. However, often, the real value lies in an organization’s structured data. Making that data reliable and available for LLMs involves leveraging the components that dbt provides: - Automatic trust frameworks, such as [data testing](https://www.getdbt.com/blog/data-testing) and [data lineage](https://www.getdbt.com/blog/what-is-data-lineage), ensure data quality and traceability. - [Writing documentation](https://docs.getdbt.com/docs/build/documentation) alongside your transformation models gives LLMs the context they need to better understand and use your data. - [Semantic models](https://docs.getdbt.com/docs/build/semantic-models) make data available to both consumers and AI applications in a consistent manner, using the language of the business in lieu of technical jargon. ### Meeting users where they are dbt accomplishes all of this in three ways. The first is through improving data quality and trust. The second is by driving efficiency, which is critical for processing all that data that’s still sitting on-premises. And third, and perhaps most important, is meeting users where they are. Your organization likely has people with varying levels of technical expertise, from seasoned SQL developers to business analysts who understand the data but may not be comfortable with complex coding environments. dbt addresses this through multiple interfaces: - A [VS Code extension](https://docs.getdbt.com/docs/about-dbt-extension) for your most technical team members - A browser-based [Studio IDE](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/develop-in-the-cloud) for those comfortable with SQL but new to local development environments - [dbt Canvas](https://docs.getdbt.com/docs/cloud/canvas), a new low-code development environment for analysts who have business context but may need support with SQL complexity. The result is that everyone can contribute safely and productively to the same projects, backed by the governance and safety features that enterprise organizations require. Our recently launched [dbt Fusion](https://www.getdbt.com/product/fusion) engine takes this further, delivering significant performance gains and intelligent cost optimization. Currently in public beta for [Snowflake](https://snowflake.com/) and [Databricks](https://www.databricks.com/), Fusion provides real-time feedback while writing code, providing a developer experience that’s orders of magnitude faster than anything before it. For organizations concerned about cloud costs, Fusion's intelligent optimization features help ensure you're not overspending on compute resources. ## Migration strategy: beyond lift-and-shift We’ve talked about the why. Now it’s time to talk about the how - i.e., migration. The statistics are sobering: only one in seven data platform migrations succeeds. The reasons usually fall into two categories: organizations either attempt **massive manual rewrites** or rely on fully **automated lift-and-shift** approaches. **Manual rewrites** are appealing in theory but problematic in practice. You'll spend years recreating functionality, learning from past mistakes the hard way, and trying to maintain business continuity while building everything from scratch. Even with AI assistance, the complexity and time requirements (typically years) make this approach risky and expensive. **Full automation** seems like the obvious alternative, but most automated migration tools are essentially **lift-and-shift** solutions. They convert your legacy code to run on modern platforms without any insight into how to optimize it for cloud architectures. You end up with the same patterns designed for on-premises systems, often resulting in higher compute costs and missed opportunities for improvement. As with most things in life, the truth often lies in the middle. The smart approach combines the best of both worlds. This is what [Flowline](https://infinitelambda.com/flowline-legacy-data-migration/) from Infinite Lambda does. Flowline combines an initial lift-and-shift with a set of baked-in best practices, including post-migration refactoring and data validation. dbt integrates seamlessly with Flowline to provide best-in-class migration for porting from legacy systems onto a modern data architecture. Leveraging the best attributes of both systems, Flowline achieves over 95% automated code conversion from platforms like Informatica and Talend to native dbt code. The technical process is thorough: - First, deterministic code migration converts your existing ETL jobs to dbt models with established best practices built in. - In the testing phase, every data output is compared between your legacy system and the new dbt implementation. Any differences are identified and explained, whether they're due to platform differences, data type changes, or even bug fixes in the legacy system. - Finally, the refactoring phase optimizes your new dbt models for performance and cost efficiency. This may involve consolidating multiple legacy transformations into a single model, creating reusable macros, or restructuring data flows to leverage modern warehouse capabilities. The results speak for themselves. Flowline’s approach delivers approximately 10 times the improvement over manual migration timelines, ensuring you end up with an AI-ready data stack rather than just a translated version of your legacy system. It does this at a fixed cost with a predictable timeline, eliminating the number one reason why data migrations tend to fail. ## Real-world success stories Consider [Macif](https://www.macif.fr/), a leading French insurance provider with over six million members. They were facing the classic legacy ETL challenges: pipelines taking hours to complete, massive Informatica license costs, and a planned two-year modernization effort that seemed overwhelming. Using Flowline’s software-accelerated approach, what Macif thought would take two years was completed in just three months. More importantly, the results were dramatic: pipelines that previously took two hours now complete in five minutes, enabling their operations team to finish work during business hours rather than waiting for overnight batch processes. They reduced their ETL licensing costs from €1 million annually to €200,000—an 80% reduction that freed up budget for strategic initiatives. [AstraZeneca's](https://www.astrazeneca.com/) journey illustrates the broader strategic value. As a global pharmaceutical company, they were dealing with increasingly complex legacy transformations across Informatica, Talend, and even SSIS. These systems were not just expensive to maintain. They were actively slowing down their ability to develop new reports and dashboards, taking weeks for what should have been simple data products. More critically for a healthcare organization, their legacy infrastructure left no capacity for AI innovation. After migrating to a modern data stack with dbt, they now have an AI-ready platform that supports generative AI use cases while delivering significantly faster development cycles and reduced costs. Perhaps most importantly, their teams now have the bandwidth to focus on advanced AI applications in healthcare rather than maintaining legacy infrastructure. ## Getting started: Your migration roadmap If any of this resonates with your current situation, your next step is assessment. Understanding your current state—the complexity of your existing ETL jobs, the volume of data you're processing, your performance requirements, and your strategic objectives—is crucial for planning a successful migration. We recommend starting with Infinite Lambda’s [comprehensive migration readiness assessment](https://migrationassessment.infinitelambda.com/). This evaluation, which takes about five minutes to complete, will help you understand the key areas that need attention in any migration project. You'll receive a detailed migration guide that goes beyond high-level strategy to provide practical, actionable guidance for your specific situation. The assessment covers critical factors like your current ETL complexity, data volumes, performance requirements, team skills, and strategic timeline. Based on your responses, you'll get targeted recommendations for addressing potential challenges before they become roadblocks. If the assessment reveals that your migration needs are complex—and most enterprise migrations are—you don't have to tackle this alone. Professional migration services can provide the end-to-end support you need, from initial assessment through production deployment and team training. These services typically offer fixed-cost engagements that eliminate the budget uncertainty that has derailed so many migration projects. ## The time to act is now The AI era isn't coming—it's here. While you're wrestling with legacy ETL maintenance and escalating license costs, your competitors may already be leveraging modern data stacks to power AI initiatives that will define the next decade of competitive advantage. Your legacy systems served you well in their time, but they're now actively limiting your potential. The migration path exists, the tools are proven, and organizations across industries have demonstrated that dramatic improvements in performance, cost, and capability are achievable. Used together, Flowline and dbt rapidly accelerate your transition from legacy ETL systems such as Talend and Informatica. With Flowline, you’re ensured a consistent, high-quality migration in less time than a manual or automated lift-and-shift effort. The end result is a native dbt implementation - a [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) that’s flexible and cross-platform, collaborative, and produces trustworthy outputs. If you’re ready to see how to make this shift, without the pitfalls of manual rewrites or lift-and-shift dead ends, watch our [on-demand webinar](https://infinitelambda.com/business-events/life-after-talend-informatica-v2/?utm_source=partner&utm_medium=dbt) with Infinite Lambda. You'll learn how teams like yours are moving beyond Talend and Informatica to build an AI-ready data foundation with Flowline and dbt. --- --- title: "Guide to AI data products" description: "This guide covers what AI data products are, why they matter, and how to design them for business impact." url: "https://www.getdbt.com/blog/guide-to-ai-data-products" date: "2025-07-07" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Guide to AI data products AI is no longer a bolt-on feature—it’s becoming the foundation of modern data products. This guide breaks down what AI data products are, why they matter, and how they evolve from simple tools into production-grade assets. Whether you’re a data leader, product owner, or executive, you’ll learn how to harness AI to build smarter, more scalable systems that turn data into a true competitive advantage. ## What are AI data products? AI data products combine data management capabilities—leveraging tools like [dbt](https://www.getdbt.com/product/dbt)—with artificial intelligence to deliver specific business outcomes. Unlike traditional tools, they incorporate machine learning, natural language processing, and other AI techniques to provide sophisticated insights, automate workflows, and enable predictions. These products represent the evolution from static reporting to dynamic, intelligent systems that learn and adapt over time. They transform raw data into actionable intelligence that drives better decision-making across the organization. AI data products solve real business problems by finding patterns humans might miss and scaling analysis beyond what traditional methods can achieve. They turn data from a passive resource into an active business driver. The true power of these products lies in their ability to continuously improve. As they process more data, their outputs become more accurate and valuable, creating a virtuous cycle of increasing returns. ## The business value of AI data products ### Delivering trusted data AI data products increase the reliability and traceability of business data. They create a foundation of trust that allows everyone to confidently base decisions on data outputs. For example, a financial services company might develop an AI-powered risk assessment product that analyzes multiple factors to provide precise, consistent risk evaluations. This trust has tangible business impacts. Companies spend less time reconciling conflicting data sources. They save resources previously devoted to manual validation. Their brand equity grows through consistently accurate insights. [Trust also accelerates decision-making](https://www.getdbt.com/product/build-trust-in-data-and-data-teams). When executives know they can rely on the data, they act more quickly and decisively. This speed creates competitive advantages in fast-moving markets. Data trust also extends beyond the organization. Customers, partners, and regulators gain confidence in companies that demonstrate data mastery through AI products. ### Accelerating development cycles AI solutions drastically compress development timelines. They automate routine coding tasks, generate documentation automatically, and enable reusable components that eliminate redundant work. A retail analytics team might use AI assistance to build a customer segmentation model in days rather than weeks. This speed means marketing initiatives launch much faster than with conventional development approaches. Faster development also means organizations can test more ideas. Teams try different approaches, learn from the results, and refine their products quickly. This iteration leads to better outcomes and more innovation. The acceleration extends beyond initial development to maintenance and updates. AI helps teams adapt existing products to new requirements with minimal effort, ensuring data products remain relevant as business needs change. ### Increasing return on investment Modern AI data products offer compelling ROI advantages. They lower maintenance costs through automated optimization. They reduce the need for specialized technical expertise for basic tasks. They deliver faster time-to-value for data initiatives. They scale without proportional cost increases. A manufacturing company using AI-powered predictive maintenance might improve production efficiency by 15-20% while reducing data engineering resources. The ROI compounds as the system prevents costly failures and extends equipment life. AI data products also create new revenue opportunities. Companies monetize insights through new offerings or use AI-enhanced data to differentiate existing products. These revenue streams often have higher margins than traditional business lines. The most significant ROI often comes from opportunity costs avoided. Better forecasting prevents inventory issues. Smarter risk models reduce losses. Improved customer insights increase retention. These benefits, while harder to measure directly, often exceed the explicit cost savings. ## Key components of AI data products ### The semantic layer [The semantic layer](https://www.getdbt.com/product/semantic-layer) serves as the critical interface between raw data and business users. It translates technical data structures into business concepts that stakeholders understand. In AI data products, the semantic layer defines standardized metrics and dimensions. It ensures consistent interpretation of data across applications. It enables natural language interactions with complex datasets. It provides governance and access controls. A healthcare organization might implement a semantic layer defining standard metrics like "readmission rate" that can be consistently referenced across different AI applications. This standardization prevents the confusion of multiple definitions for the same business concept. The semantic layer also protects AI systems from changes in underlying data structures. When source systems change, only the semantic layer needs updating, not every downstream application. This abstraction creates resilience in the AI data ecosystem. ### AI-assisted development Modern AI data products use AI throughout the development lifecycle. AI converts natural language requirements into code. It automatically creates tests for data pipelines. It generates clear documentation from existing code. It suggests performance improvements for queries and data models. A marketing team might use AI to rapidly transform a request like "Show me customer acquisition cost by channel over the last three quarters" into a complete, tested data model. This assistance makes data professionals more productive and lets business users create simple data products without coding. AI assistance also improves code quality. It applies best practices consistently, reduces errors, and suggests optimizations humans might miss. The resulting data products perform better and require less maintenance. The real power comes when AI helps data teams focus on high-value work. By handling routine tasks, AI frees skilled professionals to solve complex problems and create innovative solutions that drive business value. **** ### Data pipeline automation Reliable data pipelines form the backbone of effective AI data products. Modern tools automate much of this pipeline work with features like automatic notifications for pipeline breaks, visible data lineage, and scheduled refresh processes. Automation reduces errors and improves consistency. Manual processes inevitably create mistakes, while automated pipelines run the same way every time. This reliability is essential for AI systems that depend on clean, consistent data. Automated pipelines also adapt to changing conditions. They handle volume spikes, detect anomalies, and recover from failures without human intervention. This resilience keeps data flowing even when problems occur. Pipeline automation creates transparency. Teams see exactly where data comes from, how it's transformed, and where it's used. This visibility builds trust in the final outputs and makes troubleshooting easier when issues arise. ## Implementation lifecycle for AI data products ### Phase 0: Exploratory analysis The product development journey begins with exploration to understand the business problem and possible data-driven solutions. Analysts and data scientists experiment with different approaches. Speed and flexibility take precedence over governance. Various hypotheses are tested against available data. For example, a telecommunications company might explore customer churn patterns by examining dozens of variables across different time periods before determining which factors truly predict churn. This exploration phase is messy but essential for discovery. The best exploration happens when business and technical teams work together. Business experts bring domain knowledge that helps focus the analysis on meaningful questions. Technical teams bring data skills that extract insights from complex information sources. Successful exploration balances open-ended discovery with practical constraints. While teams should have freedom to investigate widely, they also need to maintain focus on solving real business problems rather than pursuing academic interests. ### Phase 1: Personal reporting As exploratory insights prove useful, they evolve into personal reporting tools. These typically have limited distribution, minimal governance requirements, and practical utility for specific business questions. A sales analyst might create a personal dashboard tracking key accounts' purchasing patterns. Though simple, these personal tools validate the business value of the underlying data. They prove concepts before investing in more robust development. Personal reporting creates advocates for data-driven approaches. When individuals experience the benefits firsthand, they champion expansion to their teams and departments. This organic growth builds momentum for broader AI data initiatives. These early tools often reveal data quality issues or gaps in available information. Addressing these problems early makes later development phases more successful and prevents building sophisticated products on faulty foundations. ### Phase 2: Shared reporting When reports prove valuable enough to share with others, they need enhanced capabilities. They require access controls to ensure appropriate data visibility. They need more robust testing to prevent errors. They need change tracking and version control. They need consistent definitions aligned with the semantic layer. That sales dashboard might evolve to become a shared tool for the entire sales organization, requiring standardized definitions of metrics like "qualified opportunity" and "sales cycle length". This standardization ensures everyone makes decisions based on the same information. Shared reports need clearer documentation and intuitive interfaces. While a creator understands their own work implicitly, others need context and guidance to use the tool effectively. This user-focused design becomes increasingly important as the audience grows. At this stage, feedback loops become critical. Regular input from users helps refine the product and ensures it continues to meet business needs as they evolve. This ongoing improvement transforms a static report into a dynamic business tool. ### Phase 3: Production artifact The most mature AI data products operate as critical business infrastructure. They have formal SLAs for performance and availability. They include comprehensive documentation. They follow regular maintenance cycles. They integrate with other enterprise systems. A fully developed version of the sales analysis tool might become an enterprise-wide solution that not only reports on sales patterns but also predicts outcomes, recommends actions, and integrates with CRM systems. At this stage, the product becomes essential to how the business operates. Production artifacts require rigorous governance and security. Teams implement controls that protect sensitive data while making insights available to authorized users. This balance between protection and access is essential for enterprise-grade products. The transition to production status often involves organizational changes. Teams establish clear ownership for maintenance and enhancements. They create support processes for users. They implement monitoring to ensure reliability. These operational elements are as important as the technical features. ## The impact of AI on data product development ### AI as the new EDA interface AI interfaces are becoming more effective for exploratory data analysis than traditional tools. They allow natural language questions about data. They generate analytical code faster than humans can write it. They combine semantic understanding with technical execution. They enable rapid iteration without constantly switching contexts. A financial analyst might simply ask, "What's driving the variance in Q3 profitability across our European markets?" and receive interactive visualizations and analyses without writing code. This natural interaction removes technical barriers to data exploration. AI interfaces democratize analysis. Business users ask sophisticated questions without learning SQL or programming. This broader access to insights creates more data-driven decision-making throughout the organization. The conversational nature of AI interfaces also changes how teams work with data. Analysis becomes more iterative and intuitive. Users follow their curiosity through a series of questions rather than defining all requirements upfront. This flexibility leads to unexpected discoveries. ### Enhancing the "conveyor belt" model AI systems are disrupting the traditional BI "conveyor belt" that moves data artifacts from exploration to production. They provide superior exploratory capabilities outside traditional BI tools. They create new integration patterns between AI interfaces and governance systems. They force BI vendors to rethink their value proposition. Organizations increasingly adopt hybrid approaches where exploration happens in AI-powered interfaces while governance and presentation remain in traditional platforms. This combination leverages the strengths of each system type. The new model creates challenges for data teams. They must ensure consistent results across different tools and maintain governance without blocking innovation. Solving these challenges requires both technical solutions and process changes. Despite these complications, the benefits of AI-enhanced development are compelling. Teams create more valuable data products faster. Business users get more direct access to insights. Organizations become more responsive to changing conditions and opportunities. ### The role of context protocols Context protocols like MCP (Model Context Protocol) are fundamentally changing how AI data products operate. They allow AI systems to access enterprise metadata and data models. They enable accurate answers without compromising governance. They create interoperability between different systems. They ensure AI responses reflect current, accurate business context. A marketing executive using an AI assistant can get accurate, governed answers about campaign performance because the underlying AI has access to the organization's semantic layer through context protocols. This contextual awareness makes AI tools genuinely useful for business decisions. Context protocols solve a critical problem for enterprise AI: ensuring responses reflect official company data rather than generic or outdated information. They connect conversational interfaces to trusted data sources while maintaining security and governance. These protocols also extend the value of existing data investments. Organizations connect their carefully built semantic layers and data warehouses to new AI interfaces without starting over. This integration preserves past work while adding new capabilities. ## Best practices for AI data products ### Prioritize data quality AI systems amplify both the benefits of good data and the problems of poor data. Implement comprehensive data quality monitoring. Establish clear ownership for critical data elements. Create feedback loops to continuously improve quality. Document data lineage thoroughly. Poor data quality creates a negative feedback loop with AI systems. Bad data leads to incorrect outputs, which cause users to lose trust and stop using the system. This neglect further degrades quality. Breaking this cycle requires relentless focus on data excellence. Quality matters most for the data that drives key business decisions. Rather than trying to perfect all data, focus on the critical elements that power your most important AI products. This targeted approach delivers better returns on quality investments. Data quality is not just a technical issue. It requires clear ownership and accountability. Assign data stewards who understand both the business context and technical aspects of key data domains. These stewards become the guardians of quality throughout the data lifecycle. ### Design for reusability Maximize efficiency by creating modular, reusable components. Build a comprehensive semantic layer before extensive AI product development. Create standardized data models that serve multiple use cases. Design transformation logic that can be repurposed. Document components thoroughly to encourage reuse. Reusability accelerates development exponentially. Each reusable component saves time not just once, but every time it's used. This compound effect dramatically increases team productivity and ensures consistency across products. The semantic layer becomes the foundation for reuse. By defining business concepts once and using them many times, organizations create coherent data products that speak the same language. This consistency improves user understanding and trust. Documentation is essential for reuse. Even the best components won't be reused if people don't know they exist or how to use them. Create clear documentation that explains the purpose, usage, and limitations of each reusable element. ### Implement progressive governance Apply governance appropriate to each product's maturity stage. Use minimal governance for exploratory work to encourage innovation. Increase governance progressively as products move toward production. Integrate automated testing into development workflows. Create clear access control policies tied to data sensitivity. Progressive governance balances innovation and control. Too much governance early stifles creativity and slows discovery. Too little governance late creates risk and quality problems. The right approach applies controls appropriate to each stage. Automation makes governance sustainable. Manual approval processes create bottlenecks and frustration. Automated testing, validation, and documentation keep products compliant without slowing development. The best governance is invisible to users. It works behind the scenes to ensure quality and compliance without creating friction. Teams that achieve this balance deliver both innovation and reliability. ### Embrace continuous learning Both AI systems and the teams that build them should continuously improve. Monitor AI product performance against business objectives. Collect user feedback systematically. Retrain models with fresh data. Stay current with evolving AI capabilities and best practices. Learning happens by measuring outcomes. Define clear metrics that show whether AI products are delivering business value. Use these metrics to guide improvement efforts and justify further investment. User feedback provides essential insights for improvement. Create simple ways for users to report issues and suggest enhancements. Act on this feedback quickly to show users their input matters. The AI field evolves rapidly. Dedicate time for teams to learn about new capabilities and techniques. This ongoing education helps organizations stay competitive and make the most of emerging opportunities. ## The future of AI data products As AI capabilities advance, several trends will shape the future of AI data products. AI will become embedded throughout the data lifecycle rather than applied as a separate layer. More business users will create sophisticated data products with minimal technical expertise. AI systems will increasingly optimize themselves, adjusting to changing data patterns and business requirements automatically. This autonomous operation will reduce maintenance needs and help products adapt to changing conditions. Natural language interfaces will become the primary way most users interact with data. These conversational experiences will make insights accessible to everyone, regardless of technical skills. Complex analysis will be as simple as asking a question. The boundaries between different types of data products will blur. Reporting, analysis, prediction, and automation will merge into unified experiences that adapt to user needs. This convergence will create more powerful and flexible tools. ## Conclusion AI data products represent a fundamental shift in how organizations derive value from their data assets. By combining robust data management with cutting-edge AI capabilities, these products deliver faster insights, more accurate predictions, and more accessible analytics than traditional approaches. The organizations that succeed will be those that embrace AI as a core component of their data strategy rather than treating it as a separate initiative. They will build on a foundation of high-quality, well-governed data and leverage AI throughout the development lifecycle. As you embark on your own AI data product journey, remember that technology is only one piece of the puzzle. Equally important are the people, processes, and governance frameworks that ensure these powerful tools deliver reliable, ethical, and valuable business outcomes. The journey might seem challenging, but the rewards are substantial. Organizations that master AI data products will make better decisions, operate more efficiently, and create new value for customers. In today's data-driven economy, these capabilities aren't just advantages—they're necessities. --- --- title: "Tracking data transformations: Best practices & tools" description: "Gain visibility into your transformation layer to build trust and scalability in your analytics pipeline." url: "https://www.getdbt.com/blog/tracking-data-transformations" date: "2025-07-03" authors: ["Joey Gault"] categories: ["Pulse"] --- # Tracking data transformations: Best practices & tools Data transformation tracking encompasses several interconnected elements that form the backbone of reliable analytics operations. At its core, you need visibility into the transformation logic itself—the SQL and Python code that converts raw data into analysis-ready datasets. This includes understanding what transformations are being applied, when they run, and how they relate to one another in your data pipeline. Beyond the code, tracking involves monitoring the quality and consistency of your transformed data. This means establishing tests that verify data integrity, implementing validation rules that catch anomalies, and maintaining documentation that explains the business logic embedded in your transformations. The goal is creating a comprehensive view of your data transformation landscape that enables both technical teams and business stakeholders to understand and trust the data they're working with. [Version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) emerges as a fundamental requirement in this context. Just as software development relies on tracking code changes over time, data transformations require the same level of rigor. Every modification to transformation logic should be documented, reviewed, and deployed through controlled processes that maintain data quality and prevent unexpected downstream impacts. ## Building systematic tracking approaches Effective transformation tracking begins with establishing clear conventions and standards across your data team. This involves defining naming conventions for datasets, standardizing SQL practices, and implementing consistent testing protocols. Without these foundational elements, tracking becomes exponentially more difficult as your data operations scale. The challenge of consistency across multiple datasets cannot be overstated. Teams often struggle with ensuring that similar data follows standardized formats, that timezone handling remains uniform, and that primary keys maintain consistent naming patterns. These seemingly minor inconsistencies compound over time, creating confusion and reducing the reliability of downstream analytics. [Data modeling conventions](https://www.getdbt.com/blog/modular-data-modeling-techniques) play a crucial role in systematic tracking. Before transformation work begins, teams need established style guides that cover everything from column naming to code commenting standards. This standardization ensures that transformations remain readable and maintainable across different team members, reducing the barrier for contribution and making tracking more manageable. The standardization of core KPIs represents another critical tracking challenge. Key business metrics should be version-controlled, defined in code, and accessible within business intelligence tools. When different teams generate conflicting reports due to inconsistent metric definitions, the entire data organization suffers. Proper tracking ensures that there's one authoritative source for each critical business metric. ## Implementing modern tracking solutions The evolution from traditional ETL to [modern ELT](https://www.getdbt.com/blog/extract-load-transform) approaches has fundamentally changed how organizations should think about transformation tracking. In the ELT paradigm, where data is loaded into warehouses before transformation, tracking becomes both more complex and more important. The flexibility of transforming data within the warehouse creates opportunities for better tracking, but also requires more sophisticated approaches to manage the increased complexity. Modern data transformation tools address these challenges by providing integrated tracking capabilities. dbt exemplifies this approach by treating transformations as code, enabling version control, automated testing, and comprehensive documentation generation. This transforms tracking from a manual, error-prone process into an automated system that scales with your data operations. The modular transformation logic that tools like [dbt](https://www.getdbt.com/product/what-is-dbt) enable creates natural tracking boundaries. Each transformation becomes a discrete, testable unit with clear inputs and outputs. This modularity makes it easier to understand data lineage, identify the impact of changes, and maintain comprehensive documentation of your transformation landscape. [Automated documentation generation](https://docs.getdbt.com/docs/build/documentation) represents a significant advancement in transformation tracking. Rather than relying on manually maintained documentation that quickly becomes outdated, modern tools can automatically generate and update documentation based on the transformation code itself. This ensures that tracking information remains current and accessible to all stakeholders. ## Establishing governance and quality controls Tracking data transformations effectively requires robust governance frameworks that ensure consistency and quality across your entire data organization. This involves [implementing continuous integration and continuous deployment (CI/CD)](https://docs.getdbt.com/guides/custom-cicd-pipelines?step=1) pipelines specifically designed for data work. These pipelines automate the testing and deployment of transformation changes, providing systematic tracking of what changes are made, when they're deployed, and how they impact downstream systems. Data testing emerges as a critical component of transformation tracking. Automated tests that verify data quality, check for anomalies, and validate business rules provide ongoing visibility into the health of your transformations. These tests serve dual purposes: they catch issues before they impact business users, and they create a historical record of data quality over time. The integration of testing with version control creates powerful tracking capabilities. Every change to transformation logic can be automatically tested against historical data patterns, ensuring that modifications don't introduce unexpected behaviors. This systematic approach to change management provides the confidence needed to evolve transformations while maintaining data reliability. Monitoring and alerting systems complement these governance frameworks by providing real-time visibility into transformation performance. When transformations fail, take longer than expected, or produce unexpected results, automated alerting ensures that issues are identified and addressed quickly. This proactive approach to tracking prevents small problems from becoming major data quality incidents. ## Scaling tracking across organizations As data organizations grow, the complexity of tracking transformations increases exponentially. What works for a small team manually managing a handful of transformations breaks down when dealing with hundreds of models, multiple data sources, and diverse stakeholder requirements. Scalable tracking requires architectural approaches that can handle this complexity without overwhelming data teams. The concept of a data transformation layer becomes crucial at scale. This layer provides a centralized approach to managing transformations, ensuring consistency across different projects and teams. Rather than having scattered transformation logic across various systems, a unified transformation layer creates a single source of truth for how data is processed and prepared for analysis. Collaboration features become increasingly important as teams scale. Multiple data engineers, analysts, and business stakeholders need to work together on transformation development, and tracking systems must support this collaborative workflow. This includes providing shared development environments, facilitating code reviews, and ensuring that changes are properly communicated across teams. The infrastructure management aspects of transformation tracking also become more complex at scale. Teams need systems that can handle increasing data volumes, support multiple environments for development and testing, and provide the performance required for timely data delivery. Cloud-based solutions often provide the scalability needed for enterprise-scale transformation tracking. ## Measuring success and continuous improvement Effective transformation tracking isn't a one-time implementation; it requires ongoing measurement and optimization. Organizations need metrics that help them understand the health of their transformation tracking systems and identify areas for improvement. This includes monitoring data quality trends, tracking transformation performance, and measuring user satisfaction with data products. The cost and resource efficiency of tracking systems becomes a key consideration as organizations mature. While comprehensive tracking provides significant benefits, it also requires investment in tools, processes, and personnel. Successful organizations find ways to optimize these investments, focusing tracking efforts on the most critical transformations and automating routine monitoring tasks. Continuous optimization ensures that tracking systems evolve with changing business needs. As new data sources are added, business requirements change, and teams grow, tracking approaches must adapt accordingly. This requires regular review of tracking practices and willingness to invest in improvements that enhance data reliability and team productivity. ## The strategic impact of comprehensive tracking Organizations that implement comprehensive transformation tracking see significant strategic benefits beyond just operational improvements. The visibility and control that effective tracking provides enables faster innovation, reduces risk, and improves the overall return on investment from data initiatives. Companies like [Condé Nast](https://www.getdbt.com/case-studies/conde-nast), [Nasdaq](https://www.getdbt.com/case-studies/nasdaq), and [Rocket Money](https://www.getdbt.com/case-studies/rocket-money) have demonstrated how systematic transformation tracking can streamline operations, reduce engineering bottlenecks, and improve collaboration across teams. These organizations leveraged tools like [dbt to centralize and automate their transformation workflows](https://www.getdbt.com/product/dbt), resulting in faster delivery of business-critical insights and more reliable data-driven decision making. The competitive advantage that comes from reliable, well-tracked data transformations cannot be overstated. In an environment where data drives business decisions, organizations with superior transformation tracking capabilities can move faster, make better decisions, and respond more effectively to market changes. Ultimately, the question of how to track data transformations comes down to implementing systematic approaches that scale with your organization's needs. This requires investment in modern tools, establishment of clear processes, and commitment to continuous improvement. The organizations that get this right will find themselves with a significant competitive advantage in an increasingly data-driven business environment. ## Data transformation tracking FAQs **What is data transformation tracking?** Data transformation tracking encompasses monitoring and documenting the transformation logic, data quality, and pipeline relationships in your analytics operations. It involves maintaining visibility into SQL and Python code that converts raw data into analysis-ready datasets, understanding when transformations run, monitoring data quality and consistency through validation rules and tests, and maintaining comprehensive documentation that explains the business logic. The goal is creating a complete view of your data transformation landscape that enables both technical teams and business stakeholders to understand and trust the data they're working with. **Why do businesses need systematic transformation tracking?** Businesses need systematic transformation tracking to ensure data reliability, maintain consistency across teams, and scale their data operations effectively. Without proper tracking, organizations face challenges like inconsistent data formats, conflicting business metrics definitions, and difficulty troubleshooting data quality issues. Tracking enables faster innovation, reduces operational risk, and improves return on investment from data initiatives. Companies with comprehensive tracking can move faster, make better decisions, and respond more effectively to market changes, providing a significant competitive advantage in data-driven business environments. **Does the tool have version control features?** Modern data transformation tools like dbt treat transformations as code, enabling comprehensive version control capabilities. Every modification to transformation logic is documented, reviewed, and deployed through controlled processes that maintain data quality and prevent unexpected downstream impacts. Version control integration with automated testing creates powerful tracking capabilities where every change can be tested against historical data patterns. This systematic approach to change management provides the confidence needed to evolve transformations while maintaining data reliability, similar to software development practices. --- --- title: "dbt Labs on dbt: Building the habit of cost-aware data development" description: "How the dbt Labs team uses dbt to build a habit of cost-aware development and cut data platform spend." url: "https://www.getdbt.com/blog/building-the-habit-of-cost-aware-data-development" date: "2025-07-03" authors: ["Brandon Thomson"] categories: ["Insights"] --- # dbt Labs on dbt: Building the habit of cost-aware data development **** One of the perks of my job as an analytics engineering manager at dbt Labs? I get to dogfood the latest and greatest from the dbt platform. That means playing with new features before they hit general availability, and seeing firsthand how small changes can add up to big wins. Lately, that’s included a trio I’m pretty excited about: - dbt’s new [**cost management dashboard**](https://docs.getdbt.com/docs/cloud/cost-management) - The [**dbt Fusion engine**](https://docs.getdbt.com/blog/dbt-fusion-engine) - And Fusion powered [**state-aware orchestration**](https://www.getdbt.com/product/dbt-state) Together, they’ve helped us build the habit of cost-aware data development: a mindset where optimization isn’t an afterthought, it’s just how we work. This post explores how we’re using the dbt platform to free up **25% of our annual spend**, speed up workflows, and make cost optimization part of our everyday work. ## It starts with visibility ![Cost management dashboard](https://cdn.sanity.io/images/wl0ndo6t/main/332dbd27b00c2e49357f19eddeb0f0cdf0e36c25-2880x1800.png) With [cost monitoring](https://docs.getdbt.com/docs/cloud/cost-management) directly in the dbt platform, we finally have a clear view into which projects, environments, and models are using the most compute, and which ones to focus on. Before this dashboard, we were flying in a fog, only optimizing the obvious big models. We used to rely on query tagging and build our own data pipelines just to track costs; essentially incurring overhead and costs just to manage our costs. Now that entire burden is shifted to the dbt platform, eliminating the need for these extra models and pipelines.. We uncover hidden inefficiencies and prioritize models that truly matter. Instead of trying to optimize everything, we can see exactly which transformations drive the most value: core metrics, key dashboards, AI inputs, and which ones were silently consuming resources without delivering proportional value. This visibility into our cost drivers was a game-changer. Visibility shifts the game. We stop guessing where the cost lives. Instead, we spot issues we wouldn’t otherwise notice: long-tail jobs, legacy models, or transformations we haven’t touched in months. Without that signal, they quietly run every day and quietly burn compute. Now, we don’t need to wait for a fire drill to investigate. We review cost impact regularly, prioritize what matters, and keep the entire pipeline healthy. I saw things I couldn’t unsee, and I had to fix them. **** ## **Remediation: How do we fix the opportunities identified?** With better visibility, we can focus on real opportunities for efficiency. In April, we made a lot of changes based on what the dashboard was telling us. I have to admit, this part is fun: identify + fix is a fun workflow. I’ll give three examples of areas that really moved the needle for us - these remediations brought down our spend by 8% annually with just a few hours of work. ### Model tuning [Materializations](https://docs.getdbt.com/docs/build/materializations) represent one of the most significant cost drivers, with incremental materializations in particular emerging as a big expense bucket for us. When we examined our data platform costs through the cost management dashboard, we discovered that these operations were consuming a disproportionate amount of compute resources, especially for large event-based models. The fix in many of these turned out to be a fairly deep concept in dbt called [incremental predicates](https://docs.getdbt.com/docs/build/incremental-strategy#about-incremental_predicates). Incremental predicates are a dbt configuration you add to incremental models. They allow you to limit the portion of the existing target table that gets scanned during a merge operation, based on conditions like only considering rows from the last week. We use incremental predicates to scan only the rows we need. Same results, less compute and less build time. Just a few of these provided ~3.5% annual savings ### Depreciations It’s difficult to understand ROI without visibility into both productions costs and consumption; another feature of the cost management dashboard is the ability to see consumption queries: The total number of queries of a given resource across all usage in the warehouse (includes BI/analytics tools, query consoles, etc.) ![Consumption queries](https://cdn.sanity.io/images/wl0ndo6t/main/d4b0cc41c40ac9992c3150234cf40c0738a5411f-1060x816.png) This visibility, combined with understanding downstream lineage, allowed us to find some very clear deprecations: expensive models with little to no consumption metrics and no downstream dependencies. By examining both consumption metrics and lineage analysis, we could confidently identify models that were costly but provided no value. Simply disabling or deleting these unused models in the DAG provided an additional ~2% in our yearly warehouse costs with zero impact on business operations. ### Optimized testing Another area that is relatively expensive is testing, particularly for huge event tables. Luckily, again, the dashboard makes it easy to identify some expensive culprits. ![Test selection](https://cdn.sanity.io/images/wl0ndo6t/main/04881110ff4dd7a2c454d4baff43c61c968bb499-1068x739.png) We ended up finding around 15 tests that were costing a lot unnecessarily. Another 3% annually saved. The fix? You guessed it, another obscure but powerful feature: the `WHERE` clause in tests. For large, incremental event tables, we don’t need to test every single row, on every single run - we can just test the last few days of data. This can be achieved with the [where test property](https://docs.getdbt.com/reference/resource-configs/where). Fixing existing issues is great, but what if there were a way to avoid excess costs and inefficiencies altogether? ### Avoidance: Fusion powered state-aware orchestration If this sounds like rocket science, it [sort of is](https://docs.getdbt.com/blog/the-levels-of-sql-comprehension). Lucky for me, this feature is as simple as flipping a setting. [State-aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-about) is a feature powered by the dbt Fusion engine that intelligently determines which models to build based on detecting changes in code or data. It significantly reduces compute costs and runtime by only building models that will actually change. ![State-aware orchestration](https://cdn.sanity.io/images/wl0ndo6t/main/f5ee18d1acf7079c65215e9152af89dc263480d3-742x428.gif) Key principles: - **Real-time shared state:** Jobs write to a shared model-level state, allowing dbt to rebuild only changed models across all jobs - **Model-level queueing:** Prevents "collisions" by queueing at the model level, avoiding unnecessary rebuilds - **Flexible support:** Works with both dynamic (state-aware) and explicit (state-agnostic) job building approaches - **Sensible defaults:** Works out-of-the-box with optional advanced configurations This feature is currently in beta and available to Enterprise and Enterprise+ customers using the [dbt Fusion engine](https://docs.getdbt.com/docs/fusion/about-fusion). For us, turning it on will save us 9%+ of our annual dbt warehouse workload costs and will result in 25% fewer excess models built. There’s a lot more we can do with state-aware orchestration. Next up, we will start to use the [advanced configurations](https://docs.getdbt.com/docs/deploy/state-aware-setup#advanced-configurations) to set SLAs, and allow us to do just a single `dbt build` command that intelligently orchestrates only what it needs to, this would reduce the number of jobs we have and significantly reduce maintenance. ### **What’s next** Building the habit of cost-aware engineering isn’t just good for your budget: it’s good for your workflows. It leads to faster builds, leaner DAGs, and fewer surprises. With the dbt platform, we’re no longer guessing where inefficiencies live. We’re acting on real signals and building with intention. If you’re on an eligible data platform and plan, here are a few steps to take: - **Join the Preview** for the cost management dashboard - **[Migrate](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-fusion) to the dbt Fusion engine** to lay the foundations for smarter orchestration and better performance - **Ask to join the Beta for Fusion powered state-aware orchestration** to stop running what doesn’t need to be run Each step gets you closer to a more efficient, scalable data platform and gives your team more time to focus on what actually matters. There’s a lot more coming in this area of the platform, so stay tuned. Cost awareness isn’t a side quest. It’s just good engineering. --- --- title: "The dbt Fusion engine public beta is now available on BigQuery" description: "Launching more platform and feature support on the path towards GA." url: "https://www.getdbt.com/blog/dbt-fusion-engine-public-beta-bigquery" date: "2025-07-02" authors: ["Jeff Mills"] categories: ["Product"] --- # The dbt Fusion engine public beta is now available on BigQuery Exciting news on the [dbt Fusion engine](https://www.getdbt.com/product/fusion) front: teams using BigQuery can now participate in the public beta of Fusion. Since the [initial beta launch of Fusion](https://www.getdbt.com/blog/get-to-know-the-new-dbt-fusion-engine-and-vs-code-extension) in late May, which included support for Snowflake, we’ve expanded data platform support to include [Databricks](https://www.getdbt.com/blog/databricks-users-get-ready-to-experience-the-new-dbt-fusion-engine), and now BigQuery. This is an important milestone in our [path towards GA](https://docs.getdbt.com/blog/dbt-fusion-engine-path-to-ga), as BigQuery is one of the most-used adapters across all of dbt, and now these users will be able to experience the power of Fusion. [Watch video](https://youtu.be/NiNkdThkKAI) Fusion isn’t just an incremental improvement to dbt—it’s a fundamental shift in speed, intelligence, and developer experience for anyone building with dbt, and we’re excited to extend these benefits to the BigQuery ecosystem. [Read the docs](https://docs.getdbt.com/docs/fusion/about-fusion) to learn more and get started with Fusion. ## The dbt Fusion engine now on BigQuery The dbt Fusion engine represents the technological foundation for a [new era of analytics engineering](https://www.getdbt.com/blog/where-we-re-headed-with-the-dbt-fusion-engine). Born from a complete ground-up rewrite of dbt in Rust with SQL comprehension built in, Fusion is designed for performance at scale, pairing blazing-fast execution with deep understanding of your analytics code. For BigQuery users, this means: - Improved developer experience: Fusion doesn’t just execute your SQL; it understands it. SQL comprehension gives developers real-time feedback and error checking as they write SQL, catching mistakes and providing suggestions instantly, without executing queries against the warehouse. This now precisely understands BigQuery SQL and provides all the same benefits to this ecosystem. - Significant performance gains: Project parsing and compilation are as much as 30x faster than with dbt Core, dramatically improving velocity and reducing turnaround time from code change to insight. ![Parsing 10k](https://cdn.sanity.io/images/wl0ndo6t/main/2b2250b9b2d7eef12e96e6cfc32a5b1d2dba960c-700x184.gif) - Automatic cost optimization: Fusion is state-aware, meaning it has an always up-to-date understanding of what's in the project code or what exists in the database, bringing a new level of intelligence to dbt and the platform experiences we deliver. For example, with [state-aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-about), dbt now knows what’s changed in your DAG and only runs models with new data. Not only does this improve data velocity and eliminate redundant workloads, it helps organizations reduce cloud compute spending by 10% or more…just by turning it on. ![Automatic cost op](https://cdn.sanity.io/images/wl0ndo6t/main/afb62fbf472e7cf9c573f32fc2326d1a8d8f5a5e-1600x995.png) - More detailed governance (coming soon): Precise column-level lineage and richer metadata will soon provide enhanced data governance, making it easier to manage risk, enforce policies, and lay the foundation for safe AI applications downstream. **** Fusion is a game-changer: it empowers analytics engineers to move faster, iterate with confidence, and focus more on business logic than debugging or firefighting performance issues. ## Who benefits? - Data engineers and analytics engineers running transformations at scale on BigQuery will see dramatically reduced parse times of dbt projects, lower cloud spend, and richer metadata for downstream consumers. - Organizations requiring strict governance and/or auditability gain robust lineage, improved cost management, and the freedom to scale beyond a single data warehouse. This is especially valuable for companies orchestrating complex pipelines, handling sensitive data, or operating in hybrid/multi-cloud environments. ## Getting started with Fusion **Fusion is now in public beta for projects on Snowflake, Databricks, and BigQuery. **Stay tuned for more upcoming platform support and feature coverage in [our path towards GA](https://docs.getdbt.com/blog/dbt-fusion-engine-path-to-ga). You can get started with Fusion today with the freely available and permissively licensed source code and binary, or try it locally using the [dbt VS Code Extension](http://docs.getdbt.com/docs/install-dbt-extension). If you’re a managed dbt customer and your project(s) are eligible, you can enable Fusion with just a few clicks and unlock premium features like state-aware orchestration, cost dashboards, and enterprise governance. Ready to dive in?** **[Explore the docs](https://docs.getdbt.com/docs/fusion/about-fusion) or [watch a demo](https://www.getdbt.com/product/fusion). --- --- title: "Data product management: Best practices" description: "Explore best practices for data product management—and how dbt helps teams create governed, scalable, and trusted data assets." url: "https://www.getdbt.com/blog/data-product-management" date: "2025-07-02" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data product management: Best practices The modern data stack delivers scale, but clarity demands discipline and rigor. [Data products](https://www.getdbt.com/blog/key-components-of-data-mesh-creating-and-managing-data-products) are a method of packaging polished datasets, making them easier to discover, secure, and govern. The [Gartner Survey 2024](https://www.gartner.com/en/documents/5272363) reported that 50 percent of organizations already implement data products, and 29 percent of them are currently considering it. Despite this growing adoption, the majority of organizations struggle to make data products work in practice. The issue stems from vague ownership, patchwork processes, and disconnected teams. Spreadsheets, reports, and dashboards accumulate without a common standard or ownership. What would happen if you could treat data as you treat product development? Think of versioned, documented, and reliable datasets that support decision-making instead of ad hoc requests. That’s the essence of data product management. It turns one-off tasks into consistent, scalable resources that enable the business to act decisively. In this article, we’ll outline best practices in data product management, and how dbt enables teams to apply software engineering best practices to analytics workflows. ## Core best practices for data product management Data product management requires more than tools or talent. Teams must have a defined group of practices to create dependable, trusted outputs. Here are a few of them: ### Goals, metrics, and user-centered design Base data products on quantifiable results. Teams must align all data products directly with a business goal. Setting clear [SMART](https://www.atlassian.com/blog/productivity/how-to-write-smart-goals) (Specific, Measurable, Achievable, Relevant, Time-bound) objectives ensures everyone understands how to measure progress and track success. For example: - **Reduce manual work**: Lower the time required to create monthly financial reports by eight hours (10 to two) in three months. - **Increase usage**: Achieve at least 70 percent of campaign managers using the new marketing dashboard actively within six weeks. To effectively monitor progress, think of metrics that indicate user engagement and the general health of your data products, which include: - **Adoption rate:** Proportion of planned users who are actively searching the data product. - **Data freshness:** The gap between the availability of data in the source system and the product. - **User satisfaction:** Review survey data or NPS scores to learn how users rate ease of use and trust. Interview stakeholders, analysts, data scientists, and business teams directly to uncover any pain points. Take the time to observe how they operate with the data on a daily basis, so you can identify friction that may not be apparent during interviews. Create lightweight, rapid prototypes, such as sample dbt models or simple dashboards, to obtain early feedback. Continue on a small-scale cycle to narrow the scope and ensure alignment with your objectives before expanding further. ### Defining, documenting, and discovering data products An effective data product ought to be: - Discoverable - Addressable - Self-describing - Interoperable - Secure - Properly documented Ensure data products are discoverable by maintaining a central catalog or marketplace. The catalog can enumerate all models, sources, and exposures. Metadata tags such as domain, update cadence, and owner allow users to find what they need fast. #### Interoperability and documentation Addressability and interoperability require the provision of a stable, human-readable identification for each product, using a consistent naming convention. Outputs must be in standardized formats (such as Parquet, CSV, or materialized views) to promote frictionless integration with downstream tools and teams, eliminating the need for further reformatting. Self-describing data products have well-defined metadata that conveys information related to the business logic, ownership, purpose, and frequency of updates. Such metadata makes everyone aware of how the data was generated and how to use it effectively. Additionally, good documentation clearly explains the purpose of the data, the transformations applied, and the intended decision context, encouraging understanding and trust. #### Security and automation To ensure security, use role-based access controls on your data warehouse to assign roles. Implement row-level security to restrict access to specific records as needed. dbt assists in automating many of these tasks. Its in-built documentation generator creates a [browsable catalog](https://docs.getdbt.com/docs/build/exposures), containing lineage graphs, model descriptions, and metadata annotations. The models can be enriched with tags and classifications, allowing data products to be discovered, audited, and trusted more easily throughout the organization. ### Data quality and governance Reliable data products rely on stable data quality and governance. Establish explicit data quality criteria, such as ensuring that there are no more than 0.1 percent of duplicate records or fewer than 0.01 percent of null values in key dimensions. #### Testing and checks Automated testing plays a significant role in ensuring you meet these standards and maintain strong governance. Standard validations cover uniqueness checks to prevent duplicate keys, not-null constraints for required fields, and relationship tests to ensure referential integrity. Most teams can go further and add complex assertions. For example, add checks to validate that numbers fall within a valid range. These checks act as a gateway in the CI/CD pipeline, preventing critical errors from being deployed to production. #### Lineage and stewardship Lineage and traceability facilitate the easier comprehension of how data flows as it originates from raw sources and appears in downstream reports. Teams can trace errors back to their origin and understand how upstream changes impact dependent datasets and reports. To strengthen accountability, organizations appoint data stewards for each domain. Stewards can monitor data quality, enforce compliance, and manage schema changes. Meanwhile, data owners troubleshoot failed tests, update policies as requirements change, and keep stakeholders informed of any updates. ### Cross-team collaboration and roles Coordinate across business and technical functions to make data products manageable. The key to success is a common playbook, which clarifies responsibilities, organizes communication, and fosters feedback loops. **Establish roles** The first step is to define clear roles so that everyone understands their responsibilities, roles, and contribution to the data product lifecycle. - **Product owner**: Designs the roadmap and concentrates on the most important features to the business. - **Analytics engineer**: Constructs and maintains data transformations, tests, and documentation to maintain data trustworthiness. - **Data governance lead:** In charge of compliance, access control, and quality standards in the areas. - **Data consumer:** Provides data product usage, confidence, and adoption feedback. **Build communication connections** Second, establish formal communication networks to ensure teams are in sync and remain up-to-date with the latest developments. Here are some of the ways you can create a communication channel: - **Frequent stand-ups:** Daily or weekly meetings ensure that analytics, engineering, and business teams are aligned on priorities, blockers, and progress. - **Office hours and ad-hoc support:** Special hours to exchange knowledge, technical support, and onboarding. - **Feedback loops:** Surveys and retrospectives conducted after every sprint will enable the collection of feedback, identification of problems, and process improvements over time. **The role of dbt** dbt enables collaborative workflows using modern engineering. - **Git integration:** Teams can integrate branching [strategies](https://docs.getdbt.com/best-practices/best-practice-workflows) and pull requests to review and merge changes securely. - **Multi-environment deployments:** Prevent the unexpected impact of developing, testing, and promoting models in various [environments](https://docs.getdbt.com/docs/deploy/deploy-environments). - **Slack notifications:** dbt Cloud sends real-time [notifications](https://docs.getdbt.com/docs/deploy/job-notifications) of the build status and lineage updates, so that stakeholders always know what has shipped. Implementing these practices and tools establishes a solid foundation that fosters accountability, transparency, and teamwork, ultimately delivering data products that people can trust. ### Agile iteration and prioritization An agile mindset enables data teams to rapidly produce actionable insights and data products that can be trusted to drive smarter decisions and accelerate the return on investment (ROI). Rather than shipping all the information at once, ship the smallest data asset that fulfills a business need, such as a single fact table or a simple dimension. #### Early delivery Early delivery allows teams to elicit feedback, correct logic, and demonstrate value before scaling. Sprints and a prioritized backlog make development organized and predictable. Sprints of two to four weeks offer a well-understood cadence of planning, building, and reviewing work. The backlog tracks all potential improvements and corrections, ranked by impact and urgency. This model helps the team stay focused on business goals and improves in phases. #### Feature flags Introduce new functionality under feature flags or small-scale beta releases to monitor usage and verify outcomes without impacting all users simultaneously. Even rapid changes must undergo automated testing and adhere to your governance criteria before being deployed in production. The method enables teams to ship with confidence, iterate more quickly, and ensure data quality. #### Modular design dbt supports this approach through its modular design and [incremental models](https://docs.getdbt.com/docs/build/python-models). Teams can develop, test, and review small subsets of models in isolation before merging them into the main. This makes it easier to ship frequent, controlled updates with clear impact analysis and minimal risk. **** ### Scalability, performance, and maintenance Scalability and long-term maintenance of data products make them reliable as demand increases. The initial strategy is to de-modularize pipelines. The small-scale reusable prototypes are also easier to test, debug, and build without affecting the other parts of the system. #### Performance tracking The secret to pipeline efficiency is performance tracking. Observing query performance and resource utilization will help identify slow transformations or bottlenecks. After that, you can further optimize by rewriting some of the logic, adjusting indexing strategies, or changing how data is materialized to fit your workloads and query patterns. For example, transitioning a view-based model to an incremental table can provide enormous performance improvements as your data grows. Organizations must also scale beyond technical optimization as pipelines and data products expand. #### Domain ownership Moving to a domain ownership model enables every business unit to own its data products, create local expertise, and minimize the dependency on a central team. This model helps to share the workload more evenly, enhances responsibility, and accelerates the iteration process. Periodically review and refactor or deprecate old models to maintain a clean and manageable environment in the long run. **How dbt helps** dbt allows flexible [materializations](https://docs.getdbt.com/best-practices/materializations/2-available-materializations) such as views, tables, and increments to design at scale and allows teams to customize performance strategies to different datasets. Its modular design makes it easy to break pipelines into manageable components. Integrations with [orchestration tools](https://docs.getdbt.com/docs/deploy/deployment-tools) like Airflow or Prefect reliably coordinate complex workflows as your environment scales. ### Monitoring, observability, and feedback Data products are living systems that require constant improvement and monitoring. Monitoring usage and freshness helps keep data relevant and trusted. Checking API calls and query logs provides insight into the frequency of each product and the team accessing it. Data latency can be avoided by performing freshness checks. They help detect outdated data before it impacts the decision-making process. #### Robust monitoring Solid monitoring practices can help maintain system reliability. Setting up alerts for test failures, sudden spikes in resource usage, or unexpected drops in data volume enables teams to catch problems early and keep systems running reliably. Monitoring platforms such as [Prometheus](https://prometheus.io/) and [Datadog](https://www.datadoghq.com/) offer a clear view of how pipelines and infrastructure are performing, which simplifies the process of identifying and resolving issues. #### User feedback Maintaining a strong connection with data users is crucial. Gathering regular input through embedded feedback tools or scheduled discussions ensures that their queries are heard and addressed. Over time, this steady feedback loop builds trust and encourages broader adoption, as people see their suggestions driving real improvements. ## Ready to get started? Data product management enables teams to consolidate disparate data processes into stable, custom-built resources. A clear focus on business goals, robust governance, and iterative development helps create reliable and trusted data products. With [dbt](https://www.getdbt.com/product/dbt), you can forge [a unified analytics development process](https://www.getdbt.com/resources/the-analytics-development-lifecycle). For example, you can: - Set up development, staging, and production environments that are isolated from each other, allowing tests to be performed without impacting live data. - Create a minimum viable product in dbt by [modeling fundamental](https://www.getdbt.com/blog/guide-to-dimensional-modeling) tables, including YAML-based tests of data quality, and snapshots to show changes over time. - Iterate on pull requests, using CI jobs that test only the affected models, and merge when all quality checks succeed. Take the free [dbt Fundamentals](https://learn.getdbt.com/courses/dbt-fundamentals) course to learn how to model, test, document, and deploy in dbt. To gain a more in-depth understanding, consider more advanced training on [data mesh](https://learn.getdbt.com/courses/dbt-mesh), [CI/CD optimizations](https://docs.getdbt.com/docs/deploy/continuous-integration), and [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot) integrations. --- --- title: "Empowering analysts with dbt: Who they are and how we help" description: "See how dbt empowers analysts of all skill levels to turn ad hoc work into trusted, production-ready data products." url: "https://www.getdbt.com/blog/empowering-analysts-with-dbt" date: "2025-07-01" authors: ["Patrick Barch"] categories: ["Insights"] --- # Empowering analysts with dbt: Who they are and how we help The term "data analyst" has always been broad, but today it's evolving faster than ever. As data roles diversify and generative AI reshapes workflows, traditional definitions no longer capture the full scope of what analysts do or what they need. At dbt Labs, we believe it’s time to rethink the role of the analyst and, more importantly, to build tools that meet them where they are. In this post, we’ll share how we think about [building for the analyst ](https://www.getdbt.com/product/analyst)at dbt Labs, why this perspective matters, and how our product strategy is intentionally designed to empower them across a wide spectrum of technical skill levels. ## Analysts are not just dashboard builders Many tools still treat analysts as lightweight BI users, clicking through dashboards, downloading CSVs, or handing off requests to engineers. But that view is outdated. Today’s analysts debug models, define metrics, investigate data lineage, and increasingly contribute to the foundations of trustworthy analytics and AI. Some write SQL daily. Others work in visual tools or prompt AI assistants to get the answers they need. Regardless of how they work, the expectations placed on them have grown. We see analysts as critical participants in the modern data workflow, and we’re building dbt to reflect that reality. ## No one definition of “analyst” The label "data analyst" covers a wide range of responsibilities, depending on the company, the team, and the stack. In one organization, a “data analyst” might be responsible for building ETL pipelines and maintaining data infrastructure (tasks that other companies might label “data engineering”). In another, a technically inclined “data analyst” could be closer to what some firms call a “business analyst,” primarily interpreting dashboards and writing non-technical reports. The reality is that data practitioners operate within a broad ecosystem and bring varied technical skill sets to different jobs to be done. Titles mask the actual scope of work: two people both called “data analyst” may have dramatically different day-to-day responsibilities, tooling needs, and comfort levels with code. ## The spectrum of technical proficiency When we build new analyst-friendly features, we think in terms of _capabilities_ and _comfort levels_ across three key dimensions: ### Coding proficiency (SQL, Python, etc.) Comfort working with code-based tools, ranging from writing advanced SQL and Python to preferring visual interfaces or natural language interactions. ### Familiarity with code management practices Experience working with version control systems like Git, including an understanding of branching, pull requests, merge conflicts, and collaborative development workflows. ### Familiarity with the Analytics Development Lifecycle (ADLC) best practices Knowledge of how testing, documentation, observability, code reviews, and rollback strategies apply to analytics and data workflows.  Some practitioners may only sporadically apply these practices; others bake them into every project. This view helps us design tools that meet people where they are, and where they might want to grow. **** ## Bringing best-in-class data practices to analysts At dbt Labs, we’re designing our products to support analysts who are shaping data products, not just consuming insights. While the broad array of tools available to analysts today have certainly empowered them to take data work into their own hands, it’s left them without the governance, testing, and SQL-first foundation the modern data stack requires. Our new capabilities–[dbt Canvas](https://docs.getdbt.com/docs/cloud/canvas), [dbt Insights](https://docs.getdbt.com/docs/explore/dbt-insights), and [dbt Catalog](https://docs.getdbt.com/docs/explore/explore-projects)–provide a clear path to production, combining speed with trusted best practices. To make this more concrete, with the new analyst capabilities in dbt: 1. Users who might work in dashboarding tools, like PowerBI or Sigma, instead have a path to promote their business logic into reusable dbt models with dbt Canvas, turning one-off analyses into production-ready reusable data products. 2. Users in legacy drag-and-drop tooling can get the same fast ad-hoc workflow in dbt Canva, but with added visibility and governance. 3. Users running ad-hoc queries can explore governed data assets and validate their work with dbt Insights, speeding up discovery while staying aligned with standards. 4. Users navigating complex data environments can easily find, understand, and trust their data with dbt Catalog, reducing duplication and improving collaboration. Our latest analyst-focused investments aren’t a shift in direction; they’re a natural extension of our mission to serve more people working with data. We’re building on our foundation to serve even more analysts, more effectively. We’re focused on meeting the needs of analysts who: - Regularly build dashboards or respond to ad hoc data requests- answering “that one quick question” from a boss, peer, or stakeholder - but who want their work to be maintainable, not just one-off assets. - Want to use context-aware, governed AI to more efficiently answer questions and build data products - Occasionally need to operationalize new data assets into production pipelines without being engineers by trade. - Seek guardrails to safely contribute to their organization’s dbt project. These analysts often sit in a gap between lightweight tools and heavy engineering processes. Our new capabilities give them clarity, guardrails, and flexibility to move faster, without being overwhelmed. **** ## Technical users benefit too Building on dbt’s foundation as the tool for helping data teams work like software engineers, our analyst features benefit even technical users. Visualization of dbt models, intelligent query interfaces, and context-aware AI help streamline development, accelerate ramp time, and improve cross-functional collaboration. The analyst suite complements and expands on the code these teams have already written, making dbt even more powerful for our existing experienced dbt analytics engineering and data analyst users. We’ve seen how AI makes data analysts more and more technical. Through our work on the [**dbt**](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) [**MCP server**](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) and AI-powered chat interfaces, users can now engage with dbt models through natural language. As these conversational experiences evolve, they’ll enable secure, self-service access to insights without requiring SQL or navigating developer tools. Our analyst offering isn’t about narrowing focus, it’s about meeting more users where they are and giving more analysts the tools they need to turn data into impact. ## Key use cases our new analyst offerings support today Our [new suite of analyst capabilities](https://www.getdbt.com/blog/dbt-canvas-is-ga), including dbt Canvas, dbt Insights, and dbt Catalog, enable analysts to build, explore, and share data products with speed and confidence, all within a governed environment that scales with their needs. ### Exploratory analytics, aided by context-aware AI Analysts often need to investigate metrics, trends, or anomalies on the fly. dbt Insights provides a streamlined environment for running ad hoc queries, with AI assistance that understands the structure and context of the analyst's dbt project. This enables analysts to generate more accurate queries, explore data faster, and get to insight more efficiently, whether they’re writing SQL themselves or starting from natural language. **** ### Visual, low-code/no-code transformations For analysts who want to contribute new models but prefer to avoid raw SQL, dbt Canvas provides a visual, drag-and-drop interface for building transformations. Behind the scenes, this generates production-ready SQL that can be audited and version controlled, but the user interacts with an intuitive UI: selecting tables, applying joins, filters, aggregations, and so on. This lowers the barrier to creating new data assets while still producing artifacts that are governed by version control, testing, and deployment checks. ### Discovery and trust Analysts can browse and search for datasets, view lineage, and inspect metadata such as freshness, owners, and descriptions with dbt Catalog. This helps analysts confidently choose and use the right data without needing to escalate to engineering teams. ### Streamlined collaboration between data developers and analysts Hardcore data developers build core data models, transformations, and pipelines. Analysts often bring deep domain knowledge, understanding the nuances of business metrics, customer behavior, or product usage patterns. dbt enables collaboration between these roles through built-in workflows like pull requests, code reviews, and shared documentation. This allows analysts to prototype or suggest model adjustments and then collaborate via version control and review processes. While developers contribute engineering rigor, analysts contribute domain insight. By aligning workflows in a shared environment, dbt ensures that both technical excellence and business context shape the data products teams create together. These workflows reflect how dbt helps analysts of all skill levels contribute more meaningfully to analytics, from first exploration to fully governed data products. ## Looking ahead The [role of the data analyst](https://www.getdbt.com/blog/data-analyst-closer-to-data) is fluid and evolving. As AI reshapes how insights are generated and shared, and as embedded and self-service analytics become more common, dbt Labs is focused on enabling analysts to do more with confidence and clarity. With our new analyst-focused offerings, **dbt Canvas**, **dbt Insights**, and **dbt Catalog, **we’re addressing key friction points head-on including, making it **easier to discover** and trust the right data, **reducing the need to switch between disconnected tools** for analysis and presentation, and helping analysts **move from exploration to production seamlessly**, all while staying within a governed environment. We also recognize that not every data practitioner fits neatly into a single role or title. While we use personas like "analyst" or "developer" to guide our product decisions, these are flexible starting points, not fixed definitions. Our analyst offerings are designed to support a broad range of use cases, and our focus remains on helping them turn curiosity into impact by removing friction, supporting collaboration, and maintaining the standards that data teams depend on. --- --- title: "Empowering the enterprise for the next era of AI and BI" description: "Unlock faster and smarter data workflows with the new dbt Fusion engine, built for AI-native development and enterprise scale." url: "https://www.getdbt.com/blog/empowering-enterprise-ai-bi" date: "2025-06-25" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Empowering the enterprise for the next era of AI and BI dbt emerged on the scene nine years ago. It has since become the standard in data transformation, introducing best practices such as modularity, testing, and documentation to analytics workflows. However, a lot’s changed in that decade. AI workflows are now a reality and are here to stay. Technologies like Apache Iceberg see us slowly inching closer to the commoditization of compute. To stay ahead of the demands of modern data teams, dbt needs an update. It needs something that keeps developers in the flow, speeds up their development loops, and integrates natively with AI systems. That’s why [dbt Labs acquired SDF](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs) - and why we embarked on a five-month journey to rewrite dbt entirely from the ground up to incorporate it. The result is the **dbt Fusion engine**, a new state-of-the-art data transformation engine and SQL compiler that powers a smarter, faster developer experience and optimizes data engineering costs. Let’s take a look at how the new dbt Fusion engine improves the developer experience, enables AI-native development, and provides enterprises more confidence in their data through real-time validation and impact analysis. ## dbt’s innovations and limitations For the past nine years, dbt Labs has released features that make it possible to institute a mature [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) around data projects. However, there’s a disconnect between dbt and the systems against which it operates that slows developers down. At its heart, dbt takes [Jinja templates](https://jinja.palletsprojects.com/en/stable/) and resolves them down to raw SQL strings. It defers understanding of these strings to data storage systems like [Snowflake](https://snowflake.com/) or [Databricks](https://databricks.com/) that are hosted in the cloud. This means a developer doesn’t know if their dbt code will compile as expected or not as they’re writing it. They have to run it against their target system, which responds with a status on the code. That introduces network latency as well as compute cost. All of this leads to slow development loops, as data developers are dependent on warehouse validation for even basic syntax checking. ## How Fusion changes the game Fusion is a ground-up rewrite of dbt that incorporates the technology from SDF, a high-performance SQL compiler that can parse every major SQL dialect (Snowflake, [Redshift](https://aws.amazon.com/redshift/), [BigQuery](https://cloud.google.com/bigquery?hl=en), etc.) and validate it locally. With Fusion, the entirety of your data warehouse is fully defined and statically analyzed as code. Fusion provides a host of benefits over dbt Core: ### Blazing fast performance SDF originally built its engine around dbt’s Jinja syntax. However, when we sought to integrate the two platforms, we ran into a major issue: dbt was written in Python, while SDF was originally written in Rust. To address this, the (now larger) dbt Labs engineering team undertook a five-month migration of the dbt engine from Python to Rust. This spells huge performance gains: dbt Fusion brings 30x faster project parsing and compilation to the table. The move to Rust also simplified dbt Fusion usage. Since Rust compiles down to binary executables, there’s no need to install Python, run pip, debug Python library dependency conflicts, etc. Fusion just works out of the box. ### Deep SQL comprehension capabilities Fusion takes dbt’s SQL models, resolves them to logical plans, and validates each query in the model according to its representation of the data warehouse’s SQL dialect. That means it can validate the correctness of SQL in the code editor, before it even hits the data warehouse. It also means it can statically analyze that plan to produce column-level lineage and propagate metadata. That gives data developers deep insights into how a given code change affects their existing environment. ### State awareness and optimization Fusion understands the freshness of data every time it runs. That means it can skip reruns of a model if the upstream data hasn’t changed. That simplifies orchestration, as it means you can run your pipelines at whatever interval and rest assured that they won’t do any unnecessary work. This optimization saves both time and money. Currently, we’re seeing early estimates of 10% cost savings. We expect to get this up to 20 to 30% with additional cost optimization capabilities we plan on launching soon. ## dbt Fusion engine + IDE: a supercharged developer experience Fusion’s power really shines through when paired with a first-class Integrated Developer Environment (IDE). You can leverage these advanced features using the [dbt Studio IDE](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/develop-in-the-cloud), Canvas, or our new [VS Code extension](https://marketplace.visualstudio.com/items?itemName=dbtLabsInc.dbt). Let’s dive into how these features can supercharge data development teams’ work using the VS Code extension as an example. (Note: You can see an interactive demo of these features in our corresponding YouTube video.) [Watch video](https://www.youtube.com/watch?v=2_YHbePC42g]) ### Validating SQL and viewing data lineage locally Assume you have a simple model that tracks baseball players and related data, such as teams and salaries, that’s stored in Microsoft SQL Server. Let’s say you have a simple dbt model, stg_players.sql, which is part of your model’s [staging layer](https://docs.getdbt.com/best-practices/how-we-structure/2-staging). You can run your model right in VS Code using the new dbt extension and see the results on the Query Results pane. More powerfully, however, you can dig into the data lineage for a specific column - let’s look at FULL_NAME - and see where the data comes from and where it’s going. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/3c2c9c6842710961de7a09003f14e11757c9e5a6-1600x893.png) Note that this is all happening locally. You could be offline, and this would run in the exact same fashion. You can really see the power of the Fusion compiler by looking at the developer experience. Say you have a block of Common Table Expressions (CTEs) in your pitcher_salary_analysis.sql model for various intermediate data, such as salary. Let’s say you wanted to add a median annual salary for the pitchers on teams. So you add a yearly_median_salary CTE…but forget to include the trailing comma at the end, making this invalid SQL. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/ce7fd91f9a0578f14b97c7451d944f792edecce2-1600x835.png) As soon as you save the file, you get a syntax error. dbt Fusion detected this error locally without pinging your data warehouse. Instead of waiting to detect this after you run it against a development copy of our warehouse or kick off a lengthy CI/CD deployment process, you get instantaneous feedback as you type and work. Finally, dbt Fusion in VS Code supports previewing a single CTE. In the past, if you wanted to do this, it required commenting out the rest of the file. Now, you can accomplish this by simply clicking the **Preview CTE** command in the editor and viewing the output. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/69b5c3dfb0793973968fab1a3c269970166071d3-1600x839.png) ### Intellisense and Copilot feedback loops The VS Code editor also brings powerful IntelliSense capabilities. If you start adding median() function to your code, for example, you can see the editor will pull up a list of functions available to us in Microsoft SQL Server. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/68b7f37b83f61f54f3b574dca87658182362c654-1600x835.png) Let’s say you try to complete this CTE by defining a field called median_salary as: median(salary) as median_salary Attempting to save this, however, gives you another and even more fundamental error detected by the compiler. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/987683f94784004587262eae03fdb78539122c1e-1600x836.png) Fusion’s compiler detects that median() is an aggregation function in your SQL dialect. ### Real-time impact analysis and refactoring Fusion in VS Code gives you instant feedback as you code, flagging downstream errors as soon as you make a change, before it ever hits CI/CD. No more waiting to find out something broke. Let’s say you alias the power_id column and break dependencies. Fusion detects the impact immediately: errors surface directly in your editor, so you can fix them before committing anything. Behind the scenes, dbt is compiling your project and analyzing downstream references as you type. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/934389637e3fbcb0c79f4a57faeefd28b4a9e4b2-1600x827.png) Fusion also expands macros inline, so when you use ref(), you see exactly what it resolves to. Change a macro definition? You’ll see any model that relies on it light up with errors. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/71eb80757521c4f7928e57206d67a92778177b95-1600x830.png) And when you rename a model, Fusion prompts you to refactor. One click updates all references across your project, with the option to preview or revert. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/4a625e7a6092fcd8ec6253ce412c4f3f7febeee7-1600x833.png) This is the kind of developer experience that makes high-trust data work possible. When refactoring is safe and fast, you're empowered to improve your codebase instead of working around it. ## Conclusion and what’s next With Fusion, data developers can ship data-driven projects faster than ever. Its blazing-fast SQL compiler, IntelliSense, and support for embedded AI copilots mean developers can create, test, and publish models in a fraction of the time they could before. Plus, with local validation and state-aware orchestration, you can slash around 10% off of your data warehouse bill. We’re continuing to evolve Fusion to add new and useful features for faster analytics engineering. Later this year, we’ll release new data governance features that leverage Fusion’s advanced data lineage capabilities to create advanced tagging and classification workflows, such as removing the classifier tag from aggregated data. Fusion is live and available now to supercharge your team’s data development. To try the next generation of data engineering for yourself, [sign up for a free dbt account now](https://www.getdbt.com/signup) or [book a demo with us today](https://www.getdbt.com/contact). --- --- title: "Analyst autonomy vs data governance: How to have both" description: "Data teams are overloaded. Here’s how to give analysts self-service options without sacrificing security and compliance." url: "https://www.getdbt.com/blog/analysts-autonomy-data-governance" date: "2025-06-24" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Analyst autonomy vs data governance: How to have both You’re finally in the zone, cranking through real work, and suddenly you receive the dreadful ping: “Hi, can you send last quarter’s revenue numbers by city? Need it ASAP for a QBR deck.” For most data teams, these moments of fielding ad-hoc requests break momentum and drain focus. And pushing analysts to self-serve hasn't solved the problem; in fact, it often places them in a vicious cycle where analysts need answers fast. Without governed access to trusted data assets, they turn to untrustworthy sources, leading to broken dashboards, misinformed decisions, and even more ad hoc engineering requests. Meanwhile, the demand for high-quality, reliable data keeps climbing. The rise of generative AI (GenAI) has raised the stakes. These AI systems thrive on large volumes of well-structured, high-quality data, which puts even more pressure on already stretched data engineering teams. But the answer isn’t to lock things down even further. It’s not just more self-service, it’s governed self-service. High-performing teams treat these moments as signals. Instead of being the sole gatekeepers to clean, trusted data, they shift their focus from “how many questions can we answer” to “how many people can we enable to answer their own questions, safely.” That late afternoon ping becomes a cue, not just to deliver but to empower analysts, reduce chaos, earn influence, and build systems that scale. We’ll explore why governance is essential to scaling self-service, the key technologies that make it possible, and how high-performing teams empower analysts to reduce chaos and build sustainable, trusted data systems ## The real threat to data governance Many data requests - whether for new datasets, changes to existing models, or clarity on the origin or function of data - still go through a centralized data engineering team. That’s left teams with a backlog stretching weeks or even months. This creates bottlenecks that have two major downsides: - **Business slows down**: It leads to delays in getting insights into the hands of analysts and decision-makers who need to make critical business decisions - **Engineering progress stalls:** It means data teams are firefighting requests, preventing them from making needed improvements to data architecture that benefit everyone Analysts are feeling the pressure, too. When data solutions are delayed, those eager to keep pushing the business forward take matters into their own hands. Since many existing data pipelines are scattered across multiple data storage systems and data transformation solutions, they resort to rolling their own solutions, grabbing data from wherever they can, and building data assets that sit entirely outside the governed data environment. This is the _real_ threat to data governance. These on-the-fly data pipelines exist outside of governance structures. And when analysts are blocked from trusted systems, they create untraceable workarounds. This leads to: - **Invisible pipelines: **As data code isn’t checked in, and there’s no single location to find production-ready datasets with no version control, lineage, or audit trail - **Unreviewed transformations**: Logic errors, inconsistent definitions, and a decrease in data quality, as data transformations aren’t subject to review - **Unsecured data access**: As ungoverned data can’t be properly secured, classified, and managed according to your data governance policies, this puts compliance and privacy policies at risk This isn’t a self-service problem, it’s a governance gap. And it happens when analysts aren’t invited into the right environment to build safely. ## Providing governed, self-service access for analytics The solution is not to tighten the gates; it’s to make it easy for analysts and anyone else to find, contribute, and work with data safely, no matter where it lives in your organization, without compromising governance. What’s needed, in other words, is a data control plane. A [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) acts as a single abstraction layer across your data stack, providing shared workflows for data transformation, testing, observability, lineage, cataloging, semantic consistency, and more. It manages data movement, enforces configuration and policies, standardizes workflows, and runs queries against data, ensuring every contributor works within a governed framework. By giving analysts access to a governed data control plane, teams can increase efficiency and reduce costs. Simultaneously, it improves data quality and increases trust in data across the business by ensuring all data transformation and data access is approved, monitored, and transparent. ## How to empower analysts with governed self-service analytics using dbt If only the data engineering team has access to the data control plane, governance ends the moment work is handed off to the analyst team. Logic becomes siloed, data products go undocumented, and trust starts to erode. But when analysts are brought into the same governed workflows, complete with version control, testing, and lineage, teams break the cycle of ad hoc requests and build scalable systems together. That’s where dbt comes in. dbt is a data control plane designed for this new model of decentralized, governed collaboration. It gives analysts the tools to explore, transform, document, and query trusted data, without breaking pipelines or introducing risk. It enables governed self-service analytics through six key dimensions: - Trusted data exploration - Governed transformation - Ad-hoc analysis - Metadata and documentation - Context-aware AI - Semantic consistency This is what modern data engineering success looks like: Empowering analysts to build data products in a governed, collaborative environment—so engineers can focus on infrastructure, ML, optimization, and long-term value instead of chasing down broken dashboards. **** ### Trusted data exploration Data can’t drive insight if it isn’t trusted. And trust doesn’t come from access alone, it comes from context. Analysts may be able to find data assets in a standard catalog, but without clear metadata, lineage, and documentation, they’re left guessing about what the data means, how it was created, or whether it’s safe to use. [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) fills that gap by providing a governed interface for exploring both dbt models and upstream platform assets, like Snowflake tables and views. It surfaces rich, automatically generated metadata for each asset, including transformation logic, freshness, ownership, and lineage. Analysts can find and explore data assets and experiment with their data in context to derive new business-trusted insights. Administrators [can use role-based access control (RBAC)](https://docs.getdbt.com/docs/mesh/govern/model-access) to limit what analysts can see based on their roles and their need to access sensitive data. ### Governed transformation Exploration is just the first step. When analysts find the data they need, they also need a safe, governed way to act on it. [dbt Canvas](https://docs.getdbt.com/docs/cloud/canvas) enables analysts to take this even further with a visual, drag-and-drop tool to build and modify data models in a visual interface. Analysts can create new data transformations or improve existing ones visually, and their changes are automatically compiled into dbt-compatible SQL. These changes can be materialized in production via [dbt orchestration](https://docs.getdbt.com/docs/deploy/deployments). All new analytics code goes through quality checkpoints (testing, versioning, and orchestration via the dbt platform) to ensure proposed changes are thoroughly vetted before going live. This allows analysts to contribute meaningfully to data development while keeping every transformation aligned with governance, quality, and team standards. ### Metadata and documentation Metadata - the data about our data - is indispensable for understanding and building trust, reducing duplication, and scaling data use. It provides context around where data comes from, who owns it, and when it was last refreshed. [Data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage), which is one form of metadata, shows how data travels across your company. This enables analysts to verify the origins and deduce the meaning of data in a self-service manner. dbt generates lineage automatically with every model build and makes it available to analysts via dbt Catalog. Documentation is another invaluable form of metadata. Historically, there’s been no uniform, built-in way to document the meanings of tables and fields in data. With dbt models, engineers, analysts, and decision-makers [can easily collaborate on documentation](https://docs.getdbt.com/docs/build/view-documentation), embedding docs directly into data transformation models. All docs are compiled and discoverable via dbt Catalog with every push to production. ​​This shared context helps analysts confidently self-serve, while keeping the broader team aligned. ### Ad-hoc analysis Ad hoc analysis is where governance often breaks down. Analysts need quick answers—but without visibility into trusted logic or model usage, they’re left with two inefficient paths: They either spin up siloed query consoles and work outside the system, or submit tickets to the data team for requests that don’t justify the time or overhead. **[dbt Insights ](https://docs.getdbt.com/docs/explore/dbt-insights)(now in Preview) **changes that by providing a single, governed environment for ad hoc analysis. Analysts can query, validate, understand, and visualize data directly from production-grade dbt models, all with built-in context like freshness, lineage, usage, and ownership. Instead of guessing what model to use or pulling logic from outdated dashboards, they get a complete picture in one place. And if they want to move even faster, they can use **context-aware AI through [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot)** to generate SQL based on natural language, directly grounded in their dbt project. This means analysts no longer need to guess what logic to use or rely on outdated queries. They can self-serve with speed and confidence, without leaving the boundaries of governance. For data teams, it’s not just fewer tickets, it’s how they scale. ### Context-aware AI [Data from Stack Overflow](https://survey.stackoverflow.co/2024/ai) shows that a majority of engineers across disciplines are leveraging AI to accelerate shipping new software solutions. That same phenomenon [is transforming how we do analytics](https://www.getdbt.com/blog/how-ai-will-disrupt-data-engineering). Our own data from [the recent State of Analytics Engineering Report](https://www.getdbt.com/resources/state-of-analytics-engineering-2025) shows 70% of respondents are also using AI for code development; another 50% are using it to assist with documentation. Analysts and data teams can use [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot) to create simple and complex SQL queries using natural language. But this isn’t generic AI, it’s context-aware, meaning dbt Copilot understands your dbt project’s models, metrics, relationships, and documentation so your queries are grounded in governed, production-ready logic, not guesses. This enables analysts to find the data they need regardless of their comfort level with SQL. Analysts can also leverage dbt Copilot to assist engineers in writing data documentation or even in making changes to data transformation models in dbt Canvas. By embedding context-aware AI into a structured environment, dbt enables analysts to be more self-sufficient, regardless of their SQL expertise. This reduces queries and load on the most constrained data source of all: the data engineering team. ### Semantic layer Definitions of core metrics and business logic vary between teams or tools. That often leads to miscommunication and confusion, where trust breaks down. A [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) reduces this risk by creating a unified governed layer that predefines key metrics and logic and makes them easily accessible to all data team members. The [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) allows teams to centrally define metrics like "monthly active users" or "net revenue," then make them accessible across BI tools, embedded applications, and LLMs via APIs and built-in integrations. This ensures everyone, from analysts to dashboards to AI agents, is working from the same consistent, trusted definitions. ## Balancing governance and speed A common misconception is that governance slows you down. In reality, poor governance is what slows teams down, creating rework, data debt, and mistrust. dbt is built on the belief that you don’t have to choose between moving fast and doing things right. Features like testing, lineage, and orchestration are baked into the development lifecycle, making quality the default. [The new dbt Fusion engine](https://docs.getdbt.com/docs/fusion/about-fusion) takes this belief in a bold new direction. dbt Fusion immediately catches errors in SQL, previews expressions inline, and traces model and column definitions across your dbt project, greatly accelerating model development while maintaining quality. It enables overwhelmed data teams to deliver high-quality data more quickly than ever before. And it’s not just about speed. dbt Fusion also unlocks powerful new capabilities for data governance. Soon, you’ll be able to use Fusion to show an audit-ready view of your footprint of personally identifiable information (PII) across your entire data landscape. Using dbt as your data control plane, you enable analysts to move independently within a trusted framework. You unlock an array of governed self-service features that analysts can use to find, inspect, verify, learn about, and glean insights from trusted data. Data engineers shouldn’t fear bringing analysts closer to the transformation layer. High-performing teams see this not as a control risk, but as a chance to scale their impact. By empowering analysts within a governed framework like dbt, they reduce chaos, reclaim engineering time, and build trust across the organization. As a result, analysts move faster, data teams stay focused, and the business runs on data that’s accurate, aligned, and accountable. Want to see how high-performing teams are empowering analysts and refocusing engineering time with dbt? [Ask us for a demo today](https://www.google.com/search?q=site%3Agetdbt.com+fusion&rlz=1C1TKQJ_jaJP1074JP1074&oq=site%3Agetdbt.com+fusion&gs_lcrp=EgZjaHJvbWUyBggAEEUYOTIGCAEQRRg60gEIMjc4M2oxajSoAgCwAgE&sourceid=chrome&ie=UTF-8). --- --- title: "From Docker to Dagger" description: "Solomon Hykes, the creator of Docker, on how containers changed everything." url: "https://www.getdbt.com/blog/from-docker-to-dagger" date: "2025-06-22" authors: ["Daniel Poppy"] categories: ["Insights"] --- # From Docker to Dagger _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/from-docker-to-dagger-w-solomon-hykes). _ In this season of the Analytics Engineering podcast, Tristan is digging deep into the world of developer tools and databases. There are few more widely used developer tools than Docker. From its launch back in 2013, Docker has completely changed how developers ship applications. In this episode, Tristan talks to Solomon Hykes, the founder and creator of [Docker](https://www.docker.com/). They trace Docker’s rise from startup obscurity to becoming foundational infrastructure in modern software development. Solomon explains the technical underpinnings of containerization, the pivotal shift from platform-as-a-service to open-source engine, and why Docker’s developer experience was so revolutionary. The conversation also dives into his next venture [Dagger](https://dagger.io/), and how it aims to solve the messy, overlooked workflows of software delivery. Bonus: Solomon shares how AI agents are reshaping how CI/CD gets done and why the next revolution in DevOps might already be here. _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways **Tristan Handy: I want to get you to give a little background on yourself, where you've been, what you've been up to for the last couple decades. I think many people will know you as the person who kicked off an avalanche that changed how we interact with compute environments by inventing Docker?** **Solomon Hykes:** Docker is the thing I'm known for. Pre-Docker, I grew up in France. I studied programming in a French school called EpiTech. It was a brand-new, unconventional school where you learned through nonstop programming, which I loved. Eventually, I got exposed to startups, despite being a complete outsider. I met someone who told me about them, and it stuck in my mind. Still in France at the time, I moved into my mom's house in the suburbs of Paris and worked out of the basement. By complete luck, I got into an early version of Y Combinator in 2010. That got us on the path to what would become Docker three years later. In 2013, we pivoted to Docker from our previous company, dotCloud. **Tristan Handy: The original thing was called dotCloud, right?** **Solomon Hykes:** Yep. It was about container technology and its potential, but we didn't quite know how to take it to market. DotCloud was about deploying and hosting people's apps—platform as a service—competing with Heroku and many clones. **Tristan Handy: When did Heroku become a thing?** **Solomon Hykes:** I became aware of it in 2009. Just as I was struggling in France with container tech. When we joined YC in 2010, we packaged that tech into dotCloud, our hosting platform. Our differentiator was using containers under the hood when others didn’t. That let us support many language stacks and even run databases in containers—which was unheard of at the time. Platform as a service was a tough business. Most startups went out of business or got acquired early. Eventually, we pivoted from selling the car to building an ecosystem around the engine—that became Docker. **Tristan Handy: Did you pivot because selling the car wasn't working? Or because people kept pointing at the engine saying, “Give me that”?** **Solomon Hykes:** Both. It was hard to market platforms. Developers expected free hosting, and hosting costs money. Margins were tight because of AWS. It always felt like pushing a boulder uphill. Meanwhile, people wanted to run things locally. There was no good ecosystem for that. Docker provided transparency, flexibility, and portability. **Tristan Handy: Can you define Docker and containerization, and how it differs from virtualization?** **Solomon Hykes:** Sure. Virtualization splits a physical machine into virtual ones using VMs—each with its own memory, compute, and storage. It gives flexibility, but with overhead. Containerization does something similar but at the operating system level. Instead of virtualizing the machine, you split the OS itself. It’s mostly done with Linux, which can subdivide itself into isolated units. Containers are more lightweight, letting you run hundreds or thousands, unlike VMs where you might manage a handful before hitting limits. Docker didn’t invent this, but we solved new problems with it. **Tristan Handy: I remember creating my first Docker container around 2015. I expected a slow boot-up like a VM, but it was instantaneous. Where is the OS in that setup?** **Solomon Hykes:** Great question. Docker relies on Linux. When you're on a Mac, it runs Linux behind the scenes—today via virtualization. Back then, we used lots of early, rough tools and kernel patches to make Linux containers work. Docker put all the pieces together in a coherent way. **Tristan Handy: So containerization wasn’t new, but Docker made it accessible?** **Solomon Hykes:** Exactly. The Linux kernel had features like namespaces and cgroups—building blocks for containers. But they weren’t user-friendly. We made a developer-centric abstraction on top of those tools. And Linux provided a massive compatibility layer. Unlike Java, which required writing your app in Java, Docker containers could wrap apps written in any language, as long as they ran on Linux. **Tristan Handy: So Docker is like infrastructure as code—a primitive that enables the whole concept?** **Solomon Hykes:** Yes! And because we wanted ubiquity, we avoided pushing too many opinions. We let developers build on top of it in many different ways. That’s what helped Docker become a de facto standard. **Tristan Handy: How fragmented is the Linux world under the hood? Did you have to do much abstraction work?** **Solomon Hykes:** We were lucky. The Linux kernel is extremely stable and consistent. But everything above it—distros, package managers, tooling—was chaotic. That chaos created the opportunity for Docker to provide a consistent experience. **Tristan Handy: Were there any drawbacks? Like “Docker sprawl” the way VMware saw VM sprawl?** **Solomon Hykes:** Definitely. With power comes chaos. Teams would run dozens of Docker containers, each configured differently. Docker doesn’t enforce opinions—by design. **Tristan Handy: And what happened when you left Docker in 2018?** **Solomon Hykes:** I took time off, became a full-time dad. But I also realized how many unsolved problems remained. Especially around CI/CD pipelines and software delivery—what we now call the software factory. That led me to start Dagger. **Tristan Handy: So Dagger is like “containers for pipelines”?** **Solomon Hykes:** Yes. Just as Docker standardized app deployment, Dagger aims to standardize and containerize software delivery. CI/CD pipelines today are often duct-taped together with YAML and bash scripts. We’re bringing consistency and modularity to that space. **Tristan Handy: Will there be a “Daggerfile” like there’s a Dockerfile?** **Solomon Hykes:** Sort of. But this time, we’re opinionated. Dagger is narrowly focused on CI/CD. That lets us provide APIs, SDKs, and a deeper abstraction stack. We give platform engineers a DAG-based system to define repeatable, containerized steps. **Tristan Handy: And what’s the role of AI and agents in all this?** **Solomon Hykes:** Great question. We didn’t plan for it, but our community showed us the way. People started building AI agents that run in Dagger pipelines—automating things like writing tests, submitting PRs, and optimizing builds. That blew our minds. Agents blur the line between development and delivery. They need programmable environments. Dagger is becoming an ideal platform for that. ## Chapters **01:30 – Early Days: From France to dotCloud** Solomon shares how his early programming experience and startup journey led to the creation of dotCloud. **04:00 – The PaaS Struggle and Birth of Docker** The team pivots from platform-as-a-service to focusing on the container engine itself—what would become Docker. **07:00 – What Is a Container, Really?** Solomon explains containerization vs. virtualization in plain terms and why it changed the game for developers. **11:00 – The Developer Experience That Won the World** The magic of fast, lightweight Docker containers—and how that first “wow” moment felt. **14:00 – Building a Ubiquitous Standard** Why Docker stayed narrow by design, resisting feature bloat to maximize compatibility. **18:00 – DevOps Before DevOps** How Docker avoided language tribalism and achieved mass developer adoption by choosing Go and CLI-first tooling. **21:00 – Complexity and Container Sprawl** Docker made infrastructure easy—but created new operational challenges at scale. **24:30 – Why CI/CD Pipelines Are Still Broken** Solomon outlines the gap Docker never got to fix: modern software delivery remains brittle and ad hoc. **27:00 – Enter Dagger: DevOps for the Modern Age** How Solomon’s new company is treating pipelines as composable software, not brittle scripts. **30:00 – Building an OS for the Software Factory** Dagger helps platform teams manage the complexity of software delivery with reusable, testable components. **33:00 – Agent-Native Workflows: A Surprise Use Case** AI agents begin using Dagger to reason about pipelines, generate code, and submit pull requests autonomously. **37:00 – Reimagining the Dev Loop with AI** Why the boundary between development and CI/CD is collapsing—and how Dagger fits the agent-powered future. **41:00 – Scaling Trust in Delivery** Tristan and Solomon reflect on how developer tooling evolves and what a stable, fast delivery layer enables. **45:00 – Final Thoughts: What’s Next for DevOps** The conversation closes with predictions on intelligent automation, composability, and the future of platform engineering. --- --- title: "How to move from manual to test-driven analytics" description: "Explore how dbt helps data teams adopt test-driven analytics with integrated testing, CI workflows, and data quality enforcement." url: "https://www.getdbt.com/blog/test-driven-analytics-workflow" date: "2025-06-20" authors: ["Joey Gault"] categories: ["Pulse"] --- # How to move from manual to test-driven analytics The foundation for test-driven analytics lies in adopting a structured approach to data development. The Analytics Development Lifecycle (ADLC) provides this framework by adapting proven software engineering practices for analytics work. Unlike traditional approaches where testing happens as an afterthought, the ADLC integrates testing as a core component of every development cycle. In the ADLC, testing serves as a quality gate that validates assumptions about data and analytics code before changes reach production. This proactive approach identifies issues early in the development process, preventing expensive rework and system downtime. The testing phase encompasses three distinct but complementary activities: writing tests for every data asset, running tests before merging changes to production, and continuously monitoring production data for anomalies. This systematic approach addresses a common problem in analytics organizations: the tendency to skip testing due to time pressure or competing priorities. When testing becomes integrated into the standard workflow rather than an optional add-on, teams naturally develop better habits around data quality assurance. ## The three pillars of analytics testing Effective test-driven analytics relies on three types of tests, each serving a specific purpose in ensuring data quality and system reliability. Understanding when and how to use each type is crucial for building a comprehensive testing strategy. Unit tests form the first pillar, focusing on validating the logic within individual data models and transformations. These tests work with small sets of static input data to verify that SQL logic produces expected results. Unit tests are particularly valuable for custom business logic, edge cases, and high-criticality models where defects could have widespread impact. In dbt, unit tests can be defined alongside SQL models and executed on demand, providing rapid feedback during development. For example, a unit test might verify that an email validation routine correctly handles malformed addresses, missing domain components, or invalid top-level domains. By testing these scenarios with controlled inputs, developers can ensure their logic works correctly before applying it to production data volumes. Data tests represent the second pillar, validating that transformations work correctly against actual data. These tests verify data freshness, model soundness, and transformation accuracy. They typically begin with basic assumptions about unique identifiers, non-null fields, and acceptable value ranges, then expand to more sophisticated domain-specific validations. Data tests catch issues that unit tests might miss, such as unexpected data distributions or violations of business rules that only become apparent with real data. Integration tests complete the testing framework by ensuring that changes work correctly within the broader system context. While unit tests focus on individual components and data tests examine single datasets, integration tests validate end-to-end functionality. In analytics environments, integration tests are particularly important for validating reusable packages and ensuring that changes don't break downstream dependencies. ## Implementing testing workflows The transition to test-driven analytics requires establishing clear workflows that define when tests run and who is responsible for creating and maintaining them. The most effective approach integrates testing into existing development processes rather than treating it as a separate activity. Developers working on analytics code should run unit and data tests locally before submitting changes. This immediate feedback loop helps catch issues early and reduces the time spent debugging problems later. When ready to deploy changes, developers should create pull requests that automatically trigger comprehensive test suites against non-production data in isolated environments. dbt supports this workflow by automatically running tests when pull requests are opened or updated. Test results appear both in the dbt dashboard and directly on the pull request page, making it easy for reviewers to assess the impact of proposed changes. Failed tests should block merges to production, ensuring that quality gates are enforced consistently. This automated approach addresses one of the biggest challenges in establishing testing culture: ensuring that tests actually run. When testing is manual and optional, it's easy for teams to skip it under pressure. Automated testing removes this temptation by making test execution a prerequisite for code deployment. ## Building a testing culture Technical implementation alone isn't sufficient for successful test-driven analytics. Organizations must also cultivate a culture that values testing and holds team members accountable for maintaining quality standards. This cultural shift often represents the most challenging aspect of the transition. Leadership plays a crucial role in establishing testing expectations. When code reviews consistently check for adequate test coverage and pull requests are rejected for insufficient testing, team members quickly understand that testing is a priority. Setting specific test coverage targets (typically 70-80% of analytics code should have associated tests) provides concrete goals for teams to work toward. The scope of individual changes also impacts testing effectiveness. Large, complex changes are inherently harder to test thoroughly than small, focused modifications. Training teams to break work into smaller, well-defined pieces makes testing more manageable and reduces the likelihood of introducing defects. This approach aligns with the ADLC's emphasis on frequent, iterative development cycles. Monitoring test coverage over time helps ensure that testing practices improve rather than degrade. dbt provides built-in reporting on test coverage, showing the percentage of models with defined tests. Regular review of these metrics helps identify areas where additional testing might be needed and tracks progress toward coverage goals. ## Managing test reliability One of the most significant threats to testing culture is the presence of unreliable or "flaky" tests. These tests fail intermittently due to environmental conditions, timing issues, or poorly written logic. When teams encounter frequent false positives, they begin to ignore test failures, undermining the entire testing system. Addressing flaky tests requires a systematic approach. Teams should investigate the root causes of intermittent failures and either fix the underlying issues or remove unreliable tests from the suite. It's better to have fewer, reliable tests than many tests that generate noise and confusion. Test maintenance should be treated as an ongoing responsibility rather than a one-time activity. As business requirements evolve and data sources change, tests need to be updated to reflect new realities. Establishing clear ownership for test maintenance ensures that this work doesn't fall through the cracks. ## Measuring success and continuous improvement The transition to test-driven analytics should be measured not just by the number of tests written, but by the impact on data quality and team productivity. Key metrics include the time to detect data quality issues, the frequency of production incidents, and the speed of development cycles. Teams often find that initial investment in testing pays dividends over time. While writing tests requires upfront effort, the reduction in debugging time, production incidents, and rework typically more than compensates for this investment. Additionally, well-tested systems are easier to modify and extend, enabling faster development of new features and capabilities. Regular retrospectives can help teams identify areas where testing practices can be improved. Common topics include test execution speed, coverage gaps, and the effectiveness of different test types. These discussions help teams refine their approach and adapt to changing requirements. ## The path forward Moving from manual to test-driven analytics represents a significant maturation in how organizations approach data quality. While the technical aspects of implementing testing frameworks are important, the cultural and process changes are equally critical for success. Organizations that successfully make this transition typically see improvements in data reliability, development velocity, and team confidence. They're able to make changes to their analytics systems with greater assurance and respond more quickly to business requirements. Most importantly, they build trust with stakeholders who depend on accurate, timely data for decision making. The journey toward test-driven analytics is iterative. Teams should start with basic testing practices and gradually expand their coverage and sophistication over time. dbt provides the technical foundation for this evolution, but success ultimately depends on organizational commitment to quality and continuous improvement. As the analytics engineering discipline continues to mature, test-driven development practices will become increasingly standard. Organizations that invest in these capabilities now will be better positioned to scale their analytics operations and deliver reliable insights to support business growth. ## Test-driven analytics FAQs **What is test-driven development? ** Test-driven development in analytics is a structured approach where testing is integrated as a core component of every development cycle, rather than being treated as an afterthought. It involves writing tests for every data asset, running tests before merging changes to production, and continuously monitoring production data for anomalies. This proactive approach identifies issues early in the development process, preventing expensive rework and system downtime by validating assumptions about data and analytics code before changes reach production. --- --- title: "The role of data governance frameworks in modern organizations" description: "Explore how governance frameworks help large organizations manage data quality, security, and compliance at scale." url: "https://www.getdbt.com/blog/data-governance-framework" date: "2025-06-19" authors: ["Joey Gault"] categories: ["Pulse"] --- # The role of data governance frameworks in modern organizations A data governance framework represents a comprehensive system of policies, processes, and standards that organizations use to manage their data assets throughout their entire lifecycle. The [Data Governance Institute](https://datagovernance.com/) defines data governance as "a system of decision rights and accountabilities for information-related processes, executed according to agreed-upon models, which describe who can take what actions with what information, and when, under what circumstances, and using what methods." While this definition may seem complex, every data governance program fundamentally serves three core objectives: establishing company-wide rules for data collection, storage, and usage; monitoring the global data estate to ensure compliance with governance standards; and resolving data-related issues while supporting end users across the organization. These objectives become increasingly critical as organizations scale. Smaller companies might manage data governance through informal processes and ad-hoc communications, but enterprise organizations require systematic approaches to handle their distributed and complex data assets effectively. ## The four pillars of enterprise data governance Effective data governance frameworks rest on four fundamental pillars that work together to ensure comprehensive data management. The first pillar, [data quality and trust](https://www.getdbt.com/product/build-trust-in-data-and-data-teams), focuses on maintaining accuracy, completeness, and consistency across all organizational data assets, regardless of size or distribution. This pillar ensures that decision-makers can rely on the information they're using to drive business outcomes. [Data stewardship](https://madison-schott.medium.com/what-is-data-stewardship-606594820e3d) forms the second pillar, recognizing that quality doesn't emerge spontaneously. This involves creating clear roles and responsibilities for individuals who manage, monitor, and ensure data quality within the organization. Data stewards serve as the front line of governance programs, defining and documenting data assets, promoting effective data sharing, and acting as liaisons between teams to resolve issues. They also ensure that governance policies are implemented correctly across all organizational touchpoints. The third pillar encompasses [data protection and compliance](https://www.getdbt.com/security), which has become increasingly complex in today's regulatory environment. This pillar includes security measures to prevent unauthorized access, privacy protections for sensitive and personal data, and processes to ensure compliance with applicable regulations and contractual requirements. With regulations like GDPR affecting billions of people worldwide and industry-specific requirements like FINRA and HIPAA governing financial and healthcare sectors, this pillar has become non-negotiable for enterprise organizations. [Data management visibility](https://www.getdbt.com/product/dbt-catalog) rounds out the four pillars, encompassing the processes and procedures for effectively storing, accessing, and manipulating data. This includes metadata management, data lifecycle management, and data integration: essentially how data is structured, stored, and linked across different systems and databases throughout the organization. ## When organizations need enterprise data governance As data systems and teams scale, informal practices lead to divergent logic, local workarounds, and unclear ownership. A control‑plane approach brings order to this growth by standardizing development, surfacing lineage and trust signals, and coordinating collaboration across many teams and projects so people can ship and use trusted data faster. Left unchecked, this divergence produces silos and fragmented definitions. Teams reinvent models, duplicate effort, and make breaking changes without shared contracts, eroding confidence in downstream insights. At the same time, many enterprises reach a maturity point where auditors, customers, or regulators require demonstrable controls and traceability—moving governance from “nice to have” to non‑negotiable. These conditions are strong signals to formalize governance: multiple domains contributing to shared datasets, frequent cross‑team dependencies, unclear data owners, increasing incident or rework rates, and emerging compliance obligations. Formalizing at this stage reduces cost and risk by making quality, ownership, and change management explicit within everyday analytics workflows. In short, the combination of scale, complexity, and accountability needs makes an enterprise governance framework essential—not as a separate bureaucracy, but as a control plane that embeds standards, visibility, and guardrails into how analytics work actually gets done. ## Implementing enterprise data governance frameworks Enterprise data governance frameworks are sophisticated platforms designed to streamline and automate the complex process of managing, organizing, and protecting data across large, potentially globally distributed organizations. While it's possible to develop governance frameworks from scratch, the "why build when you can buy" principle strongly applies here. Starting with pre-built frameworks based on proven principles significantly reduces the time required to design and implement new programs. Modern data governance tools can dramatically accelerate framework implementation. Platforms like dbt Cloud enable organizations to create data control planes that manage data uniformly within standardized frameworks. This approach allows everyone in the organization (from data engineers to business analysts to executive leaders) to work efficiently with trusted data in scalable, cost-effective ways. An enterprise-quality framework standardizes teams on terminology and concepts most important to the organization while building collaborative bridges across the full enterprise. This standardization enables business, technical, and compliance stakeholders to communicate effectively, exchanging data-driven information and ideas seamlessly. The result is an empowered organization where every person can extract value from data assets while managing cost and complexity. ## Selecting the right framework [Enterprise data governance](https://www.getdbt.com/blog/what-is-enterprise-data-governance) frameworks are sophisticated platforms designed to streamline and automate the complex process of managing, organizing, and protecting data across large, potentially globally distributed organizations. While it's possible to develop governance frameworks from scratch, the "why build when you can buy" principle strongly applies here. Starting with pre-built frameworks based on proven principles significantly reduces the time required to design and implement new programs. Modern data governance tools can dramatically accelerate framework implementation. Platforms like [dbt](https://www.getdbt.com/product/what-is-dbt) enable organizations to create data control planes that manage data uniformly within standardized frameworks. This approach allows everyone in the organization (from data engineers to business analysts to executive leaders) to work efficiently with trusted data in scalable, cost-effective ways. An [enterprise-quality framework](https://www.getdbt.com/blog/enterprise-data-governance-strategy-elements) standardizes teams on terminology and concepts most important to the organization while building collaborative bridges across the full enterprise. This standardization enables business, technical, and compliance stakeholders to communicate effectively, exchanging data-driven information and ideas seamlessly. The result is an empowered organization where every person can extract value from data assets while managing cost and complexity. ## Selecting the right framework Choosing an appropriate enterprise data governance framework requires careful consideration of several critical components. Mature, enterprise-ready platforms should include data stewardship capabilities, quality control mechanisms, cataloging features, lineage tracking, security controls, compliance management, and data visualization tools. However, not every solution offers all these capabilities in a single, integrated platform. When evaluating frameworks, organizations should prioritize solutions that manage data complexity in modular, scalable, repeatable, and governed ways directly within their data platforms rather than scattered across multiple business intelligence and technical platforms. Vendor-agnostic solutions that integrate seamlessly with major data cloud platforms like [Snowflake](https://www.getdbt.com/data-platforms/snowflake), [Databricks](https://www.getdbt.com/data-platforms/databricks), and [BigQuery](https://www.getdbt.com/data-platforms/bigquery) provide the flexibility needed for diverse enterprise environments. Effective frameworks should utilize centralized, reusable models that foster collaboration, reduce duplication, and ensure consistent data definitions across teams. Robust audit logging and access control features are essential for safeguarding data integrity, while support for software development best practices (including portability, CI/CD, observability, and documentation) enables the creation of production-grade analytics pipelines that scale with organizational workloads. The framework should also deliver accessible, easy-to-understand data models that integrate with BI tools, LLMs, and APIs, ensuring stakeholders have accurate data when and where they need it. dbt exemplifies these requirements, providing a scalable, enterprise-ready platform that standardizes data transformation processes, increases data quality and transparency through lineage tracking, and automates documentation while providing comprehensive visibility across organizational data assets. ## The strategic impact of data governance frameworks The right data governance framework instills confidence in organizational data, enabling everyone to make accurate, informed decisions regardless of their role. Enterprise organizations benefit from [enhanced data quality and consistency](https://www.getdbt.com/product/build-trust-in-data-and-data-teams), [reduced data management costs](https://www.getdbt.com/product/cost-optimization), and [accelerated insights from trusted data sources](https://www.getdbt.com/product/analyst). When frameworks automate data flow traceability and process transparency, organizations can focus on optimizing operations, improving performance, and achieving strategic goals while minimizing data security and privacy risks. [Modern data governance frameworks also play crucial roles in AI and machine learning initiatives](https://www.getdbt.com/blog/data-governance-frameworks-ai), where data quality directly impacts model performance and outcomes. As organizations increasingly rely on AI-driven insights, governance frameworks ensure that training data meets the high standards necessary for reliable, unbiased results. This becomes particularly important as AI-specific regulations emerge and organizations need to demonstrate responsible AI practices. Furthermore, governance frameworks enable data democratization by providing the guardrails necessary for safe, self-service data access. Rather than creating bottlenecks through centralized control, well-designed frameworks empower users across the organization to access and analyze data confidently, knowing that appropriate protections and quality controls are in place. ## Conclusion Data governance frameworks serve as the backbone of modern data-driven organizations, providing the structure necessary to transform data from a potential liability into a strategic asset. These frameworks establish the policies, processes, and technologies needed to ensure data quality, security, and usability while enabling organizations to scale their data operations effectively. As data volumes continue to grow and regulatory requirements become more complex, the role of governance frameworks becomes increasingly critical. Organizations that invest in robust governance frameworks position themselves to extract maximum value from their data assets while managing risks and maintaining compliance. The framework becomes not just a protective measure, but a competitive advantage that enables faster, more confident decision-making across the entire organization. The evolution toward AI-driven business processes only amplifies the importance of strong governance frameworks. As organizations continue to integrate artificial intelligence and machine learning into their operations, the quality and governance of underlying data becomes paramount to success. In this context, data governance frameworks represent not just operational necessity, but strategic imperative for organizations seeking to thrive in an increasingly data-driven world. ## Data governance framework FAQs **What is a data governance framework?** A governance program defines how data is produced, changed, and trusted. The dbt platform acts as the data control plane for analytics, embedding governance guardrails into daily work while integrating with your broader security and compliance stack. **What are the pillars of a data governance framework? ** Quality and trust, federated stewardship (Mesh), protection and auditability (RBAC, audit logs, environment permissions), and data visibility (data lineage, semantic metrics). **How to implement a data governance framework? ** Implementation involves using sophisticated platforms designed to streamline and automate the complex process of managing, organizing, and protecting data across large organizations. Rather than building from scratch, organizations should leverage pre-built frameworks based on proven principles to reduce implementation time. Modern data governance tools can dramatically accelerate framework implementation by creating data control planes that manage data uniformly within standardized frameworks. The framework should standardize teams on terminology and concepts while building collaborative bridges across the enterprise, enabling business, technical, and compliance stakeholders to communicate effectively and work with trusted data in scalable, cost-effective ways. --- --- title: "Data science vs. data engineering: Defining the difference" description: "Data science and data engineering serve different roles—see how dbt connects them through shared standards, testing, and lineage." url: "https://www.getdbt.com/blog/data-science-vs-data-engineering" date: "2025-06-19" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data science vs. data engineering: Defining the difference Data flows through pipelines, warehouses, and real-time systems. But volume doesn’t create value. Value requires a trustworthy data infrastructure combined with analytical real-time decisions, predictive modeling, and operational intelligence. Data engineering and data science bridge this gap, transforming raw data into reliable analytics using orchestrated pipelines and sophisticated modeling methods. Data engineering involves the flow, transformation, and storage of data between systems. Meanwhile, data science uses statistical and computational methods to identify patterns, conduct hypothesis tests, and inform decision-making. Teams can fall out of alignment, experience delays, and unreliable results when the boundaries between the two are not clear. In this article, we’ll investigate the similarities and differences between data science and data engineering, and look at tools that facilitate aligning both workflows. We’ll also see how these two disciplines collaborate to unlock the full potential of data. ## Technical foundations and core workflows Although both fields help extract value from data, one focuses on infrastructure and flow, while the other concentrates on analysis and insight. Let’s look at their workflows and technical details. ### Data engineering Data engineering is the practice of developing systems that move, organize, prepare, and store raw data. It supports a range of data types, including tables in databases, JSON from an API, or text and images. Data moves through a series of engineering steps to ensure it is clean and ready for downstream use. Let’s have a look at these steps: - **Data ingestion**: The first step involves pulling raw data from numerous data sources, including internal databases, third-party APIs, log files, and event logs. This process can operate in real-time or in batches, depending on the use case. - **Preprocessing**: After ingestion, data engineers cleanse and standardize data using automated pipelines. Common preprocessing operations include standard [data quality checks](https://www.getdbt.com/blog/data-quality-checks), such as formatting, deduplication, schema checks, and eliminating null values. This step ensures that downstream systems get quality data. - **Transformation**: After preprocessing, data pipelines reshape and enrich data. This includes joins, aggregations, filters, or calculations to prepare data sets for analysis and interpretation. Engineers commonly version and modularize transformations for reuse and traceability. - **Storage**: After transformation, data pipelines transfer data to the storage layer, storing data in systems such as data warehouses or lakehouses..The result is that data stakeholders can find and use high-quality, queryable datasets in their analytics, data applications, and AI solutions. ### Data science Data science utilizes processed data to analyze and perform tasks such as predictive modeling and analytics. Typical data science tasks are: - [**Exploratory data analysis (EDA)**](https://www.ibm.com/think/topics/exploratory-data-analysis): EDA involves examining distributions, correlations, and anomalies to gain a deeper understanding of the data. It recommends appropriate models or transformations and facilitates early detection of data quality problems. - **Feature engineering**: During the feature engineering process, raw data are converted into more interpretable inputs to enhance the accuracy of a predictive model. This can be done by generating ratios, binary flags, time lagging, or domain-specific indicators. - **Statistical analysis and testing**: Statistical techniques such as hypothesis testing and confidence interval estimation can be used to evaluate the nature of relationships and variability in data. This brings scientific rigor to analysis and aids data-driven conclusions. - **Visualization**: Results are frequently condensed into charts, dashboards, or visual narratives to convey the insights. Plotting libraries, such as [Matplotlib](https://matplotlib.org/), [Plotly](https://plotly.com/), and [Streamlit](https://streamlit.io/), make it easy to explore and share rich results. - [**Machine learning (ML)**](https://developers.google.com/machine-learning/crash-course): ML models learn directly based on data when pattern recognition and predictive automation are required. ML can provide scalable intelligence based on historical and streaming data, in the form of supervised learning, unsupervised discovery, or deep learning architectures. ## Tools and technologies The tools used by data science and data engineering are indicative of their respective and different objectives. Every domain has a specific tech stack tailored to its unique processes and priorities. Let’s find out in the table below: ## The challenges facing both data scientists and data engineers Data engineering and data science have distinct objectives, but they often encounter similar challenges due to architectural dependencies, evolving data sources, and organizational silos. Let’s have a look at some of the major problems these workflows encounter: **Schema drift**. Source systems may change unexpectedly by renaming fields, changing data types, or dropping columns. These changes may disrupt pipelines, undermine downstream models, and distort aggregated measures. **Distributed data inputs**. Integrating data from systems like CRM (customer relationship management) systems and payment platforms often introduces issues such as key mismatches, uneven update intervals, and format conflicts. Without proper entity resolution and event time alignment, data science models risk producing inaccurate or misleading outputs. **Latency limits**. Latency constraints create friction in various stages of data engineering and data science processes. In engineering, delays in ingestion, processing, or delivery of data may lead to downstream failures of dependent systems and delay real-time analytics. However, in data science, this means models can be trained on stale or incomplete data, limiting their capacity to make accurate and instant predictions. To address these challenges, both workflows must be designed for low-latency performance at scale. Data engineering workflows can use stream processing, incremental transformation logic, and event-driven orchestration to minimize lag. Workflows in data science can separate feature engineering, inference, and use continual learning techniques to adapt to changes in data without requiring retraining. **Lack of visibility**. When teams don’t track lineage and metadata, they lose visibility into how data flows through the pipeline. It becomes difficult to trace errors, understand dependencies, or explain how a number was calculated. Use end-to-end observability tools that track data lineage, monitor pipelines and manage metadata to identify failures. Use clear, documented datasets and track how features are created to make models easier to explain and reproduce. **Siloed definitions and fragmented ownership**. Without centralized documentation or modeling layers, inconsistencies grow as different teams independently define metrics. Data contracts, common semantic layers, and domain-based ownership (as found in data mesh architectures) can facilitate the standardization of business logic across workflows. ## How dbt brings data science and data engineering together Many of these problems occur because scientists, engineers, and other data stakeholders use different platforms and tools to work with data: - Scientists and engineers use different platforms for data storage and different languages to transform raw data into datasets that yield actionable insights - Analysts don’t have one place where they can find high-quality datasets or verify the quality of data - Every team uses different names and calculations for critical business metrics, resulting in confusion and a loss of trust in data dbt is a [data control plane](https://www.getdbt.com/blog/data-control-plane-why) for data that’s natively operable across various data and cloud platforms. It provides a common platform that all data stakeholders can use to collaborate on data products as part of a mature [analytics development lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle). Using dbt, data science and data engineering can bridge the chasms that have historically kept them divided. The result is higher-quality datasets delivered in less time and with measurable business impact. ### In data engineering Data platforms are growing, and data engineering requires flexible tools. dbt addresses this by bringing version control, testing, and documentation to the transformation layer. - **Scalable SQL pipelines: **dbt allows data engineers to write [SQL transformation logic ](https://www.getdbt.com/blog/data-transformation)as discrete, reusable models, making maintenance easier. These models also yield a clean DAG, ensuring proper build dependencies in complex pipelines. - **Fast development and testing**. With [dbt Fusion](https://www.getdbt.com/product/fusion), data engineers can debug and test analytics code locally, without setting up a remote data warehouse or checking code into source control. This reduces development lag time and shortens time to release for data product changes. **** - **Operational excellence in data pipelines: **dbt projects are stored in [Git](https://git-scm.com/) and integrated with CI/CD [pipelines](https://docs.getdbt.com/best-practices), allowing them to be peer-reviewed and deployed in a version-controlled manner. It supports a robust testing framework with built-in support for [generic tests](https://www.getdbt.com/blog/data-engineering) (distinct, not null, relationships) and custom assertion tests to identify data issues before they reach production. - **Pipeline reliability and idempotency: **Pipelines can be re-run without creating duplicates. dbt provides deduplication with stable keys and audit logs to track and troubleshoot runs. This promotes the use of separate [development ](https://docs.getdbt.com/best-practices/best-practice-workflows)and production targets, allowing for safe testing and staging environments prior to production release. - **Governance and project structure: **dbt supports industry [best practices](https://docs.getdbt.com/best-practices/how-we-structure/2-staging), such as dividing include staging, intermediate, and mart structures, which make data pipelines more modular and understandable. Engineers can centralize metadata and decouple model logic from schema details by declaring tables as sources in dbt. Updating the source config ensures all dependent models stay aligned when upstream changes occur. ### In data science Tools like dbt are becoming increasingly important as data engineering and data science merge. Data scientists can obtain clean, well-documented, and production-ready data thanks to dbt's automated testing, version control, and clear data lineage. - **Reliable data**: dbt allows automating the [data testing](https://www.getdbt.com/blog/data-quality-checks), ensuring that datasets are accurate and trusted prior to analysis. This enables data scientists to concentrate on modeling instead of tedious data cleaning. - **Feature engineering collaboration**: dbt enables rapidly building, testing, and sharing canonical [feature sets](https://www.getdbt.com/blog/snowflake-dbt), such as user activity aggregations. This ensures a consistent understanding of data definitions used for modeling. - **Data lineage visibility**: dbt automatically creates a DAG [lineage graph](https://www.getdbt.com/blog/what-is-data-lineage) through reference and source functions, making it easy to see how data was transformed and where it originated. This visibility contributes to the explainability, impact analysis, and auditability of model inputs. ## Conclusion When data engineering and data science are closely connected, it creates a strong, end-to-end data ecosystem. The common denominator between data science and data engineering is the need for common standards, observability, and scalability. [dbt](https://www.getdbt.com/product/dbt) supports these needs by centralizing transformation logic, enforcing testing and documentation, and aligning workflows across the data lifecycle. The result is better, higher-quality data products delivered in less time. --- --- title: "How to design a scalable cloud data architecture" description: "Build a cloud data architecture that scales—modular, centralized, and designed for change. Here’s how to do it right." url: "https://www.getdbt.com/blog/designing-scalable-cloud-architecture" date: "2025-06-18" authors: ["Daniel Poppy"] categories: ["Learn"] --- # How to design a scalable cloud data architecture A cloud data architecture defines how data is ingested, stored, transformed, and accessed within a cloud environment. Unlike traditional on-premises solutions, cloud architectures offer flexibility, scalability, and cost efficiency by using managed services that eliminate much of the infrastructure maintenance burden. The modern data stack typically consists of several layers: data sources (applications, databases, APIs), data ingestion tools, data storage solutions, transformation processes, and data serving tools like dashboards and applications. Each component must be designed with growth in mind to create a truly scalable architecture. For most organizations, the transition to cloud data architecture represents a shift from fixed capacity systems to environments that can expand or contract based on actual needs. This flexibility allows companies to start small and scale up as data volumes and business requirements grow, without massive upfront investments. The real power of cloud data architecture comes from its ability to separate storage from compute resources. This separation means you can store massive amounts of data cost-effectively while only paying for the processing power you need when you actually run analyses or transformations. ## Key components of a scalable cloud data architecture ### Centralized data management A centralized approach to data management is foundational to scalability. Rather than building siloed data solutions, organizations should establish a single source of truth where business logic resides. Centralization helps avoid conflicting reports and metrics that arise when different teams build their own data solutions. It also makes maintenance more efficient, as updates to business logic can be made in one place rather than across multiple systems. ### Modular design patterns Breaking down your data architecture into modular components enables independent scaling and easier maintenance. A well-designed modular system allows you to swap out one data source or analytics tool without rebuilding the entire architecture. The key to effective modularity is defining clear interfaces between components, including standard data formats, naming conventions, and access patterns. ### Scalable data storage solutions Cloud data platforms provide virtually unlimited storage capacity, but designing for efficient access patterns is crucial for performance and cost management. Use appropriate clustering and partitioning strategies, implement tiered storage, and choose columnar formats suited to analytical workloads. Match your storage strategy to query patterns and budget constraints. ### Transformation layer design The transformation layer is where raw data becomes business-ready information. Tools like [dbt](https://www.getdbt.com/product/dbt) help teams define transformations in code, supporting incremental loading, modular models, and version-controlled development. Plan for scalability early by using patterns like incremental processing, which reduce processing time while supporting growing data volumes. ## Implementation approaches for different business needs ### For startups and small teams Startups need to move quickly while setting the foundation for future growth. Focus on implementing a simplified architecture covering core business metrics. Choose managed services that require minimal operational overhead, and use a modular approach that allows for incremental expansion. A fintech startup might begin with a straightforward architecture using pre-built connectors to extract data from their application database and third-party services into a cloud warehouse, with a transformation tool handling the business logic. This approach requires minimal maintenance while providing a solid foundation that can scale. Small teams should avoid the temptation to build complex, custom data infrastructure. The goal should be to establish reliable data flows that answer critical business questions while laying groundwork that won't need to be completely rebuilt as the company grows. The most successful small team implementations tend to focus on getting clean, reliable data to business users quickly, rather than building perfect systems. Pick battles carefully and solve immediate needs while keeping an eye on future scalability. ### For midsize organizations Midsize organizations typically need to balance existing investments with scalability needs. Develop standardized data modeling practices across teams and implement data testing and documentation to ensure quality as complexity grows. Consider hybrid approaches that leverage both existing systems and cloud-native services. A midsize retailer with an existing on-premises data warehouse might implement a hybrid architecture. They maintain their operational data store on-premises while gradually migrating analytical workloads to a cloud data warehouse. They use a consistent transformation layer that works across both environments, enabling a gradual transition without disrupting business users. Midsize companies often face the challenge of legacy systems that can't simply be replaced overnight. The key is to design integration points that allow for gradual migration rather than risky "big bang" approaches that try to change everything at once. Focus on building a consistent data model that can span both old and new environments, making the transition invisible to end users while progressively moving workloads to more scalable platforms. ### For enterprise organizations Enterprises face unique challenges with scale, compliance, and organizational complexity. Design multi-tenant architectures that support different business units while implementing robust governance and security controls. Build for international operations with region-specific considerations where needed. A global industrial company might design a multi-region data architecture that keeps certain data within specific geographic boundaries for compliance reasons. They could implement a federated approach where common data models are defined centrally but deployed regionally, ensuring consistency across regions while respecting local regulations. Enterprise implementations require careful attention to organizational dynamics as well as technical considerations. Success often depends on balancing centralized standards with the flexibility needed by different business units. The most effective enterprise data architectures create clear boundaries between shared, governed data assets and areas where teams can innovate independently. This balanced approach prevents both the chaos of complete decentralization and the bottlenecks of overly rigid central control. ## Best practices for scalable transformation Data transformation is where raw data becomes valuable business information. A code-based approach to transformation brings software engineering best practices to data work, making complex transformation processes more manageable and reliable. Structure your transformation projects with scalability in mind by implementing a clear organization that separates staging, intermediate, and final models. Use subdirectories to group related models by business domain, and create reusable components for common transformation patterns. **** A healthcare analytics team might structure their transformation project with staging models that clean raw data from each source system, intermediate models that implement business logic for specific domains, and final presentation-layer models that serve specific use cases. This structure allows them to onboard new data sources without disrupting existing workflows. For large datasets, incremental processing drastically reduces computation time by transforming only new or changed records rather than reprocessing entire datasets. This approach can reduce daily processing time from hours to minutes, ensuring business users have up-to-date data without excessive processing costs. As your data models grow in complexity, comprehensive testing and documentation become essential. Write tests that validate key assumptions about your data, document model relationships, and create data dictionaries for business users. ## Integration with the modern data stack A truly scalable architecture integrates smoothly with other components of the modern data stack. Connect your transformation layer with reliable data ingestion by implementing quality checks at ingestion points and designing for idempotent processing that can safely re-process data if needed. A retail analytics team might integrate their transformation workflow with a pipeline that loads data from point-of-sale systems, inventory management, and e-commerce platforms. They could implement quality checks that verify data completeness before triggering transformation jobs, ensuring that incomplete data loads don't produce misleading analytics. Consider how transformed data will be consumed by designing models that align with how business users think about the business. Implement appropriate security controls and create semantic layers that abstract complexity from end users. The most effective integrations create clear handoffs between different components of the data stack. This means establishing conventions for when data moves from one stage to another and implementing appropriate checks to ensure data quality at each transition point. ## Future-proofing your data architecture Technology evolves rapidly, and today's scalable architecture must accommodate tomorrow's requirements. Plan for structured, semi-structured, and unstructured data by implementing flexible transformation patterns that can adapt to new data forms. An insurance company initially focused on structured policy and claims data might design their architecture to also accommodate semi-structured data from mobile apps and IoT devices, and unstructured data like claims documents. This forward-looking design allows them to quickly incorporate new data sources as business needs evolve. For large organizations, centralized approaches may eventually hit scaling limits. Consider domain-oriented ownership of data products, implement self-service capabilities for domain experts, and establish clear interfaces between domains. The most future-proof architectures focus on principles rather than specific technologies. By establishing clear data contracts, ownership boundaries, and quality standards, you create a foundation that can adapt to new tools and techniques as they emerge. ## Conclusion Building a scalable cloud data architecture requires thoughtful design across multiple dimensions: storage, processing, transformation, and serving. By implementing centralized yet modular approaches, choosing appropriate technologies, and following software engineering best practices, organizations can create data ecosystems that grow with their business. The key to success lies in starting with a solid foundation of well-organized, transformed data, which provides a consistent layer regardless of the underlying data platform. This approach ensures that as your data needs grow—whether in volume, complexity, or business coverage—your architecture can scale to meet those needs without requiring major redesigns. By treating data as a product and applying software engineering best practices like testing, version control, and documentation, teams can deliver higher quality data that business users can trust, even as the organization and its data needs grow exponentially. ## Data cloud architecture FAQs **What is the data architecture of the cloud?** Cloud data architecture defines how data is ingested, stored, transformed, and accessed within a cloud environment. It typically consists of several layers: data sources (applications, databases, APIs), data ingestion tools, data storage solutions (warehouses, lakes, or lake houses), data transformation processes, and data serving components like BI tools and dashboards. A well-designed cloud data architecture offers flexibility, scalability, and cost efficiency by leveraging managed services that reduce infrastructure maintenance burdens. **What are the 4 types of cloud architecture?** The four main types of cloud architecture are: 1. **Public Cloud**: Infrastructure owned and operated by third-party providers (AWS, Azure, Google Cloud) that deliver services over the internet. 2. **Private Cloud**: Infrastructure dedicated solely to a single organization, either hosted on-premises or by a third-party provider. 3. **Hybrid Cloud**: A combination of public and private cloud environments that operate independently but are connected to allow data and application portability. 4. **Multi-Cloud**: Using cloud services from multiple providers simultaneously, often to leverage specific strengths of different platforms or avoid vendor lock-in. **What are the 5 pillars of cloud architecture?** The 5 pillars of cloud architecture are: 1. **Operational Excellence**: Running and monitoring systems to deliver business value and continually improving processes and procedures. 2. **Security**: Protecting information, systems, and assets while delivering business value. 3. **Reliability**: Ensuring a system can recover from infrastructure or service disruptions and dynamically acquire computing resources to meet demand. 4. **Performance Efficiency**: Using computing resources efficiently and maintaining that efficiency as demand changes and technologies evolve. 5. **Cost Optimization**: Running systems to deliver business value at the lowest price point, avoiding unnecessary costs. **What is cloud architecture?** Cloud architecture is the way various technology components are organized to build a cloud computing system. It encompasses all the cloud components—software, hardware, virtual machines, networks, storage systems—and how they interact with each other to create a functional cloud environment. Cloud architecture defines the structure that supports the delivery of computing services over the internet, providing on-demand access to resources without direct active management by the user. It enables organizations to scale resources as needed, pay only for what they use, and access advanced technologies without having to build and maintain physical infrastructure. --- --- title: "dbt Labs on dbt: Why this analyst is obsessed with dbt Insights" description: "Why dbt Insights is this analyst’s go-to for faster, in-context data discovery—no more tab hopping required." url: "https://www.getdbt.com/blog/why-this-analyst-is-obsessed-with-dbt-insights" date: "2025-06-17" authors: ["Rachael Gilbert"] categories: ["Learn"] --- # dbt Labs on dbt: Why this analyst is obsessed with dbt Insights I’m Rachael, one of the data analysts at dbt. While I’m always excited when we ship new features, I’m not always on the frontlines of the folks dogfooding them. With our fantastic team of analytics engineers, it’s rare for me to get in all the weeds of goodies like [snapshot improvements](https://www.getdbt.com/blog/whats-new-in-dbt-cloud-january-2025) or [the latest on iceberg](https://www.getdbt.com/blog/iceberg-give-it-a-rest). As an analyst, I’m often coming in **after** the models are built, usually navigating what's available and stitching together what I need **to leverage for business decisions**. But even with the best models, discovering the right data and insights is time consuming. Gaining confidence that I have the specific data I need for a business decision involves the following: - Identifying which **models** I need - Identifying which **fields** in the models I need - Identifying which **values** in those fields I need Finding the answer has always been a multi-faceted workflow across multiple tabs at every company I’ve worked at: ping ponging between combing documentation in one place (a catalog, a Confluence Wiki, comments in past reporting, etc.) and writing queries to dig in deeper somewhere else (Hex, Mode, Dbeaver, etc.). After a while, you take the multi-tab existence for granted. …Until now. ## Could we be talking data discovery and querying, all in one place?! Indeed, we are! One of the new features dbt has built is dbt **Insights**: a place to actually query data from within dbt's governed workspace, without tab hopping. It goes beyond a standard SQL interface with previews and charts to tie your data exploration to all the metadata, semantic layer logic, and documentation, in the same environment as your engineering team. Likely, you already have a place you are comfortable doing your data querying, but here are a couple dbt Insights tricks that have been leveling up my workflows: - ➡️ **From dbt Catalog to queries, instantly**: I can hop directly from a model’s documentation to live querying that model in the blink of an eye…no switching tabs and typing out needed! [Watch video](https://youtu.be/YigqN-el_b4) - ➡️ **Everything in-line**: Likewise, I can hop back to dbt Catalog just as easily, or even explore it side by side (my personal preference) to ensure I'm querying trusted, governed models with full context in view: metadata, documentation, and lineage. ![Everything in-line](https://cdn.sanity.io/images/wl0ndo6t/main/db4cc33e58b3e8a76d9b23b225f3e4590357f746-3004x1568.png) ## The dbt Insights features that have transformed my workflow My dbt Insights love letter is about to get sappier: - ➡️ **AI efficiency boosts:** As a staff analyst, I’m often jumping into new business contexts. When I’m working with an unfamiliar area and not as sure on where to begin, I can chat with our context-aware AI, dbt Copilot, to get some ideas of where to start looking across both dbt Catalog and saved queries. This has proven much faster and customized than my prior catalog search process of throwing out search terms and hoping to strike gold: [Watch video](https://youtu.be/TDunJTsEIho) - ➡️ **Pre-populated dbt Semantic Layer syntax**: Another big time saver is that I no longer have to look up or try to remember the nuances of how to query the dbt Semantic Layer, which frankly I’m not in everyday and don’t have memorized. I can just look up metrics, like MAU, and dbt Insights will auto-populate a good starting point for the syntax automatically. Yes please. [Watch video](https://youtu.be/s1XqGiIAqvw) - ➡️ **Understanding what SQL the Semantic Layer syntax actually represents**: I always have trouble remembering what I need to do to see what is behind a semantic model. The product team knows I have mentioned this before! 😂 With dbt Insights, I finally get my pretty out-of-the-box view of what SQL my metric queries are exactly compiling to. No more guessing or bugging my team. [Watch video](https://youtu.be/-TuQwU3Hrq8) - ➡️ **Some extra Fusion goodness coming next**: You may have heard of [all the Fusion perks](https://www.getdbt.com/blog/building-the-next-gen-dbt-engine) we’re adding to the dbt engine, like CTE previews. Plans are in the works to enrich dbt Insights with some of this as well. ## Try dbt Insights: a better way to query and explore trusted data Above, I highlighted some of the reasons dbt Insights has won me over for discovering and querying data (beyond just the fact that I happen to work at the company that makes it). If you’re already working in dbt, dbt Catalog and Insights can help you move faster and feel more confident, without ever stepping outside the trusted, governed layer your data team has built. It makes it easier for more analysts to explore models, run queries, and contribute without relying on messy workflows or one-off help. If you’ve been the “go-to” dbt person on your team, this may help open up dbt to the rest of your team to move faster, together. This is just my personal take, and I’m always looking to learn more about other analysts’ workflows! Ping me on the [Community Slack](https://www.getdbt.com/community/join-the-community) with your thoughts if you give it a spin, or let me know what your querying tips and tricks are if you’ve found productivity elsewhere. --- --- title: "Best practices for data engineering" description: "Learn how to build scalable, resilient data pipelines with modern data engineering tools, workflows, and best practices." url: "https://www.getdbt.com/blog/data-engineering" date: "2025-06-17" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Best practices for data engineering Every day, over [400 million terabytes](https://www.techbusinessnews.com.au/blog/402-74-million-terrabytes-of-data-is-created-every-day/) of data are generated across digital platforms, powering everything from mobile apps and e-commerce to internal dashboards and reporting systems. But raw data, by itself, isn’t immediately useful. It needs to be cleaned, structured, and shaped into something teams can work with. That’s where data engineering comes in. By designing systems that collect and transform scattered inputs, data engineering helps teams produce trusted, analysis-ready datasets. These systems ensure that the data is consistent, accurate, and easily accessible, making them a core foundation for analytics, machine learning, and day-to-day operations. Let’s examine the practices that experienced data engineers rely on to build scalable, resilient, and easy-to-maintain pipelines as data and business needs evolve. ## What is data engineering and why does it matter? A predictive model is only as good as its input data, and teams can't make smart decisions if dashboards show conflicting or incomplete information. Data engineering is the behind-the-scenes work that shapes raw data into something teams can trust. This means designing systems that collect, process, and organize information so it’s ready to power everything from daily reports to machine learning models. Unlike data scientists (who develop models), data analysts (who examine trends), and DevOps engineers (who manage software infrastructures), data engineers focus on creating and maintaining [data pipelines](https://www.getdbt.com/blog/data-pipelines-snowflake-dbt) that deliver reliable, relevant, and up-to-date information. This involves: - Pulling data from APIs, event streams, and internal databases - Validating and cleaning inputs to remove errors or inconsistencies - Storing data in platforms built for scale and speed - Creating access layers so data can flow easily into tools and workflows When done well, data engineering turns scattered inputs into a single source of truth, ready for whatever the business needs next. ### Why data engineering is critical in modern business Each day, clicks, swipes, transactions, and sensor readings generate massive amounts of raw data. Without structure, that information is difficult to use. Data engineers step in to organize and build pipelines that move and transform data into a format that teams can actually work with. Their work supports everything from real-time dashboards to machine learning models. As data volumes increase, they help ensure the systems behind them scale smoothly and deliver reliable results. ### The cost of poor data engineering When data engineering breaks down, everything downstream is affected.The impact is hard to miss: - Teams end up with disconnected, overlapping datasets - Confidence in reporting and analytics drops - Pipelines fail more often and take longer to fix - Decisions slow down as teams question what data to trust - Projects miss deadlines or go unused - Engineers spend more time patching problems than improving the system ## Top strategies for modern data engineering Building reliable pipelines starts with clear priorities and creating systems that deliver accurate, trusted data. This requires you to develop workflows that are easy to scale, simple to maintain, and flexible enough to handle change. Here’s how you can do this: ### Choose the right tech stack A modern data stack typically spans four key layers. Understanding what each layer does and what to consider when selecting tools can help you make better choices from the start. Here's a quick breakdown: **Tip:** If you're just starting out, prioritize tools that are easy to learn and quick to implement. ### Assess the business need Before writing a single line of code, ask questions like: - What is the data meant to accomplish? - Who’s going to use it? - How often does it need to be updated? - What decisions will it influence? Analysts, product managers, and business leads often have the clearest view of how data is used on a day-to-day basis. Their input can help clarify the pipeline’s goals, identify where flexibility is acceptable, and establish how success should be measured. Using a method such as the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) ensures that this indispensable stakeholder input is integrated into data engineering from day one. It’s also important to agree on which metrics are most valuable to your team. For some, that may mean tracking daily revenue or user growth. For others, it could be product usage or retention. Aligning early helps ensure all the stakeholders are working toward the same outcomes. ### Design modular and scalable data pipelines Designing pipelines around smaller, focused tasks makes it easier to manage growing data volume and changing business requirements. Each task can be tested and maintained independently, which helps teams troubleshoot issues more quickly and improve system performance over time. Instead of repeating logic across different models, many teams use SQL snippets, helper functions, or shared [dbt macros](https://learn.getdbt.com/courses/jinja-macros-and-packages). This maintains development consistency while reducing unnecessary rework. Another key consideration is how data moves through your system. Some workflows require real-time updates, while others only need data to be refreshed on a schedule. The choice often comes down to what the data supports and how quickly it needs to be available. The table below outlines how batch and streaming options compare: Once you've chosen a processing method, reliability comes next. Pipelines should run multiple times without causing duplicates or corruption. Use stable identifiers like primary keys to deduplicate records. Most importantly, track each run using job states or audit tables so you know exactly what’s been processed and where issues might occur. ### Automate with CI/CD for data pipelines [CI/CD (Continuous Integration and Continuous Delivery)](https://www.redhat.com/en/topics/devops/what-is-ci-cd) provides a reliable way to test, validate, and deploy changes automatically. It provides teams with the structure to identify issues early, before they reach users or disrupt key workflows. Automating deployment also saves time and reduces manual effort. Once new code is committed, the CI pipeline can handle the rest: running tests, checking for schema changes, and promoting updates through development, staging, and production. ### Version your data and transformations As your data pipelines evolve, [version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) plays a critical role in ensuring consistency, auditability, and collaboration across teams. Start by storing all transformation logic, configuration, and scheduling metadata in Git. This gives you a clear history of what changed, when, and why. It also allows safe rollbacks when needed. But versioning shouldn't stop at code. Your raw data, intermediate outputs, and model artefacts can (and should) be versioned too, especially when accuracy and reproducibility matter. Lastly, use consistent naming, commit messages, and tagging for both code and data assets. Wherever possible, version datasets and logic together to keep them in sync. ### Use reliable orchestration and scheduling Orchestration plays a key role in managing when and how pipeline tasks run. It keeps processes on schedule, enforces the right order of operations, and helps recover from failures. #### Pick the right orchestrator Choosing the right orchestration tool depends on the size of your workflows and your team’s experience. Look for a system that makes it easy to build modular pipelines, track task progress, and connect smoothly with the rest of your data stack. You can use tools like [Apache Airflow](https://airflow.apache.org/use-cases/etl_analytics/), [Prefect](https://www.prefect.io/data-engineering), and [dbt](https://docs.getdbt.com/docs/deploy/deployment-tools). Some teams also build in-house orchestration tools tailored to their specific infrastructure and workflow needs. The downside is that these can be time-consuming to maintain and scale. #### Handle dependencies and failures Build pipelines that recover automatically. Set up retries to automatically handle temporary issues and configure alerts to notify teams of critical failures. Triggers or sensors can help track when a file arrives or an upstream job finishes, so the pipeline knows when to move forward. Keep tasks modular so failures can be isolated and resolved without disrupting the entire pipeline. ### Build robust data models Reliable analytics starts with a solid data model. Without a clear structure, even the most advanced tools can produce slow queries, inconsistent results, and user confusion. The table below describes approaches that can fit your needs: ### Monitor, log, and alert effectively Active monitoring helps catch failures early and ensures pipelines stay reliable as they scale. #### Set up observability Track key metrics such as run time, error counts, throughput, data freshness, and missed SLAs. These signals reveal bottlenecks and help prevent silent failures. #### Centralized logging Aggregate logs in one place and include relevant metadata, such as task IDs or timestamps. This makes it easier to trace issues and debug quickly. #### Use alerts strategically Define alert thresholds for critical issues like missing data, schema drift, or job failures. Route alerts to the relevant teams with sufficient context to enable prompt action. ### Data processing and transformation Data is just noise without structure and business logic. Transformation is what gives it meaning, making it ready for reporting, automation, and operational use. Transformation usually comes down to two options: - [ETL (Extract, Transform, and Load)](https://www.getdbt.com/blog/extract-transform-load) - [ELT (Extract, Load, and Transform)](https://www.getdbt.com/blog/extract-load-transform) They sound similar, but the difference in approach can have a big impact on performance, flexibility, and overall pipeline design. Here’s a quick side-by-side to help make the distinction clear: ### Secure data and maintain compliance Keeping data secure is critical as it moves through your pipeline. A combination of access control, encryption, and regulatory safeguards helps minimize risk and protect both users and the business. [Role-based access control (RBAC)](https://www.youtube.com/watch?v=2h-ak4P5n3o) restricts what each user can see or modify based on their responsibilities. Separating permissions across environments adds an extra layer of protection and makes auditing more straightforward. Encryption protects sensitive data both at rest and in transit. To maintain its effectiveness, encryption keys should be rotated regularly and stored securely, ideally in a secrets manager instead of being hardcoded in scripts. Many teams also need to meet legal and regulatory requirements. Depending on your industry and region, this may include the [GDPR](https://gdpr.eu/what-is-gdpr/), [HIPAA](https://www.cdc.gov/phlp/php/resources/health-insurance-portability-and-accountability-act-of-1996-hipaa.html), and [CCPA](https://oag.ca.gov/privacy/ccpa). Techniques such as masking, pseudonymization, and consent tracking help support these standards and mitigate risk throughout the pipeline. ### Document everything Focus your documentation on the details that help teams work efficiently and avoid confusion: - Explain each pipeline step so anyone can understand how data is processed. - List data sources and outputs to show where data comes from and where it is delivered. - Keep documentation up to date whenever code or logic changes. - Use data catalogs to show lineage and clarify which team owns each dataset. - Add metadata to support cross-team collaboration and help non-experts understand the data. Create shared channels and review routines to enable teams to raise issues, ask questions, and stay aligned. ## Elevate your data engineering with dbt With a strong foundation like modular pipelines, tested transformations, and clear documentation in place, your next step is making these practices scalable and repeatable. That requires a system that can act as a single [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) for all the data across your organization. dbt helps teams manage transformations with the same discipline used in software development. Using dbt, you can: - Collaborate across distributed teams to create [data transformation models](https://docs.getdbt.com/docs/build/models) in SQL and Python in a flexible, cross-platform fashion - Author [tests](https://docs.getdbt.com/docs/build/data-tests) to verify data transformation logic and [documentation](https://docs.getdbt.com/docs/build/documentation) to help data consumers understand and use the data - Ship changes safely to production with [an automated CI/CD process](https://docs.getdbt.com/docs/deploy/continuous-integration) - Enable stakeholders [to find and understand data quickly](https://www.getdbt.com/product/dbt-catalog) One of the easiest ways to begin is by modeling a single domain, writing basic tests, and tying those changes into a continuous integration workflow. As these practices take hold, they can gradually be applied to other parts of the stack. Over time, this approach helps improve pipeline quality while reducing the time spent fixing issues after deployment. [Get started with dbt](https://www.getdbt.com/signup) and turn well-structured pipelines into reliable, production-grade systems. ## Data engineer FAQs **What does a data engineer do?** A data engineer designs, builds, and maintains systems that collect, process, and organize data so it’s accurate, consistent, and accessible. Typical responsibilities include ingesting data from APIs, event streams, and databases; validating and cleaning inputs; modeling data for analytics; orchestrating pipelines; enforcing quality with tests; implementing CI/CD and version control; monitoring and alerting; securing data with RBAC and encryption; and documenting lineage, ownership, and usage. **How do data engineers design data pipelines to ingest, transform, and serve data while maintaining data quality and reliability?** - Start with the business need: clarify purpose, users, refresh cadence, and decisions supported. - Choose a modular architecture across ingestion, storage, transformation, and orchestration layers. - Pick the right processing mode: batch for scheduled/bulk jobs; streaming for low-latency use cases. - Build for idempotency: use stable keys, deduplicate, and track runs with audit tables or job states. - Enforce quality with tests (schema, freshness, uniqueness), validation, and contract checks. - Automate with CI/CD to run tests, detect schema changes, and promote changes through environments. - Add observability: metrics (latency, throughput, failures, SLAs), centralized logs, and actionable alerts. - Secure and govern: RBAC, encryption in transit/at rest, secrets management, and compliance controls. - Document pipelines, lineage, and ownership; keep docs updated with code changes. **What are the key differences between ETL and ELT pipelines, and when should each approach be used?** - ETL (Extract, Transform, Load): Transform data before loading it into the warehouse, often in an external compute layer. Best when you need tight control on data quality upfront, have storage constraints, or operate under rigid schemas and compliance requirements with moderate data volumes. - ELT (Extract, Load, Transform): Load raw data into the warehouse first, then transform in place. Best for cloud-native stacks that can leverage warehouse compute, analytics-heavy workloads, large-scale data, and teams that value faster iteration and simpler architectures. **** **How do workflow management systems like Apache Airflow use DAGs to orchestrate and monitor data pipelines?** Airflow represents pipelines as Directed Acyclic Graphs (DAGs), where nodes are tasks and edges define dependencies. The scheduler executes tasks in the right order, supports time-based schedules or event-driven sensors, and handles retries with backoff. Operators integrate with common services, while task isolation and modular design make failures easier to contain. The UI provides run history, logs, SLA tracking, and alerting to monitor performance and diagnose issues. Similar concepts apply in other orchestrators (e.g., Prefect, dbt) even if they use different abstractions. --- --- title: "Understanding MCP: The missing glue between governed data and AI agents" description: "The dbt MCP Server gives AI agents governed access to metrics, lineage, and models—bridging the gap between LLMs and trusted data." url: "https://www.getdbt.com/blog/mcp" date: "2025-06-16" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Understanding MCP: The missing glue between governed data and AI agents Imagine asking your best data analyst to make a decision—without giving them access to the source data or its documentation. That’s how most AI agents operate today: disconnected from the governed, trusted data that real business decisions rely on. **The result? Misleading outputs and missed insights—not because the model lacks intelligence, but because it lacks context.** Enter the **Model Context Protocol (MCP)**: a new open standard designed to close this gap. In this article, we’ll explore how dbt’s integration with MCP helps AI agents access structured, governed data and metadata—bringing transparency, accuracy, and trust to AI-powered workflows. ## What is Model Context Protocol (MCP)? **_The Model Context Protocol (MCP) is an open-source standard that enables AI agents to access structured, trusted data in real-time without requiring custom connections._** MCP was released by Anthropic in November 2024. It replaces fragile, one-off data integrations with a shared method for connecting AI to any data source, such as databases, APIs, or business tools. ![AI Agents With vs. Without MCP. Without MCP: hardcoded, brittle integrations per tool; no awareness of data models or relationships; hallucinated metrics or inconsistent queries; heavy human oversight for validation. With MCP: one standard interface for all data sources; discovers model relationships, lineage, and metadata; uses governed metrics from dbt’s semantic layer; autonomous, accurate, and auditable responses.](https://cdn.sanity.io/images/wl0ndo6t/main/40d80aefc8f4ca9c798dd217624c8dbe0e9a11dc-1536x1024.jpg) MCP solves model isolation from siloed data and legacy systems. It creates a standardized protocol with discovery, semantic, and execution tools. These tools connect AI agents to data sources. By using these tools, AI agents pull live data from databases, APIs, and business applications via MCP servers. These tools enable AI agents to explore data models, understand relationships, and perform analytics tasks. This provides accurate context without relying on hardcoded connections. Given its utility, MCP has garnered strong industry support, with major players such as [Google](https://google.com/), [Microsoft](https://microsoft.com/), and [OpenAI](https://openai.com/) backing the protocol. ## Why AI needs access to governed data Without structured, validated data, AI agents generate unreliable outputs or require labor-intensive manual oversight. Common challenges include: - **Custom pipelines: **These are hand-coded data flows, such as Python scripts or ETL jobs, that connect systems. They’re fragile; a small schema change can break them, causing silent errors or inaccurate results for AI systems. - **LLM hallucinations: **When AI models don’t have access to clearly defined, trusted metrics, they guess. For example, _"monthly revenue"_ might be interpreted in different ways, sometimes including tax, discounts, or refunds. This creates inconsistent and unreliable answers. - **Operational friction: **In many teams, AI agents depend on data specialists to explain models, clarify terms, or write SQL queries. This creates delays in delivering insights and increases the risk of errors, since every step requires manual involvement. **_This raises a critical challenge: How can AI agents reliably understand and use the rich, governed context of enterprise data?_** Moreover, agents need to do so without relying on brittle, custom connections. That’s where a MCP server, like [the dbt MCP Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server), functions as the missing glue between governed data and AI. The dbt MCP Server gives AI agents direct, structured access to your dbt [project](https://docs.getdbt.com/docs/build/projects). A set of tools built into the dbt MCP Server enables a Large Language Model (LLM) to gain a deep understanding of your data models, documentation, and underlying metadata. With these capabilities, MCP addresses the above challenges. It gives AI agents direct access to [dbt’s Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl), lineage, and documentation. This enables AI agents to understand the structure and business logic of your data without requiring manual intervention. AI agents gain structured visibility into how your business logic is defined and implemented. For example, an LLM can query monthly revenue by region, using the exact definition in dbt’s revenue metric, ensuring that AI outputs align with organizational truth. ## Inside the dbt MCP Server architecture The dbt MCP Server translates LLM requests into dbt-native operations and returns structured context. Its architecture rests on three pillars: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/fefe9ded24b20607382c1b40b87539dc59c72a2c-1600x803.jpg) ### Discovery tools This layer includes methods like `get_all_models`, `get_model_details`, `get_model_parents`, and `get_mart_models`. These tools enable autonomous metadata ingestion, allowing AI agents to automatically explore project structure, relationships, and column-level lineage. This provides an entirely contextual understanding of your data without manual effort. ### Semantic layer The Semantic Layer provides governed analytics. It uses tools such as `list_metrics`, [`get_dimensions`](https://docs.getdbt.com/docs/build/dimensions), and `query_metrics`. AI systems can query validated business metrics and dimensions, like monthly revenue_._ They query directly from the source of truth defined in dbt’s Semantic Layer. This eliminates guesswork and reduces hallucinations. ### Execution engine The Execution engine drives operational orchestration. It uses commands like dbt run, dbt test, dbt compile, and dbt build. These tools enable seamless automation, allowing AI agents to run dbt pipeline runs, tests, or compile operations directly through conversational interfaces. ### How it works These pillars activate in a coordinated workflow when requests arrive: #### Request ingestion & tool selection A user or AI agent asks a question, like "Run tests for customer models," in an MCP-enabled client. The request structures into a standardized protocol message. This message routes to the dbt MCP Server. #### Context-aware routing The server’s protocol layer acts as an intelligent dispatcher. It analyzes the request to select the optimal tool based on intent: - For metric-focused questions, such as _"_Show last quarter’s revenue trends_,"_ it activates the Semantic Layer. It then uses the `query_metrics` tool to ensure that the results are governed. - For dependency investigations, such as _"_What’s upstream of the customer table_?"_, it uses Discovery Tools like `get_model_parents` to map lineage. - For operational commands, including _"_Test payment models_,"_ it triggers the Execution Engine to run [dbt test](https://docs.getdbt.com/reference/commands/test). #### Instant context hydration The server loads pre-compiled project knowledge, such as lineage graphs and metric definitions, from memory before execution. This eliminates slow data warehouse queries. #### Permission -bound execution Every action runs within strict guardrails: - Metric queries run through [dbt Cloud APIs](https://docs.getdbt.com/dbt-cloud/api-v2#/) for certified results. - Commands like dbt run operate in isolated sandboxes. - Read-only access is enforced by default. #### Real-time result streaming The dbt MCP Server streams outcomes incrementally. Instead of static reports, it delivers insights: - Pipeline test logs appear line-by-line, showing failures during execution. - Metric results populate column-by-column, helping detect patterns before queries finish. - Lineage maps expand node-by-node, visually tracing dependencies from core models to upstream sources. ## Real-world use cases powered by the dbt MCP Server These use cases demonstrate how the MCP Server transforms dbt from a static data tool into a dynamic control plane for AI-driven operations. ### Self-service analytics Non-technical users can now explore dbt projects conversationally using natural language interfaces. For example, a business stakeholder can ask, _“_What models do we have for customer behavior_?”_ They instantly retrieve a list of relevant models, their descriptions, and lineage context. This works through Discovery Tools like `get_all_models` and `get_model_details`. These tools make dbt metadata accessible and explorable without SQL knowledge. ### AI agent workflows Agentic systems can autonomously discover and map model relationships using tools such as get_model_parents and get_mart_models. These tools help understand how models connect to each other and which ones are used for reporting and business decisions. This allows them to dynamically reason through dbt projects and understand dependencies. Based on this context, AI agents can take informed actions, such as investigating upstream changes or mapping column-level lineage, without human prompting. ### AI-accelerated dbt development Beyond simple queries, the dbt MCP Server enables advanced AI-driven development workflows. An AI agent can proactively identify and refactor dbt models, ensuring models align with best practices. For example, an agent can reference intermediate data instead of raw staging tables. The agent can automatically generate new models or update existing ones, taking into account the project's dependencies. If the agent runs a model and detects errors, it can automatically analyze the issues, such as join logic failures. It can then suggest or apply fixes directly. This capability transforms dbt into a dynamic partner in your [analytics development lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle). ### Trusted data analysis LLMs often hallucinate when disconnected from ground truth. The MCP Server enables AI systems to query canonical metrics and dimensions like monthly revenue or active users_._ The server uses dbt's Semantic Layer tools to ensure that outputs strictly reflect definitions within your dbt project, enhancing reliability and trust. ### Accelerated dbt operations AI agents can now drive operational workflows using familiar dbt [command line interface (CLI) commands](https://docs.getdbt.com/reference/dbt-commands): run, test, compile, and build. The MCP Server acts as a secure bridge between prompt-driven interactions and backend operations. For example, an agent can be triggered to _“run daily models”_ or _“compile staging layer,”_ running safely within scoped environments. That means they run in isolated, permission-controlled contexts that protect production systems. These use cases show how the dbt MCP Server transforms dbt into a dynamic control plane for AI-driven operations. It ensures: - **Governance at scale:** AI agents operate using standardized semantic definitions, maintaining consistency across reporting workflows. - **Operational efficiency:** Automating builds, tests, and queries removes the need for manual intervention, saving time and reducing human error. - **Collaboration:** Both humans and AI systems interact with the same trusted data context, eliminating silos and improving decision-making. Together, this multi-pronged value proposition illustrates why the MCP Server is indeed the “missing glue” connecting governed data with AI agents in a trusted and scalable manner. ## Limitations and security in using the dbt MCP Server As an experimental release, the dbt MCP Server presents certain limitations and requires careful security considerations for effective and safe deployment. ### Tool selection requires human oversight During testing, AI models occasionally cycle through unnecessary tools. They might call `get_all_models` before narrowing to `get_mart_models`, or pick incorrect tools for requests. While users can correct this with specific prompts, it currently undermines fully autonomous operation. This behavior stems from the protocol's immaturity. However, it will improve through community feedback and updates. ### Uncontrolled SQL execution poses risks Freeform SQL in the MCP Server supports flexible exploration. However, it bypasses dbt’s semantic safeguards. Uncontrolled use can lead to incorrect results and costly warehouse queries. dbt Labs recommends limiting SQL tools to sandbox environments. For trusted, production-grade insights, always prefer Semantic Layer tools like `query_metrics` that enforce certified logic. ### Prototype before scaling The current dbt MCP Server is experimental. Begin with limited, low-risk use cases, such as metadata queries, sandboxed metric checks, or lineage tool testing in development environments. Only scale after proving value. **_This isn’t production-ready—prototype first, then expand._** ### Strict permission scoping is essential Running tools like dbt run carry inherent risks. Mitigate them through: - **Command disabling:** Block risky operations via flags like `DISABLE_DBT_CLI=true`. - **Ephemeral environments:** Isolate runs in containers that self-destruct post-run. - **Granular role-based access control (RBAC):** Restrict dbt tokens to specific projects and environments. - **Read-first policy:** Enforce read-only access until safety is proven. Always audit tool call logs to detect prompt injection or misuse. ## How the dbt MCP Server shapes the future of AI-driven data access The dbt MCP Server is the missing glue that will fundamentally change how AI interacts with your data. It paves the way for AI to drive both business intelligence and data engineering directly. This means AI will deeply understand your dbt projects, transforming how data is used and built. Data teams will increasingly focus on creating this rich, governed context that feeds into the server. This shifts their role towards making data highly understandable and trustworthy for AI. Giving AI systems access to structured context through dbt establishes a solid and lasting part of the modern data stack. A key promise of the dbt MCP Server is enabling safe and reliable data access. It offers built-in security features and access controls. Ultimately, dbt is becoming the central data control plane for AI. This ensures AI agents can access structured data reliably, making every AI-driven insight consistent and aligned with organizational truth. **** **_Ready to Get Started with the dbt MCP Server?_** The dbt MCP Server is an experimental release that shapes the future of AI-driven data access. It’s available now on GitHub for prototyping AI agents that will benefit from a deep understanding of how your data is structured and used. To get started: - Explore the repository on [GitHub](https://github.com/dbt-labs/dbt-mcp) for installation instructions. - Connect your dbt project to begin experimenting with AI-powered data workflows. - Join the conversation in the [dbt Community Slack's](https://www.getdbt.com/community/join-the-community) #tools-dbt-mcp channel to share your findings. --- --- title: "The dbt Fusion engine shows up at 2025 Databricks Data + AI Summit" description: "Fusion for Databricks: faster development, lower costs, and smarter AI-ready data workflows with dbt" url: "https://www.getdbt.com/blog/the-dbt-fusion-engine-shows-up-at-2025-databricks-data-ai-summit" date: "2025-06-13" authors: ["Jeff Mills"] categories: ["Partnerships"] --- # The dbt Fusion engine shows up at 2025 Databricks Data + AI Summit ![The dbt Labs team at 2025 Databricks Data + AI Summit](https://cdn.sanity.io/images/wl0ndo6t/main/00a291df97bb2263b8caa3889e73f3629ab424d2-4032x3024.jpg) This week, we joined over 20,000 data professionals at the 2025 Databricks Data + AI Conference in San Francisco. We wanted to show how dbt Labs and Databricks are working together to improve the quality, trust, and cost of data and AI development. It was a week packed with product announcements, breakout sessions, booth conversations and a whole bunch of fun. ![Databricks Data + AI Summit attendees watching a dbt Fusion engine demo at the dbt Labs booth.](https://cdn.sanity.io/images/wl0ndo6t/main/f24d3987595fdd21ef683a57b12b37b8ebda5f19-5712x4284.jpg) Just in time for Databricks Summit, [we announced beta availability](https://www.getdbt.com/blog/databricks-users-get-ready-to-experience-the-new-dbt-fusion-engine) of the dbt Fusion engine for anyone using Databricks. We’re continuing to bring the speed and cost savings of our new engine to Databricks. For our Databricks users, this means: - **Significant performance gains**: Project parsing and compilation is as much as 30x faster than with dbt Core, dramatically reducing turnaround time from code change to insight. - **Developer intelligence**: Fusion’s native SQL comprehension provides real-time feedback and error checking as you write SQL, catching mistakes and providing suggestions instantly—without executing queries against your warehouse. The engine understands Databricks SQL and provides all the same benefits to this ecosystem. - **Cost control**: [Fusion’s](https://www.getdbt.com/product/fusion) state-aware orchestration understands what’s changed in your data and code, so only necessary models run. This eliminates redundant workloads and helps organizations reduce cloud compute spending by 10% or more. - **Better governance (coming later this year)**: Precise column-level lineage and richer metadata will soon provide enhanced data governance, making it easier to manage risk, enforce policies, and lay the foundation for safe AI applications downstream. ## Databricks’ new capabilities ![Lakebase launch partner logos](https://cdn.sanity.io/images/wl0ndo6t/main/8fd9be69909e07e4949d94414e50ed42a0cdb9d7-800x447.jpg) This week Databricks launched Lakebase, their separation of storage and compute to bring together OLTP data structures for fast processing, while supporting new and emerging AI workloads.Its bones are built on Postgres, and we have a long history of supporting Postgres. dbt Labs is proud to be a launch partner for this innovation. Check out [the announcement blog](https://www.databricks.com/blog/announcing-lakebase-public-preview) that includes a quote from our CPO, Ryan Segar. ## **Riot Games was a riot** ![Riot Games' Marco Garcia presenting at the 2025 Databricks Data + AI Summit](https://cdn.sanity.io/images/wl0ndo6t/main/78bd2192f5a97694c48c1e62ecf2204aff1c9c65-5712x4284.jpg) [Marco Garcia](https://www.linkedin.com/in/marcoantoniogarcia/) from Riot Games joined us this week in San Francisco to share how Riot Games is building reliable, scalable data products with dbt on top of Databricks. They turned to dbt to streamline their development, increase collaboration and leverage built-in testing to deliver trusted data to the business. [Check out Riot Games' presentation at Coalesce 2024](https://www.getdbt.com/resources/coalesce-on-demand/coalesce-2024-how-riot-games-is-building-player-first-gaming-experiences-with-databricks-and-dbt) to share how their data platform team paired with analytics engineering, machine learning, and insights teams to integrate Databricks Data Intelligence Platform and dbt Cloud to significantly mature its data capabilities. ## We also had a ton of fun ![dbt customers playing skee-ball at 2025 Databricks Data + AI Summit](https://cdn.sanity.io/images/wl0ndo6t/main/d27b6b32561fa486285181181eaf98d0b98a740f-4032x3024.jpg) On Tuesday night, we were joined by our partners Thoughtspot, Indicium, and Atlan along with 300 of our closest friends for a relaxed evening at Thriller Arcade. Well, it was relaxed unless you joined the Skee-Ball tournament. Congrats to the winners. Thank you to our sponsors for making a memorable night. I’ll always remember beating my boss, Clarke, at Skee-Ball. ![dbt customers playing basketball at 2025 Databricks Data + AI Summit](https://cdn.sanity.io/images/wl0ndo6t/main/2edccea244ab384e1ac043b25db4bedc4159c406-4032x3024.jpg) ## What’s next for dbt and Databricks? We can’t wait to see what you all take advantage of with Fusion at Databricks. Try it, get used to the blazing fast speed, and tell us what you want next. If you want to see how to deliver AI to your company - join us for our [webinar](https://www.getdbt.com/resources/webinars/empowering-data-analysts-showcase-series-part-one) on June 25 & 26 to learn more. And lastly, join us in October for [Coalesce](https://coalesce.getdbt.com/event/21662b38-2c17-4c10-9dd7-964fd652ab44/summary) where you can join thousands of enthusiastic data pros and leaders at the analytics engineering event of the year. --- --- title: "Risks of a poorly designed semantic layer — and how to avoid them" description: "A poorly designed semantic layer can slow performance, cause metric chaos, and bring governance headaches. Learn what to avoid." url: "https://www.getdbt.com/blog/semantic-layer-pitfalls" date: "2025-06-13" authors: ["Joey Gault"] categories: ["Pulse"] --- # Risks of a poorly designed semantic layer — and how to avoid them The promise of a [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) is simple: consistent metrics, streamlined governance, and easier access to data for business users. But like any layer of abstraction, its value hinges entirely on how well it’s designed and implemented. When built thoughtfully, a semantic layer can be a force multiplier for your analytics team. When built poorly, it can introduce new bottlenecks, increase complexity, and erode trust in data. In this article, we’ll explore the most common pitfalls of semantic layer design—and why data teams need to treat it like core infrastructure, not an afterthought. ## Performance degradation and scalability issues One of the most immediate impacts of a poorly designed semantic layer is performance degradation. When semantic models lack proper optimization, query response times can become unacceptably slow, particularly as data volumes and user bases grow. This often occurs when the underlying data models are not properly normalized or when the semantic layer attempts to join too many tables dynamically without considering the computational overhead. Poor caching strategies compound these performance issues. Without intelligent caching mechanisms that store frequently accessed metrics and pre-calculated results, every query forces the system to recalculate from raw data. This not only increases response times but also drives up compute costs significantly. Organizations may find themselves paying substantially more for warehouse resources while delivering a frustrating user experience. Scalability problems emerge when the semantic layer architecture cannot handle increasing numbers of concurrent users or growing data complexity. A poorly designed system may work adequately with a small team but fail catastrophically when rolled out organization-wide. This scalability ceiling often becomes apparent only after significant investment in implementation and user training. ## Inconsistent metric definitions and data quality issues Paradoxically, a poorly implemented semantic layer can actually worsen the metric consistency problems it was designed to solve. When business logic is incorrectly encoded in the semantic layer, these errors propagate across all downstream tools and reports. Unlike isolated errors in individual dashboards, semantic layer mistakes affect every consumer of that metric, amplifying the impact of any data quality issues. Inadequate metadata management creates confusion about metric definitions and calculations. When users cannot understand how metrics are calculated or what assumptions underlie the data, they lose confidence in the results. This lack of transparency can lead to shadow analytics, where teams revert to building their own calculations outside the semantic layer, defeating its primary purpose. Version control problems in semantic models can create additional inconsistencies. When changes to business logic are not properly managed, different versions of the same metric may exist simultaneously across various tools and reports. This creates the exact problem the semantic layer was meant to eliminate: multiple versions of truth within the organization. ## Governance and security vulnerabilities A poorly designed semantic layer can create significant [governance](https://www.getdbt.com/blog/data-governance) gaps. When access controls are not properly implemented, sensitive data may be exposed to unauthorized users. Unlike traditional database security models where access is controlled at the table level, semantic layers require more nuanced permission systems that understand business context and user roles. [Data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) becomes obscured in poorly implemented systems, making it difficult to trace how metrics are calculated or identify the source of data quality issues. When problems arise, data teams struggle to diagnose root causes, leading to longer resolution times and decreased trust in the data platform. Compliance requirements become harder to meet when the semantic layer lacks proper audit trails or cannot demonstrate how sensitive data is being accessed and used. Organizations in regulated industries may find themselves unable to satisfy regulatory requirements, creating legal and financial risks. ## Increased complexity and maintenance burden Rather than simplifying data access, a poorly designed semantic layer can add unnecessary complexity to the data stack. When the abstraction layer is overly complicated or poorly documented, it becomes a bottleneck rather than an enabler. Data teams may spend more time maintaining the semantic layer than they would have spent managing individual data marts and reports. The learning curve for poorly designed systems can be steep, requiring extensive training for both technical and business users. When the semantic layer interface is not intuitive or when error messages are unclear, user adoption suffers. Teams may abandon the semantic layer in favor of familiar but less optimal approaches. Maintenance overhead increases when the semantic layer is not well-integrated with existing data workflows. If the system requires separate tooling, different deployment processes, or specialized skills, it becomes an operational burden rather than a productivity enhancer. ## Integration and compatability challenges A poorly designed semantic layer may not integrate well with existing BI tools and analytics platforms. When integrations are incomplete or unreliable, users cannot access semantic layer metrics from their preferred tools, limiting adoption and value realization. This forces organizations to either change their tooling or maintain parallel systems, both of which are costly and inefficient. API limitations can restrict how external systems interact with the semantic layer. When APIs are poorly designed, slow, or unreliable, they become bottlenecks that limit the semantic layer's utility. Applications that depend on semantic layer data may experience timeouts, errors, or inconsistent responses. [Data freshness](https://docs.getdbt.com/reference/resource-properties/freshness) issues arise when the semantic layer cannot keep pace with underlying data updates. If metrics become stale or if there are delays in reflecting new data, users may lose confidence in the system's reliability. This is particularly problematic for operational use cases that require near real-time data. ## Business impact and user adoption problems Perhaps the most significant downside of a poorly designed semantic layer is its impact on business decision-making. When users cannot trust the data or when the system is too slow or complex to use effectively, they may make decisions based on incomplete or incorrect information. This can lead to poor strategic choices, missed opportunities, and operational inefficiencies. User adoption challenges emerge when the semantic layer fails to deliver on its promises. If business users find the system difficult to use or if they encounter frequent errors, they will likely revert to previous methods of accessing data. This not only wastes the investment in the semantic layer but also perpetuates the data silos and inconsistencies it was meant to address. Training and change management costs can escalate when the semantic layer is poorly designed. Organizations may need to invest heavily in user education and support, only to find that adoption remains low due to fundamental usability issues. ## Technical debt and long-term consequences A poorly implemented semantic layer can create significant technical debt that becomes increasingly expensive to address over time. When the initial design is flawed, making corrections often requires substantial rework of semantic models, metric definitions, and integration points. This technical debt can limit the organization's ability to adapt to changing business requirements or take advantage of new technologies. Migration challenges become more severe when organizations need to move away from a poorly designed semantic layer. The interconnected nature of semantic layer implementations means that changes can have far-reaching impacts across the data ecosystem. Organizations may find themselves locked into suboptimal solutions due to the cost and complexity of migration. ## Mitigation strategies and best practices Understanding these potential downsides highlights the importance of careful planning and design when implementing a semantic layer. Organizations should invest in proper requirements gathering, stakeholder alignment, and technical architecture before beginning implementation. Starting with a narrow, high-impact use case allows teams to validate their approach before scaling organization-wide. Regular performance monitoring and optimization should be built into the semantic layer operations from the beginning. This includes implementing proper caching strategies, monitoring query performance, and establishing clear escalation procedures for performance issues. Strong governance frameworks must be established early, including clear ownership models, change management processes, and security protocols. Documentation and training programs should be comprehensive and continuously updated to support user adoption and system maintenance. The key to avoiding these pitfalls lies in treating the semantic layer as a critical infrastructure component that requires the same level of planning, testing, and operational rigor as any other core system. When properly implemented, semantic layers deliver tremendous value, but the consequences of poor design can be severe and long-lasting. Data engineering leaders must carefully weigh these risks and invest appropriately in design, implementation, and ongoing operations to realize the full potential of semantic layer architectures. **** ## Semantic layer FAQs **What is a semantic layer?** A semantic layer is an abstraction layer that sits between raw data and business users, providing a unified view of data with consistent metric definitions and business logic. It translates complex database structures into business-friendly terms and ensures that metrics are calculated consistently across all downstream tools and reports. **Why use a semantic layer?** A semantic layer eliminates data silos and metric inconsistencies by providing a single source of truth for business definitions. It simplifies data access for business users, reduces the need for technical expertise to query data, and ensures that all reports and dashboards use the same calculations and business logic, leading to more reliable decision-making. **When does a BI truly become scalable? ** BI becomes truly scalable when it can handle increasing numbers of concurrent users and growing data complexity without performance degradation. This requires proper optimization of semantic models, intelligent caching strategies, and an architecture that can accommodate organization-wide rollouts while maintaining consistent performance and data quality across all users and use cases. --- --- title: "Snowflake data transformation architecture & performance" description: "How to design and optimize transformation workloads in Snowflake: materializations, cost control, governance, and modern tooling." url: "https://www.getdbt.com/blog/snowflake-data-transformation-architecture" date: "2025-06-12" authors: ["Joey Gault"] categories: ["Pulse"] --- # Snowflake data transformation architecture & performance [Snowflake's](https://www.snowflake.com/) unique architecture separates compute from storage, fundamentally changing how data transformation workloads should be designed. Unlike traditional data warehouses, Snowflake allows multiple compute clusters to access the same data simultaneously without contention. This architecture enables parallel processing of transformation workloads, but it also requires careful consideration of how transformations are structured and executed. The elastic nature of Snowflake compute means that transformation jobs can scale up or down based on workload demands. However, this flexibility comes with the responsibility of managing compute costs effectively. Transformation logic that works well on smaller datasets may become prohibitively expensive when scaled to production volumes without proper optimization. [Snowflake's columnar storage and automatic clustering capabilities](https://docs.snowflake.com/en/user-guide/tables-micro-partitions) can significantly impact transformation performance. Understanding how data is physically organized and accessed becomes crucial when designing transformation logic. Queries that leverage Snowflake's metadata and pruning capabilities will perform better than those that require full table scans, making the design of transformation models a critical performance consideration. ## Materialization strategies The choice of how to materialize transformed data in Snowflake directly impacts both performance and cost. Views offer the lowest storage cost but require computation on every query. Tables provide fast query performance but consume storage and require regular refreshes. Incremental models process only changed data, reducing compute costs for large datasets while maintaining reasonable query performance. [Snowflake's Dynamic Tables](https://docs.snowflake.com/en/user-guide/dynamic-tables-about) feature adds another materialization option that automatically manages refresh schedules and dependencies. This can simplify operational overhead but requires careful consideration of refresh frequency and resource allocation. The choice between traditional materialization approaches and Dynamic Tables depends on factors like data freshness requirements, query patterns, and operational complexity tolerance. Materialization decisions should align with downstream usage patterns. Frequently accessed data that powers critical dashboards may justify table materialization for performance, while exploratory datasets might be better served as views to minimize storage costs. The ability to change materialization strategies as requirements evolve provides flexibility but requires ongoing monitoring and optimization. ## Cost optimization approaches Snowflake's consumption-based pricing model makes cost optimization a continuous consideration in transformation design. Compute costs are driven by warehouse size and runtime, making efficient SQL and appropriate warehouse sizing critical factors. Transformation jobs that can complete quickly on smaller warehouses often cost less than those requiring larger warehouses for extended periods. The timing of transformation jobs affects costs through Snowflake's per-second billing model. Jobs that can be batched and run during off-peak hours may benefit from larger warehouses that complete work faster, while real-time transformations might require smaller, continuously running warehouses. Understanding these trade-offs helps optimize the total cost of transformation workloads. Storage costs in Snowflake are relatively low, but they accumulate over time, especially with frequent table refreshes that create multiple versions of data. Implementing appropriate data retention policies and leveraging Snowflake's time travel features judiciously helps manage storage costs. The choice between storing intermediate transformation results versus recomputing them on demand involves balancing storage costs against compute costs. ## Data quality and testing frameworkeworkss Ensuring data quality in Snowflake transformations requires systematic testing approaches that can scale with data volume and complexity. Traditional [data quality checks](https://www.getdbt.com/blog/data-quality-checks) may not be sufficient for the scale and speed of modern data pipelines. Implementing automated testing that validates data integrity, business rules, and transformation logic becomes essential for maintaining trust in transformed data. Snowflake's ability to process large datasets quickly enables comprehensive data quality testing that might be impractical on other platforms. However, the cost of running extensive tests must be balanced against the value they provide. Designing efficient test suites that provide maximum coverage with minimal compute overhead requires careful consideration of test design and execution strategies. The integration of data quality testing into transformation workflows affects both development velocity and operational reliability. Tests that run too frequently may slow development cycles, while insufficient testing may allow data quality issues to reach production. Finding the right balance requires understanding the specific quality requirements of each transformation and its downstream consumers. ## Governance and lineage tracking Data governance in Snowflake transformations extends beyond traditional access controls to include transformation logic documentation, lineage tracking, and change management. As transformation complexity grows, maintaining visibility into how data flows through the system becomes increasingly challenging. Implementing systematic approaches to governance helps maintain control and understanding of transformation processes. Lineage tracking becomes particularly important in Snowflake environments where data can be easily shared and accessed across different workloads. Understanding the dependencies between transformed datasets and their downstream consumers helps assess the impact of changes and ensures appropriate communication when modifications are necessary. This visibility is crucial for maintaining data trust and operational stability. Change management processes must account for Snowflake's ability to rapidly deploy and scale transformations. While this agility enables faster development cycles, it also increases the risk of unintended consequences from poorly tested changes. Implementing appropriate review processes and deployment controls helps balance development speed with operational stability. ## Integration with the broader data ecosystem Snowflake transformations rarely exist in isolation but must integrate with broader data ecosystems that include ingestion tools, orchestration platforms, and consumption applications. Understanding how transformation processes fit into these larger workflows affects design decisions and operational procedures. The choice of transformation tools and approaches should consider compatibility with existing infrastructure and future scalability requirements. The integration between transformation tools and Snowflake's native features requires careful consideration. While Snowflake provides powerful built-in capabilities for data processing, external transformation tools like [dbt](https://www.getdbt.com/product/what-is-dbt) offer additional features for development workflow, testing, and documentation. Determining the right balance between leveraging Snowflake's native capabilities and external tooling depends on team skills, operational requirements, and long-term strategic goals. Data sharing capabilities in Snowflake create opportunities for transformation results to be consumed across organizational boundaries. This capability requires additional consideration of data governance, security, and performance implications. Transformations that will be shared externally may require different design approaches than those consumed only within the organization. ## Operational monitoring and maintenance Monitoring Snowflake transformations requires understanding both the technical performance metrics and the business impact of transformation processes. Traditional database monitoring approaches may not capture the full picture of transformation health in a cloud-native environment. Implementing comprehensive monitoring that covers compute utilization, data freshness, quality metrics, and business KPIs provides the visibility needed for effective operations. The elastic nature of Snowflake compute means that performance issues may manifest differently than in traditional environments. A transformation that performs well under normal conditions might experience significant degradation during peak usage periods or when processing larger data volumes. Designing monitoring systems that can detect and alert on these variable conditions helps maintain consistent performance. Maintenance procedures for Snowflake transformations must account for the platform's continuous evolution and feature updates. Snowflake regularly introduces new capabilities that may benefit existing transformations, but adopting these features requires careful evaluation and testing. Establishing processes for evaluating and incorporating new Snowflake features helps teams optimize their transformation processes over time. ## Team skills and organizational readiness Successfully implementing data transformation on Snowflake requires teams with appropriate skills in SQL optimization, cloud data architecture, and modern development practices. The shift from traditional data warehousing approaches to cloud-native transformation patterns may require significant learning and adaptation. Assessing current team capabilities and identifying skill gaps helps inform training and hiring decisions. The collaborative nature of modern transformation development, particularly when using tools like [dbt](https://www.getdbt.com/data-platforms/snowflake), requires teams to adopt software engineering practices like version control, code review, and automated testing. Organizations accustomed to traditional BI development approaches may need to invest in process changes and tooling to support these new workflows effectively. Change management extends beyond technical considerations to include organizational culture and processes. Teams must be prepared to embrace iterative development, continuous improvement, and data-driven decision making. The speed and flexibility of Snowflake transformation capabilities can only be fully realized when supported by appropriate organizational practices and mindset. The considerations outlined above represent the key areas that data engineering leaders must evaluate when implementing transformation processes on Snowflake. Success requires balancing technical capabilities with cost considerations, operational requirements, and organizational readiness. By carefully considering these factors, teams can build transformation processes that leverage Snowflake's strengths while avoiding common pitfalls and ensuring long-term success. ## Snowflake data transformation FAQs **How does Snowflake's unique architecture impact data transformation performance?** Snowflake's architecture separates compute from storage, allowing multiple compute clusters to access the same data simultaneously without contention. This enables parallel processing of transformation workloads and elastic scaling based on demand. The columnar storage and automatic clustering capabilities significantly impact performance, with queries that leverage metadata and pruning capabilities performing better than those requiring full table scans. However, this flexibility requires careful consideration of compute costs and proper optimization of transformation logic. **What are the different materialization strategies available in Snowflake and when should each be used?** Snowflake offers several materialization options: Views provide the lowest storage cost but require computation on every query, making them suitable for exploratory datasets. Tables offer fast query performance but consume storage and require regular refreshes, making them ideal for frequently accessed data powering critical dashboards. Incremental models process only changed data, reducing compute costs for large datasets. Dynamic Tables automatically manage refresh schedules and dependencies, simplifying operational overhead but requiring careful consideration of refresh frequency and resource allocation. **How can teams optimize costs when running data transformations in Snowflake?** Cost optimization in Snowflake's consumption-based pricing model requires efficient SQL and appropriate warehouse sizing, as compute costs are driven by warehouse size and runtime. Jobs that complete quickly on smaller warehouses often cost less than those requiring larger warehouses for extended periods. Batching transformation jobs during off-peak hours and using larger warehouses that complete work faster can reduce total costs. Additionally, implementing appropriate data retention policies and balancing storage costs against compute costs when deciding whether to store intermediate results or recompute them on demand helps manage overall expenses. --- --- title: "Data transformation on Databricks: Best practices with dbt" description: "Databricks and dbt combine for scalable, governed transformation." url: "https://www.getdbt.com/blog/data-transformation-dbt-databricks" date: "2025-06-10" authors: ["Joey Gault"] categories: ["Pulse"] --- # Data transformation on Databricks: Best practices with dbt When implementing data transformation on [Databricks](https://www.databricks.com/), teams typically encounter several architectural patterns, each with distinct trade-offs. The most straightforward approach involves running transformations directly within Databricks notebooks or jobs, utilizing Spark SQL or PySpark for data processing. This pattern offers tight integration with the Databricks ecosystem and can be effective for teams with strong Spark expertise. However, many organizations find that pure Databricks-native transformation approaches can become difficult to manage as complexity grows. SQL code scattered across notebooks lacks the modularity and testing capabilities that modern data teams expect. Version control becomes challenging, and collaboration between team members can suffer when transformation logic is embedded within notebook environments. The integration of [dbt with Databricks](https://www.getdbt.com/data-platforms/databricks) addresses many of these limitations by providing a structured framework for transformation development while still leveraging Databricks' computational capabilities. dbt compiles transformation logic into SQL that executes on Databricks SQL warehouses, creating a clear separation between transformation logic and execution infrastructure. This pattern has gained significant traction because it allows teams to maintain the benefits of the lakehouse architecture while adopting software engineering best practices for their transformation code. Recent developments have further enhanced this integration pattern. The introduction of dbt Platform Task Types in Databricks Lakeflow Jobs enables teams to orchestrate dbt transformations directly within Databricks workflows. This capability eliminates the need for separate orchestration tools and provides end-to-end pipeline management within a single interface. Organizations can now manage raw data ingestion, transformation, and downstream analytics processes through unified Databricks orchestration while maintaining the development experience and capabilities that dbt provides. ## Performance and cost optimization Performance considerations play a crucial role in Databricks transformation implementations. The platform's flexible compute model allows teams to scale resources based on workload requirements, but this flexibility requires careful management to avoid unexpected costs. Databricks SQL warehouses provide optimized query execution for analytical workloads, with recent improvements delivering significant performance gains through intelligent workload management and automatic optimization features. The choice of [materialization strategy](https://docs.getdbt.com/reference/resource-configs/databricks-configs) becomes particularly important in the Databricks context. Views provide cost-effective development and testing environments but may not deliver adequate performance for complex transformations or high-concurrency scenarios. Tables offer better query performance but consume more storage and require periodic rebuilds to maintain data freshness. Incremental models can provide an optimal balance, processing only changed data to minimize compute costs while maintaining acceptable performance levels. [Liquid clustering](https://docs.databricks.com/aws/en/delta/clustering), a Databricks feature that automatically manages data layout without manual intervention, can significantly improve query performance for dbt models. This capability eliminates the need for manual partitioning strategies and adapts automatically to changing query patterns. Teams implementing dbt on Databricks should consider how their transformation patterns align with these optimization features to maximize both performance and cost efficiency. Storage optimization also requires attention in Databricks environments. The platform's integration with cloud storage services like Azure Data Lake Storage provides cost-effective data storage, but transformation patterns can impact storage efficiency. Frequent full table rebuilds can lead to storage bloat, while poorly designed incremental strategies may result in small file problems that degrade query performance. ## Governance and security frameworks Data governance takes on particular importance in Databricks environments due to the platform's broad capabilities and multi-persona access patterns. Unity Catalog provides centralized governance and lineage tracking across the entire Databricks ecosystem, but [implementing effective governance](https://www.getdbt.com/product/governance) requires careful consideration of how transformation processes interact with these capabilities. When dbt transformations execute on Databricks, metadata flows seamlessly between the two systems, providing comprehensive lineage tracking from raw data sources through final analytical outputs. This integration enables data teams to understand data dependencies and impact analysis when making changes to transformation logic. However, organizations must establish clear governance policies around who can modify transformation code, how changes are reviewed and approved, and how production deployments are managed. Security considerations extend beyond traditional access controls to include compute resource management and data exposure patterns. Databricks service principals and personal access tokens require careful management, particularly when integrating with external tools like dbt. Organizations should establish clear policies around credential management, rotation schedules, and access scope limitations. The shared nature of Databricks compute resources also introduces considerations around workload isolation and resource allocation. Transformation jobs can impact the performance of other workloads running on the same clusters, requiring careful planning around compute resource allocation and scheduling. Teams should consider dedicated compute resources for production transformation workloads to ensure predictable performance and avoid resource contention issues. ## Development workflow considerations The development experience for data transformation on Databricks varies significantly depending on the chosen tooling approach. Native Databricks development through notebooks provides immediate access to the full platform capabilities but can create challenges around code organization, version control, and collaborative development practices. [dbt](https://www.getdbt.com/product/what-is-dbt) integration addresses many of these development workflow challenges by providing familiar software development patterns. Teams can leverage Git-based version control, pull request workflows, and automated testing practices. The [dbt](https://www.getdbt.com/product/dbt) development environment supports local development, cloud-based IDEs, and integration with popular development tools like VS Code. This flexibility allows teams to adopt development practices that align with their existing workflows and preferences. However, the multi-tool nature of dbt and Databricks integration can introduce complexity in development workflows. Developers need to understand both the [dbt compilation process](https://docs.getdbt.com/reference/commands/compile) and Databricks execution environment. Debugging failed transformations may require investigation across multiple systems, and performance optimization requires understanding both dbt's compilation patterns and Databricks' execution characteristics. The recent introduction of enhanced IDE integration and debugging capabilities helps address some of these challenges. [dbt's integration with Databricks provides better visibility into query execution plans and performance metrics](https://www.getdbt.com/data-platforms/databricks), enabling developers to optimize their transformation logic more effectively. Additionally, the unified orchestration capabilities reduce the complexity of managing separate scheduling and monitoring systems. ## Scalability and operational considerations As transformation workloads grow in complexity and volume, operational considerations become increasingly important. Databricks' auto-scaling capabilities can help manage variable workload demands, but teams need to establish appropriate scaling policies and monitoring practices to ensure reliable operation. The combination of dbt and Databricks provides several scalability advantages. dbt's compilation process enables efficient dependency resolution and parallel execution of transformation tasks. Databricks' distributed computing capabilities can handle large-scale data processing requirements. However, realizing these benefits requires careful attention to transformation design patterns and execution strategies. Monitoring and alerting become critical as transformation pipelines grow in complexity. Teams need visibility into both transformation logic execution and underlying infrastructure performance. Databricks provides comprehensive monitoring capabilities for compute resource utilization and query performance, while dbt offers transformation-specific monitoring around model execution, test results, and data quality metrics. The operational overhead of managing multiple systems can be significant, particularly for smaller teams. Organizations should consider their operational capabilities and resource constraints when evaluating the complexity of multi-tool architectures. The unified orchestration capabilities introduced through Databricks Lakeflow Jobs can help reduce this operational burden by consolidating pipeline management within a single platform. ## Strategic implementation considerations Successfully implementing data transformation on Databricks requires alignment between technical capabilities and organizational needs. Teams with strong SQL skills and analytics engineering practices may find dbt integration provides immediate productivity benefits. Organizations with primarily data engineering backgrounds might prefer Databricks-native approaches that leverage existing Spark expertise. The choice of implementation approach should consider long-term organizational goals around data team structure, skill development, and technology standardization. dbt's growing adoption across the industry can provide advantages in talent acquisition and knowledge sharing. However, organizations heavily invested in Databricks ecosystems might benefit from deeper platform integration through native tooling approaches. Change management becomes particularly important when introducing new transformation patterns. Teams accustomed to notebook-based development may require training and support to adopt dbt's model-based approach. Conversely, teams with strong dbt experience may need to develop understanding of Databricks-specific optimization techniques and operational practices. The evolving nature of both Databricks and dbt capabilities means that implementation decisions made today may need revisiting as new features become available. Organizations should maintain flexibility in their architectural approaches and stay informed about platform developments that might impact their transformation strategies. Ultimately, the success of Databricks data transformation implementations depends on careful consideration of technical requirements, organizational capabilities, and strategic objectives. The combination of Databricks' powerful lakehouse platform with structured transformation approaches like dbt offers compelling benefits, but realizing these benefits requires thoughtful planning and execution across multiple dimensions of the data platform architecture. ## Data transformation on Databricks FAQs **What is data transformation on Databricks? ** Data transformation on Databricks involves processing and converting raw data into analytical-ready formats using the platform's computational capabilities. Teams can implement transformations through various approaches, including running transformations directly within Databricks notebooks or jobs using Spark SQL or PySpark, or by integrating with structured frameworks like dbt that compile transformation logic into SQL executed on Databricks SQL warehouses. The platform's flexible compute model allows scaling resources based on workload requirements while leveraging features like Unity Catalog for governance and Liquid clustering for performance optimization. **Why is data transformation important? ** Data transformation is crucial for converting raw data into formats suitable for analytics and business intelligence. It enables organizations to clean, structure, and enrich their data while maintaining data quality through testing and validation processes. Effective transformation processes provide comprehensive lineage tracking, enable impact analysis when making changes, and support governance requirements through centralized metadata management. Additionally, proper transformation strategies help optimize both performance and costs by implementing appropriate materialization strategies and leveraging platform-specific optimization features. **When should you use declarative data transformation with Delta Live Tables (DLT) versus procedural transformations with Apache Spark code on Databricks?** The choice between declarative and procedural approaches depends on your team's expertise, complexity requirements, and operational preferences. Procedural transformations with Apache Spark code offer tight integration with the Databricks ecosystem and work well for teams with strong Spark expertise, but can become difficult to manage as complexity grows due to challenges with modularity, testing, and version control. Declarative approaches like DLT or dbt integration provide structured frameworks with software engineering best practices, better collaboration capabilities, and clearer separation between transformation logic and execution infrastructure, making them more suitable for teams seeking maintainable, scalable transformation processes with comprehensive governance and lineage tracking. --- --- title: "How dbt enhances your Snowflake data stack" description: "Already using Snowflake? Here’s why dbt adds the structure, testing, and workflow automation needed to scale with speed and trust." url: "https://www.getdbt.com/blog/snowflake-dbt" date: "2025-06-10" authors: ["Daniel Poppy"] categories: ["Learn"] --- # How dbt enhances your Snowflake data stack Snowflake provides a powerful cloud data platform with elastic compute, separate storage and compute layers, and strong security out of the box. Its zero-maintenance infrastructure, pay-as-you-go model, and native cloud architecture have made it a go-to choice for data teams across industries. Snowflake supports both structured and semi-structured data, delivers high-performance analytics, and integrates easily with a broad data ecosystem—making it a strong foundation for modern data platforms. dbt is a transformation framework that works [natively within Snowflake](https://app.snowflake.com/marketplace/listing/GZTYZSRT2UA/dbt-labs-dbt). It brings structure, governance, and workflow automation to your transformation layer. Together, [Snowflake and dbt](https://www.getdbt.com/data-platforms/snowflake) combine computational power with modern development practices—so data teams can scale analytics with confidence, speed, and trust. ## Development structure and modular design dbt introduces a transformation framework that runs directly in Snowflake, allowing teams to build reusable, scalable data models. Foundational models, like a customers model, can serve multiple use cases across the business, from churn analysis to customer lifetime value, without duplicating logic. With dbt, teams define transformations once and reuse them across the stack. When business rules change (e.g., how “active customers” are defined), the update is made in one place and automatically cascades to downstream models. dbt’s dependency management ensures everything runs in the correct order, leveraging Snowflake’s performance without sacrificing transparency or control. Tools like [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) give teams full visibility into their data ecosystem, showing how models are connected, what transformations exist, and where logic is being reused. This makes it easier to discover, extend, and govern your dbt projects — especially as your Snowflake footprint grows. ## Collaboration and version control [dbt applies software engineering best practices to analytics workflows](https://www.getdbt.com/resources/the-analytics-development-lifecycle) — bringing version control, code reviews, and collaborative development into the data stack. All transformation logic is written as code and stored in Git repositories, enabling workflows like branching, pull requests, and peer reviews. Analysts and engineers can safely develop in isolated branches, test changes, and merge them after approval — without impacting production models. These workflows run seamlessly on Snowflake’s compute layer, combining analytical performance with structured governance. [Version control ](https://docs.getdbt.com/docs/cloud/git/version-control-basics)also creates a full audit trail of every change to your data models. This adds transparency and accountability as your team and project complexity grow. Multiple contributors can work in parallel, with Git coordinating efforts and reducing time spent on coordination. Beyond technical workflows, dbt fosters a shared language across teams. Analysts, engineers, and business users align around documented models and consistent definitions, improving clarity and speeding up onboarding. New team members can quickly understand how models are built and how Snowflake is being used — by exploring the dbt project structure itself. ## Data quality and testing dbt introduces built-in [data testing](https://docs.getdbt.com/docs/build/data-tests) to ensure the reliability of your Snowflake transformations — before they power dashboards, reports, or ML models. While Snowflake handles execution, dbt enforces that data meets your team’s technical and business requirements. You can define tests for: - Uniqueness (e.g., no duplicate IDs) - Field validity (e.g., acceptable values in categorical columns) - Referential integrity (e.g., foreign key relationships) - Metric thresholds (e.g., revenue must be non-negative) These tests run automatically during development and deployment, catching issues early in the pipeline. If a test fails, dbt flags the issue before it reaches downstream consumers — reducing fire drills and increasing confidence in your data. As your Snowflake environment scales, so does dbt’s testing safety net. It grows alongside your complexity, safeguarding logic assumptions even as data sources or business definitions evolve. This makes dbt a critical quality layer in production-grade analytics and data science workflows. ## Cost optimization dbt helps you control Snowflake compute costs by optimizing how and when transformations run. With strategic use of materializations — views, tables, and incremental models — teams can balance performance with cost efficiency. - [Views](https://docs.getdbt.com/reference/resource-configs/snowflake-configs#secure-views) are ideal for infrequently accessed data, minimizing storage costs. - [Tables](https://docs.getdbt.com/reference/resource-configs/snowflake-configs) offer fast query performance for frequently used datasets. - [Incremental models](https://docs.getdbt.com/reference/resource-configs/snowflake-configs#merge-behavior-incremental-models) process only new or changed records, significantly reducing compute time for large datasets. For example, a media company analyzing user engagement can transform just the latest interaction data instead of reprocessing historical events. dbt’s dependency management ensures only the necessary transformations run when source data changes, avoiding redundant compute. As your Snowflake environment grows, dbt helps teams make smarter decisions about: - How often to refresh data - What compute resources to allocate - Which models need to run (and when) This approach allows you to [scale usage without runaway costs](https://www.getdbt.com/product/cost-optimization), helping teams extract full value from their Snowflake investment while keeping operations lean. ## Documentation and knowledge management dbt embeds documentation directly into your Snowflake transformation workflows — creating a [living, searchable data catalog](https://docs.getdbt.com/guides/sl-snowflake-qs?step=1). Teams can document everything from business definitions and update frequency to ownership and usage guidelines, all within the same environment they use for development. Because documentation lives alongside code, updates happen as part of the same workflow. When a model is added or updated in Snowflake, dbt ensures its documentation reflects those changes. This prevents drift between data logic and its explanations—making it easier for both technical and business users to trust the data. Over time, this builds a shared knowledge base that includes: - Column-level metadata - Model-level descriptions - Data lineage and dependencies - Business logic explanations New team members can quickly ramp up, and existing teams can collaborate more efficiently, with a clear understanding of what data exists, how it was created, and how it’s used across the organization. ## Orchestration and workflow management With dbt, teams can [orchestrate Snowflake transformations directly](https://docs.getdbt.com/docs/cloud-integrations/set-up-snowflake-native-app) — no third-party schedulers required. You can define jobs that run on a schedule or trigger based on events, with built-in support for dependency management, resource control, notifications, and failure handling. This allows data teams to power everything from hourly dashboards to weekly reporting jobs — without switching platforms or writing custom orchestration logic. You can manage production-grade workflows for Snowflake in the same environment where you develop and test models. As project complexity grows, so do the capabilities: - Define job dependencies and conditional logic - Enable parallel execution to reduce processing time - Manage orchestration across multiple projects or environments By integrating orchestration with development and testing, dbt ensures your Snowflake transformations run reliably and efficiently at scale. ## AI and Machine Learning support For organizations implementing AI and machine learning solutions, [dbt and Snowflake](https://www.getdbt.com/data-platforms/snowflake) create a strong foundation for data science workflows. dbt’s transformation framework provides data scientists with clean, tested, and well-documented feature sets derived from Snowflake. This reduces time spent on data preparation and validation, accelerating model development. With [column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage), data scientists can trace how features were derived—improving model transparency and trust. dbt’s [testing framework](https://docs.getdbt.com/docs/build/tests) adds confidence that the data powering models meets quality standards, even as underlying sources or business logic evolve. Teams can use dbt to build AI-ready feature tables in Snowflake, implement tests for accuracy and completeness, document business logic, and track dependencies. Integration with [Snowflake Cortex AI](https://docs.snowflake.com/en/guides-overview-ai-features) enables teams to embed AI directly into their data pipelines—combining dbt’s governance and transformation strengths with Snowflake’s scalable AI capabilities. This joint foundation supports both the technical and governance needs of production AI systems. Models receive consistent inputs, and teams maintain a clear understanding of data provenance and transformation logic — critical for regulated and high-trust AI use cases. ## Implementation examples Leading organizations across industries are using dbt and Snowflake together to power scalable, performant analytics workflows. [**McDonald’s Nordics**](https://www.getdbt.com/case-studies/mcdonalds-nordics) established a standardized Data Vault structure to track €1.3 billion in sales across multiple channels. By centralizing data from four distinct markets with different tech stacks, and applying dbt to drive consistent transformation logic, the team focused more on business modeling than technical rework. [**Siemens**](https://www.getdbt.com/case-studies/siemens) achieved a 93% reduction in daily data load times—cutting from 6 hours to just 25 minutes. They also decreased dashboard maintenance costs by 90%, thanks to dbt’s efficient incremental processing and transformation patterns on top of Snowflake’s compute power. [**Reforge**](https://www.getdbt.com/case-studies/reforge) tripled the size of their data team and saved 18 hours per week by improving development workflows and onboarding. With dbt’s modular structure and built-in documentation, new team members ramped faster and collaborated more effectively—while Snowflake provided a reliable, high-performance data backbone. ## dbt and Snowflake FAQs **What exactly does dbt add to my existing Snowflake setup?** dbt brings a structured transformation framework into your Snowflake environment. While Snowflake handles data storage and processing, dbt adds software engineering best practices like version control, modular development, automated testing, documentation, and orchestration. The result: You can develop, test, and deploy SQL transformations more efficiently—without changing how you use Snowflake. dbt doesn’t replace Snowflake functionality; it enhances how your team builds and manages data workflows on top of it. **How does using dbt with Snowflake affect my costs?** dbt can reduce your Snowflake compute costs by optimizing how data is transformed and materialized. With dbt, you can: - Use views for rarely accessed models - Materialize high-use models as tables - Create incremental models that process only new data Its dependency-aware execution ensures that Snowflake only runs what’s needed—avoiding redundant processing. For example, Siemens cut dashboard maintenance costs by 90% by applying dbt’s optimization techniques on Snowflake. **Do I need to migrate away from my current Snowflake workflows to use dbt?** Not at all. dbt integrates directly with your existing Snowflake environment. You don’t need to move data, change permissions, or rebuild your architecture. You can start small—by converting a few key SQL transformations into dbt models—and expand from there. Many teams adopt dbt incrementally, gaining benefits like versioning and testing without disrupting current workflows. --- --- title: "Databricks users, get ready to experience the new dbt Fusion engine" description: "Fusion is here for Databricks: welcome to the new era of analytics engineering." url: "https://www.getdbt.com/blog/databricks-users-get-ready-to-experience-the-new-dbt-fusion-engine" date: "2025-06-10" authors: ["David Tishgart"] categories: ["Partnerships"] --- # Databricks users, get ready to experience the new dbt Fusion engine Exciting news from the Databricks Data + AI Summit in San Francisco as today we’re expanding access for the public beta of the [dbt Fusion engine](https://www.getdbt.com/product/fusion) to teams using Databricks as their data platform. For customers familiar with dbt, this upgrade isn’t just an incremental improvement on the dbt Core engine—it’s a fundamental shift in speed, intelligence, and developer experience for anyone building with dbt in the lakehouse. ## The dbt Fusion engine experience on Databricks The dbt Fusion engine represents a [new era of analytics engineering](https://www.getdbt.com/blog/where-we-re-headed-with-the-dbt-fusion-engine). Born from a complete ground-up rewrite of dbt in Rust, Fusion is designed for performance at scale, pairing blazing-fast execution with deep understanding of your analytics code. For Databricks users, this means: - **Significant performance gains:** Project parsing and compilation is as much as 30x faster than with dbt Core, dramatically reducing turnaround time from code change to insight. - **Developer intelligence:** Fusion’s native SQL comprehension provides real-time feedback and error checking as you write SQL, catching mistakes and providing suggestions instantly—without executing queries against your warehouse. This now precisely understands Databricks SQL and provides all the same benefits to this ecosystem. - **Cost control:** Fusion’s “state-aware orchestration” understands what’s changed in your data and code, so only necessary models run—eliminating redundant workloads and helping organizations reduce cloud compute spending by 10% or more. - **Better governance (coming later this year):** Precise column-level lineage and richer metadata will soon provide enhanced data governance, making it easier to manage risk, enforce policies, and lay the foundation for safe AI applications downstream. For teams used to the original Python-based dbt Core engine, Fusion’s upgrade is a game-changer. It empowers analytics engineers to move faster, iterate with confidence, and focus more on business logic than debugging or firefighting performance issues. **** ## Who benefits? - **Data engineers and analytics engineers** running transformations at scale on Databricks will see dramatically reduced compile and parse times of dbt projects, lower cloud spend, and richer metadata for downstream consumers. - **Organizations** requiring strict governance, auditability, or cross-platform flexibility gain robust lineage, improved cost management, and the freedom to scale beyond a single data warehouse. This is especially valuable for companies orchestrating complex pipelines, handling sensitive data, or operating in hybrid/multi-cloud environments. ## **The new industry standard** dbt is regarded as the [industry standard for AI on structured data](https://www.getdbt.com/product/dbt). The Fusion engine, with deep SQL comprehension, will power the next generation of dbt use cases. Following the public beta, the dbt Fusion engine will be available for Databricks users through dbt or the [VS Code extension](https://docs.getdbt.com/docs/install-dbt-extension). If you’re a dbt managed customer, you can enable Fusion for your projects with just a few clicks. For those working with Databricks on-premises or in hybrid setups, Fusion is also available as a source-available package under the Elastic License v2 (ELv2). This gives organizations the flexibility to install, use, and put dbt Fusion into production locally. For a premium developer experience—including features like state-aware orchestration, cost management dashboards, and enterprise governance—dbt unlocks Fusion’s full value. With the Fusion engine now powering Databricks transformations, we are setting the new standard for scalable, intelligent, and cloud-agnostic analytics workflows. And of course, we're not done yet—look out for [support for more data platforms and dbt projects in the coming weeks](https://docs.getdbt.com/blog/dbt-fusion-engine-path-to-ga). **Ready to experience it? dbt Fusion is currently in Public Beta, but you [can review the docs](https://docs.getdbt.com/docs/fusion/about-fusion) or [check out a demo](https://www.getdbt.com/product/fusion).** We’ll be hosting Fusion demos all week at the Databricks Data + AI Summit. If you’re attending the show, [swing by our booth](https://www.getdbt.com/events/summit/databricks-data-ai-summit-2025) to grab some great swag and experience the new super-charged dbt. --- --- title: "What are the key roles in data governance?" description: "Learn who’s responsible for data governance—from CDOs and stewards to engineers and business stakeholders." url: "https://www.getdbt.com/blog/data-governance-key-roles" date: "2025-06-09" authors: ["Joey Gault"] categories: ["Pulse"] --- # What are the key roles in data governance? Strong data governance isn’t just about checking boxes or avoiding compliance issues. It’s about making sure the right people can access the right data, trust what they’re using, and understand how it was created. That’s hard to do without clear roles and responsibilities. As data stacks grow more complex and AI raises the stakes, getting governance right means knowing who’s doing what — and making sure everyone works together. In this article, we’ll walk through the key roles that make data governance work, from executive sponsors to frontline data stewards to the engineers building governance into everyday workflows. Each plays a different part, but they all share one goal: making data safer, more usable, and more trustworthy. ## Data governance executive: setting strategic direction At the executive level, many organizations are appointing Chief Data Officers (CDOs) to provide strategic oversight of the entire data estate. While not every company requires a formal CDO, every organization needs someone at the executive level who is responsible for achieving data governance goals and connecting them to broader business objectives. This executive role, whether it's a dedicated CDO or another senior leader, serves as the "data czar" who champions governance initiatives across the organization. They are responsible for securing resources, removing organizational barriers, and ensuring that data governance aligns with business strategy. This person translates technical governance concepts into business value, helping stakeholders understand why governance matters for their specific roles and objectives. The executive sponsor also plays a crucial role in establishing governance as an organization-wide priority rather than just a technical concern. They communicate the connection between data governance and company goals, making it clear that everyone who works with data shares responsibility for its accuracy and security. This top-down support is essential for creating the cultural shift necessary for governance success. ## Data stewards: the front line of governance Data stewards represent the operational backbone of any governance program. These individuals serve as the front line for governance initiatives because they are collectively responsible for defining and documenting the organization's data assets, ensuring data quality, and promoting effective data sharing across teams. The stewardship role encompasses several critical responsibilities. Data stewards define business rules and data standards, monitor data quality metrics, and work to resolve data issues as they arise. They act as liaisons between different teams, helping to bridge the gap between technical and business stakeholders when data problems need resolution. Perhaps most importantly, data stewards ensure that governance policies and procedures are implemented correctly in day-to-day operations. They translate high-level governance frameworks into practical, actionable guidelines that teams can follow. This role requires both technical understanding and business acumen, as stewards must understand how data flows through systems while also grasping its business context and usage. Data stewards also play a vital role in data discovery and cataloging. They help identify and classify data assets, ensuring that metadata is accurate and comprehensive. This work makes data more findable and understandable across the organization, enabling broader self-service capabilities while maintaining governance standards. ## Data owners: accountability at the source Data owners are typically the individuals or teams closest to where data is created and initially managed. They have primary accountability for specific datasets and are responsible for making decisions about data access, usage policies, and quality standards for their domain. The data owner role carries significant responsibility for ensuring that data meets quality standards and complies with relevant policies. They work closely with data stewards and other governance roles to establish and maintain data quality processes. Data owners also make decisions about who should have access to their data and under what circumstances. This role is particularly important because data owners understand the business context and operational requirements that shape how data should be managed. They can provide insights into [data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage), business rules, and quality requirements that might not be apparent to those further removed from the data's source. Their domain expertise is essential for making informed decisions about data governance policies and procedures. ## Data managers: coordination and communication Data managers provide guidance and oversight to data owners while handling communication with leadership and other teams. They coordinate governance activities across different domains and ensure that governance standards are consistently applied throughout the organization. This role involves significant coordination responsibilities. Data managers work to align governance practices across different business units, resolve conflicts between competing data needs, and ensure that governance policies are practical and implementable. They often serve as the primary interface between the governance program and senior leadership, providing updates on governance metrics and initiatives. Data managers also play a crucial role in continuous improvement of governance practices. They gather feedback from data owners and stewards, identify areas where governance processes can be streamlined or improved, and help evolve governance practices as the organization's data landscape changes. ## Technical governance roles: engineering and architecture While governance involves significant organizational and process components, technical roles remain essential for implementing and maintaining the systems that enable governance. Data engineers build and maintain the infrastructure that supports governance, including data pipelines, quality monitoring systems, and [access controls](https://docs.getdbt.com/best-practices/dbt-unity-catalog-best-practices#access-control). These technical roles are responsible for implementing the technical aspects of governance policies. They build automated data quality checks, establish monitoring and alerting systems, and create the technical infrastructure that enables data lineage tracking and metadata management. Tools like [dbt](https://www.getdbt.com/product/what-is-dbt) help these teams implement [governance practices](https://docs.getdbt.com/docs/fusion/supported-features#features-and-capabilities) directly within data transformation workflows, making governance a natural part of the development process rather than an external constraint. Technical governance roles also include data architects who design systems with governance principles in mind. They ensure that data architecture supports governance requirements like auditability, security, and quality monitoring. These roles require deep technical expertise combined with understanding of governance principles and business requirements. ## Compliance and security specialists As regulatory requirements become more complex and data security threats evolve, specialized roles focused on compliance and security have become increasingly important. These specialists ensure that governance practices meet regulatory requirements and protect sensitive data from unauthorized access or misuse. Compliance specialists stay current with evolving regulations like GDPR, HIPAA, and [emerging AI-specific requirements](https://www.getdbt.com/blog/understanding-data-governance-ai). They translate regulatory requirements into practical governance policies and help ensure that data handling practices meet compliance obligations. This role requires both legal/regulatory knowledge and understanding of how data flows through technical systems. Security specialists focus on protecting data throughout its lifecycle. They implement access controls, monitor for security threats, and ensure that governance practices include appropriate security measures. With the rise of AI and machine learning, these roles increasingly need to understand new types of security risks like model inversion attacks or prompt injection. ## The collaborative nature of modern governance Modern data governance succeeds when these roles work collaboratively rather than in isolation. The most effective governance programs break down silos between technical and business teams, creating shared workflows where everyone contributes to governance within their area of expertise. This collaborative approach is particularly important as organizations adopt modern data tools that enable broader participation in data work. When [analysts](https://www.getdbt.com/product/analyst) and other business users can contribute to data transformation and analysis within governed frameworks, the traditional boundaries between governance roles become more fluid. The key is ensuring that collaboration happens within appropriate guardrails. [Role-based access controls](https://docs.getdbt.com/docs/cloud/manage-access/about-user-access#role-based-access-control-), automated testing, and clear approval processes allow different roles to contribute while maintaining governance standards. This approach scales governance by distributing responsibility appropriately rather than creating bottlenecks around a small number of gatekeepers. ## Evolving roles in the age of AI As artificial intelligence becomes more prevalent in data work, governance roles are evolving to address new challenges and opportunities. AI introduces new considerations around bias, explainability, and model governance that traditional data governance roles weren't designed to handle. Some organizations are creating new roles specifically focused on [AI governance](https://www.ibm.com/think/topics/ai-governance#:~:text=Organizations%20can%20use%20several%20frameworks,%2C%20privacy%2C%20security%20and%20safety.), including AI ethics officers and model risk managers. These roles work alongside traditional data governance roles to ensure that AI systems are built on well-governed data and operate according to ethical and regulatory standards. The integration of AI also changes how existing governance roles operate. Data stewards increasingly need to understand how AI systems use data and what governance practices are needed to ensure responsible AI outcomes. Technical roles must implement new types of monitoring and controls designed for AI systems. ## Building sustainable governance organizations Successful governance programs recognize that roles and responsibilities must evolve as organizations grow and data landscapes become more complex. The most effective approach is to start with core roles and responsibilities, then adapt and expand as needed based on organizational maturity and requirements. The key is maintaining clarity about accountability while allowing flexibility in how roles are implemented. Smaller organizations might have individuals wearing multiple governance hats, while larger enterprises might have dedicated teams for each role. What matters is ensuring that all essential governance functions are covered and that everyone understands their responsibilities. Regular evaluation and adjustment of governance roles helps ensure that the program remains effective as the organization evolves. This includes gathering feedback from role holders, measuring governance outcomes, and adjusting responsibilities as needed to address gaps or inefficiencies. Data governance succeeds when it becomes embedded in how organizations naturally work with data rather than existing as a separate, parallel process. By establishing clear roles that work together effectively, data engineering leaders can build governance programs that enable innovation while managing risk, creating sustainable competitive advantages through trusted, well-managed data assets. ## Data governance roles FAQs **Who is responsible for data governance? ** Data governance responsibility is distributed across multiple roles within an organization. At the executive level, Chief Data Officers (CDOs) or senior leaders serve as "data czars" who provide strategic oversight and connect governance to business objectives. Data stewards act as the operational backbone, handling day-to-day governance activities like defining data standards and monitoring quality. Data owners have primary accountability for specific datasets and make decisions about access and usage policies. Additionally, data managers coordinate activities across domains, while technical roles like data engineers and architects implement the systems that enable governance. **What are the key differences in accountability and day-to-day responsibilities between Data Owners and Data Stewards?** Data owners have primary accountability for specific datasets and focus on decision-making authority, while data stewards handle the operational implementation of governance policies. Data owners are typically closest to where data is created and make decisions about data access, usage policies, and quality standards for their domain. They understand the business context and provide insights into data lineage and business rules. Data stewards, on the other hand, serve as the front line for governance initiatives, defining and documenting data assets, monitoring data quality metrics, resolving data issues, and translating high-level governance frameworks into practical guidelines that teams can follow. **What skills do data governance roles require?** Data governance roles require a combination of technical and business skills. Data stewards need both technical understanding and business acumen to comprehend how data flows through systems while grasping business context and usage. Technical governance roles like data engineers require deep technical expertise combined with understanding of governance principles and business requirements. Compliance specialists need legal/regulatory knowledge along with understanding of data flows through technical systems. Security specialists must understand traditional security measures plus emerging risks like AI-specific threats. Executive-level roles require the ability to translate technical governance concepts into business value and communicate effectively with stakeholders across the organization. --- --- title: "The history and future of the data ecosystem" description: "Mainframes, relational databases, ETL, Hadoop, the cloud, and all of it with Lonne Jaffe of Insight Partners." url: "https://www.getdbt.com/blog/the-history-and-future-of-the-data-ecosystem" date: "2025-06-08" authors: ["Daniel Poppy"] categories: ["Insights"] --- # The history and future of the data ecosystem _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/the-history-and-future-of-the-data). _ In this decades-spanning episode, Tristan talks with Lonne Jaffe, Managing Director at Insight Partners and former CEO of Syncsort (now Precisely), to trace the history of the data ecosystem—from its mainframe origins to its AI-infused future. Lonne reflects on the evolution of ETL, the unexpected staying power of legacy tech, and why AI may finally erode the switching costs that have long protected incumbents. The future of the AI and standards era is bright. _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Episode chapters **00:46 – Meet Lonne Jaffe: background & career jurney** Lonne shares his career highlights from Insight Partners, Syncsort/Precisely, and IBM, including major acquisitions and tech focus areas. **04:20 – The origins of Syncsort & sorting in mainframes** Discussion on why sorting was a critical early problem in hierarchical databases and how early systems like IMS worked. **07:00 – M&A as innovation strategy** How Syncsort used inorganic growth to modernize its platform, including an early example of migrating data from IMS to DB2 without rewriting apps. **09:35 – Technical vs. strategic experience** Tristan probes Lonne’s technical depth despite his business titles; Lonne shares his background in programming and a fun fact about juggling. **11:55 – Why this history matters** Tristan sets up the key question: what lessons from 1970s-2000s ETL tooling still shape the modern data stack? **13:00 – Proto-ETL: The real OGs** Lonne traces the origins of ETL to 1970s CDC, JCL, and early IBM tools. Prism Solutions in 1988 gets credit as the first real ETL startup. **15:40 – Rise of the ETL market (1990s)** From Prism to Informatica and DataStage—early 90s vendors brought visual development to what was once COBOL-heavy backend work. **18:00 – Why people offloaded Teradata to Hadoop** Exploring how cost, contention, and capacity drove ETL out of the warehouse and into Hadoop in the 2000s. **20:00 – Performance vs. price: Jevons Paradox in ETL** Why lower compute and storage costs led to _more_ ETL, not less—and how parallelization changed the game. **22:30 – Evolution of data management suites** How ETL expanded into app-to-app integration, catalogs, metadata management, and why these bundles got bloated. **25:00 – Rise of data prep & self-service analytics** Tools like Kettle, Pentaho, and Tableau mirrored ETL for business users—spawning a whole “data prep” category. **27:30 – Clickstream, logs & big data chaos** How clickstream and log data changed the ETL landscape, and the hope (and letdown) of zero-copy analytics. **29:10 – Why is old software so sticky?** Tristan and Lonne explore the economics of switching costs, the illusion of freedom, and whether GenAI could break the lock-in. **33:30 – Are old tools actually… good?** Defending mainframes and 30-year-old databases like Cache. Sometimes the mature option is better—just not sexy. **36:00 – The new vs. the durable** Modern tools must prove themselves against decades of reliability and robustness in finance, healthcare, and compliance. **38:20 – GenAI in data: The early movers** Lonne highlights why companies like Atlan and dbt Labs are in the best position to win—distribution, trust, and product maturity. **41:00 – TAM and the Jevons Paradox, again** Revisiting how price drops expand TAM. Some categories vanish, others explode—depending on elasticity of demand. **43:15 – Unlocking new personas with LLMs** Structured data access for non-technical users is finally viable, but “it has to be right”—trust and quality remain the barrier. **46:00 – Real-world examples: dbt’s MCP server win** Tristan shares how dbt’s Metadata API became a catalog replacement for a traditional financial institution—an unplanned AI GTM success. **48:30 – Agents, not interfaces** New pattern: LLMs as agents interacting directly with infrastructure via APIs. Tool use is becoming table stakes for AI integration. **50:30 – Are LLMs birthright tools yet?** Discussion around adoption of ChatGPT Enterprise, Claude, etc. Lonne suggests adoption is accelerating fast—and the usage model matters. **52:00 – Looking ahead** The conversation ends with a reflection on GenAI’s near future in data workflows, TAM expansion, and what the next episode might tackle. **** ## Key takeaways from this episode **Tristan Handy: You've had a long career in tech. Maybe start by giving us the 30,000-foot view of what you've been up to over the last couple decades?** **Lonne Jaffe:** I’ve been at Insight Partners for about eight years now, working mostly on deep tech investments—AI infrastructure companies like Run AI and [deci.ai](http://deci.ai/), both acquired by Nvidia. I’ve also done work with data infrastructure companies like SingleStore. Before Insight, I was CEO of a portfolio company called Syncsort, now Precisely. It was founded in 1968. Prior to that, I was at IBM for 13 years, working in middleware and mainframe technologies. Products like WebSphere, CICS, and TPF—foundational systems for enterprise computing. **Tristan Handy: And Syncsort's origin was in sorting, right? Literally sorting files?** **Lonne Jaffe:** Exactly. In the early days of computing, sorting was a huge part of what you did. Much of the data was hierarchical—stored in IMS—and had to be flattened into files to process. The algorithms were optimized to run in extremely resource-constrained environments. **Tristan Handy: Fascinating. And I assume as compute and storage improved, the data integration landscape evolved?** **Lonne Jaffe:** Yes. We saw a move from hierarchical to relational databases, then toward ETL tools in the 80s and 90s. The first real ETL startup was probably Prism Solutions in 1988. Informatica and DataStage showed up in the early 90s, followed by Talend and others. **Tristan Handy: It seems like we got a whole bundle of tools over time—ETL, CDC, app integration, metadata, and so on.** **Lonne Jaffe:** Yes, often bundled together, even though data prep and app integration were treated separately. That persisted for longer than you'd expect. At Syncsort, we acquired a company with a "transparency" solution that allowed IMS applications to use data stored in DB2 without rewriting code—a clever way to manage switching costs. **Tristan Handy: Speaking of switching costs—why are these legacy tools so sticky?** **Lonne Jaffe:** Great question. In many cases, no customer loves the product. They’d switch in a heartbeat—if it were easy. But rewriting jobs and ensuring reliability is a heavy lift. The best outcome is a new system that replicates old functionality. And for many organizations, that’s not worth the risk. **Tristan Handy: But if generative AI could reduce those switching costs?** **Lonne Jaffe:** That’s the potential. Code generation, agents that explore and iterate—those could erode the moat that’s protected these incumbents for decades. Not tomorrow, but it’s a real possibility. **Tristan Handy: It also seems like some of these systems are more robust than people give them credit for.** **Lonne Jaffe:** Absolutely. Mainframes are IO supercomputers. Products like InterSystems Cache, used by Epic, are incredibly performant. But new systems must match or exceed those capabilities in reliability and scale, which is a high bar. **Tristan Handy: As you look at the evolution of the modern data stack, how do you think about its impact on the market?** **Lonne Jaffe:** In the 2010s, we saw disaggregation—tools like Fivetran, dbt, and Snowflake each tackled a slice of the old enterprise bundle. But the TAM isn’t infinite. Some categories may compress or vanish entirely if price drops aren’t offset by new demand. **Tristan Handy: Do you think AI expands or compresses the data stack?** **Lonne Jaffe:** It depends. High elasticity of demand—like with dashboards or analytics—can drive massive TAM expansion. But some categories, like logo redesign or simple data movement, might get commoditized. For more complex workflows, AI agents accessing platforms like dbt or Atlan could dramatically increase value by automating common tasks and enabling new personas. **Tristan Handy: We’ve seen an example already—a customer replaced their data catalog with our dbt Cloud metadata server and AI interface.** **Lonne Jaffe:** That’s telling. If AI interfaces can connect to tools like dbt and generate value—self-service, documentation, lineage—it changes the game. Especially for organizations already standardized on those platforms. **Tristan Handy: What’s your view on how these AI interfaces get distributed?** **Lonne Jaffe:** ChatGPT Enterprise, Claude, and others are spreading fast. Eventually, you’ll want those tools to search files, access internal metadata, and interact with your data stack—not just answer questions from the open web. **Tristan Handy: It makes a lot of sense. If AI is going to serve enterprise users, it needs access to the real data. Otherwise, it’s just a toy.** **Lonne Jaffe:** Exactly. A model that can’t query or verify against your actual environment won’t be reliable. And data quality and observability—something dbt Cloud is already good at—become foundational. --- --- title: "Fresh pow: Snowflake Summit 2025 was as satisfying as first tracks" description: "Snowflake Summit attendees got a first look at the new dbt Fusion engine, which will power Snowflake’s dbt Projects." url: "https://www.getdbt.com/blog/snowflake-summit-2025-recap" date: "2025-06-06" authors: ["Jeff Mills"] categories: ["Partnerships"] --- # Fresh pow: Snowflake Summit 2025 was as satisfying as first tracks What an amazing week in San Francisco. We met with customers and partners to listen, learn, and share in the excitement of the strengthening dbt Labs and Snowflake partnership. Our booth was busy with talks, demos, and s’mores. What was clear was that last week’s [dbt Launch Showcase](https://www.getdbt.com/blog/dbt-launch-showcase-2025-recap)—where we announced the dbt Fusion engine, the GA of dbt Canvas, and upcoming cost management features - resonated with the Snowflake community. ![Team photo](https://cdn.sanity.io/images/wl0ndo6t/main/b2bd12ca7fc3f43bf6665a4da9fdb74be91baf6d-2487x1399.png) ## dbt Projects powered by the dbt Fusion engine [The Fusion engine](https://www.getdbt.com/product/fusion) allows everyone to build faster, more efficiently in dbt. It’s a game-changer for dbt Labs and our customers that delivers lightning-fast performance, with parse times up to 30x faster than dbt Core and native SQL comprehension that can provide real-time validation of code—without the need to query your warehouse. This week Snowflake announced [dbt Projects](https://www.snowflake.com/en/news/press-releases/snowflake-openflow-unlocks-full-data-interoperability-accelerating-data-movement-for-ai-innovation/) available natively inside of Snowflake will be powered by dbt Fusion. We’re excited to get the Fusion engine into the hands of as many dbt developers as possible, and Snowflake is a great distribution channel. That said, dbt Fusion delivers superior features when paired with dbt’s data control plane: - dbt Fusion is the most efficient way to run dbt, and paired with [Cost Optimization](https://www.getdbt.com/product/cost-optimization) and [state-aware orchestration](https://www.getdbt.com/product/fusion) (only available from dbt) you’ll ensure that your spend is the most efficient and effective on Snowflake. - dbt is platform agnostic. For organizations that run Snowflake and another data platform, dbt is the best way to manage cross-platform data pipelines and projects. - dbt Fusion improves dbt-specific features like [dbt Canvas](https://www.getdbt.com/product/develop) experience for analysts and the AI-powered [dbt Copilot](https://www.getdbt.com/product/dbt-copilot). Whether you build in Snowflake or with us, we’re thrilled you’re on dbt and can’t wait to see what you build. ![Snowflake Summit stage](https://cdn.sanity.io/images/wl0ndo6t/main/6cc5a958f3a3ed7e7acc0222cab7d2944a21ef02-2016x1512.jpg) **** ## Customers shared their successes Shoutout to [Whoop](https://www.getdbt.com/case-studies/whoop) for its featured spot during the Builder Keynote and highlighting how dbt and Snowflake Cortex are bringing AI and data together to produce better health outcomes for Whoop-device wearers. It was super cool to see. ![WHOOP presentation](https://cdn.sanity.io/images/wl0ndo6t/main/fd0dfa5b119a6b4a77aef9556c476e1489485943-2016x1512.jpg) [Shanna Anderson](https://www.linkedin.com/in/shanna-anderson-545a8010/) and [Katherine Long](https://www.linkedin.com/in/katherineglong/) of Fifth Third Bank presented on how they’re building an amazing assortment of data products for their business to use. ![Fifth Third Bank](https://cdn.sanity.io/images/wl0ndo6t/main/d721600cfd6d6b7f16fd4103536e8289954229ac-4032x2268.jpg) [Pooja Crahen](https://www.linkedin.com/in/drexelpooja/), from Okta, shared how the company moved from siloed, fragmented data processes to one that was centralized, optimized, and governed with dbt and Snowflake. ![Okta](https://cdn.sanity.io/images/wl0ndo6t/main/aef5b674be4b7aa2f054c319f56ebeabb8b4605d-4032x2268.jpg) [Rahavan (Ra) Raman](https://www.linkedin.com/in/rahavan-raman/) from Zscaler talked about how the data team speeds data model development in a governed framework, highlighting features like column-level lineage, dashboard health tiles, and CI/CD process with dbt. ![ZScaler](https://cdn.sanity.io/images/wl0ndo6t/main/d421ea531938c78ab773163f37fec75058ff2c44-2253x2254.jpg) ## The booth was slammed Attendees couldn’t get enough of dbt Fusion, Canvas, Catalog, and the socks. Did I mention we had USA-made socks in the booth? ![dbt Labs booth](https://cdn.sanity.io/images/wl0ndo6t/main/2220e4a5e607e37896da53c980d3f4884b030246-4032x3024.jpg) ![Drew demo](https://cdn.sanity.io/images/wl0ndo6t/main/d1fdf722005db490e46f4ad8a3b7eb45cc98260a-3024x4032.jpg) ![dbt socks](https://cdn.sanity.io/images/wl0ndo6t/main/d98806e56cc6becf4b07ff52558910b84925d3d0-4032x3024.jpg) ## The parties were epic It wasn’t just all work this week. We hosted our customers and partners at many after-hours get-togethers. Thanks to our customers, partners, and employees for making this week an amazing experience. ![Tristan talk](https://cdn.sanity.io/images/wl0ndo6t/main/2fcec26425e3d06542b84b8ed4d91ff6fa67551f-2500x1663.jpg) ![Ancillary event](https://cdn.sanity.io/images/wl0ndo6t/main/92337bb3a94b9b10ed54bd2da8a0569bc590186c-2500x1663.jpg) ## What’s next? The summer is just getting started. Please join us for a webinar on June 25th or 26th, depending on which day and time works best for you, to learn more about the new, Fusion-powered dbt and Snowflake. We’ll also highlight exciting new analyst features, including dbt Canvas, dbt Insights, and the dbt Catalog, which now extends to Snowflake assets. [Register here](https://www.getdbt.com/resources/webinars/empowering-data-analysts-showcase-series-part-one). --- --- title: "The changing role of the analyst: Getting closer to the data source" description: "Data analysts are being asked to do more than ever. Here’s how they can deliver higher-quality, well-governed data faster." url: "https://www.getdbt.com/blog/data-analyst-closer-to-data" date: "2025-06-05" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # The changing role of the analyst: Getting closer to the data source As the demand for high-quality data accelerates, driven in large part by generative AI, data teams face mounting pressure to deliver more, faster. To do this, a wide range of analysts - from technical data experts to business-focused users who prefer visual interfaces - need a faster, more reliable path from question to insight. However, that’s creating tension between speed and good data governance. In any given company, there are usually multiple data analysts per data engineer. This makes it unrealistic for every data request to flow through the data engineering team. With long request queues and competing priorities, engineers simply don’t have the bandwidth to support every ad-hoc query or model analysts need. This means more analysts are using self-serve tooling, and in many cases, spinning up their own data marts to work around engineering bottlenecks. But when these workflows live outside of governed pipelines, relying on inconsistent logic and unstructured data assets, they introduce serious risks, including duplicated work, data security concerns, and rising cloud costs from redundant or unmanaged assets. The solution isn't more dashboards. It's empowering the right analysts with the right self-service tools, without sacrificing governance. We'll look at how the role of the analyst is changing in response to this demand, and how companies can use governed collaboration to increase data velocity without compromising on quality. ## The evolving role of the analyst Data analysts are increasingly taking a more active role in shaping data within their companies. This is driven by both business demands for data and changes in the underlying technology. As a result, analysts are: - Getting closer to raw and modeled data sources - Becoming more familiar with data tooling - Incorporating AI Let’s take a look at each of these areas in detail and what’s driving them. ### Getting closer to raw and modeled data sources Data quality is still a leading concern across industries. [In dbt Labs’ 2024 State of Engineering Analytics report](https://www.getdbt.com/blog/the-2024-state-of-analytics-engineering-report), we found 57% of data professionals citing data quality as the largest data-related issue. That’s up from 41% in 2022. Companies are expecting analysts to be more than just passive users of data. Analysts are increasingly expected to have the skills and tools to verify that the data sets they have are accurate, up to date, and have been properly cleaned for business use. The need for more high-quality data is also driving analysts to seek out useful data sources they can incorporate into their work. Data silos, islands of data that are independent from and often incompatible with more governed and highly structured data, are still a vexing issue plaguing most companies. Data analysts play a pivotal role in helping to find and transform this data to ensure that it's compatible with the company's governed datasets. ### Becoming more familiar with data tooling In the past, data pipelines were solely the province of data engineering teams. They were often written in different languages, hidden away in stored procedure code in a database or data warehouse. They were as hard to find as they were to use and manage. Today, [with tools like dbt](https://www.getdbt.com/blog/what-exactly-is-dbt), anyone with knowledge of SQL or Python can contribute to data transformation code. dbt provides a common and governed approach to data transformation backed by software development best practices like documentation, version control, and testing. As a result, analysts are becoming more familiar with the technical tools required to create and maintain data pipelines, including source control systems such as Git. That enables data engineers and analysts to collaborate on analytics code, data tests, documentation, and data metrics in ways that weren’t previously possible. In other words, analysts, [who were always quite technical](https://www.getdbt.com/blog/there-is-no-such-thing-as-a-non-technical-data-analyst), are becoming increasingly more comfortable with more technical tools and workflows, blurring the lines between business and data roles. ### Incorporating AI dbt Labs co-founder Tristan Handy has noted [how AI is disrupting the way we do data engineering](https://www.getdbt.com/blog/how-ai-will-disrupt-data-engineering). The advent of GenAI means that analysts can do more and do it more quickly than ever before: - Beginner analysts can query data using natural language prompts to a [large language model (LLM)](https://www.ibm.com/think/topics/large-language-models), which can translate their requests into SQL and run the results for them against the source systems - Experienced analysts can use AI to help them automatically edit, develop, or understand complex queries that otherwise might take time to develop and debug fully - All analysts can leverage AI to generate boilerplate code for new data pipelines and tests, as well as base documentation for data models Of course, AI doesn’t replace the analyst, it supports them. High-quality reports and data products still require human judgment, context, and oversight. AI simply accelerates the work, keeping a skilled analyst in the loop every step of the way. ## The challenges that analysts face All this means that, more than ever, data analysts can dive headlong into data and find the answers they need without waiting on an already overtaxed data engineering team. However, analysts also run multiple risks when dealing directly with ungoverned and unstructured data: **No mechanisms to ensure data quality**. Data stored in multiple systems often isn’t rationalized or harmonized. It may exist in different formats across different data stores. Key data values - e.g., revenue - may even differ from system to system, leading to doubts around which system is the “source of truth.” **Missing (or unavailable) metadata**. Ungoverned data often lacks appropriate or complete [metadata](https://www.ibm.com/think/topics/metadata) - data about data. This can include **technical metadata** (tables, columns, data types, relationships, last update time, upstream source) as well as **business metadata** (owner, description, method of calculation, business meaning, and usage). Without metadata, it can be difficult to tell who’s responsible for a given dataset or how certain values were calculated. **Documentation is light or nonexistent**. A critical form of metadata is documentation about the meaning and purpose of a given dataset. Documentation is critical for collaborating across roles. However, without a tool that supports documenting datasets in a data model, such rich metadata might be impossible to capture. ## The tools that data analysts can use to collaborate on governed data In the end, data analysts are concerned primarily with delivering high-quality data and insights to their stakeholders as quickly as possible. It’s the job of a company’s data engineering and central governance teams to set standards and monitor data to ensure that this data is well-governed, secure, and compliant. With the right tools, data analysts can play a more active role, contributing to structured, governed data by building on shared models, documenting usage, and working within trusted workflows. Together, analysts and governance teams can raise the bar for data quality, security, and compliance. dbt serves as [a data control plane for analytics and AI](https://www.getdbt.com/blog/data-control-plane-why) that centralizes your analytics workflows so that teams can ship and use trusted data, faster. With the rise of AI and self-service analytics, the definition of an “analyst” is evolving. Today’s analysts span a wide spectrum, from SQL-fluent data experts to business users and data scientists who rely on visual tools or natural language. dbt opens the door for all of them to contribute to high-quality, trusted data products—without compromising on governance, speed, or security. Using dbt, analysts and data teams can collaborate on creating reliable, well-documented data sets, all within one central, governed environment by leveraging: **Easy model building**. [dbt Canvas](https://docs.getdbt.com/docs/cloud/canvas) is a visual tool that any analyst can use to contribute to data models. Using dbt Canvas’ visual, drag-and-drop experience and built-in context-aware AI powered by dbt Copilot, analysts can create model changes that compile to production-ready SQL with all the benefits of dbt, including version control, [orchestration](https://docs.getdbt.com/docs/deploy/deployments), and discovery. **Collaborative discovery**. dbt provides access to a company’s data transformation models and associated metadata via [dbt Catalog](https://docs.getdbt.com/docs/explore/explore-projects). This feature provides a full view of your data estate, including non-dbt data objects in Snowflake. Engineers, analysts, and business decision-makers can collaborate on code and documentation as part of [a single collaborative workflow](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). dbt supports writing documentation as an intrinsic part of each data model. Once a data pipeline is pushed to production, analysts can find a governed dataset, examine its metadata, and read its associated documentation before putting it to use. **Frictionless data insights**. With [dbt Insights](https://docs.getdbt.com/docs/explore/dbt-insights), analysts can freely query, validate, visualize, and share trusted data. Analysts can write SQL queries from scratch or use context aware AI, powered by dbt Copilot, to generate new queries using natural language prompts. This means analysts, regardless of technical skill, can explore data, uncover insights, and make decisions, all within a secure, governed environment built for trusted, self-service. dbt Insights is available in Preview. **Data lineage**. [Data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) provides a visual map of how data flows across your models, generating a visual lineage graph that helps teams understand upstream sources and downstream dependencies. Using these data lineage maps, analysts can trace issues to the origin without filing a support ticket or digging through disconnected tooling. Analysts are empowered to detect an issue in the data, report it, and data engineers can quickly assess the impact and resolve the problem. **AI-powered workflows**. [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot) is our context-aware AI-powered solution that supports engineers, analysts, and business users at every step of the data lifecycle. It understands the structure, logic, and metadata of your dbt project, so it can help analysts generate SQL, create data tests, and draft documentation with accuracy and speed. Engineers can use it to accelerate development and enforce standards. Analysts can safely explore models with guardrails in place. As more people engage with analytics, dbt Copilot helps maintain governance, quality, and consistency, by keeping everyone aligned on trusted, project-specific context. To learn more about how dbt brings analysts of all technical skill levels into one central, governed workflow powering faster, more trusted analytics for tomorrow’s data solutions, [request a demo today](https://www.getdbt.com/contact). --- --- title: "Aligning analytics initiatives to broader business strategies" description: "Analytics success begins with business objectives. Align your work, embed in domains, and move from insights to impact." url: "https://www.getdbt.com/blog/align-analytics-to-business-strategies" date: "2025-06-04" authors: ["Joey Gault"] categories: ["Pulse"] --- # Aligning analytics initiatives to broader business strategies The most successful analytics initiatives begin not with technical requirements but with clear business impact objectives. This represents a significant shift from traditional approaches where data teams focused primarily on data collection and left interpretation to others. Modern analytics organizations embed their technical capabilities directly within business domains, creating strategic partnerships that ensure relevance and actionability. When data teams are embedded within specific business units, they develop intimate knowledge of the problems that matter most. This proximity enables them to identify which data to collect, how to transform it effectively, and how to present insights that directly address business needs. The result is analytics work that drives measurable outcomes rather than simply producing reports. This business-first approach requires data engineering leaders to structure their teams accordingly. Rather than organizing around technical capabilities alone, successful organizations create cross-functional teams that combine technical expertise with domain knowledge. These teams can move from identifying a business problem to delivering a solution without navigating complex handoffs between departments. The embedded model also changes how success is measured. Instead of focusing solely on technical metrics like data pipeline reliability or query performance, teams can track business outcomes like revenue impact, operational efficiency gains, or customer satisfaction improvements. This shift in measurement creates natural alignment between analytics work and organizational priorities. ## Leveraging the versatility of analytics engineers [Analytics engineers](https://www.getdbt.com/blog/what-is-analytics-engineering#what-is-an-analytics-engineer) represent a crucial bridge between technical capability and business understanding. Their unique combination of analytical thinking and engineering skills makes them particularly valuable for organizations seeking to align technical initiatives with business strategy. Data engineering leaders who recognize and leverage this versatility can build more effective analytics organizations. The role of analytics engineers extends beyond traditional boundaries. They possess the business context of analysts while maintaining the technical skills necessary to build robust, scalable solutions. This dual capability allows them to understand both the "what" and the "how" of analytics initiatives, ensuring that technical implementations serve business needs effectively. As organizations scale their analytics capabilities, analytics engineers can flex into infrastructure challenges that might traditionally require specialized DevOps resources. Their understanding of both the business context and technical requirements enables them to solve infrastructure problems in ways that directly support analytical workflows. This versatility becomes particularly valuable when teams need to move quickly or when specialized resources are constrained. The key insight for data engineering leaders is that analytics engineers can serve as force multipliers across the organization. By empowering them to work across traditional role boundaries, teams can maintain alignment between technical work and business objectives while building more resilient and adaptable analytics capabilities. ## Implementing DevOps principles for alignment Successful alignment between analytics initiatives and business strategy requires more than good intentions; it demands systematic approaches to collaboration and delivery. [DevOps principles](https://www.browserstack.com/guide/devops-lifecycle), adapted for analytics workflows, provide a proven framework for maintaining this alignment at scale. The foundation of effective analytics DevOps lies not in technology but in people and processes. Organizations with scattered teams using different methodologies and working on different timelines struggle to maintain strategic alignment. The solution involves standardizing ways of working across all analytics teams, regardless of their specific technical stacks or business domains. This standardization begins with establishing common development practices. Teams working in waterfall methodologies cannot easily collaborate with those using agile approaches. Similarly, teams with different sprint cadences or release cycles create friction that impedes strategic alignment. Successful organizations invest significant effort in bringing all analytics teams onto consistent operational rhythms. The cultural shift required for effective analytics DevOps cannot be understated. Teams must embrace practices like continuous integration and continuous deployment not just as technical necessities but as enablers of business alignment. When analytics teams can deploy production-grade solutions every two weeks rather than every six months, they can respond more effectively to changing business needs and maintain closer alignment with strategic objectives. Technology choices should support these cultural and process changes rather than drive them. Code-based approaches enable the collaboration and reliability necessary for strategic alignment, but only when supported by appropriate organizational practices. The combination of standardized processes, collaborative culture, and enabling technology creates the foundation for analytics initiatives that remain aligned with business strategy over time. ## Building scalable analytics workflows Strategic alignment requires analytics capabilities that can scale with business needs while maintaining quality and reliability. [The Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) provides a framework for building such capabilities by applying software engineering best practices to analytics workflows. ![nfinity loop diagram showing the Analytics Development Lifecycle (ADLC) stages: Plan, Develop, Test, Deploy, Operate, Observe, Discover, and Analyze—blending Data and Ops.](https://cdn.sanity.io/images/wl0ndo6t/main/6a2eca8c4be8bc95297a2c67c14b3a4aa8e862a4-1850x906.png) The ADLC recognizes that analytics systems are fundamentally software systems and should be treated as such. This perspective enables data engineering leaders to leverage decades of software engineering experience in building reliable, scalable systems. The eight-stage lifecycle (Plan, Develop, Test, Deploy, Operate, Observe, Discover, and Analyze) creates a structured approach to analytics development that maintains business alignment throughout. The planning phase becomes particularly critical for strategic alignment. Every analytics initiative should begin with a clear business case that articulates expected outcomes and success metrics. This business case serves as the foundation for all subsequent technical decisions, ensuring that development work remains focused on delivering business value. The iterative nature of the ADLC supports strategic alignment by enabling rapid feedback cycles between analytics teams and business stakeholders. Rather than building large, monolithic solutions that may miss the mark, teams can deliver smaller increments that can be validated against business objectives and adjusted as needed. Quality assurance throughout the ADLC ensures that analytics initiatives can scale from experimental prototypes to mission-critical business systems without requiring complete rebuilds. This capability is essential for maintaining strategic alignment as business needs evolve and analytics requirements become more demanding. ## Enabling collaboration through data products Strategic alignment requires effective collaboration between data producers and consumers across the organization. [Data products](https://www.getdbt.com/blog/build-trust-in-data-products) provide a framework for structuring this collaboration by treating data sets like software releases, with versioned contracts and clear interfaces. The data product approach enables analytics teams to serve multiple business constituencies simultaneously while maintaining consistency and quality. A single data product might support both business intelligence applications and machine learning initiatives, with each consumer accessing the data through well-defined interfaces that abstract away implementation complexity. This abstraction is crucial for strategic alignment because it allows business teams to focus on outcomes rather than technical details. When data products provide reliable, well-documented interfaces, business users can build applications and analyses without needing deep technical knowledge of underlying data systems. Data products also enable better governance and compliance, which becomes increasingly important as analytics initiatives scale and handle more sensitive data. By building governance capabilities directly into data products, organizations can ensure that strategic analytics initiatives meet regulatory requirements without sacrificing agility or innovation. The collaborative aspects of data products extend beyond technical interfaces to include feedback mechanisms that inform future development. When business users can easily provide feedback on data products, this information flows back into the planning phase of the ADLC, maintaining alignment between technical capabilities and business needs. ## Preparing for AI and advanced analytics The emergence of generative AI and advanced analytics capabilities creates new opportunities for strategic alignment while introducing additional complexity. Data engineering leaders must prepare their organizations to leverage these capabilities while maintaining the governance and quality standards necessary for business-critical applications. The foundation for AI-enabled analytics remains high-quality, well-governed data. Organizations that have invested in mature analytics workflows and data products are better positioned to take advantage of AI capabilities because they have the data infrastructure necessary to support these more demanding applications. However, AI initiatives also require new approaches to collaboration and governance. The non-deterministic nature of many AI systems creates challenges for traditional testing and validation approaches. Organizations must develop new quality assurance practices that can handle the uncertainty inherent in AI-generated outputs while maintaining the reliability necessary for business applications. The role of data engineering leaders in AI initiatives extends beyond providing data infrastructure. They must also help organizations navigate the governance challenges associated with AI, including bias detection, explainability requirements, and compliance with emerging regulations. This requires close collaboration with business stakeholders to understand acceptable risk levels and appropriate use cases. Strategic alignment becomes even more critical in AI initiatives because the potential for both positive and negative business impact is amplified. Organizations that maintain strong alignment between AI capabilities and business strategy are more likely to realize the benefits while avoiding the pitfalls associated with poorly implemented AI systems. ## Measuring and maintaining alignment Successful alignment between analytics initiatives and business strategy requires ongoing measurement and adjustment. Data engineering leaders must establish metrics that capture both technical performance and business impact, creating feedback loops that enable continuous improvement. Technical metrics remain important for ensuring system reliability and performance, but they must be balanced with business outcome measures. Organizations should track metrics like time-to-insight, decision-making velocity, and business impact alongside traditional measures like data quality and system uptime. The measurement framework should also capture the health of collaboration between analytics teams and business stakeholders. Metrics like stakeholder satisfaction, request fulfillment time, and cross-functional project success rates provide insights into how well analytics initiatives are serving business needs. Regular review cycles ensure that alignment is maintained as business priorities evolve. Quarterly business reviews that examine both technical performance and business outcomes create opportunities to adjust analytics initiatives based on changing strategic priorities. These reviews should involve both technical and business leadership to ensure that all perspectives are considered. The goal is not perfect alignment; business needs will always evolve faster than technical capabilities can adapt. Instead, the goal is responsive alignment that enables analytics initiatives to adjust quickly when business priorities change while maintaining the technical excellence necessary for reliable operations. Strategic alignment between analytics initiatives and business objectives represents both a significant challenge and a tremendous opportunity for data engineering leaders. Organizations that successfully achieve this alignment can leverage their data capabilities as true competitive advantages, driving business outcomes while building technical capabilities that scale with organizational needs. The key lies in treating alignment not as a one-time achievement but as an ongoing practice that requires attention to people, processes, and technology in equal measure. ## Analytics and Business strategy FAQs **How will the data and analytics strategy help achieve stakeholders' required business outcomes?** A business-first analytics strategy achieves stakeholder outcomes by beginning with clear business impact objectives rather than technical requirements. When data teams are embedded within specific business units, they develop intimate knowledge of the most critical problems and can identify which data to collect, how to transform it effectively, and how to present insights that directly address business needs. Success is measured through business outcomes like revenue impact, operational efficiency gains, and customer satisfaction improvements rather than solely technical metrics, creating natural alignment between analytics work and organizational priorities. **How do you create a data strategy roadmap?** Creating an effective data strategy roadmap requires implementing the Analytics Development Lifecycle (ADLC), an eight-stage framework that includes Plan, Develop, Test, Deploy, Operate, Observe, Discover, and Analyze phases. The planning phase becomes critical for strategic alignment, with every analytics initiative beginning with a clear business case that articulates expected outcomes and success metrics. The iterative nature of this approach enables rapid feedback cycles between analytics teams and business stakeholders, allowing teams to deliver smaller increments that can be validated against business objectives and adjusted as needed, while maintaining quality assurance throughout the process. --- --- title: "dbt Labs Named Snowflake Monetization Data Cloud Product Partner of the Year" description: "Snowflake selected dbt Labs as a Snowflake partner award winner for the third consecutive year" url: "https://www.getdbt.com/blog/dbt-labs-named-snowflake-monetization-data-cloud-product-partner-of-the-year" date: "2025-06-03" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Named Snowflake Monetization Data Cloud Product Partner of the Year **SAN FRANCISCO, June 3, 2025 — **dbt Labs, the leader in standards for AI-ready structured data, today announced that it has been named the 2025 Snowflake Monetization Data Cloud Product Partner of the Year by [Snowflake](https://www.snowflake.com/), the AI Data Cloud company. The award was announced at Snowflake’s annual user conference, [Snowflake Summit 2025](https://www.snowflake.com/summit/). dbt Labs was recognized for its achievements as part of the Snowflake AI Data Cloud ecosystem, helping joint customers integrate dbt as their data foundation, faster and easier than before. Enterprises are empowered to discover, procure and deploy dbt on Snowflake to maximize the value of their data and AI investments, accelerate productivity across the Analytics Development Lifecycle ([ADLC](https://www.getdbt.com/resources/the-analytics-development-lifecycle)), and position their organizations to scale analytics in the age of AI. “The Snowflake Marketplace is a powerful resource to connect enterprises to dbt, the standard for AI-ready structured data, and its [latest features](https://www.getdbt.com/blog/dbt-labs-redefines-dbt-with-new-fusion-engine), all built to enable customers to bring more data stakeholders into the ADLC and capitalize on the massive opportunity AI is ushering in,” said Shawn Toldo, Vice President, WW Partner Organization at dbt Labs. “Having one of our largest customers purchase dbt Cloud on Snowflake Marketplace last year is among our team's key achievements, and it’s a testament to the strength of our partnership. We are honored to accept this year’s Snowflake Monetization Data Cloud Product Partner of the Year award and are eager to continue our collective work, delivering significant value to our current and future joint customers." This is the third consecutive year that Snowflake selected dbt Labs as a Snowflake partner award winner, illustrating the continued investment in and impact of this longstanding partnership. Together, [dbt and Snowflake](https://www.getdbt.com/data-platforms/snowflake) are committed to helping customers cost-effectively build AI-powered insights and data assets, ultimately improving organizational trust in data and data teams. "We're proud to name dbt Labs as Snowflake's 2025 Monetization Partner of the Year," said Kieran Kennedy, VP, Data Cloud Product Partners, Snowflake. "Their team’s work with the AI Data Cloud ecosystem has delivered remarkable results for our shared customers, allowing for data-driven innovation at scale. This award underscores our mutual dedication to expanding the frontiers of data and AI capabilities." Learn more about dbt Labs** **and Snowflake [here](https://www.getdbt.com/data-platforms/snowflake), and visit the dbt Labs booth (#1808) during this week’s Snowflake Summit to explore the latest dbt innovations, including [the dbt Fusion engine](https://www.getdbt.com/blog/dbt-labs-redefines-dbt-with-new-fusion-engine).** **Check out keynotes from Snowflake Summit 2025 live or on-demand [here](https://www.snowflake.com/summit/?utm_source=pressrelease&utm_medium=partner&utm_campaign=--en-&utm_content=-evv-) and stay on top of the latest news and announcements from Snowflake on [LinkedIn](https://www.linkedin.com/company/3653845/) and [Twitter/X](https://x.com/Snowflake). **About dbt Labs** Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 60,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at [getdbt.com](http://getdbt.com), and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "Houston Food Bank improves fundraising and donor outreach by modernizing its data layer" description: "How Houston Food Bank scaled impact and unlocked insights with dbt—modernizing data to better serve 18 counties." url: "https://www.getdbt.com/blog/houston-food-bank-dbt" date: "2025-06-03" authors: ["Hrishi Kulkarni"] categories: ["Product"] --- # Houston Food Bank improves fundraising and donor outreach by modernizing its data layer The Houston Food Bank (HFB) is the largest food bank in the U.S. Each year, it serves over 140 million meals across 18 counties in southeast Texas. To do that, it works with an extensive network of over 1,600 community partners in lockstep with their partner food banks. It’s a massive operation, and it requires a lot of data. But up until recently, HFB relied on siloed systems and manual processes to manage it. Understandably, the data team struggled to scale their efforts, surface strategic insights, or innovate beyond day-to-day operations. Let’s explore how the data team modernized their data infrastructure using dbt. We’ll see how the journey unlocked critical insights—while increasing collaboration, trust, and governance across HFB. ## The problem: unscalable, siloed systems Before adopting dbt, HFB’s data environment existed in fragments: every department tracked its KPIs in its own spreadsheets, and data transformations were stored across different tools, systems, and formats. Meanwhile, the data team was stretched thin. There were only two dedicated data professionals, who were bogged down with basic reporting. Without a unified architecture, they couldn’t offer more strategic support, like identifying high-value donors or measuring impact. When the COVID-19 pandemic hit, HFB realized it had a data problem. As demand for services surged, HFB’s executive leadership needed real-time visibility into operations, but there was no way to view the organization’s core KPIs in one place. In response, the data team rapidly built stopgap infrastructure to meet emergency needs. It was a turning point: HFB saw the value of data infrastructure and committed to a broader digital transformation. ## The solution: a unified data architecture with dbt After building integrations and a data warehouse, the team struggled with a critical gap: managing and automating data once it was in the warehouse. That’s where [dbt Core](https://www.getdbt.com/product/what-is-dbt) came in. The team migrated all of their queries to dbt Core, which allowed them to deploy and orchestrate data models alongside their integration infrastructure. “For the first time, almost all of our SQL code was under version control via GitLab,” says Herndon-Miller. “dbt Core empowered us to implement software engineering best practices for our data pipelines.” Two years later, the team officially made the jump to [dbt](https://www.getdbt.com/product/dbt-cloud). It’s been a game-changer for collaboration: now, analysts can write, test, and deploy SQL models themselves, without engineering support. Everyone can see what’s changing and why, thanks to dbt’s built-in lineage. “Today, our data team of eight oversees more than 70 reports that deliver more than 180 metrics across the organization,” says Erwin Kristel, Data Analyst at HFB. “dbt improved our ability to build trust with our stakeholders and help them make faster decisions.” ## The impact: data that changes lives More than that, the data transformation has amplified HFB’s capacity to serve. With better visibility into donor behavior, partner activity, and community needs, HFB has unlocked millions in grant funding and identified patterns for reducing food insecurity. ## A centralized dashboard for monetary donors Kristel led the effort to unify volunteer and monetary-donor data, previously spread across multiple unintegrated systems. To support the fundraising team, he transformed that data into an interactive dashboard designed to surface high-potential donors. It was a big initiative to complete this—but it paid off. In just one 30-minute meeting, the fundraising team identified nearly 20 major donors who could contribute $10K-$50K or more but hadn’t been prioritized for outreach. ## Automating essential metrics and reporting Previously, the data team had to manually compile essential metrics for things like grant reports, partner-performance tracking, and resource planning. To automate this process, the data team turned to [dbt Seeds](https://docs.getdbt.com/docs/build/seeds) and modeling. Now HFB is allocating resources more efficiently—and even uncovering new-funding opportunities. “Last year, our community-level partner metrics generated $4 million in three different grants,” shares Susan Quiros, Data Analyst at HFB. “Not every community has the same needs, and now we can create tailored funding strategies that serve them effectively.” ## Measuring the impact of food-benefits programs HFB’s Community Assistance Program (CAP) helps people apply for benefits like SNAP. It’s a critical program, but it was deeply siloed, making it difficult to understand the program’s impact. After integrating CAP data with pantry-usage data, the data team learned that 48% of new SNAP applicants reduced pantry visits within six months. It’s powerful evidence that access to these benefits has an impact on food insecurity. Armed with these insights, the government relations team is better equipped to advocate for SNAP in conversations with policymakers. ## Meeting their neighbors where they are By adopting dbt, the HFB data team has increased data accessibility across the organization. They’ve built trust in metrics at every level: their work has improved donor engagement, identified funding opportunities, and validated the impact of state programs. For HFB, their data transformation has been a force multiplier. dbt is now central to how HFB operates—helping the organization serve the community, one meal at a time. If you’re a data professional at a nonprofit looking to modernize your data stack, we’d be honored to help you build your transformation strategy. [Book a demo](https://www.getdbt.com/contact) to see how dbt works; you can also [sign up for dbt](https://www.getdbt.com/signup) to connect your data warehouse and start building. [Watch video](https://youtu.be/sJg5_kB31Ik?si=lG2_NXtPsxIIcf8v) --- --- title: "How The Philadelphia Inquirer increases productivity and enables self-service with the dbt Semantic Layer" description: "How the Philadelphia Inquirer used the dbt Semantic Layer to enable self-service, reduce errors, and build trust in their data." url: "https://www.getdbt.com/blog/philadelphia-inquirer-dbt-semantic-layer" date: "2025-06-02" authors: ["Hrishi Kulkarni", "Chakshu Mehta"] categories: ["Product"] --- # How The Philadelphia Inquirer increases productivity and enables self-service with the dbt Semantic Layer For the data team at _The Philadelphia Inquirer_, answering simple data questions fast, like, “Can I see a chart for Homepage traffic by day?” is essential for shaping editorial decisions, subscription strategies, and long-term growth initiatives. ![Slide describing a common business challenge: stakeholders frequently ask for homepage traffic dashboards, but delivering insights requires SQL skills and contextual knowledge, including the SQL query to define “Homepage traffic.”](https://cdn.sanity.io/images/wl0ndo6t/main/917e2d8b67a8315d80aa555e4bdfd05a70ba26e3-512x269.png) But getting to those answers was far from simple. To get quick insights, business users relied heavily on the data team to translate their questions into SQL to query the data warehouse. Unfortunately, this led to organizational bottlenecks, inconsistent metrics, and barriers to scaling. Rather than doubling down on custom dashboards, the data team took a new approach. Let’s explore how they used the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) and [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) to enable true data self-service—and what other data teams can learn from their journey. ## Too many dashboards, not enough clarity Years ago, the Inquirer’s data team created denormalized data marts and dashboards, each stitched together with custom SQL. The problem was, the codebase wasn’t flexible or easy to maintain. Even small changes—like adding a new column to filter data—required going back into the SQL models, modifying the code, and deploying changes. Core metrics like “Total Pageviews” or “subscriptions” could be calculated in different ways (i.e. sum vs count), leading to confusion and a loss of trust in the data. ![Slide illustrating the problem of inconsistent data calculations, showing two SQL queries using different tables and methods (fct_events and fct_sessions) to calculate “Total Pageviews,” leading to discrepancies and reduced trust in data quality.](https://cdn.sanity.io/images/wl0ndo6t/main/a0e0c3f2a77ec9b223df0c104503831633cb7cc9-512x277.png) “Our denormalized layer became the most brittle part of our system,” says Brian Waligorski, Lead Data Engineer at _The Philadelphia Inquirer_. “Copying and pasting large blocks of SQL across models opened the door to countless errors and inconsistencies.” Without a centralized source of truth, business users often resorted to exporting data into their own spreadsheets to do their calculations manually, and even added their own business logic. The result was fragmented reporting and misaligned metrics. Ultimately, the process was a drain on resources. “We’ve all seen the dreaded after-hours Slack ping to ‘pull a number real fast,’” comments Waligorski. “But every message interrupts what you’re doing to dive into SQL troubleshooting. We spent a lot of time reacting, instead of focusing on strategic work.” ## A single source of truth for scalable self-serve To break out of this cycle of ad-hoc requests and inconsistent reporting, the data team needed to centralize metrics to enable true self-service. “Self-serve doesn’t just mean ‘analysts building dashboards, faster,’” Waligorski emphasizes. “It’s gaining direct access to trusted data faster. A solution like the dbt Semantic Layer reduces organizational bottlenecks and empowers users to interact with the data directly without the wait.” By adopting the [dbt Semantic Layer](https://www.getdbt.com/blog/semantic-layer-introduction), the team built a centralized, governed single source of truth for their business metrics and logic that powers self-service across the organization tooling and systems. Here’s how: - **Centralized metric definitions.** Metrics and business logic are now defined directly in dbt, right alongside the data models they rely on. This eliminates metric and logic discrepancies while simplifying maintenance: update a metric definition once, and it’s updated everywhere. - **Standardized naming conventions.** The data team partnered with key stakeholders to define clear, consistent naming conventions and canonical metrics like “total users” vs “total viewers.” to reduce confusion and reporting discrepancies. - **Enabled self-service across every tool.** The team integrated metrics into tools that meet users where they already work—including Steep (for product managers), Hex (for analysts), Google Sheets (for the finance team), and Exports/Saved Queries (for traditional BI consumption). “Now, with the dbt Semantic Layer, we have a single, enforceable source of truth for our business metrics and logic,” says Waligorski. “It’s like a holy grail of analytics engineering, and it’s possible because of dbt.” ![Diagram showing data flow from raw data in a cloud data platform through transformations and models, then unified by the dbt Semantic Layer, enabling analytics tools like BI, AI/LLMs, notebooks, and exports to access consistent data.](https://cdn.sanity.io/images/wl0ndo6t/main/0327f97679d834efcfbd2c98b2dc7df96a74556d-512x371.png) ## Faster insights, greater trust By implementing the dbt Semantic Layer, the data team at The Philadelphia Inquirer has become more productive than ever. For example, when they get a request, analysts no longer have to rewrite SQL or reverse-engineer metric logic. They can simply pull trusted metrics from their suite of governed tools, speeding up delivery and reducing friction. “With the dbt Semantic Layer, our time-to-delivery for dashboards has gone down significantly,” reports Waligorski. “By reducing the back-and-forth between data engineers and analysts, we’ve become more efficient.” The dbt Semantic Layer has also unlocked a new level of flexibility. Because it’s tool-agnostic—with 10+ out-of-the-box integrations and a robust API—the data team can deliver metrics wherever stakeholders need them. Rather than being tied to a single BI platform, the data team can deliver insights in whatever format best suits each stakeholder, whether that’s spreadsheets, visual dashboards, or notebooks. The Inquirer’s data team also saw a significant drop in metric errors and inconsistencies. Before, different teams often reported different numbers for the same metric, causing confusion, duplication of work, and long Slack threads trying to reconcile the truth. Stakeholders can finally rely on the numbers they see, increasing confidence in decision-making, trust in the data, and alignment across teams. ## Takeaways and the road ahead After a year of working with the dbt Semantic Layer, Waligorski has four pieces of advice for data teams looking to follow a similar path: 1. **Normalize your models.** Denormalized models limit flexibility—but refactoring for modularity pays dividends in the long run. 2. **Avoid redundant joins.** Let MetricFlow handle join logic and transformations to reduce code duplication. (But allow exceptions for specific, mission-critical edge cases.) 3. **Be deliberate with naming.** Agree on metric and business terminology early. It’s an investment of cross-functional work, but it will prevent confusion and rework later. 4. **Start small.** Pilot with key use cases, prove value, then expand adoption across departments and tools. ![Slide outlining next steps: using AI chatbots to answer business questions with a set of baseline queries, and enabling metric observability and anomaly detection to monitor when key metrics spike or dip and understand why.](https://cdn.sanity.io/images/wl0ndo6t/main/7f8162c90fe0e35fbbdb12570cddc83cd31ce23b-512x248.png) Looking ahead, the team is getting ready to launch AI chatbots for real-time metric queries, making it even easier for stakeholders to get answers faster directly in their AI systems. The Inquirer is also exploring anomaly detection and metric observability to proactively understand when key metrics spike or dip and why. “Today, we can pull a strong set of metrics with low overhead with the dbt Semantic Layer,” says Waligorski. “That has freed up our capacity to focus on high-impact work and scale the business.” Check out the full story on YouTube. [Watch video](https://youtu.be/vsJOtQkmlxw?si=5DNhOMBe-9X4KCc9) Curious how the dbt Semantic Layer can empower your team to move faster and trust their metrics?, [Book a demo](https://www.getdbt.com/contact). We’d love to answer any questions about what dbt can do. You can also [sign up for dbt](https://www.getdbt.com/signup) to connect your data warehouse and start building today. --- --- title: "Unlocking analyst-driven data transformation: dbt Canvas is GA" description: "Empower analysts without sacrificing governance. Welcome to the ADLC on your terms." url: "https://www.getdbt.com/blog/dbt-canvas-is-ga" date: "2025-05-28" authors: ["Greg McKeon", "Sara Gawlinski"] categories: ["Product"] --- # Unlocking analyst-driven data transformation: dbt Canvas is GA Today, dbt Canvas, an all-new, AI-powered visual editing experience in dbt is generally available. With dbt Canvas, data analysts can finally take an active role in data transformation. They can use an intuitive drag-and-drop interface that’s seamlessly integrated into the governed, scalable [analytics development lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle) and powered by context-aware AI. It empowers teams to move faster, reduce bottlenecks, and collaborate more effectively without compromising on quality, trust, or version control. [Watch video](https://youtu.be/pO_TUnCt6es) You can [get started with dbt Canvas today](https://docs.getdbt.com/docs/cloud/canvas)—just open a project and start building. ## Making analytics a team sport with dbt Canvas At dbt Labs, we believe data is a team sport. For nearly a decade, dbt has helped organizations bring software engineering best practices like version control, testing, and CI/CD to analytics engineering. But that was just the beginning. The full promise of data can only be realized when all data practitioners, not just data and analytics engineers, can contribute to their organization’s data estate in an accessible, yet governed way. dbt Canvas, now GA, makes this promise a reality. ## Self-service analytics requires more than just access to data; it requires governance For years, data leaders have chased the promise of self-service analytics. But true self-service doesn’t mean simply handing analysts access to datasets or procuring visualization tools. It means enabling them to contribute to a mature analytics workflow—the ADLC—where data is shaped, defined, and prepared for the entire organization. Data analysts are closer to the business questions than anyone else on the team. They know the “why” behind the metrics. They understand the quirks in the data. They have a deep understanding of evolving business strategies and goals. But when it comes to contributing directly to data transformation, they’re often blocked. Why? Because data teams are overwhelmed. Engineers are buried in tickets and ad-hoc requests. Analysts wait in queues. To move faster, they turn to ungoverned low-code editors, SQL editors, BI tools, or scripting in isolation, creating shadow pipelines. We’ve heard the frustration loud and clear: Data teams need a way for analysts to safely, confidently, and effectively contribute to the data transformation process without sacrificing governance or quality. And in the AI era, analysts need trustworthy AI that's powered by the full context of their data assets to contribute faster. ## Empower analysts to build with confidence and control dbt Canvas is an AI-powered visual experience built for analysts. It enables data practitioners across a spectrum of technical skills to build and collaborate on data transformations without needing to write SQL from scratch. > > > — James Wright, InterWorks ### Build and edit dbt models visually - no SQL required With dbt Canvas, analysts start by exploring the input data that is available, whether it’s curated models or sources imported from the warehouse, or an uploaded CSV file. They can drag and drop these data sets directly onto the canvas, where they can preview the data and associated output columns. Analysts may not know the data that’s available to them. Canvas includes robust search and discovery to find existing datasets and always-on data profiling capabilities that help you understand that data. From there, analysts can apply transformations. Common operations like joins, aggregations, filters, calculated columns (formulas), sorts, and unions are all available as modular building blocks. Each transformation is visually represented as a node in the canvas, making it simple to trace logic, understand relationships, and collaborate with teammates. As analysts build out a model, they can preview the data output at every stage, seeing exactly how each transformation impacts the dataset. This step-by-step feedback loop builds confidence, reduces errors, and accelerates iteration. What makes dbt Canvas truly powerful is what’s happening under the hood. _Every _transformation, filter, and calculation built in Canvas is automatically compiled into SQL that adheres to your organization’s dbt project structure and coding standards. This means the visual logic analysts create isn’t locked into a proprietary format, it becomes production-grade dbt code that’s version-controlled, testable, and reviewable just like any other model in your dbt project. ### Built-in governance with Git and PR workflows Every model built in Canvas can be committed to your dbt project using the same version control workflows you already use in dbt Studio or the VS Code extension. Even if you’re not a git expert, Canvas makes it simple: click to commit, and submit a pull request for review. Analysts become first-class contributors to the analytics engineering workflow - with all the safeguards teams require. ### Accelerate development with dbt Copilot Canvas includes dbt Copilot, our built-in, context-aware AI assistant, which can help generate SQL expressions in formula nodes, explain specific steps in the workflow, modify existing models, or create brand-new models from scratch—just by describing what you want in natural language. Copilot captures project context, like model relationships, tested logic, and naming conventions, from your dbt project to power Canvas, helping analysts build more accurate, trusted models using natural language and visual workflows Whether you’re exploring a data source or defining a new KPI, Copilot helps you move from idea to model faster than ever. Whether you're a data analyst building your first model or a senior analytics engineer reviewing a pull request, dbt Canvas ensures that what’s created visually can be integrated directly into your governed analytics workflow. And because everything in Canvas aligns with dbt’s underlying principles—modularity, reusability, and transparency—teams get the best of both worlds: accessibility for analysts, and accountability for engineering. ## Coming soon to dbt Canvas We’re just getting started. Here’s what’s next: - Support for additional SQL operations, including window functions and advanced filter expressions - Connect and import data to the warehouse from other sources, like Google Sheets. - Define dbt tests and unit tests directly in Canvas ## Making dbt the standard for analysts dbt Canvas is part of a growing set of new AI-powered self-service capabilities we’re launching to empower analysts to participate in dbt's governed workflows with confidence and clarity. New capabilities include: - [dbt Insights](https://docs.getdbt.com/docs/explore/dbt-insights) (now in Preview) helps analysts ask any ad-hoc question and get an answer fast from governed data, all within dbt. Using SQL or natural language prompts, analysts can create queries, validate models, generate visualizations, and easily share findings, all without writing complex code. Transparent governance and full lineage ensure trust and deliver high‑quality insights in one seamless workflow. With Insights, you can: - Write, run, validate, and iterate on SQL queries without switching tools using features like dbt Copilot, syntax highlighting, and query history. - Leverage dbt metadata, trust signals, and model lineage to write better, faster queries and quickly identify issues and root causes. - Make data accessible to a wide range of users, from SQL pros to AI-assisted analysts, using both code and context-aware AI. - Go from question to model faster by validating ideas before formalizing them in dbt Canvas or the dbt Studio (IDE). - Quickly explore patterns, trends, or anomalies in your data with built-in visualizations, no need to export to external tools - New functionality in [dbt Catalog](https://docs.getdbt.com/docs/explore/explore-projects) (formerly known as dbt Explorer) allows users to [go beyond their dbt estate](https://docs.getdbt.com/docs/explore/external-metadata-ingestion), and search and explore Snowflake assets—like tables and views—directly within dbt. Now, analysts and developers can build holistic context for their data estate without switching platforms or tabs. Together, these new platform experiences reflect our belief that the future of analytics engineering is collaborative, governed, and accessible to every data practitioner. Not only are we introducing new capabilities, but we’re also making dbt Cloud more accessible than ever with our new flexible pricing. This approach is designed to make it easy to expand access to your entire analyst team. ## Get started today Check out the below resources to get your teams started with Canvas: - [Explore the docs](https://docs.getdbt.com/docs/cloud/canvas) - Take the [dbt Canvas Fundamentals](https://learn.getdbt.com/learn/course/canvas-fundamentals/welcome-to-canvas-fundamentals/welcome-20-mins?client=internal) training course to ramp up fast - Learn more about [dbt for analysts](https://www.getdbt.com/product/analyst) - [Join our webinar](https://www.getdbt.com/resources/webinars/empowering-data-analysts-showcase-series-part-one) in late June for an in-depth look at the new features tailor-made for analysts --- --- title: "Get to know the new dbt Fusion engine and VS Code Extension" description: "The all-new dbt Fusion engine and the dbt VS Code Extension are now in public beta." url: "https://www.getdbt.com/blog/get-to-know-the-new-dbt-fusion-engine-and-vs-code-extension" date: "2025-05-28" authors: ["Elias DeFaria", "Azzam Aijazi"] categories: ["Product"] --- # Get to know the new dbt Fusion engine and VS Code Extension Today, at our annual [dbt Launch Showcase](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase), we announced the public beta release of the all-new [dbt Fusion engine](https://www.getdbt.com/product/fusion), available to [eligible Snowflake projects](https://docs.getdbt.com/docs/fusion/supported-features)**.** We also released the official [dbt VS Code Extension](http://docs.getdbt.com/docs/install-dbt-extension) in public beta. Together, these represent a huge leap forward for everyone building with dbt. Fusion isn’t just another feature. It’s the [foundation for a new era of analytics engineering](https://getdbt.com/blog/where-we-re-headed-with-the-dbt-fusion-engine), ushering in a whole new dbt experience that is far faster, more intelligent, and more cost-efficient than ever before. [Watch video](https://youtu.be/NiNkdThkKAI) ### Why now? Since its launch in 2016, dbt Core has revolutionized the way data teams approach data transformation, bringing software engineering principles to analytics workflows for the first time. But the world has changed in the ensuing nine (!) years. The modern data stack is no longer novel, storage and compute are cheaper than ever, and AI is reshaping how we interact with data. Meanwhile, even as analytics has become a team sport—with developers, analysts, and business stakeholders working more closely than ever—the bar for what’s possible in a first-class data development experience has continued to rise. dbt is ready for an update. That’s why we [we acquired SDF Labs earlier this year](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs), and why we’ve incorporated SDF’s technology into the heart of dbt to create the new dbt Fusion engine. ## Fusion: A new engine for a new era Fusion is a fundamental shift in how dbt understands and executes your analytics code. It introduces **three key new technological breakthroughs** to dbt: ### Lightning-fast performance Performance isn’t just a nice-to-have. It’s the backbone of rapid, high-trust data development. Fusion is built in Rust, a systems-level language known for performance, and optimized to parse even the largest dbt projects 30x faster than dbt Core. This allows developers to iterate rapidly, stay in flow longer, and ultimately, ship high-quality data products faster. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/3b9c198a1be26f7be18a25d71a8eca2153d2f1c2-1225x1224.png) ### Native SQL comprehension Fusion represents a fundamental change to how dbt operates: dbt is no longer just _passing along_ your code to the data warehouse. Now, dbt is also able to‌ understand that code. With this understanding, dbt can emulate your cloud data warehouse locally and reason about your code with the [intelligence of a modern compiler](https://docs.getdbt.com/blog/the-levels-of-sql-comprehension). It’s a major shift that enables powerful new capabilities—like validating your code in real time, without ever needing to hit the warehouse. ### State-awareness Fusion introduces a first-of-its-kind understanding of your dbt project’s state: both what exists in your codebase, and what’s already been materialized in your warehouse. This means that dbt can now dynamically make determinations about, for example, which models truly need to be rebuilt, as opposed to which ones have already been built very recently. This results in higher velocity pipelines and built-in cost savings. ## What this means for data teams The above three technological breakthroughs translate directly into tangible benefits for your organization: faster development cycles, lower warehouse spend, and stronger confidence in the quality and governance of your data products. Here’s what that looks like in practice: ### A hyper-responsive development experience Fusion is designed to allow developers to move faster, and stay in flow for longer. With SQL comprehension built into the engine, you get real-time, context-aware assistance as you write code. **IntelliSense** offers smart autocompletion for models, columns, macros, and functions. **Hover previews** surface column types and schema details. Need to rename a model or column? **Automatic refactoring** updates references across your entire project instantly. With **go-to-definition** and **inline CTE previews**, navigating large projects is seamless and debugging is faster than ever. > > > — James Dorado, Bilt Rewards ### Built-in cost savings Fusion understands your project’s lineage and database state in real-time. This unlocks [**state-aware orchestration**](http://docs.getdbt.com/docs/deploy/state-aware-about): only run models when upstream data has changed. No more wasteful rebuilds. We expect customers to see a **~10% reduction in warehouse spend** driven by this, with even greater savings expected as new state-aware capabilities are launched in the coming months. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/6493dd677fca3d14f1cc3b6f6f0f8285774551d0-1196x543.png) > > > — Matt Karan, Obie Insurance ### Improved lineage and data governance Fusion’s native SQL comprehension will unlock true, precise column-level lineage. This enhances impact analysis, helps reduce risk of compliance issues, and lays the groundwork for **built-in PII classification and policy enforcement** (coming soon). The kind of robust data governance this will unlock is particularly pivotal to a safe, effective AI strategy. ## Available now: The official dbt VS Code Extension The dbt Fusion engine will power every development experience in the dbt platform. That said, we know many of you do your best work locally in the familiar confines of VS Code (or VS Code based IDEs such as Cursor.) That’s why we’re excited to launch an official [dbt VS Code Extension](http://docs.getdbt.com/docs/install-dbt-extension) (now in public beta)—built by dbt Labs, and optimized for dbt. The dbt VS Code Extension brings all the development experience enhancements of Fusion directly to VS Code. It’s the only way to unleash the full power of the dbt Fusion engine while developing locally. > > > — Bruno Souza de Lima, pHData Now, regardless of how your team chooses to contribute to the analytics loop in dbt—whether it’s the [dbt Studio IDE](https://www.getdbt.com/product/develop), [dbt Canvas](https://getdbt.com/blog/dbt-canvas-is-ga), [dbt Insights](https://docs.getdbt.com/docs/explore/dbt-insights), or locally with the new VS Code Extension—they can do so on one powerful, shared engine: Fusion. Note that to use the VS Code Extension, your project must be running on the Fusion engine. The extension does not support dbt Core, as the key developer experience enhancements it’s designed to enable rely on the technological foundations of ‌the new engine. ## What’s available today? ✅ **Fusion engine**: Available in public beta for [eligible projects](https://docs.getdbt.com/docs/fusion/supported-features) on Snowflake (with support for more projects and data platforms coming soon). If your projects are eligible, you'll soon see a notification when you log in to dbt. ✅ **Official** **dbt VS Code Extension**: Now available via the [VS Code Marketplace](http://docs.getdbt.com/docs/install-dbt-extension). Requires use of the dbt Fusion engine, although you do not need to be a customer of the commercial dbt platform. ✅ **State-aware orchestration**: Unlock ~10% warehouse cost savings with smarter, more efficient job runs. [Automatically enabled](http://docs.getdbt.com/docs/deploy/state-aware-about) for customers on Enterprise plans running on the dbt Fusion engine. ## How to get started If your [dbt project is eligible](https://docs.getdbt.com/docs/fusion/supported-features) and you’re a dbt customer, you’ll soon see a notification in the dbt UI to begin the process of migrating to Fusion. If your project is not yet eligible, fret not: you can [follow these steps](https://www.getdbt.com/blog/how-to-get-ready-for-the-new-dbt-engine) to get prepared in the meantime. We’re rapidly expanding support, and it shouldn’t be long before your project can tap into the Fusion engine. If you’d like to begin trying Fusion immediately, or if you’re not a customer, you can [install dbt Fusion locally](http://docs.getdbt.com/docs/fusion/install-fusion) and [use the VS Code Extension](http://docs.getdbt.com/docs/install-dbt-extension) today. You can also explore our new [Fusion Quickstart](https://docs.getdbt.com/guides/fusion) to experiment with a sample project running on Fusion. ## What’s next? Fusion is the future, and this is just the beginning. Soon, you’ll see: - **Support for more projects and more data platforms**. Our team is working diligently to make Fusion available to more teams and users as a top priority in our [approach to GA](https://docs.getdbt.com/blog/dbt-fusion-engine-path-to-ga). Hang on tight. - **Even smarter CI/CD** with column-aware change detection. Expect more cost savings in the near future, powered by the Fusion engine’s state-awareness. - **Built-in data governance** with support for tagging of PII/PHI and enforcement of data policies. But that's still just scratching the surface. The longer-term implications of the dbt Fusion engine across AI and data infrastructure go far beyond these capabilities. Read Tristan’s blog post on [where we’re headed with the new dbt Fusion engine](https://getdbt.com/blog/where-we-re-headed-with-the-dbt-fusion-engine) for more. --- --- title: "Where we're headed with the dbt Fusion engine" description: "Paving the way for the next era" url: "https://www.getdbt.com/blog/where-we-re-headed-with-the-dbt-fusion-engine" date: "2025-05-28" authors: ["Tristan Handy"] categories: ["Product"] --- # Where we're headed with the dbt Fusion engine Today, we [launched the dbt Fusion engine](https://getdbt.com/blog/get-to-know-the-new-dbt-fusion-engine-and-vs-code-extension), a complete rewrite of dbt from the ground up. If you’d like more information on what just launched check out these posts: - [What is Fusion and how does it help?](https://docs.getdbt.com/blog/dbt-fusion-engine) - [New Code, New License](https://www.getdbt.com/blog/new-code-new-license-understanding-the-new-license-for-the-dbt-fusion-engine) - [Technical Components and Licensing](https://docs.getdbt.com/blog/dbt-fusion-engine-components) - [Current Maturity and Path to GA](https://docs.getdbt.com/blog/dbt-fusion-engine-path-to-ga) In this post, though, I want to talk about the future. **What does this complete technical overhaul say about the future of dbt?** Let’s dig in. ## Ripping out an engine is hard! In the world of commercial open source infrastructure software, there is an emerging trend: rebuilding the engine at the heart of the platform. Databricks did this with [Photon](https://www.databricks.com/blog/2022/08/03/announcing-photon-engine-general-availability-on-the-databricks-lakehouse-platform.html). Photon is a complete rewrite of the Apache Spark engine in C++. Before Photon, Databricks was a really nice way to run Apache Spark in the cloud. Now, Databricks delivers meaningful benefits that are simply not possible with “vanilla Spark,” including significant price and performance gains. Confluent did this with [Kora](https://www.confluent.io/blog/cloud-native-data-streaming-kafka-engine/). Kora is a protocol-compatible rewrite of the Apache Kafka engine. Before Kora, Confluent was a really nice way to run Apache Kafka in the cloud. Now, Confluent can rightfully claim meaningful benefits that are simply not possible with “vanilla Kafka,” including significant price, performance, and reliability gains. There are others. MongoDB’s launch of Atlas comes to mind. These decisions are fascinating to me as the founder of a commercial open source business. In each of these cases, these companies had to say: “What got us here won’t get us there.” And I don’t think Databricks, Confluent, or Mongo would be the success stories they are today without these investments. But they are _incredibly hard to execute on_. The demands of growth—thousands of customers, different segments, different user profiles, a fast-moving ecosystem, etc.—require so much attention that it is incredibly hard to say “we’re going to take 1-3 years and rebuild the engine that everything else is built on.” But without making these kinds of investments, a commercial OSS business is unlikely to be successful over the long term. This thought has been bouncing around in my head for several years, and it’s been pretty clear to me that we were going to have to make this leap as well. There were simply aspects of the dbt Core code base—tracking back from 2016!—that were not able to get us to the future, that we couldn’t iterate our way through. Performance. Functionality. Etc. There was just no path towards the world we want to build without rebuilding the foundations. We were on our own internal path towards this rebuild when I originally met [Lukas](https://www.linkedin.com/in/lukas-schulte-a6b16254/) and got to hear what he and the SDF team were up to. This [match made in heaven](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs) has allowed us to move 1-2 years faster than we otherwise would’ve been able to. I feel as though I’m shipping tech today that was brought from the future into the present in a time machine. So: now that the dbt Fusion engine is live in the wild, where are we headed? What does this unlock for the future of dbt? ## What the dbt Fusion engine unlocks The below are the medium-term (12-24 month) directions that the dbt Fusion engine will allow us to innovate in. While I’m not here to commit to specific dates, you should expect that we have a direct line-of-sight to all of these themes based on the state of the underlying Fusion technology. ### Parse & compilation times **The new dbt Fusion engine does a more advanced parse and compile than dbt Core, and even so is already [around 30x faster ](https://docs.getdbt.com/blog/faster-project-parsing-with-rust)to parse and substantially faster to compile.** Parse and compilation times are incredibly important for any piece of developer tooling. The faster they are, the more useful that developer tooling becomes. Pre-Fusion, dbt’s parse times have been just barely fast enough to support traditional developer workflows. Even for this use case they are a pain point in larger projects. But imagine other use cases that have not been supportable by dbt because of parse times: - [Agent-based chat experiences with MCP](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) where an agent iteratively writes and tests code based on user requests. Current parse times are just too slow. - Better developer experiences that require recompile-on-keypress (more below). - Ever-larger dbt projects authored by ever-larger teams enabled by ever-more-accessible authoring experiences. …and more that we cannot even anticipate. In general, every time you improve the performance of a developer tool by an order of magnitude, you discover use cases that were totally unanticipated. (Developer creativity is a beautiful thing!) In the future, you should expect parse and compile times to continue to drop, as there is now significantly more headroom for optimization within Fusion. You can already see this inside of the VS Code extension, where we can incrementally recompile a single file in milliseconds. ### Improved developer experience Developer experience isn’t just about making developers happier: it is about making developers _more productive_. And Fusion, along with the new dbt Language Server and VS Code extension, delivers. Historically, dbt developers were left to actually execute dbt in order to find errors in their code. Even in the best of worlds, that is a cumbersome process, and the performance challenges of dbt Core only made the problem worse. Now errors show up as you type—and they are significantly more complete and detailed. dbt will now validate not only your dbt code, but your SQL as well—not just coarse-grained things like function parameters but fine-grained things like type checking. And dbt will now not only validate the model you’re developing in; it will, as you type, validate downstream dependencies as well and surface those errors. Add to this features like auto-refactoring, auto-renaming, go-to-source, and more and the tasks you do day-in, day-out just got a _whole lot more efficient_. ### Local execution for development environments The original thesis for dbt, authored in 2016, was that data practitioners should adopt the tooling and best practices of software engineers. And much of that has come to pass over the past almost-decade. But one of the ways that this has definitely NOT come to pass is local development environments. It is generally superior, especially in a post-Docker world, to have development environments that can run locally end-to-end. Local development environments are more amenable to building great tooling, are more customizable, reduce latency, and eliminate a source of cost. And while software engineers have been doing this for a long time—packaging Postgres locally and pointing to RDS in production—data practitioners operating in the cloud simply couldn’t do that. You can’t run BigQuery or Snowflake or Athena or Databricks (etc.) on your local machine. Until now. The dbt Fusion engine can fully emulate—with consistency guarantees—the functionality of the underlying data platform and allow all developers to execute their code locally, without ever reaching out to the remote data platform. This is not “best guess” emulation; this is Fusion fully emulating, [down to the logical plan level](https://www.getdbt.com/blog/how-logical-plan-impact-modern-data-workflows), the exact behavior of the underlying platform. When paired with [dbt’s existing data sampling functionality](https://docs.getdbt.com/docs/build/sample-flag), this will be a dramatic upgrade in the dbt experience. While we’re not ready to share when this will ship to users, this type of functionality is something we are already playing with internally. ### Cost savings The dbt Fusion engine gives us huge scope to mitigate costs for users. In fact, our goal will be that adopting the full capabilities of Fusion will be ‘cheaper than free’ due to the savings it will create. For almost all companies that use dbt, it is the single largest driver of consumption on their underlying cloud data platform. dbt makes it easy to author data pipelines, and data pipelines can be expensive to execute. Companies have historically only really had two options: a) accept this reality and do their best to optimize, or b) limit the number of humans who had access to author pipelines. Neither are good answers. With Fusion, dbt will be able to automatically optimize your pipelines and orchestration to allow them to simply cost less to execute. As of today, dbt Enterprise customers on Fusion get access to a feature we are calling [state-aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-about) which is expected to result in an average of 10% cost reduction, simply by turning the feature on and letting Fusion optimize how and when jobs are run. In a typical customer spending $1 million on their underlying platform, dbt can often be 50% of that, so this 10% of the $500k of dbt-driven data platform spend represents a $50k annual cost savings simply by standardizing on Fusion. This is only the first tranche of cost saving strategies unlocked by Fusion. In the future, we anticipate this number growing significantly beyond 10%. ### PII, governance, and lineage Tracking the flow of PII in a sufficiently-complex data ecosystem is a very challenging problem when you stare at it long enough. At first it would seem easy: tag all the source data, and make sure you have column-level lineage. As it turns out, that is not sufficient. PII is often transformed into non-PII with certain predictable transformations. Imagine running a `count(user.email)`: the resulting column is based on PII, but it is not, itself, PII. If you provide a false positive on downstream uses of the resulting column, users will lose trust in the system and begin to ignore it. This may seem like a small thing, but in environments with tens of thousands to millions of tables, the complexity of these seemingly-simple things quickly spirals out of control. The technology behind the dbt Fusion engine was originally built to serve exactly this requirement inside one of the most complex data ecosystems on the planet: Meta’s internal data warehouse. In the not-too-distant future, we’ll be offering the ability to show an audit-ready view of your PII footprint across your entire data landscape, all powered by Fusion. ### Cross-platform workload portability I do not believe that we are in an ecosystem that will ultimately be dominated by a single large player. This is not Windows in the 90’s (monopoly). It is not even Mobile in the 2010’s (duopoly). Instead, there is a group of 5-10 major vendors, and major platforms, that dominate the data ecosystem (oligopoly). Not 50-100! But 5-10. And I think this will be a persistent fact. Within those 5-10 vendors, nearly all customers will require flexibility. It creates exactly zero enterprise value to migrate code from one platform to another. No CIO / CDO wants to spend time migrating or replatforming. And so they tend to stay on older, more established technologies for longer than they would want to avoid having to pay this cost. This desire for cross-platform flexibility is exactly what is driving the massive demand for Iceberg and other open-table formats. With the dbt Fusion engine, in the future, dbt-authored pipelines will never need to be migrated again, for any data platform that Fusion supports. Just as Fusion can emulate the underlying data platform with complete type-aware fidelity to support local compute environments, it will be able to use that logical plan to perform real-time, guaranteed-faithful SQL transpilation. This allows SQL written in one dialect to be automatically ported into others—at runtime. When paired with Iceberg, this capability represents the future of enterprise cross-platform flexibility. This capability is likely a bit further away, but it is nonetheless very real and unlocked by the dbt Fusion engine. ## Just the beginning… And these benefits are just the beginning of what Fusion will bring to the dbt Community in the coming years. In a period of such intense change (both at dbt Labs and in the technology industry more broadly), predicting the future can be challenging. But I am incredibly confident that Fusion—the new technical foundation that we are building the entire dbt platform on top of—will help us accelerate into that future. --- --- title: "dbt Labs Launches AI-Powered Features to Onboard Data Analysts into dbt" description: "Analysts can now build, explore, and validate models leveraging the power of dbt." url: "https://www.getdbt.com/blog/dbt-labs-launches-ai-powered-features-to-onboard-data-analysts-into-dbt" date: "2025-05-28" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Launches AI-Powered Features to Onboard Data Analysts into dbt #### Analysts can now build, explore, and validate models leveraging the power of dbt **PHILADELPHIA, May 28, 2025** – [dbt Labs](http://www.getdbt.com), the leader in standards for AI-ready structured data, today announced a powerful new suite of AI-enhanced features that give data analysts a fast and governed way to explore data and deliver insights within dbt's workflows. These new capabilities empower analysts across a range of technical backgrounds to lean on natural language or visual interfaces to build, explore and validate data in the same version-controlled environment trusted by data teams. This release includes dbt Canvas (a visual, drag-and-drop interface for model development), dbt Insights (an AI-powered query tool for quick analysis and sharing), and an enhanced dbt Catalog (for global asset discovery). Additionally, organizations can now use the new cost management dashboard to optimize their data warehouse spend. ## Bridging the gap between self-service and governance [Gartner® predicts](https://www.gartner.com/en/data-analytics/topics/data-governance) that, “by 2027, 60% of organizations will fail to realize the full value of their AI use cases due to fragmented data governance frameworks.”* One contributing factor is the rise of ungoverned data workflows, often driven by analysts working around limited engineering support. To get the insights they need, data analysts rely on unsupported, disconnected tools and un-tested, bespoke logic to build, query, and explore data, leading to compliance risks, increased costs, and poor data quality that undermine organizational decision making. dbt’s new AI-powered capabilities are purpose-built to solve this issue by giving analysts greater autonomy while ensuring every action remains governed, version-controlled and aligned with organizational data standards. "Data teams today face a fundamental tension – analysts need speed and independence, while organizations require strong governance and security," said Tristan Handy, founder and CEO of dbt Labs. "Our new AI-powered solutions break down these traditional barriers for data analysts across any skill level and collaborate with developers in the same platform, which will have a significant, positive impact throughout the business." ## Unlocking trusted self-service for analysts with dbt The [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) is a vendor-agnostic framework that helps organizations mature how they build, maintain, and scale trusted data products. As the data control plane for the modern enterprise, dbt brings the ADLC to life, enabling version-controlled, governed workflows that power analytics across teams. dbt Labs is now making it easy for downstream analysts to participate in the ADLC with the following new capabilities: - **dbt Canvas**, a new visual editing environment in dbt, enables users more comfortable with drag-and-drop tooling to build and edit data models. Analysts can describe what they want to build in natural language using dbt Copilot, allowing teams with limited SQL knowledge to build effective data models using context-rich AI. It automatically maintains governance and quality standards, while reducing reliance on data engineers, boosting collaboration and improving productivity. dbt Canvas is now GA. - **dbt Insights**, a new AI-powered query interface that helps analysts ask questions and get answers faster, all within dbt. With full awareness of an organization's models, lineage and governance rules, it enables users to query, validate, visualize, and share findings using SQL or natural language in one seamless, governed workspace. This eliminates the need to wait on central data teams to process requests or switch tabs to get answers. dbt Insights is available in preview. - **An expanded dbt Catalog** (formerly dbt Explorer) includes a unified discovery experience that enables global search and exploration for overall Snowflake assets not managed by dbt, offering analysts a comprehensive view of their data landscape. Analysts can easily discover, understand and trust the assets they use, without switching tools. dbt Catalog is now generally available, with the ability to explore Snowflake data assets currently in preview. Integrations for additional data platforms are coming soon. > "Lowering the technical barrier to entry for data analysts has been important to Tableau from the beginning of the company," said Dan Jewett, Senior Vice President, Product Management at Tableau. "dbt’s expanded offering is a game changer for customers that are looking to reduce the sizable burden on their data engineering teams, while simultaneously enabling analysts across the business in a meaningful way. It’s a massive step forward for the future of data teams and one we’re thrilled to continue to partner on." dbt Labs customer [WHOOP](https://www.getdbt.com/case-studies/whoop) is eager to boost self-service for its analysts, while leaning on easy workflows. “As our data needs evolve, empowering analysts with seamless self-exploration becomes increasingly critical,” said William Tsu, Senior Analytics Engineer at WHOOP. “By keeping them within the familiar dbt Catalog they already use daily, dbt's new analyst offerings enhance discoverability and enable faster, more intuitive, and governed self-service.” For dbt systems integrator InterWorks, dbt Canvas is poised to remove bottlenecks and power trusted self-service analytics across the organization. > “dbt Canvas is unlocking a future where analysts can build confidently alongside engineers within the same trusted and governed workflows,” said James Wright, Chief Strategy Officer at [InterWorks](https://www.interworks.com). “We're excited about how this new development environment will help our customers unlock true self-service while maintaining the standards, security, and collaboration required to scale analytics responsibly.” ## Empowering Organizations to Manage Data Warehouse Spend dbt Labs is also providing new features that allow organizations to optimize data platform costs and ensure the long-term flexibility of their data investments. This includes a cost management dashboard that helps organizations understand data platform costs from their dbt workloads, and also view consumption and realized savings from standardizing on dbt. Powered by the dbt Fusion engine, the cost management dashboard offers visibility into costs at the project, environment, model, and test level, helping users identify and resolve cost inefficiencies. No other vendor owns the transformation workflow from development to production, allowing dbt to embed cost optimization natively rather than as an add-on. The cost management dashboard is in preview for Snowflake customers ahead of the 2025 [Snowflake Summit](https://www.snowflake.com/en/summit/), June 2-5 in San Francisco. ## A Better-than-ever Developer Experience Announced earlier today, dbt Labs launched [the new dbt Fusion engine](https://www.getdbt.com/blog/dbt-labs-redefines-dbt-with-new-fusion-engine), incorporating the technology from its [acquisition of SDF Labs this year](https://www.prnewswire.com/news-releases/dbt-labs-acquires-sdf-labs-to-introduce-robust-sql-comprehension-into-dbt-and-supercharge-developer-efficiency-302350354.html). Fusion delivers massive performance improvements and introduces features that significantly enhance the developer experience. These include next-gen data transformation capabilities that improve code quality by providing real-time feedback, lower costs by avoiding unnecessary warehouse compute, and make dbt 30x faster than dbt Core. For more information on the future of the dbt platform, visit [https://getdbt.com/blog/dbt-launch-showcase-2025-recap](https://getdbt.com/blog/dbt-launch-showcase-2025-recap). For more information on the new dbt features for analysts, visit [https://www.getdbt.com/product/analyst](https://www.getdbt.com/product/analyst). *Gartner Insights, Adopt a Data Governance Approach That Enables Business Outcomes, [https://www.gartner.com/en/data-analytics/topics/data-governance](https://www.gartner.com/en/data-analytics/topics/data-governance). GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved. ### About dbt Labs Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 60,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at [getdbt.com](http://getdbt.com), and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "New era, new engine, and new names" description: "Understanding our new, simplified naming structure" url: "https://www.getdbt.com/blog/updated-names-for-dbt-platform-and-features" date: "2025-05-28" authors: ["Tristan Handy"] categories: ["Product"] --- # New era, new engine, and new names Today is an exciting day for dbt Labs. In case you missed it, we held our [annual dbt Launch Showcase](https://getdbt.com/blog/dbt-launch-showcase-2025-recap) at which we unveiled new capabilities for how developers, analysts, and organizations leverage the power of dbt. And as we rebuilt dbt's foundations with the new [dbt Fusion engine](https://getdbt.com/blog/get-to-know-the-new-dbt-fusion-engine-and-vs-code-extension), we saw an opportunity to refactor our naming to better align with the product as it now exists. So: as part of today's launch, we're also announcing a series of name changes that reflect a more unified vision for dbt's end-to-end user experiences. Let’s dig in. ## A new engine for the future of dbt By now you’ve likely heard of the dbt Fusion engine. If you missed it, check out the blog [here](https://getdbt.com/blog/get-to-know-the-new-dbt-fusion-engine-and-vs-code-extension). The dbt Fusion engine is a fundamental shift in how dbt understands and executes your code. It introduces three key new technological breakthroughs to dbt: lightning-fast performance, native SQL comprehension and state-awareness. A literal fusing together of the [capabilities we acquired with SDF](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs) with everything you’ve known dbt to be to this point. Fusion isn't just an engine—it's a glimpse of what's ahead. As data teams seek faster feedback loops, deeper runtime insights, and smarter automation, Fusion provides the foundation for this future. It’s the future of dbt and the thing that stitches it all together. ## It’s all just dbt In 2016, when dbt Labs (then Fishtown Analytics) was founded, I described dbt as _“a command line tool that enables data analysts and engineers to transform data in their warehouses more effectively.”_ The product, for ~4 years, was _very_ basic: a Python-based command-line tool that allowed users to take code from their local machine and deploy it into a data warehouse. Needless to say, a lot has changed since then. Today’s dbt ecosystem includes: - The all-new, all-Rust dbt Fusion engine. - A widely-used suite of APIs powered by the dbt engine and metadata platform. - A sophisticated identity and auth platform. - A whole suite of products: 2 IDEs (local and web), a low-code editing experience, a data catalog, an orchestrator, a CI tool, a cost observability tool, a semantic layer, a tool for exploratory data analysis, an MCP server, and almost certainly more that I’m forgetting. - An AI copilot experience that integrates with the entire platform. Many of these experiences are local; others live in a remote server. That server can be in our cloud or it can be in yours. Some are OSS, some are source available, and some are proprietary. Some are free; some require a commercial license. **All of them come together to form a single, integrated platform for data and analytics engineering.** We are moving towards having a single CLI that controls every bit of this functionality, from executing `dbt run` to reading metadata to adding a new user account. This … is not the world of dbt in 2020, when our previous product names (dbt Core and dbt Cloud) were born. A very long way from it. The product experiences are new and the technology that underlies them is new. I announced one of my main goals for the future of dbt at Coalesce 2024: **One dbt**. I had become hyper-frustrated with the separation between our Core and Cloud products; it didn’t make any sense and it didn’t help anyone. In my head, there was only a single product—dbt—and the failure was ours for not making that product work together in a cohesive fashion. Today, we have taken significant technical steps to bring this vision to reality. The dbt Fusion engine sits at the heart of the dbt experience. It can be installed locally or used in the cloud. It can be used for free under a source available license, or it can be used under commercial terms. But wherever and however it runs, Fusion is Fusion. The APIs, products, and AI experiences I listed before? They’re all built on top of Fusion. And they’re all a part of a single platform called, simply, dbt. ## Associating product names with jobs to be done If you haven’t noticed, we like boring names. The name “dbt” was originally short for “data build tool” which ended up becoming the name when none of the original committers could think of anything more creative. Sometimes boring is good :) As a part of our effort to clarify and unify our product naming, we’re going to stick with our history of straightforward, descriptive names: - **dbt Explorer** is now…wait for it…**dbt Catalog.** - Our drag-and-drop visual editing experience is **dbt Canvas.** - **dbt Insights** is the new exploratory data analysis interface. - Our web-based IDE is called **dbt Studio.** These new names align with how data teams already think and—we hope!—will make it easier for you to describe the breadth of what dbt has to offer. ## A new era for dbt - one product, choose your adventure While these naming changes might seem subtle, they represent a fundamental shift in our vision for dbt. We're creating a more intuitive, integrated, and clear experience throughout. - A single integrated platform: **dbt** - **Descriptive product names** to better represent the roles they play in your workflow - And with **Fusion,** the foundation for an exciting future ahead Getting started is easier than ever. [Download the VSCode extension](http://docs.getdbt.com/docs/install-dbt-extension) and start experiencing Fusion first-hand. Build a model in Canvas and collaborate with your peers. Go wild. It’s all dbt. --- --- title: "dbt Labs Redefines dbt with New Fusion Engine, Built to Revolutionize Developer Experience in the Age of AI" description: "New engine enables faster analytics delivery, lower cloud costs, and trusted data pipelines built for AI at scale" url: "https://www.getdbt.com/blog/dbt-labs-redefines-dbt-with-new-fusion-engine" date: "2025-05-28" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Redefines dbt with New Fusion Engine, Built to Revolutionize Developer Experience in the Age of AI #### New engine enables faster analytics delivery, lower cloud costs, and trusted data pipelines built for AI at scale **PHILADELPHIA** – May 28, 2025 – [dbt Labs](http://getdbt.com), the leader in standards for AI-ready structured data, today unveiled the new dbt Fusion engine, a monumental evolution of the technology that powers dbt. Fusion, built on Rust and equipped with native SQL comprehension, introduces a lightning-fast developer experience that delivers productivity, data velocity, and platform intelligence to drive substantial cost savings. dbt Labs also launched its VS Code extension, unlocking broad access to the power of Fusion for local developers, and is introducing a free, source-available version of the Fusion engine with a subset of features. These foundational enhancements, along with several others [announced today](https://www.getdbt.com/blog/dbt-labs-launches-ai-powered-features-to-onboard-data-analysts-into-dbt) and tailored to bring data analysts into the dbt workflow, will empower organizations to scale analytics in the age of AI. “AI is completely changing the way we interact with data, and dbt is in a prime position to drive the next phase of innovation in the market,” said Tristan Handy, founder and CEO, dbt Labs. “Fusion is the most significant evolution of dbt in its history. It gives enterprise data teams the control, speed, and intelligence they need to scale analytics and AI responsibly while keeping costs down.” ## Fusion Engine Brings SQL Comprehension to dbt The [dbt Fusion engine](https://getdbt.com/where-we-re-headed-with-the-dbt-fusion-engine) now powers the entire dbt platform, from the CLI that is in use by over 60,000 teams today, to dbt Orchestrator, Catalog, Studio, and the other commercial products powering [dbt Labs’ rapid growth](https://www.getdbt.com/blog/dbt-labs-100m-arr-milestone). Fusion introduces powerful SQL comprehension and a host of other capabilities that collectively deliver a best-in-class developer experience, all while empowering organizations to operate with the highest quality, context-rich data while optimizing costs. New capabilities include lightning-fast parse times, up to 30x faster than dbt Core, allowing large dbt projects to execute in milliseconds instead of minutes. Instant feedback loops and live error detection now uncover and surface parse, compilation and logic errors as code is being written – and before running code against the warehouse – optimizing both developer efficiency and data platform costs. Fusion also brings state-awareness to dbt and with it, a new level of intelligence to how dbt orchestrates pipelines. With state-aware orchestration, available in beta for commercial customers running Fusion, dbt will automatically run jobs as soon as sources are fresh and limit builds to only the models that changed. This helps organizations save on data platform compute and maintain pipeline velocity. Early customer feedback indicates an average 10% cost savings as a result, with additional savings expected as Fusion matures. Organizations can validate these savings in the new cost management dashboard (in preview for Snowflake users), which offers visibility into costs at the warehouse, project, model, and environment level, helping to identify inefficiencies earlier. **Other standout functionality includes:** - **Powerful IntelliSense**, which autocompletes SQL functions, model names, columns, macros and more; - **Instant refactoring** to rename models or columns and see references update project-wide; - **Go-to-definition**, allowing users to jump to definitions in a single click, a useful feature for large projects with many models and macros; - **Hover insights**, which enable users to see context on tables, columns and functions without leaving code; - **Live CTE previews** directly inside dbt models, for faster validating and debugging; - **Rich lineage, in context**, allowing developers to see lineage at the column or table level as they develop, without breaking flow; and - **View compiled code**, which gives a live view of the SQL code built by models, alongside dbt code. > “The data team at Bilt is very excited to roll out the new dbt Fusion engine,” said James Dorodo, VP of Data Analytics at Bilt Rewards. “The improvements it brings will address many of the pain points we currently face in our development cycle, and we believe it will provide a step function increase in our velocity.” ## Multiple Paths to Fusion Engine Access Fusion is now available for eligible dbt projects on Snowflake, with support for Databricks, BigQuery, and Redshift coming soon. dbt Labs is also introducing its VS Code extension, the sole way to access the full power of the Fusion engine while developing locally. Now, wherever developers are doing their work, they can do so backed by Fusion. The VS Code extension is downloadable now from the [VS Code Marketplace](https://docs.getdbt.com/docs/install-dbt-extension). In addition, dbt Labs is making a subset of Fusion’s capabilities broadly available via a new source-available license. This will provide users in the dbt community free access to Fusion’s robust developer experience features. ## AI Adoption Drives The Need For More Quality Data and Unified Standards A major disruption is underway in analytics, driven by and in service to AI. According to the latest [State of Analytics Engineering report](https://www.getdbt.com/resources/state-of-analytics-engineering-2025), organizations rank AI at the top of their budget priority lists. As these investments grow, the pressure on data teams to deliver trustworthy, contextual data – and a scalable AI strategy – has never been greater. Yet inconsistent standards and fragmented workflows often result in a lack of a single source of truth, making it difficult to scale AI initiatives responsibly. To address this challenge, dbt Labs launched the dbt MCP server, leveraging Model Context Protocol to enable seamless, universal connectivity between AI systems and the governed, structured data in dbt. As the standard for creating governed, trustworthy datasets on top of structured data, dbt unifies models, metrics, documentation, and testing into one collaborative environment. The dbt MCP server gives business users the confidence that all AI endpoints are fueled by context-rich, reliable data, no matter how the AI stack evolves. > “dbt Labs’ continuous investment in revolutionizing the developer experience aligns well with our commitment to giving customers the very best platform for all of their data engineering needs, with low costs and accelerated performance in the AI Data Cloud,” said Chris Child, VP of Product, Data Engineering, Snowflake. “As more of our joint customers adopt AI across their businesses, we know these critical data initiatives require context to be successful. The dbt MCP server complements our investments in AI, and now, with the Fusion engine powering dbt, we’re eager to see how much more productive and successful our joint customers will be.” ## dbt Empowers Data Analysts with New Features In conjunction with the rollout of the Fusion engine, VS Code extension and dbt MCP server, dbt Labs [launched](https://www.getdbt.com/blog/dbt-labs-launches-ai-powered-features-to-onboard-data-analysts-into-dbt) a suite of new governed, accessible features designed to bring data analysts into the dbt workflow, including: - **dbt Canvas**, a new AI-powered drag-and-drop visual editing experience that enables analysts less familiar with dbt or SQL to create new and edit existing dbt models within a governed environment. - **dbt Insights**, a new AI-powered query interface that lets analysts perform ad-hoc analysis by asking questions about their data models in SQL or natural language and get answers faster, without waiting on engineering. - Extending the functionality of **dbt Catalog**, formerly known as dbt Explorer, to include search and lineage for overall Snowflake assets alongside your dbt models, making it easier to explore your full data environment. Support for other data platforms is coming soon. These features expand the impact of dbt across the analytics workflow, empowering more collaborators to build, analyze, and explore the data they need, within a governed environment. By supporting governed, scalable self-service, teams reduce engineering bottlenecks, minimize security risks, cut compute costs, and ensure high quality data across their analytics workflow. > “Fivetran and dbt have a long history of innovation and delivering on the promise of the modern data stack. With the launch of the dbt Fusion engine and dbt Labs’ acquisition of SDF Labs, dbt Labs is accelerating what’s possible with data and AI,” said Taylor Brown, COO and co-founder, Fivetran. “We’re excited to have SDF Labs co-founder Elias DeFaria join us at the Fivetran booth at Snowflake Summit next week to showcase the next wave of tooling.” These new features will be the focus of dbt Labs’ upcoming presences at [Snowflake Summit](https://www.getdbt.com/events/summit/snowflake-summit-2025) (June 2-5, booth #1808) and [Databricks Summit](https://www.getdbt.com/events/summit/databricks-data-ai-summit-2025) (June 9-12, booth #326). The dbt Labs team will be onsite at both industry events to discuss and demo the power of Fusion and these new capabilities. To learn more and book a meeting, visit [https://www.getdbt.com/events/summit/snowflake-summit-2025](https://www.getdbt.com/events/summit/snowflake-summit-2025). ### About dbt Labs Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 60,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at [getdbt.com](https://getdbt.com), and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "dbt Launch Showcase 2025 recap" description: "New features in dbt empower organizations to scale analytics for the age of AI" url: "https://www.getdbt.com/blog/dbt-launch-showcase-2025-recap" date: "2025-05-28" authors: ["Alexis Jones", "James Mayfield"] categories: ["Product"] --- # dbt Launch Showcase 2025 recap Today, at the annual [dbt Launch Showcase](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase), we released a ton of exciting new features for dbt, all purpose-built to help our customers and users tackle the next wave of analytics and AI. It was without a doubt, the biggest launch event in our company’s history. In a jam-packed 90 minutes, we had an executive keynote, new product announcements, and in-depth demos that all centered around the theme of **how dbt is empowering enterprises for the next era of analytics.** Let’s break that down 👇 ## Empowering the enterprise for the next era of analytics The next era is without a doubt driven by and in service of AI. We showcased the **dbt MCP server** and other dbt Copilot-enhanced features to help our users connect their AI systems to governed, trusted data. And in this new AI era, data has to be a team sport. That means that data tooling needs to extend beyond data developers to include the downstream analysts who need more governed inroads to participate in the data workflow. And also to include the organizational stakeholders who need to ensure that data initiatives stay within budget and that they’re future-proofing their investments. The innovations we shared are designed to empower our users—whether data developers, data analysts, or their organizational leaders—to standardize on dbt to help their organizations win with data in this new era. ![dbt 2025 Showcase event recap overview](https://cdn.sanity.io/images/wl0ndo6t/main/29c56724e16a8557c96a0fbaec0e0c3da08df8b1-1056x934.png) - For developers: The new dbt Fusion engine and VS Code extension (both in public beta) deliver **productivity** and **developer experience** improvements that make it turnkey (and delightful!) to ship high-quality data at speed and scale - For analysts: A new suite of platform features designed specifically for analysts—dbt Canvas (GA), dbt Insights (Preview), an expanded dbt Catalog (Preview for Snowflake assets), and a new flexible seat type—make it easy for data analysts to bring their business knowledge to bear and **participate in governed, self-service data development** in a dbt-tonic way - For organizations: New cost optimization and platform flexibility features help budget owners and organizational leaders **optimize every dollar** spent on data and standardize on a platform that supports flexible navigation of **cross-platform** architectures ### A new, simplified naming philosophy This is such a massive launch that not only did we have to start from scratch on the technology behind dbt, we had to re-think the way we talk about it. In the past, we primarily built two separate but interlocking products: dbt Core and dbt Cloud. But with the launch of the dbt Fusion engine, we’ve rearchitected significant parts of the technology to work more closely together. As a result, we don’t really have two products anymore…we just have one. **And it’s called dbt.** You can read a lot more about our new naming and the ethos behind it [here](https://getdbt.com/blog/updated-names-for-dbt-platform-and-features). The dbt Fusion engine is the future of dbt and is available to you whether you’re a managed dbt customer or not: you can adopt Fusion via the freely available and permissively licensed source code and binary, or via the dbt platform. [See the full post here](https://www.getdbt.com/blog/new-code-new-license-understanding-the-new-license-for-the-dbt-fusion-engine) on how we've licensed Fusion. So, back to the naming: you can choose to just use the open source / source available components (the engines), or you can use paid features on top. But either way, the experience is tightly integrated. It’s all just dbt. Let’s dive in to the specifics of our announcements. **** ## AI needs a strong data foundation ### dbt MCP Server As your AI stack evolves, your AI systems need structured context that stays consistent. And that context needs to be centralized, governed, and available across every agent, tool, and workflow. It can’t live in fragments or tribal knowledge; it has to be defined in code. [MCP](https://www.anthropic.com/news/model-context-protocol) is quickly emerging as the standard for providing context to LLMs and AI agents, allowing them to function at a high level in real world, operational scenarios. With the **dbt MCP Server** ([in beta](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) and publicly available in an [open source repo](https://github.com/dbt-labs/dbt-mcp/tree/main)), you can expose your dbt project’s trusted models, metrics, tests, and lineage to AI systems, giving agents and LLMs a structured foundation of governed context so you can trust the queries, reasoning, and actions that they inform. [Watch video](https://youtu.be/fJ-72qOA7BE) By integrating with the dbt MCP server, AI agents and LLMs can now **discover** trusted data models and metrics, **query** the semantic layer using generated SQL, and safely **execute** dbt projects, enabling structured, governed AI workflows. Together, these dbt MCP server tools make dbt not just the structured context layer, but also: - The governed control layer between your data workflows and your AI. - The bridge between the governed warehouse and your LLMs, agents, and AI-powered tools. - The standard for creating governed, trustworthy datasets so structure is defined once and reused across every AI workflow. With more tools on the way, dbt is becoming the connective layer powering enterprise AI systems with confidence. Download the repo [here](https://github.com/dbt-labs/dbt-mcp/tree/main), and learn more about how dbt can help with your AI initiatives [here](https://www.getdbt.com/product/ai). ## Empowering developers with the new Fusion engine and VS Code Extension ### New era, new engine We announced that the new [dbt Fusion engine](https://www.getdbt.com/product/fusion) is in public beta for [eligible Snowflake projects](https://docs.getdbt.com/docs/fusion/about-fusion). We also released the official [dbt VS Code Extension](http://docs.getdbt.com/docs/install-dbt-extension) in public beta. Fusion represents a huge leap forward for teams building with dbt. Fusion isn’t just another feature. It’s the [foundation for a new era of analytics engineering](https://getdbt.com/blog/where-we-re-headed-with-the-dbt-fusion-engine), ushering in a whole new dbt experience that is faster, more intelligent, and more cost-efficient than ever before. By standardizing on Fusion, teams can unlock: - Lightning-fast performance, with parse times up to 30x faster than dbt Core. - Native SQL comprehension, unlocking capabilities like real-time validation of your code—without the need to query your warehouse. - State-awareness, giving dbt a rich understanding of the state of your dbt project and what’s been materialized in your warehouse. This allows dbt to intelligently avoid unnecessary builds, resulting in higher velocity pipelines and substantial cost savings. **Early customers are already seeing ~10% reductions in warehouse spend.** > > > — Matt Karan, Obie Insurance ### A purpose-built VS Code extension, powered by the Fusion engine The official [dbt VS Code Extension](http://docs.getdbt.com/docs/install-dbt-extension) is the only way to tap into the full power of the Fusion engine when developing locally. With it, users get a hyper-responsive development experience bolstered by capabilities like: - IntelliSense for smart autocompletion for models, columns, macros, and functions. - Automatic refactoring to update references across your entire project instantly when a model or column is re-named. - Go-to-definition and inline CTE previews to aid in navigating large projects and rapid debugging. Note that to use the VS Code Extension, your project must be running on the Fusion engine. The extension does not support dbt Core, as the key developer experience enhancements it’s designed to enable rely on the technological foundations of ‌the new engine. You can learn more [here](https://docs.getdbt.com/blog/dbt-fusion-engine-components). **** ## Empowering analysts with dbt Canvas, dbt Insights, and an expanded dbt Catalog We’re launching a new suite of capabilities to [make it easier for analysts to contribute directly to transformation workflows in dbt](https://www.getdbt.com/product/analyst) without sacrificing the governance and quality teams data teams require: dbt Canvas, dbt Insights, and an expanded dbt Catalog. ### Build and edit models in a visual development environment with dbt Canvas dbt Canvas is a AI-powered visual editing experience that makes it easy for analysts to build and edit dbt models. With Canvas, analysts can work out of a drag-and-drop interface to discover trusted sources and models, apply transformations like joins, filters, and aggregations, and preview outputs step-by-step as they go. Every transformation is automatically compiled into SQL that fits seamlessly into your existing dbt project. Canvas includes always-on data profiling, context-aware AI assistance with dbt Copilot, and intuitive Git-based version control. Analysts can commit their work and open PRs without leaving the Canvas interface, making it simple to contribute while maintaining quality and oversight. Whether defining a new KPI or iterating on an existing model, analysts can work with confidence and speed inside a fully governed workflow. [dbt Canvas is now GA](https://docs.getdbt.com/docs/cloud/canvas) for Enterprise customers. For more details, check out the [dedicated Canvas launch blog post](https://getdbt.com/blog/dbt-canvas-is-ga). ### Get ad-hoc insights, fast with dbt Insights dbt Insights is a new interface for fast, governed data exploration that combines metadata, documentation, AI-assistance, and powerful querying capabilities into one unified experience. Insights supports both SQL and natural language queries, with built-in AI assistance from dbt Copilot, query history, and instant visualizations. Users can quickly validate ideas, explore trends, and iterate on queries in a context-rich environment that surfaces metadata, lineage, and trust signals from across the dbt platform. When ready, analysts can turn their exploratory work into models in dbt Canvas or Studio all within the same workflow. With tight integration across dbt, Insights helps teams collaborate, explore, and build with confidence. [dbt Insights is in Preview](https://docs.getdbt.com/docs/explore/dbt-insights) for Enterprise customers. ### Broaden data discovery with an expanded dbt Catalog The [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) (formerly dbt Explorer) now goes beyond just dbt models, and allows users to search and explore Snowflake assets, like tables and views, directly within dbt. This unified view helps teams move faster and make more informed decisions, without switching tools or duplicating efforts. By centralizing discovery, the expanded Catalog streamlines collaboration, accelerates insight, and ensures governance in dbt workflows. The ability to find and visualize Snowflake assets in dbt Catalog is now in Preview, with integrations for other data platforms coming soon. ## Empowering organizations with cost optimization and enterprise features We're introducing new features to make dbt analytics workflows more cost-efficient, flexible, and enterprise-ready. These updates help teams reduce overhead, improve visibility, and work confidently across clouds and platforms. ### Understand and optimize transformation spend with the cost management dashboard Understanding data platform costs has been a persistent challenge for data leaders. Through monitoring, remediation, and state-aware orchestration, dbt provides comprehensive cost management capabilities. While monitoring gives visibility into warehouse spend and remediation helps fix existing inefficiencies, Fusion-powered state-aware orchestration prevents unnecessary costs by intelligently avoiding redundant builds and queries in the first place. Rather than struggling with billing data or building custom pipelines, you can track spend directly in dbt. The cost management dashboard provides detailed breakdowns by account, project, environment, and model. This makes it simple to identify high-cost workloads, find optimization opportunities, and measure your ROI. With the cost management dashboard, you can: - Compare spend over time - Drill into specific areas of platform spend - Tie optimization work to real savings - Track how Fusion-powered state-aware orchestration reduces warehouse compute costs - Uncover waste with confidence We’ve been using these features internally and have already discovered $10,000+ yearly savings with less than two hours of work. This is transformative for analytics teams looking to shift from reactive clean-up to proactive optimization. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/561912e59ad09b03d961efa80e4b61cbc9a2272c-2880x1800.png) > > > — Ra Raman, Zscaler [The cost management dashboard is currently available in Preview](https://docs.getdbt.com/docs/cloud/cost-management) for Snowflake customers, with other platforms and functionality coming soon. ### Simplify user management with SCIM (GA) Manual user management doesn't scale. dbt now supports SCIM for automatic user provisioning and de-provisioning through identity providers like Okta and Entra ID. This keeps access in sync with your organization without burdening your admins. [SCIM is generally available for Okta](https://docs.getdbt.com/docs/cloud/manage-access/scim) and in Preview for other providers. ### Flexibility across catalogs and clouds dbt is expanding support for hybrid and multi-cloud data architectures. You can now integrate with [Iceberg, Unity, Polaris, and BigLake](https://docs.getdbt.com/docs/mesh/iceberg/about-catalogs) catalogs across Snowflake, Databricks, and BigQuery. This gives teams the freedom to choose the tools and platforms that best serve their needs without getting locked into a single vendor. We've also launched a new hosting option on Google Cloud that enables co-location of your data and analytics layers. This includes support for Private Service Connect and GCP Marketplace procurement (coming soon), streamlining deployment in regulated or high-security environments. GCP hosting is [available](https://docs.getdbt.com/docs/cloud/about-cloud/access-regions-ip-addresses) in North America today. ## See you at Summits We couldn’t be more energized for the future and to help our passionate and vibrant community ‌continue to thrive with dbt. Fusion is the future of dbt and we can’t wait to see what you build on top of it. We’ll be on the road in San Francisco in June for both the [Snowflake Summit](https://www.getdbt.com/events/summit/snowflake-summit-2025) and [Databricks Summit](https://www.getdbt.com/events/summit/databricks-data-ai-summit-2025) and we look forward to seeing you there! --- --- title: "New code, new license: Understanding the new license for the dbt Fusion engine" description: "The philosophy behind our new license" url: "https://www.getdbt.com/blog/new-code-new-license-understanding-the-new-license-for-the-dbt-fusion-engine" date: "2025-05-28" authors: ["Tristan Handy"] categories: ["Product"] --- # New code, new license: Understanding the new license for the dbt Fusion engine Hi all! You’ve likely read about our [massive launch today](https://getdbt.com/blog/dbt-launch-showcase-2025-recap). It is not an exaggeration to say that this is the single biggest launch event for us in many years. The [dbt Fusion engine](https://getdbt.com/blog/get-to-know-the-new-dbt-fusion-engine-and-vs-code-extension). [dbt Canvas](https://getdbt.com/blog/dbt-canvas-is-ga). Cost Management. And more. There’s one announcement that I want to zoom in on in this post: licensing for the dbt Fusion engine. This is something that I care _a lot_ about and have thought a _lot_ about, and I want to share that thinking, zero filter, with you. Here are the main three points to know at the outset: 1. **The license for dbt Core is not changing.** dbt Core remains Apache 2 OSS and we will continue to support it indefinitely. 2. Today, we introduced the brand new dbt Fusion engine to power the next era of analytics engineering. **The dbt Fusion engine contains a mixture of source-available, proprietary, and open source code.** The source-available components of the Fusion engine will be available under the Elastic Version 2 (ELv2) License. 3. At GA, the source available and open source components of the dbt Fusion engine will support all features in dbt Core _and more_. **Fusion will be a strictly better offering whether you’re paying us or not.** You can read the [Licensing FAQ](https://www.getdbt.com/licenses-faq) for full details. The upshot is: If you're a data team that uses dbt Core today, nothing about the Fusion engine’s new license (ELv2) prevents you from upgrading to Fusion and using it just like you use Core today. You don’t need to pay us anything, give us your contact info, or speak with anyone. You run it locally, and you can see (and contribute to!) the source code. You _cannot_ use it to directly compete with dbt Labs by hosting a managed service powered by the dbt Fusion engine. Let’s dive in. ## The biggest update: dbt Core and the dbt Fusion engine In January, we announced that we had [acquired a company called SDF Labs](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs). SDF had, over the course of multiple years in stealth, built a brand new engine for data transformation. It was written in Rust, blazing fast, and had meaningful new features, all stemming from its ability to both parse and compile SQL in many dialects. Since then, we have been sprinting to integrate SDF’s technology into the overall dbt experience. The main outcome of this is the [launch of Fusion](https://getdbt.com/blog/get-to-know-the-new-dbt-fusion-engine-and-vs-code-extension), a brand new engine for dbt, rewritten from the ground up incorporating the best elements of dbt and SDF into a single engine. New code, new product, more functionality. The authoring specification itself is shared—code that is working in dbt Core will also work in Fusion, with a [few exceptions](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-fusion). But under the hood - other than the adapter macros and materializations - not a single line of code between Fusion and dbt Core is shared! Upon acquisition, SDF’s source code was proprietary, and _the original plan was actually to stick to that strategy_. As you would expect, we had to pay real $$$ for the acquisition, and, given that, open sourcing all the IP seemed hard to justify. But, having some time to process, we wanted to see if we could do better. Better for the dbt Community and better for dbt Labs. And I think the approach we’ve come to is _much better for both._ Here are the specifics: 1. **Separate the code** Separate dbt Core and the dbt Fusion engine. They are two totally separate engines, with zero shared tech. Reflect that in the repository structure. 2. **Continue to support dbt Core as Apache 2.0** Continue to support dbt Core indefinitely. Bug fixes, security patches, maintaining compatibility with data platforms, etc. Keep the Apache 2.0 license on this repo. If you are currently using dbt Core and for some reason prefer not to migrate to Fusion, nothing is being taken away from you. If you are a vendor using dbt Core to compete directly with dbt Labs, your ability to do that has not changed. 3. **Apply a new license to Fusion** License the Fusion code base under ELv2 to protect dbt Labs’ commercial interests but still maximizes users’ ability to view source code, contribute, locally install, use, _and productionize_ an exclusively-better dbt experience. Ultimately, Fusion is not only significantly better for the things that users have historically come to expect from dbt—data modeling and transformation—it also presents unique opportunities to build brand-new products. We think the entire dbt Community should use it. The way we’ve licensed it enables exactly that—there are no limitations on what _users of Fusion_ can do with the code and its binaries as long as they comply with the three [rules of ELv2](https://www.elastic.co/licensing/elastic-license/faq): - You may not provide the products to others as a managed service - You may not circumvent the license key functionality or remove/obscure features protected by license keys - You may not remove or obscure any licensing, copyright, or other notices The ELv2 license does prevent other companies from using code from Fusion to directly compete with dbt Labs. Our goal is to build a long-term sustainable business, and this IP is an important part of our ability to do so. Our choice was between making this new code proprietary and making it more open. We chose to make it more open. ## Net new Apache 2 OSS While Fusion will be released under the ELv2 license, we are continuing to contribute significant net new code in Apache 2.0-licensed open source. As a part of this launch, we are committing to contribute three new code bases licensed under Apache 2.0: Our brand new dbt Fusion adapters, all based on the Apache 2 Arrow Database Connector (ADBC) framework. We're all in on Arrow and are not only open sourcing these adapters, but dbt Labs team members are also frequent contributors to Arrow and ADBC. - The ANTLR grammars that power multi-dialect SQL parsing, the result of thousands of hours of analysis across millions of SQL queries. - A Rust-based implementation of Jinja called [dbt-jinja](https://github.com/dbt-labs/dbt-fusion/tree/main/crates/dbt-jinja), an essential component for being able to parse dbt projects without a Python interpreter. We are making these libraries Apache 2 OSS because we believe there is significant Community value to build on top of these. Adapters, grammars, and the new faster standard for parsing dbt Jinja in Rust are foundational pieces of technology, and we hope the entire community finds them as useful as we do. We’ll have more to say about this code in the coming weeks. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/4d801cfa813fd46a38390c8742f7e84641b78921-2367x2517.png) ## Why now? What changed? While we are not changing the license of any existing code, I do want to acknowledge that this is an evolution of our open source strategy and I owe you an explanation. dbt Core was originally built by data practitioners, for data practitioners. There’s a reason it caught fire: the workflow it enabled resonated deeply with the folks doing data work every day. And the Apache 2 license made it easy to adopt, and to become a standard, back when dbt Labs was half a dozen people and a viewpoint. But dbt Core was built in 2016 and could never have imagined the needs of 2025. Translating between SQL dialects with complete fidelity. Local execution. Compilation times that are suitable for conversational AI interfaces. And much more. These needs just weren’t present in 2016, and they aren’t things that the original codebase can be extended to handle. I am incredibly proud of what we all, collectively, have accomplished with dbt Core: elevating the careers of a million analytics engineers, bringing software engineering best practices to data, and building (IMHO) the most inclusive, magical community in data. But one of the most challenging things in life is when you realize “what got us here won’t get us there.” Those moments show up in my personal and professional life and every single time they are hard and uncomfortable. Over my 44 years I’ve learned to lean into those moments and embrace the future—despite the challenges it often brings!—rather than getting anchored in the past. Fusion is the future of the dbt engine. And Fusion is industrial-grade. It was built by some of the most talented software engineers in all of data. And with industrial-grade quality comes industrial-grade investment. We need to build the business on top of this technology that it deserves, so that we can continue investing in it for many years to come. Choosing the Elastic License was our best answer to the question: “How do we give the entire dbt Community an exclusively-better dbt experience while preserving space for ourselves to build a business around it?” I continue to be a believer in Open Source, but I am realistic about the resources required to fund it. This is my best answer on how to balance both. If you’re a longtime dbt user and want to chat about this, ping me on dbt Slack—I’m @tristan. Or ask questions in #dbt-fusion-engine. I could not be more excited for the next decade of dbt, now powered by Fusion. --- --- title: "What is data governance?" description: "Explore the key pillars and tools driving modern data governance—from quality and compliance to AI-readiness." url: "https://www.getdbt.com/blog/data-governance" date: "2025-05-27" authors: ["Joey Gault"] categories: ["Pulse"] --- # What is data governance? ## The fundamental pillars of data governance Effective data governance rests on four essential pillars that work together to ensure comprehensive data management. The first pillar, [data quality](https://www.getdbt.com/blog/getting-started-data-quality-management), focuses on ensuring accuracy, completeness, and consistency of all data assets across the organization. This becomes particularly challenging as companies scale and data flows through multiple systems and transformations. [Data stewardship](https://www.getdbt.com/blog/data-governance-ownership) forms the second pillar, establishing clear roles and responsibilities for managing data throughout its lifecycle. Data stewards serve as the front line of governance programs, defining and documenting data assets, ensuring quality standards are met, and facilitating effective data sharing across teams. The third pillar, [data protection and compliance](https://www.getdbt.com/security), encompasses security measure that prevent unauthorized access, privacy protections for sensitive information, and processes to ensure compliance with applicable regulations. This pillar has become increasingly important as data privacy regulations like GDPR and CCPA affect billions of people worldwide, and heavily regulated industries face additional requirements from frameworks like [FINRA](https://www.getdbt.com/industry/financial-services) and [HIPAA](https://www.getdbt.com/industry/healthcare). [Data management](https://www.getdbt.com/blog/data-products-data-mesh), the fourth pillar, covers the processes and procedures for storing, accessing, and manipulating data effectively. This includes metadata management, data lifecycle management, and data integration: essentially how data is structured, stored, and linked across different systems within the organization. ## When organizations need formal data governance The transition from informal to formal data governance typically occurs when organizations reach a size where casual, ad-hoc management can no longer effectively control data-related activities across the entire company. As companies grow, the number and complexity of data systems multiply, and without structured governance frameworks, data estates quickly become siloed as teams naturally diverge in their priorities and approaches. The fragmentation makes it impossible to gain informed, enterprise-level visibility into the masses of data collected and processed throughout the organization. Teams may unknowingly duplicate efforts, create conflicting definitions for the same metrics, or implement incompatible data formats that break downstream processes. Regulatory requirements often serve as another catalyst for implementing formal governance programs. A majority of the world's population is covered by national data privacy regulations, most [enterprise-level companies eventually face compliance obligations that require documented data handling procedures](https://www.getdbt.com/blog/what-is-enterprise-data-governance), audit trails, and formal oversight mechanisms. ## The evolution toward modern data governance Traditional data governance approaches were largely static, policy-based, and top-down. A central authority would establish standards and policies, which data stewards would then work to implement and enforce at the team level. While this approach provided structure, it often proved too slow and rigid for today's fast-paced data environments. Modern [data governance, particularly in the age of AI and machine learning](https://www.getdbt.com/blog/understanding-data-governance-ai), requires a more dynamic and collaborative approach. The exponential growth of data volumes, combined with the rise of generative AI applications that demand large, high-quality datasets, has strained traditional manual governance processes beyond their breaking point. Contemporary governance strategies emphasize automation, continuous monitoring, and federated responsibility. Rather than relying solely on centralized control, modern approaches enable teams to work independently while ensuring compliance through automated tools and shared standards. This federated computational approach makes data governance a community effort where data producers, consumers, and governance experts collaborate to create and maintain high-quality datasets. ## Essential tools for governance at scale Implementing [data governance at enterprise scale](https://www.getdbt.com/blog/enterprise-data-governance-strategy-elements) requires sophisticated tooling that combines automation with human oversight. Data catalogs serve as the foundation, providing a single source of truth that describes not just data assets but all associated metadata: ownership, update frequency, quality metrics, and usage patterns. Data lineage tools complement catalogs by visualizing how data flows through systems, enabling root cause analysis when issues occur and impact assessment before making changes. These tools prove invaluable for [understanding data provenance and building trust in analytical outputs](https://www.getdbt.com/product/build-trust-in-data-and-data-teams). Data security management tools provide fine-grained access controls, ensuring that sensitive information remains protected while enabling appropriate self-service access. Data classification capabilities automatically tag datasets according to sensitivity levels, enabling automated enforcement of governance policies based on regulatory requirements. Quality management tools enable teams to implement DataOps methodologies, treating data changes as code that can be version-controlled, tested, and deployed through continuous integration pipelines. This approach catches errors before they reach production and provides audit trails for all data modifications. **** ## How dbt enables comprehensive governance With [the dbt platform](https://www.getdbt.com/product/what-is-dbt) as your [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction), teams standardize how they build, test, deploy, and discover analytics code using the ADLC. dbt unifies transformations, testing, documentation, CI, and dbt Catalog for discovery and lineage, while remaining vendor-agnostic. Through [dbt Catalog](https://www.getdbt.com/product/dbt-catalog), teams gain end-to-end visibility into data pipelines and dependencies, with column-level lineage that facilitates both troubleshooting and impact analysis. dbt's built-in testing framework enables data engineers to create comprehensive test suites that validate transformations before deployment, ensuring [data quality standards](https://www.getdbt.com/blog/data-quality-metrics) are maintained consistently across all teams. dbt's approach to governance emphasizes collaboration through shared transformation code and standardized practices. Teams can package their work as [reusable data products](https://www.getdbt.com/blog/data-product-data-as-product), making high-quality datasets discoverable and accessible to other teams while maintaining appropriate access controls. The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) further enhances governance by centralizing metric definitions, eliminating inconsistencies that arise when teams implement their own calculations. [dbt's continuous integration capabilities](https://docs.getdbt.com/docs/deploy/about-ci) enable organizations to implement rigorous change management processes. All modifications go through peer review and automated testing before reaching production, with role-based access controls ensuring that only authorized personnel can make changes to critical data models. ## Governance in the AI era The emergence of [AI and machine learning applications has introduced new governance challenges](https://www.getdbt.com/blog/understanding-data-governance-ai) that traditional approaches weren't designed to handle. AI systems require large volumes of high-quality training data, and the nature of large language models makes their outputs difficult to predict or control. These characteristics raise concerns about bias in underlying data, lack of transparency in model decision-making, and the inability to explain why models product specific outputs. AI systems are also susceptible to unique threats like data poisoning, prompt injection, and model inversion attacks that require specialized governance approaches. Modern [data governance frameworks must address these AI-specific challenges](https://www.getdbt.com/blog/data-governance-frameworks-ai) while maintaining the fundamental principles of data quality, security, and compliance. This requires continuous monitoring of both data inputs and model outputs, with automated systems that can detect anomalies and bias in real-time. [dbt supports AI governance](https://docs.getdbt.com/blog/ai-eval-in-dbt) by providing the data quality foundation that machine learning models require. Through comprehensive testing, documentation, and lineage tracking, teams can ensure that AI systems are built on trustworthy data with clear provenance. dbt's ability to create standardized, well-documented data products makes it easier to verify the origin and quality of datasets used in AI applications. ## Building competitive advantage through governance Strong data governance creates tangible competitive advantages beyond risk mitigation. Organizations with mature governance practices can move faster because teams spend less time hunting for data, resolving quality issues, or rebuilding broken pipelines. Standardized processes are automated quality checks reduce the friction associated with data projects, enabling teams to focus on generating insights rather than managing infrastructure. [Trust in data](https://www.getdbt.com/product/build-trust-in-data-and-data-teams) enables more ambitions analytics initiatives and faster decision-making. When stakeholders have confidence in data quality and understand how metrics are calculated, they're more likely to act on analytics insights. This trust becomes particularly valuable in AI applications, where explainable, well-governed models build user confidence and adoption. Governance frameworks also future-proof organizations against evolving regulatory requirements. Rather than scrambling to achieve compliance when new regulations emerge, organizations with mature governance practices can adapt their existing frameworks to meet new requirements efficiently. ## Conclusion Data governance represents far more than a compliance exercise: it's a strategic capability that enables organizations to extract maximum value from their data assets while managing associated risks. As data volumes continue to grow and AI applications become more prevalent, the organizations that thrive will be those that have invested in scalable, automated governance frameworks. The key to successful governance lies in choosing approaches and tools that enable collaboration rather than creating bottlenecks. Modern solutions like [dbt](https://www.getdbt.com/product/dbt) provide the foundation for governance that scales with organizational growth while maintaining the flexibility to adapt to changing requirements and technologies. For data engineering leaders, the question isn't whether to implement data governance, but how to build governance capabilities that accelerate rather than impede data initiatives. Teams that bake governance into their ADLC workflow and utilize catalogs to make lineage and context visible move faster with fewer incidents—and build lasting trust in data. ## Data governance FAQs **What is data governance?** **When do organizations need formal data governance?** Organizations typically need formal data governance when they reach a size where casual, ad-hoc management can no longer effectively control data-related activities across the entire company. As companies grow, data systems multiply in number and complexity, leading to siloed data estates where teams diverge in their priorities and approaches. This fragmentation prevents enterprise-level visibility and can result in duplicated efforts, conflicting metric definitions, and incompatible data formats. Regulatory requirements also serve as a catalyst, with much of the world's population covered by national data privacy regulations that require documented procedures and formal oversight mechanisms. **What are the benefits of data governance?** Strong data governance creates tangible competitive advantages by enabling organizations to move faster, as teams spend less time hunting for data, resolving quality issues, or rebuilding broken pipelines. It builds trust in data through standardized processes and automated quality checks, enabling more ambitious analytics initiatives and faster decision-making. When stakeholders have confidence in data quality and understand how metrics are calculated, they're more likely to act on analytical insights. Additionally, governance frameworks future-proof organizations against evolving regulatory requirements, allowing them to adapt existing frameworks efficiently rather than scrambling for compliance when new regulations emerge. --- --- title: "Everything terminals" description: "The universal integration layer...the command line? Tristan talks terminals with Zach Lloyd, the founder of Warp." url: "https://www.getdbt.com/blog/everything-terminals" date: "2025-05-25" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Everything terminals _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/everything-terminals-w-zach-lloyd). _ In this episode, Tristan talks with Zach Lloyd, founder of [Warp](https://www.warp.dev/)—a terminal built for the modern era, including for AI agents. They explore the history of terminals, differences between terminals and shells, and what the future might look like. In a world driven by generative AI, the terminal could once again be the control center of computer usage. _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Chapters - **01:00 – Introducing Warp and Zach Lloyd** - Zach Lloyd explains Warp's origin, mission, and initial vision. - **02:40 – Why redesign the terminal?** - Zach describes why traditional terminal UX was ripe for reinvention. - **04:43 – Enter LLMs: A new direction for Warp** - Warp evolves into a natural language interface for developer workflows. - **06:34 – What is a shell?** - Zach defines shells, how they process text, and their role in the CLI ecosystem. - **07:58 – Shells vs programs vs built-ins** - Distinguishing between shell commands and standalone programs. - **10:00 – Why do developers debate shells?** - Features, syntax, and licensing behind the Bash vs Z Shell discussion. - **12:17 – Why terminals still matter** - The enduring power of text-based computing and scripting. - **16:40 – What is a terminal, really?** - Clarifying the difference between terminal hardware, emulators, and modern terminal apps. - **20:13 – The Warp interface** - Zach breaks down Warp’s UI: input editor, output blocks, and mouse support. - **22:48 – Will Warp replace your IDE?** - The vision of AI-driven development and the convergence of terminal, editor, and chat. - **27:20 – Rethinking development interfaces** - Finding the ideal hub for AI-native software development. - **35:00 – Why the terminal has an edge** - Advantages of the terminal for cross-project, full-lifecycle developer tasks. - **37:10 – Bottom-up adoption strategy** - How Warp approaches growth: focus on individual developers, not top-down mandates. - **39:50 – Is Warp redefining the terminal?** - The challenges of innovating in a legacy-dominated space and creating a new category. - **42:45 – Developer control & context in Warp** - Customization, context-awareness, and MCP integration in Warp’s AI tooling. - **46:32 – Closing reflections** - Zach and Tristan wrap up their thoughts on the future of terminals, AI, and developer tools. ## Key takeaways from this episode **Tristan Handy: Can you tell us about Warp, where the idea came from, and where you’re at today?** **Zach Lloyd:** Warp reimagines the command line to make it more approachable, powerful, and useful for developers. I've been a software engineer for over 20 years and always used the terminal, but never understood why it worked the way it did. I used to learn the minimum I needed and rely on team members when I ran into issues. After my last startup, I looked at tools I used frequently that could have a big impact if improved. The terminal stood out. I realized better UX—like being able to use a mouse to position the cursor or select output for copy-paste—could unlock a lot of productivity. That was the initial idea about five years ago. We spent the first couple of years redesigning the interface. Today, Warp is more than a terminal—it's a natural language interface to the command line, powered by large language models (LLMs). You can use it to set up projects, write code, debug production, and more. **Tristan: I want to dig into fundamentals. Can you define what a shell is?** **Zach:** A shell is a program that parses text input, runs commands, and returns text output. You can run it interactively or through scripts. Terminals, by contrast, are the graphical layer that displays text and captures keyboard input. Shells like Bash, Z Shell, and Fish offer different features, syntaxes, and configurations. Some programs like `cp` are shell built-ins, which don’t require forking new processes. **Tristan: Why do terminals persist in a GUI-dominated world?** **Zach:** A few reasons. First, it’s easier to write command-line apps than GUI apps. Second, the interface is infinitely flexible—you can pass endless flags and parameters. Third, command-line programs interoperate cleanly via text streams. And lastly, they’re scriptable. Developers can automate repetitive workflows easily, which is powerful. **Tristan: So a terminal just runs a shell. But I never think of terminals as having features. What makes a terminal more than a simple interface?** **Zach:** Terminals emulate old hardware—keyboards and text displays. Today’s terminal apps are GUI shells that simulate this behavior. Most are "dumb terminals," just rendering characters. But they can support features like theming, control characters for advanced UI (e.g., in Vim), and even bitmap rendering. **Tristan: Warp looks very different. Can you describe it?** **Zach:** Warp looks more like a chat or notebook interface. Each command's output is grouped in a logical block instead of being dumped in a scroll. The input area behaves more like a code editor, with syntax highlighting and first-class mouse support. We're aiming for modern UX. **Tristan: So you're blending terminal, editor, and chat. Will people eventually write all their code in Warp?** **Zach:** My vision is that developers will increasingly describe what they want in natural language, and agents will do the work. Developers supervise the results. That interface needs to support managing many tasks at once. That’s what we’re building towards. It won’t even be called a terminal—it’s a new category of software. **Tristan: The boundaries between these tools are blurring. And maybe the best interface for AI-assisted development isn't an IDE or chat app—it could be the terminal.** **Zach:** The terminal spans all phases of development—from setup to deployment and debugging. It also supports cross-project work, which IDEs don’t. That’s a huge strength. **Tristan: But terminals are a personal choice. How do you think about adoption and your business model?** **Zach:** Like editors, terminals are developer-choice tools. We don’t go top-down. Our motion is bottoms-up: get individuals to love Warp, then expand into teams and enterprises for security, privacy, and data controls. **Tristan: Are you trying to reset the baseline for what a terminal is?** **Zach:** We're not open source, though we’ve considered it. It’s risky. But our focus isn’t on redefining "the terminal." It’s on building the best tool for developers to ship software. That might require a new category name. **Tristan: What’s the dev experience in Warp like? Is it customizable?** **Zach:** We support theming and shortcuts. But the most important part is AI context. Warp can use any CLI tool to gather context—GitHub CLI, GCloud, etc. We’re also implementing the Model Context Protocol (MCP) and plan to better support custom/internal tools as well. --- --- title: "How to solve data collaboration challenges at scale" description: "Learn how to scale collaboration across teams and tools by standardizing workflows and aligning producers and consumers." url: "https://www.getdbt.com/blog/solving-data-collaboration-challenges" date: "2025-05-23" authors: ["Joey Gault"] categories: ["Pulse"] --- # How to solve data collaboration challenges at scale Teams across the business rely on data to make decisions—but too often, they’re working in isolation. One team’s definition of “customer” doesn’t match another’s, or a dashboard breaks because someone upstream changed a model. Solving these challenges starts with improving how teams collaborate on data: using consistent processes, shared tools, and a clear understanding of how work flows from source to insight. ## The standardization imperative Addressing data collaboration challenges requires establishing a common framework that all stakeholders can understand and use. When different teams solve similar problems using disparate approaches, organizations lose the ability to share solutions, align quickly, and collaborate seamlessly. The solution lies in standardizing on a unified transformation framework that transcends cloud providers, data platforms, and team boundaries. This standardization enables scenarios that were previously difficult to achieve: product teams deploying [dbt](https://www.getdbt.com/product/what-is-dbt) on AWS while engineering colleagues run dbt on Azure, data science teams on Databricks directly referencing dbt projects managed by finance teams on Snowflake, and central data teams building models in CLI environments that downstream marketing operations teams can investigate and extend through visual interfaces. The key insight is that standardization doesn't mean uniformity; it means establishing a common language and set of practices that work across diverse technical environments. The benefits of this approach extend beyond technical compatibility. When organizations adopt a standard transformation framework, they create opportunities for knowledge sharing, reduce duplicate work, and accelerate onboarding of new team members. More importantly, they establish the foundation for effective collaboration between data producers and consumers. ## Implementing the Analytics Development Lifecycle Effective data collaboration requires more than just standardized tools; it demands a structured approach to analytics work. The [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) provides this structure by establishing a vendor-agnostic framework that helps organizations mature their analytics workflows regardless of size or technical complexity. The ADLC draws inspiration from the Software Development Lifecycle, which successfully broke down barriers between software engineers and IT professionals in the early 2000s. By providing a standardized, repeatable framework, the SDLC enabled cross-functional teams to work together with greater agility and velocity. The analytics industry needs a similar revolution to accelerate and harden data workflows. The eight phases of the ADLC ([Plan](https://www.getdbt.com/product/dbt-mesh), [Develop](https://www.getdbt.com/product/develop), [Test](https://www.getdbt.com/product/test-and-observe), [Deploy](https://www.getdbt.com/product/deploy), [Operate & Observe](https://www.getdbt.com/product/test-and-observe), [Discover](https://www.getdbt.com/product/dbt-catalog), and [Analyze](https://www.getdbt.com/product/semantic-layer)) create a structured approach that encourages collaboration among various stakeholders. This framework helps data producers, consumers, and business stakeholders ship and use trusted data products at speed and scale. Each phase has specific objectives and deliverables that ensure all team members understand their roles and responsibilities in the broader analytics workflow. The planning phase establishes requirements and scope, while development focuses on building transformation logic. Testing ensures data quality and reliability, and deployment moves code to production environments. The operate and observe phases monitor system performance and data health, while discover and analyze phases enable stakeholders to find and use data assets effectively. This structured approach prevents the ad hoc workflows that often lead to collaboration breakdowns. ## Establishing a data control plane While the ADLC provides the process framework, effective collaboration requires a technological foundation that unifies access to data and metadata across the organization. The modern data stack's complexity (with separate solutions for orchestration, observability, catalogs, and semantic layers) often creates the very silos that collaboration efforts aim to eliminate. A [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) addresses this challenge by sitting across the entire data stack and unifying capabilities that are typically fragmented. This centralized approach consolidates metadata across the business, providing clear signals about data estate health, freshness, cost optimization, and metric definitions. The control plane serves as a single source of truth for understanding data lineage, quality, and usage patterns. The most effective data control planes exhibit three key characteristics. First, they maintain flexibility across platforms, enabling distributed teams to work with their preferred data platforms while avoiding vendor lock-in. This flexibility is crucial for organizations with complex technical requirements or those managing costs across multiple cloud providers. Second, they democratize data access by making information streamlined, accessible, and governed for users beyond the core data engineering team. This broader accessibility is essential for true collaboration, as it enables business stakeholders to engage with data assets directly rather than relying entirely on data team intermediaries. Third, they produce trustworthy outputs by providing clear visibility into data provenance, quality metrics, and troubleshooting capabilities. Trust is fundamental to collaboration; stakeholders need confidence that the data they're using is fresh, accurate, and reliable. ## Bridging producer and consumer perspectives The most persistent collaboration challenge stems from the different perspectives that data producers and consumers bring to their work. Data engineers focus on pipeline reliability, performance optimization, and technical debt management. Business stakeholders prioritize speed of insight, data accessibility, and decision-making confidence. These different priorities often create tension and miscommunication. Successful collaboration requires tools and processes that serve both perspectives simultaneously. Data producers need visibility into how their work impacts downstream users, including which models are most heavily consumed and which dashboards depend on specific data assets. They also need efficient ways to communicate data quality issues and planned maintenance to consumers. Data consumers, meanwhile, need context about data assets without requiring deep technical knowledge. They need to understand data lineage in business terms, assess data quality through clear health indicators, and identify the right datasets for their specific use cases. Most importantly, they need confidence that the data they're using is current and accurate. The solution lies in creating shared interfaces that present information relevant to both audiences. Resource pages that combine technical metadata with business context, lineage visualizations that show both technical dependencies and business impact, and health indicators that translate technical metrics into business-relevant signals all contribute to bridging this gap. ## Enabling discovery and transparency Effective data collaboration depends on stakeholders' ability to discover relevant data assets and understand their context. In large organizations, data consumers often struggle to find existing datasets that meet their needs, leading to duplicate work and inconsistent metrics. Similarly, data producers may not understand how their work is being used downstream, making it difficult to prioritize improvements and maintenance. Discovery capabilities must go beyond simple search functionality. They need to surface data assets based on business context, usage patterns, and quality indicators. Auto-generated exposures that connect dbt models to downstream BI tools provide crucial visibility into how data flows through the organization. This connectivity helps both producers and consumers understand the full impact of data assets. Query history and usage analytics add another layer of insight by revealing which models are most heavily consumed and which may be candidates for optimization or retirement. This information helps data teams prioritize their work based on actual business impact rather than assumptions about usage patterns. Transparency extends to data quality and health monitoring. Embedding health indicators directly into the tools where data is consumed (such as dashboards and reports) ensures that stakeholders have immediate visibility into data reliability. This proactive approach to quality communication prevents the trust issues that arise when stakeholders discover data problems independently. ## Scaling collaboration with federated approaches As organizations grow, centralized data teams often become bottlenecks for analytics work. The traditional model of routing all data requests through a central team doesn't scale effectively and can slow down business decision-making. However, completely decentralized approaches risk creating inconsistent definitions, duplicated work, and governance gaps. [Federated collaboration](https://www.getdbt.com/blog/federated-data-governance) models offer a middle path that combines the benefits of distributed ownership with centralized governance. In this approach, domain teams maintain ownership of their data pipelines and can choose the data platforms that best serve their needs. Meanwhile, central data teams maintain visibility into end-to-end lineage and establish global development standards. This federated approach requires technical capabilities that support cross-project references and multi-platform integration. Teams need to seamlessly reference models from other dbt projects or data platforms to avoid duplication and streamline development. They also need shared governance frameworks that ensure consistency without stifling innovation. The key to successful federation lies in establishing clear boundaries and interfaces between teams. Shared definitions for key business metrics, standardized data quality practices, and common documentation standards ensure that distributed teams can work independently while maintaining organizational alignment. ## Measuring collaboration success Organizations implementing improved data collaboration practices need clear metrics to assess their progress and identify areas for continued improvement. Traditional metrics like pipeline uptime and query performance, while important, don't capture the full picture of collaboration effectiveness. More relevant metrics include time-to-insight for business stakeholders, the percentage of data assets with clear business context and documentation, and the frequency of cross-team data reuse. Organizations should also track the reduction in duplicate data work and the speed of resolving data quality issues. User satisfaction surveys can provide qualitative insights into collaboration effectiveness. Regular feedback from both data producers and consumers helps identify friction points and opportunities for improvement. These surveys should assess not just tool satisfaction but also confidence in data quality, ease of finding relevant datasets, and effectiveness of communication between teams. The ultimate measure of collaboration success is business impact. Organizations with effective data collaboration typically see faster decision-making, more consistent metrics across teams, and increased confidence in data-driven initiatives. While these outcomes may be harder to quantify directly, they represent the true value of solving data collaboration challenges. ## Looking ahead Data collaboration challenges will continue to evolve as organizations adopt new technologies and scale their operations. The [integration of AI and machine learning capabilities into data workflows](https://www.getdbt.com/blog/traditional-to-ai-data-engineering) presents both opportunities and challenges for collaboration. While AI can automate routine tasks and provide intelligent recommendations, it also requires new forms of collaboration around model development, validation, and monitoring. The most successful organizations will be those that establish strong collaboration foundations early and continuously adapt their practices as their needs evolve. This means investing in both technological capabilities and organizational processes that support effective collaboration across diverse stakeholder groups. By focusing on standardization, structured workflows, unified control planes, and federated governance models, data engineering leaders can build the foundation for scalable, effective data collaboration that serves their organizations' growing analytical needs. ## Data collaboration FAQs **What exactly is data collaboration?** Data collaboration is the practice of enabling different teams and stakeholders to work together effectively with shared data assets across an organization. It involves establishing common frameworks, tools, and processes that allow data producers (like data engineers) and data consumers (like business stakeholders) to seamlessly share solutions, align quickly, and work together despite using different technical environments or having different priorities. **What are some best practices for data collaboration?** Key best practices include establishing standardized transformation frameworks that work across different cloud providers and platforms, implementing structured workflows like the Analytics Development Lifecycle with clear phases for planning, development, testing, and deployment, and creating shared interfaces that serve both technical and business perspectives. Organizations should also focus on enabling discovery through auto-generated connections between data models and downstream tools, embedding health indicators directly into dashboards, and adopting federated approaches that combine distributed ownership with centralized governance. **What is a data collaboration platform?** A data collaboration platform functions as a data control plane that sits across the entire data stack and unifies capabilities that are typically fragmented across separate solutions. It consolidates metadata across the business, provides clear visibility into data lineage, quality metrics, and usage patterns, and serves as a single source of truth for understanding data health and freshness. Effective platforms maintain flexibility across different data platforms, democratize access for users beyond core data engineering teams, and produce trustworthy outputs with clear data provenance and quality indicators. --- --- title: "dbt and Google Cloud: What’s new and what’s next" description: "How dbt data engineers and data scientists can leverage BigQuery easily today - and a sneak peek of what’s to come." url: "https://www.getdbt.com/blog/dbt-google-cloud-integration" date: "2025-05-22" authors: ["Kathryn Chubb"] categories: ["Product"] --- # dbt and Google Cloud: What’s new and what’s next As the industry standard for data transformation, dbt brings software engineering practices into the analytics workflow, including version control, testing, CI/CD deployments, data lineage, and other features. As part of our commitment to our customers, we continue to evolve dbt to remain vendor-independent and work seamlessly with the major data storage and cloud providers in the industry. That’s why so many Google Cloud customers choose dbt for data transformations with Google BigQuery. Besides integrating tightly with BigQuery and Google Cloud, dbt works with a large number of third-party applications. This enables Google Cloud developers to integrate other data applications easily with the Google Cloud ecosystem. dbt and Google Cloud are working closely together to provide even better support for BigQuery from within dbt-powered data pipelines. We’ll review how dbt supports BigQuery today and what features you can expect to see in the near future. **** ## Integrating dbt and BigQuery gets easier dbt’s support for [a mature analytics development lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle) enables data and analytics engineers to produce and ship high-quality data in a collaborative, governed manner. Using standard data engineering languages, such as SQL and Python, data engineering teams can gain more control and power over their data workflows than they’ve had before. dbt is validated for Google Cloud BigQuery. That means it’s an officially recognized partner solution that has fulfilled a specific set of requirements to ensure the best possible integration and performance. We’ve worked closely with the Google team to build dbt from the ground up for Google Cloud. In addition to blazing performance, users can harness specific BigQuery-specific features such as partitioning, clustering in BigQuery ML, and more. Today, BigQuery is the second largest adapter for dbt, with tens of thousands of projects in production. Multiple Google Cloud customers - including [Rocket Money](https://www.getdbt.com/case-studies/rocket-money), [Bilt Rewards](https://www.getdbt.com/case-studies/bilt-rewards-professional-services), and [Virgin Media O2](https://www.getdbt.com/blog/virgin-media-o2-dbt) - have leveraged dbt and BigQuery for their data transformation workflows. Customers report that using dbt lowers BigQuery ramp-up times and results in increased productivity, reduced errors, and faster delivery. Getting started with dbt from Google Cloud is easy, thanks to the latest integration released by the Google Cloud and dbt teams. If you don’t have a dbt account, you can create one by navigating to the Google Cloud Partner Center. From there, you can establish a connection between dbt and BigQuery. You can authorize dbt to connect to BigQuery using either a [service account key](https://cloud.google.com/iam/docs/keys-create-delete) downloaded in a JSON file or OAuth. From there, you can create a new model in the dbt IDE and run it to see how your transformation shows up in BigQuery. dbt stores your changes in a Git repository, so if you need to roll back a change, you can do so quickly and easily. ## Using BigQuery from dbt - no upskilling required Have your Python developers ever come across scenarios where their local runtime is just not large or powerful enough to process their Python data? When processing runs into terabytes of data, a single Python runtime often isn’t enough. [BigQuery contains built-in support for DataFrames](https://cloud.google.com/bigquery/docs/use-bigquery-dataframes), providing a Pythonic DataFrame and machine learning (ML) API powered by the BigQuery engine. This gives Python developers a Pandas or Scikit-like interface that automatically transpiles to SQL for server-side execution. This means developers can now leverage the distributed power of BigQuery and the scale of BigQuery from their laptops. A data scientist can go from tens of gigabytes of data to terabytes of data without having to switch on infrastructure. BigQuery accomplishes this with several libraries that provide extensions above and beyond the default Pandas and Scikit features: - **bigframes.pandas**: Transpiles Python into BigQuery SQL - **bigframes.ml**: Transpiles Scikit code to BigQuery ML, BigQuery’s native machine learning language, which can handle featurization, training, and batch prediction - **bigframes.ml.llm**: Provides access to large language models and features such as multimodal processing So how does this come together within dbt? On top of SQL-based transformations, [dbt supports Python models directly](https://docs.getdbt.com/docs/build/python-models). BigQuery’s dbt adapter automatically converts this Python code to work with BigFrames, the BigQuery DataFrames implementation. This means that data scientists and other Python developers who use dbt can take advantage of the power of BigQuery without any upskilling or direct provisioning of infrastructure. dbt users can also use dbt features such as [incremental materialization](https://docs.getdbt.com/docs/build/incremental-models), and bring in their own custom Python code as user-defined functions (UDF). To see this in action, [check out the demo in the webinar](https://www.getdbt.com/resources/webinars/dbt-bigquery-whats-new-from-google-next) (demo begins near the 34-minute mark). You can also try it yourself using [our quickstart for BigQuery DataFrames with dbt Python models](https://docs.getdbt.com/guides/dbt-python-bigframes?step=1). ## Other dbt and BigQuery features Beyond this, there are a number of improvements coming for dbt and BigQuery integration. Google Cloud is currently rolling out support for Workload Identity Federation. This will increase security for dbt/BigQuery connections by issuing ephemeral, short-lived tokens for any deployment credentials. This eliminates the need to download a JSON credentials file, reducing [the risk of credential leaks](https://reliaquest.com/blog/service-account-abuse/). Additionally, Google Cloud is rolling out support for [Private Service Connect](https://cloud.google.com/vpc/docs/private-service-connect) for BigQuery. This means that customers can access Google Cloud managed services via private endpoints in their VPC networks, meaning traffic never traverses the public Internet. This will make your network security teams a lot happier for those workloads that require advanced security. A huge change that has both the teams at dbt and Google Cloud excited is support for the Iceberg open table format. [Apache Iceberg](https://iceberg.apache.org/) is a high-performance open table format developed for modern data lakes. [dbt currently supports Iceberg](https://docs.getdbt.com/blog/icebeg-is-an-implementation-detail), and we’re excited to bring this support to BigQuery in the next few months. Finally, down the road, we’ll be integrating cost monitoring for BigQuery and other data warehouses into dbt. A lot of customers have asked us for a way to understand how dbt is driving data warehouse costs. (We hear this in general - not just for BigQuery.) Using this information, you can optimize costs and create data governance policies around data warehouse usage. ## A thriving community Another exciting aspect of the collaboration between dbt and BigQuery is the community support. There are a ton of interesting community projects involving the two technologies that are up and coming. For example, there’s the [BigQuery ML for dbt](https://github.com/kristeligt-dagblad/dbt_ml) project that started in 202 and is currently under active development, which enables users to train, audit, and use BigQuery ML models from inside of dbt projects. We’re honored that so many BigQuery users utilize dbt for their data transformation pipelines. The dbt and Google Cloud teams are excited to continue working together on even deeper integrations to make the dbt + BigQuery experience as easy and seamless as possible. --- --- title: "Who is responsible for data governance?" description: "Governance isn’t just a team—it’s everyone’s job. Learn which roles own what parts of data governance today." url: "https://www.getdbt.com/blog/data-governance-ownership" date: "2025-05-22" authors: ["Joey Gault"] categories: ["Pulse"] --- # Who is responsible for data governance? ## The evolution from centralized to distributed governance [Traditional data governance](https://www.ibm.com/think/topics/data-governance) followed a top-down, centralized model where a dedicated governance team established policies and data stewards enforced them across the organization. While this approach provided clear accountability and consistent standards, it often created bottlenecks that couldn't keep up with the speed of modern business operations. The limitations of purely centralized governance become apparent when organizations scale. Data continues to grow exponentially year over year, and the emergence of AI applications has only increased the demand for diverse, high-quality datasets. A small central team simply cannot manage the complexity and volume of data decisions required across a large enterprise. [Modern data governance has shifted toward a federated model](https://www.getdbt.com/blog/federated-data-governance) that distributes responsibility while maintaining centralized standards and oversight. This approach recognizes that the people closest to the data often have the best understanding of its quality, usage patterns, and business context. By empowering domain teams to take ownership of their data while providing them with standardized tools and frameworks, organizations can achieve both scale and quality. This federated approach enables what's often called a "shared, community effort" where data producers, consumers, and governance experts collaborate to create high-quality datasets. Rather than governance being something imposed from above, it becomes embedded in the daily workflows of data teams across the organization. ## Executive leadership and strategic oversight At the highest level, data governance requires executive sponsorship and strategic direction. Chief Data Officers (CDOs), Chief Information Officers (CIOs), and other C-suite executives play a critical role in establishing the governance vision, securing resources, and ensuring alignment with business objectives. Executive leadership is responsible for defining the overall governance strategy, including the balance between centralized control and distributed ownership. They must also ensure that governance initiatives receive adequate funding and that governance considerations are integrated into broader business planning processes. Perhaps most importantly, executives set the tone for data culture across the organization. When leadership demonstrates commitment to data quality, security, and ethical use, it signals to the entire organization that governance is a business priority rather than just a technical requirement. Executive oversight becomes particularly important when governance decisions have cross-functional implications or when conflicts arise between different teams' data needs. Having clear executive accountability ensures that governance issues receive appropriate attention and resources for resolution. ## Data stewards and domain expertise Data stewards serve as the operational backbone of any governance program. These individuals act as the bridge between high-level governance policies and day-to-day data management activities. They are responsible for implementing governance standards within their specific domains while serving as liaisons between different teams to resolve data issues. The role of data stewards has evolved significantly in modern governance models. Rather than simply enforcing policies created by others, today's data stewards are expected to be active participants in defining governance standards based on their deep understanding of business requirements and data usage patterns. Data stewards are collectively responsible for defining and documenting data assets, ensuring data quality, and promoting effective data sharing across the organization. They also play a crucial role in ensuring that governance policies are practical and implementable rather than theoretical constructs that don't work in real-world scenarios. In organizations using tools like [dbt](https://www.getdbt.com/product/what-is-dbt), data stewards often work closely with analytics engineers to implement governance controls directly in data transformation workflows. This embedded approach ensures that [governance becomes part of the natural development process](https://www.getdbt.com/product/governance) rather than an additional burden. ## Analytics engineers and technical implementation [Analytics engineers](https://www.getdbt.com/blog/what-is-analytics-engineering) have emerged as key players in modern data governance, particularly in organizations that have adopted tools like [dbt](https://www.getdbt.com/product/dbt) for data transformation. These professionals bridge the gap between traditional data engineering and business analysis, making them ideally positioned to implement governance controls that serve both technical and business requirements. Analytics engineers are responsible for implementing data quality tests, documentation standards, and lineage tracking within data transformation workflows. They work with data stewards to translate business requirements into technical controls and ensure that governance standards are consistently applied across all data models. The role of analytics engineers in governance extends beyond just implementation. They often serve as advocates for governance best practices, helping to educate other team members about the importance of testing, documentation, and code review processes. Their technical expertise combined with business understanding makes them effective champions for governance initiatives. When using [dbt, analytics engineers can standardize governance practices](https://docs.getdbt.com/docs/mesh/govern/about-model-governance) across teams by creating reusable models, implementing consistent testing frameworks, and establishing clear documentation standards. This technical standardization supports the broader governance objectives while making it easier for teams to collaborate and share data assets. ## Data engineering teams and infrastructure Data engineering teams play a foundational role in governance by building and maintaining the infrastructure that supports governance activities. They are responsible for implementing security controls, access management systems, and the technical frameworks that enable other teams to practice good governance. Data engineers work closely with security teams to implement role-based access controls, audit logging, and other technical safeguards that protect sensitive data. They also build and maintain the data pipelines that move information throughout the organization, ensuring that these systems include appropriate monitoring and quality controls. The infrastructure decisions made by data engineering teams have far-reaching implications for governance. Choices about data storage formats, processing frameworks, and integration patterns all affect how easily governance controls can be implemented and maintained across the organization. In modern data stacks, data engineers increasingly focus on building platforms that enable self-service analytics while maintaining appropriate governance controls. This might involve implementing tools like [dbt that allow analysts to work independently](https://www.getdbt.com/product/analyst) while ensuring that all transformations go through proper testing and review processes. ## Business stakeholders and data consumers Business stakeholders and data consumers have important governance responsibilities, even though they may not think of their role in governance terms. These users are often the first to notice data quality issues, and their feedback is essential for maintaining and improving governance standards. Business users are responsible for understanding and following data usage policies, particularly those related to privacy and security. They also play a crucial role in validating that governance controls are working effectively by reporting issues when data doesn't meet their expectations. The relationship between business stakeholders and governance teams has become more collaborative in recent years. Rather than simply consuming data that others have prepared, business users are increasingly involved in defining data requirements, validating data quality, and participating in governance decisions that affect their work. This increased involvement requires business stakeholders to develop a better understanding of [data governance concepts](https://www.databricks.com/discover/data-governance) and their role in maintaining data quality. Organizations that invest in governance education for business users often see better outcomes from their governance programs. ## Compliance and legal teams [Compliance](https://docs.getdbt.com/faqs/Accounts/account-specific-features#bring-your-own-key-byok-) and legal teams have become increasingly important in data governance as regulations like GDPR, CCPA, and industry-specific requirements have proliferated. These teams are responsible for ensuring that governance programs meet all applicable regulatory requirements and that data handling practices align with legal obligations. Legal teams work with technical teams to translate regulatory requirements into specific technical controls and processes. They also provide guidance on data retention policies, privacy requirements, and cross-border data transfer restrictions that must be built into governance frameworks. The role of compliance teams extends beyond just ensuring regulatory adherence. They also help organizations understand the risk implications of different governance decisions and provide guidance on balancing business needs with compliance requirements. As AI applications become more prevalent, compliance and legal teams are taking on additional responsibilities related to algorithmic fairness, bias detection, and explainability requirements. These new challenges require close collaboration with technical teams to ensure that governance frameworks can support responsible AI development. ## The role of modern tools in distributed governance Modern data governance tools play a crucial role in enabling distributed responsibility models. Tools like [dbt](https://www.getdbt.com/product/dbt) allow organizations to embed governance controls directly into data development workflows, making it easier for individual contributors to practice good governance without requiring extensive oversight. These tools support governance at scale by combining automation with human oversight. They enable data producers and consumers to manage large volumes of data effectively while ensuring compliance with governance standards. Features like automated testing, documentation generation, and lineage tracking reduce the manual effort required to maintain governance standards. The availability of sophisticated governance tools has changed the skill requirements for governance roles. Rather than needing dedicated governance specialists for every domain, organizations can now train analytics engineers and other technical contributors to implement governance controls as part of their regular work. This democratization of governance capabilities allows organizations to scale their governance programs without proportionally increasing their governance headcount. However, it also requires investment in training and change management to ensure that distributed teams understand and embrace their governance responsibilities. ## Building accountability in distributed models Successfully distributing governance responsibility requires clear accountability structures that define who is responsible for what aspects of governance. This includes establishing clear roles and responsibilities, defining escalation paths for governance issues, and creating metrics that track governance effectiveness across different domains. Organizations need to balance autonomy with accountability, giving teams the freedom to make governance decisions within their domains while ensuring that these decisions align with broader organizational standards. This often involves creating [governance frameworks](https://www.getdbt.com/blog/data-governance-frameworks-ai) that provide clear guidelines while allowing for domain-specific adaptations. Regular communication and coordination between different governance stakeholders is essential for maintaining alignment and addressing cross-functional governance challenges. This might involve regular governance committee meetings, cross-team reviews of governance practices, or shared dashboards that provide visibility into governance metrics across the organization. The most successful distributed governance models create a culture where governance is seen as everyone's responsibility rather than something that belongs to a specific team. This cultural shift requires ongoing investment in education, communication, and recognition of good governance practices across the organization. ## Conclusion The responsibility for data governance has evolved from a centralized, top-down model to a distributed approach that engages stakeholders across the organization. While executive leadership provides strategic direction and data stewards coordinate implementation, the day-to-day practice of governance increasingly relies on analytics engineers, data engineers, and even business users who understand their role in maintaining data quality and security. Modern tools like [dbt](https://www.getdbt.com/product/dbt) have made this distributed model more practical by embedding governance controls directly into data development workflows. This allows organizations to scale their governance programs while maintaining high standards for data quality, security, and compliance. Success in this distributed model requires clear accountability structures, ongoing education, and a culture that values governance as a shared responsibility. Organizations that can effectively balance centralized oversight with distributed ownership will be best positioned to meet the growing demands for high-quality, well-governed data in an AI-driven business environment. The question of who is responsible for data governance doesn't have a simple answer, but the most effective approach involves everyone who touches data taking ownership of governance within their domain while working together toward common standards and objectives. ## Data governance FAQs **What is data governance, and how does it differ from broader data management?** Data governance is a strategic framework that establishes policies, roles, and responsibilities for managing data quality, security, and compliance across an organization. Unlike broader data management, which focuses on the technical aspects of storing, processing, and moving data, governance specifically addresses the "who, what, when, and how" of data decision-making. It involves creating accountability structures, defining standards, and ensuring that data practices align with business objectives and regulatory requirements. While data management handles the infrastructure and technical implementation, governance provides the oversight, policies, and cultural framework that guides how data is used responsibly throughout the organization. **What roles, responsibilities, and policies should a data governance framework include to ensure data quality, security, and compliance?** A comprehensive data governance framework should include executive leadership (CDOs, CIOs) who provide strategic direction and secure resources, data stewards who serve as the operational backbone implementing policies within specific domains, and analytics engineers who embed governance controls into data workflows. Technical roles include data engineering teams responsible for infrastructure and security controls, while business stakeholders and compliance teams ensure practical usability and regulatory adherence. Key policies should cover data quality standards, access controls, privacy requirements, documentation standards, and clear escalation paths for governance issues. The framework must also establish accountability structures that balance centralized oversight with distributed ownership, enabling domain teams to make governance decisions while maintaining organizational alignment. **How can organizations enforce data governance across hybrid and multicloud environments while balancing self-service access with privacy and security?** Organizations can enforce governance across distributed environments by implementing federated governance models that combine centralized standards with distributed responsibility. This involves using modern tools that embed governance controls directly into data development workflows, enabling automated testing, documentation, and lineage tracking without requiring extensive manual oversight. Technical implementation includes role-based access controls, audit logging, and standardized frameworks that work across different cloud platforms. The key is creating governance platforms that enable self-service analytics while maintaining appropriate security controls, such as implementing tools that allow analysts to work independently while ensuring all transformations go through proper testing and review processes. Success requires clear accountability structures, ongoing education, and building a culture where governance is viewed as everyone's responsibility rather than a barrier to data access. --- --- title: "How to build reliable data pipelines with data quality checks" description: "A guide to implementing data quality checks across pipelines, plus how dbt can help automate testing, alerting, and documentation." url: "https://www.getdbt.com/blog/data-pipeline-quality-checks" date: "2025-05-19" authors: ["Daniel Poppy"] categories: ["Learn"] --- # How to build reliable data pipelines with data quality checks High-quality data is the foundation of trustworthy analytics. But without automated checks, even modern pipelines can quietly deliver incomplete, stale, or incorrect information—leading to costly business decisions. This post walks through how to implement practical, scalable data quality checks in your pipelines. You’ll learn the key dimensions of data quality, which tests matter most, where to place them in development and production, and how tools like dbt can help you enforce standards with automation and CI/CD. ## Understanding data quality dimensions Before you can enforce data quality, you need to define what “quality” means in context. Here are the core dimensions that form a practical framework: - **Accuracy (Correctness):** Does the data reflect real-world values? For example, a product’s listed price should match its actual sale price. - **Completeness:** Are all required fields populated? Missing IDs, emails, or timestamps can break downstream processes. - **Validity:** Does the data conform to expected formats, ranges, or business rules? Think of dates, enum values, or transaction types. - **Consistency:** Are values uniform across systems and datasets? Inconsistent naming or duplication introduces conflict and confusion. - **Freshness (Timeliness):** Is the data current enough to support decisions? Delayed updates make reports misleading. - **Uniqueness:** Are entities (like order IDs or user accounts) represented only once? Duplicates lead to inflated metrics or failed joins. Together, these dimensions help you define what “good” looks like for your data. By embedding checks that enforce them throughout your pipelines, you build trust—not just in your data, but in the decisions it drives. **** ## Essential data quality checks for your pipelines Implementing a few strategic tests can catch most issues before they impact business decisions. Here are the foundational data quality checks every team should include: ### Uniqueness tests Duplicate data distorts analysis and leads to incorrect business insights. Uniqueness tests make sure values in key columns appear only once in your dataset. For example, in a sales system, order IDs should never be duplicated, as this could cause revenue to be counted twice. These tests are straightforward to implement in most data tools. When a uniqueness test fails, it immediately signals potential data corruption or process issues that require attention. Regular uniqueness checks prevent downstream reporting errors. Implementing uniqueness checks early in your pipeline catches problems before they spread through your data ecosystem. These checks form a basic but critical part of your data quality system. ### Non-null tests Missing values in critical fields can break processes and create incomplete insights. Non-null tests verify that essential data is always present. For instance, in customer records, fields like ID, email, and signup date typically must contain values. These tests help catch data entry issues or system failures that lead to incomplete records. Data with proper completion rates is more reliable for analysis and operational use. Non-nullness tests are simple to implement but deliver significant value. By ensuring critical fields are always populated, you prevent many common data problems before they affect business operations. ### Accepted values tests Data often needs to fall within specific categories or ranges to be meaningful. Accepted values tests enforce these boundaries. For example, a financial system might need all transaction types to be one of "deposit", "withdrawal", or "transfer" – any other value would indicate a problem. These tests catch both technical failures and user input errors. They help maintain data consistency across systems and prevent nonsensical analysis. When combined with business rules, they ensure data matches operational reality. Implementing accepted values tests creates guardrails that keep your data aligned with business expectations and technical requirements. ### Referential integrity tests As data moves through transformations, relationships between tables must remain intact. Referential integrity tests verify that foreign keys in one table exist as primary keys in related tables. For instance, every product ID in a sales table should exist in the products master table. When these relationships break, reports can show incomplete information or fail entirely. Maintaining proper connections between data entities ensures accurate joins and aggregations. These tests help prevent the "missing data" problems that often puzzle end users. Regular referential integrity checks maintain the connectedness of your data model, making all downstream analysis more reliable. ### Freshness and recency tests Outdated information can lead to poor decisions. Freshness tests verify that data is updated on schedule and remains current. For instance, a sales dashboard becomes useless if yesterday's transactions haven't loaded properly. These tests often check timestamps to ensure recent updates have occurred. They can trigger alerts when data flows stop or slow down. Freshness checks are particularly important for time-sensitive business processes. By monitoring the timeliness of your data, you ensure business users always have current information for decision-making. ## Where to implement data quality checks ### During development Quality starts with good design. Implementing checks during development helps catch issues before they hit production. Developers should explore raw source data to understand its baseline quality, then test transformations to ensure they preserve or improve that quality. Test-driven development (TDD) works well for data. Write tests before building transformations to clarify expected outcomes and catch logic errors early. **** Early checks create a solid foundation for reliable downstream analytics. ### During pull requests Code review is a critical gate for quality. Running automated tests during pull requests (PRs) ensures changes won’t break existing models or dashboards. PRs are an ideal time to validate assumptions, review test coverage, and ensure new logic integrates cleanly with the existing project. With dbt Cloud, you can [enable CI jobs](https://docs.getdbt.com/docs/deploy/continuous-integration) to run your tests on every pull request—automatically. Making tests part of your merge criteria helps shift data quality left and builds a stronger team culture around trust. ### In production Even with perfect development workflows, things can break in production. Scheduled tests in your production environment catch issues from upstream schema changes, unexpected edge cases, or pipeline failures. Production tests should run after every data refresh and before dashboards or reports go live. Use alerts to notify the right people when issues arise—and block data from flowing downstream if needed. Ongoing production validation ensures your data stays accurate as systems evolve. ## Best practices for data quality implementation ### Start simple and expand Start with high-impact tests—like uniqueness and non-null constraints—on your most critical tables. These catch common issues quickly and deliver immediate value. As your framework matures, layer in more sophisticated checks (e.g., accepted values, referential integrity). This incremental approach keeps the implementation manageable and builds team momentum. Early wins help build trust in the process and secure buy-in for deeper investment in quality. ### Automate testing Manual checks don’t scale. Automating your data quality tests ensures consistency across datasets and frees up engineers to focus on higher-value work. Most modern data platforms—including [dbt](https://www.getdbt.com/product/dbt)—offer native testing frameworks that integrate seamlessly with your transformation workflows. Automation creates a safety net that works even as teams, tools, or data change. ### Alert on failures Testing is only useful if someone sees the results. Set up alerts to notify the right stakeholders when issues arise—whether through email, Slack, or incident response tools. Tailor alerts based on severity. Some failures might trigger automated data blocks; others might open tickets for triage. Proactive alerting turns passive monitoring into active data reliability management. ### Document your tests Clear documentation helps everyone understand what tests exist and why they matter. For each quality check, document what it verifies, why it's important, and what action to take if it fails. This documentation serves both immediate operational needs and onboarding of new team members. It preserves knowledge about data expectations even as teams change. Good documentation also helps business users understand the quality measures protecting their data. Treating test documentation as a first-class deliverable improves team alignment and operational response to issues. ## Common challenges and solutions Implementing data quality checks isn’t without its challenges. Some of the most common issues include: **Handling exceptions.** Not all data that fails a rule is wrong—edge cases, legacy formats, or evolving business logic may require flexibility. Instead of weakening standards across the board, consider scoped exceptions or custom tests that handle these cases without sacrificing trust. **Balancing coverage with performance.** Comprehensive testing can slow data delivery. Focus frequent checks on critical fields and logic, while running heavier validations during off-peak hours or staging environments. **Adapting to change.** As your business evolves, so do your schemas and expectations. Design your test framework to be modular and easy to update, so your quality coverage evolves with your data. **** **** **Fostering a culture of quality.** Perhaps the hardest challenge is cultural. When teams see testing as overhead, adoption suffers. The shift happens when quality is clearly tied to business outcomes—like faster delivery, fewer dashboard errors, and higher trust in data. ## Conclusions Robust data quality checks are essential to building trust in your analytics and driving confident business decisions. By focusing on key quality dimensions—like accuracy, completeness, and freshness—and placing checks strategically throughout your pipelines, you can significantly improve data reliability. Start simple, automate where possible, and set up alerting so issues are surfaced and resolved quickly. As your systems evolve, so should your tests. Treating quality as a continuous process, not a one-time task, ensures long-term resilience. The payoff? Less time firefighting, faster insights, and stronger trust across your organization. 🔗 _Want to start testing your data today? [Explore how dbt automates testing and monitoring](https://docs.getdbt.com/docs/build/tests)._ ## Data quality FAQs **What is a data quality check?** A data quality check is a process or test implemented within data pipelines to verify that data meets specified quality standards. These checks ensure that data is reliable for analytics and business decisions by validating various aspects such as uniqueness, completeness, validity, consistency, and freshness. Implementing these checks at various stages of data processing helps catch issues early, preventing inaccurate information from reaching stakeholders and eroding trust in data teams. **What are the 4 C's of data quality?** The 4 C's of data quality are: 1. Correctness - ensuring the data is accurate and matches reality 2. Completeness - verifying all expected data is present 3. Consistency - confirming the data is consistent across different systems and datasets 4. Currency (often referred to as freshness) - checking how up-to-date the data is These dimensions form a framework for evaluating and maintaining high-quality data throughout organizations. **What are the 5 elements of data quality?** The 5 elements of data quality are: 1. Correctness (or accuracy) - Is the data accurate and does it match reality? 2. Completeness - Is all the expected data present? 3. Validity - Does the data conform to defined formats and rules? 4. Consistency - Is the data consistent across different systems and datasets? 5. Freshness (or timeliness) - How up-to-date is the data? These dimensions help organizations address different aspects of quality in their data management practices. **What are the 6 measures of data quality?** The 6 measures of data quality are: 1. Accuracy - Does the data correctly represent the real-world entity or event it describes? 2. Completeness - Is all necessary data present without gaps? 3. Consistency - Is data uniform across different datasets and systems? 4. Timeliness - Is the data current and updated at appropriate intervals? 5. Validity - Does the data conform to required formats, ranges, and business rules? 6. Uniqueness - Is each entity represented once without duplication? These measures provide a comprehensive framework for assessing data quality and implementing appropriate checks throughout data pipelines. --- --- title: "Introducing the New Cloud Architect Certification" description: "dbt Labs has launched its latest certification exam, the Cloud Architect Certification." url: "https://www.getdbt.com/blog/introducing-the-new-cloud-architect-certification" date: "2025-05-16" authors: ["Stephen Robb"] categories: ["Learn"] --- # Introducing the New Cloud Architect Certification Today, we're launching the newest addition to the dbt Certification Program: the dbt Cloud Architect Certification. This new certification is designed for experienced data professionals ready to take their dbt skills beyond model building and into full-platform architecture — enabling scalable, secure, and governed data operations using dbt Cloud. This is the first credential specifically focused on the architecture and operationalization of dbt in enterprise environments. Whether designing workflows across teams, managing environments and permissions, or scaling dbt in production, this certification validates the knowledge and skills you need to do it confidently. Earning this certification is a signal to your current team, and the next one, that you take the craft of analytics engineering seriously and are committed and qualified to build data the right way. ## Why get dbt Cloud Architect Certified? Over the past few years, dbt Cloud has become the central control plane of modern data teams. It powers CI/CD pipelines, orchestrates transformations, and integrates with the entire data and analytics ecosystem. With that growth has come a new set of roles and responsibilities: from managing job schedules and development lifecycles to integrating dbt with tools like GitHub, BigQuery, and Snowflake. The dbt Cloud Architect Certification helps establish a standard set of best practices — and recognizes professionals who can confidently design and operate robust dbt Cloud deployments at scale. As a dbt Cloud Architect, you’ll be able to: - **Demonstrate platform mastery**, from environment management and access control to production job design and governance. - **Showcase your expertise** in deploying dbt Cloud as a mission-critical tool across teams and domains. - **Stand out to employers and clients** looking for experienced dbt practitioners who can lead architecture decisions and scale their data development workflows. - **Earn a digital badge** backed by dbt Labs to share on LinkedIn, resumes, and in your org’s internal learning programs. ## Who is the dbt Cloud Architect Certification for? This certification is designed for: - Analytics engineers and data platform engineers who manage dbt Cloud deployments - Data team leads responsible for productionizing dbt workflows - Consultants and solution architects guiding organizations through dbt Cloud adoption - Anyone who wants to prove their ability to scale dbt Cloud in the real world ## Where this certification fits in Think of the current certification landscape like a specialized toolkit. You might already have tools focused on data analysis, cloud fundamentals, or specific database technologies. The dbt Cloud Architect Certification serves as a unique, trusted, credible, and invaluable tool for demonstrating your deep expertise. ## Ready to get certified? You can register now to take the dbt Cloud Architect Certification Exam via online proctoring, with appointments starting now! We’ve also put together a full study guide, practice questions, learning path, and exam overview to help you prepare. We can’t wait to see what you build with dbt Cloud — and we’re proud to recognize the architects leading the way. 👉 [Register for the dbt Cloud Architect Exam](https://pages.talview.com/dbtlabs/certifications/) 👉 [View the Study Guide](https://www.getdbt.com/dbt-assets/dbt-certificate-study-guide-for-cloud-architect) 👉 [Get Started on the Learning Path](https://www.getdbt.com/certifications/dbt-architect-certification-exam) --- --- title: "The state of analytics engineering in 2025: A summary" description: "Discover key trends shaping analytics engineering in 2025—from AI adoption to data trust and evolving team roles." url: "https://www.getdbt.com/blog/state-of-analytics-engineering-2025-summary" date: "2025-05-16" authors: ["Daniel Poppy"] categories: ["Insights"] --- # The state of analytics engineering in 2025: A summary Every year, dbt Labs releases our [State of Analytics Engineering report](https://www.getdbt.com/resources/state-of-analytics-engineering-2025). In our latest 2025 report, we identified a number of key trends, including: - A growth in data investment after a period of caution; - The use of AI to augment (not replace) data teams; and - A continued concern over data quality and an emphasis on building trust in data. We also wanted to get some thoughts from leaders in the industry on what they’re seeing beyond these numbers. We held a webinar roundtable recently featuring our own Senior Manager, Developer Experience, Jason Ganz, as well as Yannick Misteli, Head of Engineering, Go-to-Marke at Roche and Jenna Jordan, Senior Data Management Consultant at Analytics 8. [You can watch the full webinar yourself](https://www.getdbt.com/resources/webinars/2025-state-of-analytics-engineering-virtual-event). Here are some of the main takeaways from our participants. ## The biggest changes in analytics engineering In all, 2024 was a transition period for analytics. Instead of any major new developments, the industry mostly laid the groundwork for the next transformations to come. Given how overwhelming sudden change can be, this stability was refreshing. For the most part, organizations in 2024 focused on developing fully mature data pipelines. Observability and data quality became more prominent concerns across the board. Others were excited by some of the changes announced by dbt Labs over the past year. In particular, more large organizations took advantage of [dbt Mesh](https://docs.getdbt.com/best-practices/how-we-mesh/mesh-1-intro). Enterprise found that new features incorporated into Mesh enabled them to manage a growing number of data transformation models, and introduce the additional flexibility and distributed approach to data [that a data mesh architecture offers](https://www.getdbt.com/blog/what-is-data-mesh). There was also excitement [around dbt Labs’ acquisition of SDF](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs). SDF is a high-performance toolchain that can represent various SQL dialects. That means SDF can faithfully emulate all popular data warehouses locally, providing data pipeline developers with immediate feedback before they even run a single line of SQL. ## The analytics engineer role in 2025 [Analytics engineering](https://www.getdbt.com/blog/what-is-analytics-engineering) is a fairly new industry whose practitioners focus on providing clean data sets to end users, using [software engineering best practices](https://www.getdbt.com/resources/the-analytics-development-lifecycle) to maintain a clean analytics codebase. Originally, this role took on some of the scope of data analysts and some of the scope of data engineers. Different organizations have chopped this up differently. The result is we’re seeing more of a blurring of lines between these three roles. An analytics engineer in one org, for example, may continue to move left and take on more traditional data engineering tasks. Others, however, may move more to the right and start interfacing more often with data stakeholders, bringing more of the business perspective into their work. We’ve seen similar movements in other roles as well. Some data analysts, for example, have been drifting more leftward, taking on dbt modeling and some of the other tasks typically done by data and analytics engineers. At larger organizations, however, we’re seeing a clearer delineation between these roles. At these companies, analytics engineers are becoming more focused on data modeling. We’re also seeing the emergence of visual analytics engineers who excel at turning data into reports in tools such as [Qlik](https://www.qlik.com/us), [ThoughtSpot](https://www.thoughtspot.com/), and [Tableau](https://www.tableau.com/). Typically, at such organizations, we see a 3:2:1 ratio between analysts, analytics engineers, and visual analytics engineers. **** ## AI's impact on analytics engineering AI isn’t just changing how we do business. It’s changing analytics engineering as well. Our latest report found that 80% of data practitioners are using AI in some way as part of their workflows. AI is impacting all roles in the data lifecycle - data and analytics engineers, analysts, and business stakeholders. Participants have found they can use AI to: - Better define requirements - Improve code quality and efficiency - Connect BI tools to data using text-to-SQL capabilities The primary use of AI currently in data workflows is to reduce drudgery and generate code (dbt YAML files, for example) for repetitive tasks. Offloading the drudgery to tools like [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot) has given analytics engineers and data engineers more time to focus on the creative side of their work. This can involve an increased focus on data handling, for example, or on revising the company’s data architecture. Additionally, AI is empowering more people who are just starting out on their journey with dbt or even SQL. Even if someone doesn’t know exactly what goes into a dbt project YAML file, for example, they can use AI to give them a boost and move up the stack, enabling them to contribute to the company’s dbt projects more directly. ### The importance of context in AI applications On the other hand, in many ways, we’re not there yet. We can ask LLMs any questions we may have about our data. But they don’t always return the correct answers. We continue to see accuracy issues with AI when LLMs and AI agents don’t have full access to our codebases or standards. Passing data context to LLMs - e.g., using technologies such as [Model Context Protocol (MCP)](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) - will become increasingly important over the next year. [The semantic layer](https://www.getdbt.com/product/semantic-layer) also has a role to play in improving AI results over your data. In the past, different departments in an organization have had their own, slightly divergent calculations for key data points, such as sales volume or revenue. A semantic layer standardizes vocabulary and meaning for these key metrics across teams. Connecting LLMs to this layer is one way of increasing the likelihood that, when we ask them a question, we’ll get the right answers back. Finally, good documentation will become more important than ever. Writing out detailed descriptions of models, fields, etc. provides more context to a query engine on how to query that data. Analytics 8’s Jordan said she’s not a huge fan of using AI for documentation. In many cases, however, she says data teams can get a limited Ouroboros effect going by using AI itself to stub out some of the initial documentation (which you then refine and supplement with your own human knowledge). Roche’s Misteli concurred by citing his team’s experience in learning the importance of documentation firsthand. Roche built a chatbot on top of its technical documentation. Engineers panned it, however, saying it wasn’t useful. After analyzing why, Roche came to a simple conclusion: its documentation wasn’t up to snuff. That led to an effort to clean up the docs and make them more valuable - for both humans and LLMs. “If your documentation isn’t good,” Misteli concluded, “your chatbot won’t be, either.” ## Data quality remains a top issue Surprising to no one who deals with data for a living, data quality still remains a top issue. 56% of survey respondents identified data quality as a problem. Participants emphasized that data quality isn’t just a technology problem - it’s a process problem. You need [a mature analytics workflow](https://www.getdbt.com/blog/adlc-plan) in place to find and fix problems early in order to minimize their business impact. Jordan from Analytics 8 agreed and said the key is “bringing your analytics systems closer to your operational systems and organizing by domain.” That includes using data mesh architectures, data contracts, data quality testing, and defined Service Level Agreements (SLAs) for data. It’s also important to get all data stakeholders in the room early on - i.e., when the data contracts are being written. ## Predictions for the future In terms of what’s coming up in the near future, the fusion of AI agents with semantic web technology holds a lot of promise in realizing the true value of AI agents. Connecting these agents to the ontologies that exist inherently in our web-based data could provide agents with true semantics and reasoning capabilities. On the analytics workflow front, improvements such as SDF and the Visual Studio add-on for dbt promise to cut down dev time for analytics code changes even further. We should also expect to see LLMs take on a larger role through the analytics development lifecycle, particularly via ChatGPT-style interfaces for analytics. No one can predict the future, of course. Given current trends, though, we can expect that by the end of 2025, it’ll be easier to transform, publish, find, and utilize data than we ever previously thought possible. --- --- title: "Automating your ETL: A guide to improved efficiency" description: "ETL automation can streamline data workflows, improve accuracy, and support scalable data infrastructure in modern organizations." url: "https://www.getdbt.com/blog/automating-your-etl" date: "2025-05-15" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Automating your ETL: A guide to improved efficiency Most modern organizations depend on robust ETL workflows for all data-related tasks. These pipelines streamline data integration and apply necessary transformations to create a standardized dataset. However, manually operating the pipeline becomes tedious and prone to human error as data grows. Automating the ETL process ensures that the data reaches its destination when required without any human intervention. It offers numerous benefits, such as improved process efficiency, better data quality, and consistency between data refreshes. In this article, we’ll address what you need to know about automating your organization's ETL workflows, including the fundamentals, the challenges, and how to streamline your ETL processes. ## What is ETL? Extract, Transform, and Load ([ETL](https://www.getdbt.com/blog/etl-pipeline-best-practices)) is a data integration process that collects and delivers data from various sources to a central location. Data engineers build pipelines that collect data from locations like relational databases, cloud storage, or streaming platforms. Once extracted, the data is [transformed](https://www.getdbt.com/blog/analytics-engineering-transformation) to meet the needs of the target system. The transformation logics may include data cleaning, validation, or aggregations, depending on the business requirement. Finally, the processed data is loaded into a destination such as a [data warehouse](https://www.ibm.com/think/topics/data-warehouse) or [data lake](https://www.databricks.com/discover/data-lakes). ETL is essential for ensuring data consistency, quality, and accessibility across an organization. Without it, data might remain impossible to find or unusable in its existing format. ## Why automation? The ETL process is run periodically to capture newly generated data and bring it to the designated warehouse. Manual operation of these pipelines requires data engineers to execute the steps one by one while ensuring error-free processing in between. Although this may work for smaller workloads, manual operation becomes a hassle as data operations grow. The increasing workload also introduces human errors and impacts data quality. According to the Data Science Council of America, [developers spend 80% of their time](https://www.dasca.org/world-of-data-science/article/the-art-of-data-wrangling-in-2024-techniques-and-trends?utm_source=chatgpt.com) cleaning and processing data. ETL automation brings several benefits for organizations and users. These benefits are discussed in detail below. ### Streamlined data processes An automated ETL process ensures that fresh and accurate data is delivered promptly to all stakeholders. The automated pipeline includes all necessary data checks and runs on a pre-defined schedule. This removes any unnecessary human intervention and ensures fresh data is always available. Automation removes bottlenecks from human operations and reduces time to market for data products. ### Improved data quality and reduced human errors A core part of ETL automation is implementing data transformation steps and quality checks. Whenever the pipeline runs, it automatically checks for any [data quality](https://www.getdbt.com/blog/data-quality-best-practices) issues and notifies developers if any issues are found. This ensures that data remains of the highest quality across all refreshes. Moreover, since the process is automated, the same steps are followed during each refresh, ensuring consistency in results and removing human error. ### Scalability and flexibility Increasing user activity and uneven workloads are the biggest challenges for modern organizations. ETL automation tackles this issue by deploying transformation processes on scalable cloud hardware. These cloud platforms provision hardware according to requirements and can increase or decrease resources based on workload. In a cloud deployment, resources can scale [vertically and horizontally](https://www.geeksforgeeks.org/horizontal-and-vertical-scaling-in-databases/) to cater to larger workloads, decommissioning compute capacity when demand returns to baseline. These scalable solutions ensure minimum downtimes and smooth operations at all times. ### Cost efficiency Errors in data pipelines can be costly as they require time to locate and fix. Moreover, poor data quality leads to inaccurate reporting, which can cause financial losses to the company. An automated ETL solution improves data quality by fixing issues once, making the correct procedure a permanent, automated part of the data pipeline. This reduces the costs associated with time delays and bug fixes. ## Fundamentals of ETL automation Automating the ETL process can be an overwhelming task that requires careful consideration. Below are some crucial points that can help design a robust, scalable, and maintainable ETL pipeline. ### Architecture design The first must always be to gather requirements and build an end-to-end development road map. The requirements must be business and technical and cover all aspects of the pipeline. This will include: - Exploring the available data sources and types - Quantifying the workload - Deciding on the quality checks and transformations to include - Deciding on the destination Having an architecture design beforehand will help prevent any surprises during development and ensure pipeline robustness and data quality. ### Selecting the correct automation tool Your architecture design may involve a single or multiple tools for complete ETL automation. However, selecting the right tools is a critical factor as it will determine the user experience. To make the decision easier, consider the following factors: - **Existing infrastructure**: Your selected tool should integrate well with your existing infrastructure. It must support the various data types present in your database and provide connectors for the different platforms you use. - **Ease of use**: An easy-to-use tool will take less time to deploy and for user onboarding. This will ensure fewer errors during development and faster time-to-market. - **Automation features**: During selection, consider key features like continuous integration/continuous deployment (CI/CD) for easy changes, deployment, and job scheduling. - **Testing**: Check if the tool includes data and pipeline testing features. These can include automated duplicate removals, schema validation, data drift detection, etc. These tests can improve the data quality standards and ensure consistent and correct results. ### Implementation and monitoring Once the tools are decided, the implementation process begins, following the roadmap. The automation process is implemented in steps, and each module is thoroughly tested for accuracy, robustness, and consistency before proceeding to the next. Once the automated pipeline is complete, it is tested end-to-end by running dummy jobs. Data passes through the pipeline, is monitored at every stage, and is validated for errors and bugs. The testing process also involves putting the pipeline through varying workloads to test its scalability and ensure smooth performance under different circumstances. ### Implement logging and debugging Despite extensive testing, it is vital to implement logging and debugging mechanisms. These mechanisms help identify and trace data errors that occur during production. Logging must be implemented at every step to ensure a detailed traceback is generated during each run. These traces allow you to validate the logic and output at every step and identify where the mismatches start to occur. ## ETL automation challenges As helpful as it may be, the automation process accompanies various challenges that must be addressed. Identifying these challenges beforehand allows you to build resolution strategies and make the process smoother. Here are some key challenges you will encounter: ### Architecture complexity Larger organizations often deal with disparate data sources and types. Integrating all these sources may require using multiple tools and complex integration pipelines. This can complicate the architecture and make the overall process overwhelming. These complications can cause errors, mismanagement, and time delays during development. ### Schema changes The ETL automation process relies on static data sources and schemas. The pipeline expects the source data to be in a certain format, and any changes to this format can break the entire pipeline. Adapting the pipeline to such changes can involve changing entire modules or logics, which is counter-productive. Moreover, implementing automation can be a big challenge in use cases where data is expected to change constantly. ### Error detection and resolution In any data-related project, errors in logic are not always apparent. Flaws in the transformation logic or calculations can lead to incorrect information. These errors are sometimes subtle and go unnoticed, especially if data validation is not handled carefully. If errors are left in an automated pipeline, they continue to poison the data so long asthe pipeline is active. Moreover, when the errors are finally identified, their resolution can take time and cause downtimes. ### Data security and compliance Automated workflows often require access to sensitive information, making it essential to implement strict access controls and encryption protocols. Without proper safeguards, there's a risk of data breaches or violations of regulatory requirements, which can have serious legal and reputational consequences. ### Performance and scalability Automated pipelines must be carefully tuned to maintain high performance without overloading system resources. Inefficient processing or poorly optimized queries can slow data flow, delay insights, and increase infrastructure costs. Additionally, without dynamic resource allocation and monitoring, these pipelines can become expensive to run at scale, making it crucial to balance speed, efficiency, and cost-effectiveness. ## Streamlining ETL with dbt Automating your ETL process streamlines data workflows, making them more efficient, reliable, and scalable. With automation, your dashboards, reports, and analytics are consistently powered by fresh, up-to-date data, eliminating manual effort and reducing the risk of human error. It also enhances data quality through built-in validation, testing, and error-handling mechanisms that catch issues early and improve trust in your data. [This is where dbt shines](https://www.getdbt.com/product/dbt). dbt automates the transformation layer of your pipeline by allowing you to define modular SQL models, enforce data quality through built-in tests, and track lineage across your data warehouse. With support for version control, documentation, and seamless integration with third-party tools, dbt empowers data teams to scale confidently, maintain transparency, and deliver insights faster. Interested in taking your data pipelines to the next level? Book a [demo](https://www.getdbt.com/contact) today. --- --- title: "Do you really need ETL tools for data transformation?" description: "You don’t need ETL tools to transform data. Explore SQL, scripts, and modern ELT approaches for flexible transformation." url: "https://www.getdbt.com/blog/data-transformation-without-etl" date: "2025-05-14" authors: ["Joey Gault"] categories: ["Pulse"] --- # Do you really need ETL tools for data transformation? Modern data teams have more options than ever when it comes to shaping raw data into something useful. For decades, ETL tools were the default way to clean, format, and prepare data before loading it into a warehouse. But the shift to [cloud-native platforms and ELT architectures](https://www.getdbt.com/blog/best-elt-tools) has changed that assumption. Today, teams can transform data using SQL, scripts, in‑warehouse functions, or specialized transformation frameworks—no traditional ETL tools required. The real question isn’t _whether_ transformation can happen without ETL, but _which approach gives teams the flexibility, governance, and scalability they need_. This article breaks down how data transformation works with and without ETL tools, and what modern teams should consider when choosing their approach. ## Understanding the relationship between data transformation and ETL [Data transformation](https://www.getdbt.com/blog/data-transformation) is the process engineers use to take data from one state to another, utilizing query systems like SQL or transformation packages for programming languages like Python. This process involves writing code that converts one table of data into another table, set of tables, or views on the original data. It's the mechanism that turns raw data into meaningful analytics useful for decision-making. [ETL (Extract, Transform, Load)](https://www.getdbt.com/blog/extract-transform-load) represents a sequential pipeline where data transformation serves as the middle step. In this traditional approach, data is extracted from source systems, transformed according to business requirements, and then loaded into a target system like a data warehouse. The transformation step within ETL handles the cleaning, formatting, and structuring of data before it reaches its final destination. However, data transformation itself is not inherently dependent on ETL tools. The transformation process can occur using various methods and at different stages of the data pipeline, depending on the architectural approach and available infrastructure. ## The evolution from ETL to ELT architectures The emergence of cloud computing has fundamentally changed how organizations approach data transformation. Traditional ETL architectures were designed decades ago when client-side storage and bandwidth were limited and costly. This approach made sense when organizations needed to minimize data movement across networks and storage requirements. Modern cloud environments have shifted this paradigm toward ELT (Extract, Load, Transform) architectures. In ELT, data is first extracted from source systems and loaded into a data warehouse in its raw form, where transformations occur afterward. This approach leverages the scalability and computational power of cloud-native data platforms like Snowflake, BigQuery, and Redshift. The key advantage of ELT lies in its flexibility. Raw data becomes immediately available for analysis, while transformations can be applied iteratively as business needs evolve. This eliminates the rigid preprocessing requirements of traditional ETL and allows data teams to work more responsively with their datasets. ## Data transformation without traditional ETL tools Organizations can absolutely perform data transformation without relying on traditional ETL tools. Several approaches enable this capability: **Direct SQL transformation within data warehouses** represents the most straightforward method. Modern cloud data warehouses provide powerful SQL engines capable of handling complex transformations directly within the platform. Data engineers can write SQL queries that clean, aggregate, and restructure data without requiring external transformation tools. **Custom scripting solutions** offer another path forward. Teams can develop transformation logic using programming languages like Python, R, or Scala, creating custom scripts that process data according to specific business requirements. These scripts can be scheduled and automated using workflow orchestration tools or cloud-native scheduling services. **Cloud-native transformation services** provide managed solutions for data transformation. Services like AWS Glue, Azure Data Factory, or Google Cloud Dataflow offer transformation capabilities without requiring traditional ETL tool installations or management. **In-database transformation capabilities** leverage the computational power of modern data platforms. Many cloud warehouses include built-in functions for data cleaning, statistical analysis, and complex aggregations that can handle sophisticated transformation requirements. ## The role of modern transformation tools While data transformation is possible without traditional ETL tools, modern transformation-focused tools like dbt have emerged to address the limitations of both traditional ETL and ad-hoc transformation approaches. dbt represents a SQL-first transformation workflow that operates within the ELT paradigm, focusing specifically on the transformation layer. [dbt](https://www.getdbt.com/product/what-is-dbt) enables teams to manage transformations as code, providing version control, automated testing, and comprehensive documentation. It automatically tracks dependencies between different transformed tables, eliminating the need to manually manage complex transformation relationships. The tool also promotes data quality through configurable tests that automatically validate new datasets against defined standards. These capabilities address common challenges that arise when performing transformations without dedicated tools: lack of documentation, difficulty tracking data lineage, inconsistent transformation logic across teams, and limited collaboration capabilities. ## Challenges of transformation without ETL tools Organizations attempting data transformation without proper tooling often encounter several significant challenges. Consistency becomes a major issue when different teams develop their own transformation approaches. Without standardized processes, similar cleaning and preparation work gets repeated across the organization, leading to inefficient resource utilization. Documentation and lineage tracking present ongoing difficulties. When transformations are scattered across various scripts and systems, understanding data flow and dependencies becomes increasingly complex. This lack of visibility makes debugging and maintenance significantly more challenging as data systems scale. Collaboration suffers when transformation logic exists in isolated scripts or individual SQL files. Team members struggle to share knowledge, review each other's work, or maintain consistent standards across projects. This fragmentation often leads to conflicting metrics and definitions across different business units. Quality assurance becomes more difficult without systematic testing frameworks. Ad-hoc transformation approaches often lack consistent validation processes, increasing the risk of data quality issues that can undermine business decision-making. ## Best practices for transformation without traditional ETL Organizations choosing to implement data transformation without traditional ETL tools should adopt several key practices to ensure success. Establishing clear data modeling conventions before beginning transformation work helps maintain consistency across all engineering efforts. Teams should define style guides, naming conventions, and SQL best practices that all contributors follow. Version control becomes critical when managing transformation logic outside of dedicated ETL platforms. All transformation code should be stored in centralized repositories with proper branching strategies and code review processes. This ensures changes are tracked, tested, and can be rolled back if necessary. Implementing automated testing frameworks helps maintain data quality even without built-in ETL tool capabilities. Teams can develop custom testing scripts that validate data accuracy, completeness, and consistency across transformation outputs. Standardization of core KPIs and business metrics prevents the confusion that arises when different teams generate conflicting reports. Key business metrics should be defined in code, version-controlled, and accessible within BI tools to ensure organizational alignment. ## The modern data transformation landscape Today's data transformation landscape offers multiple viable paths for organizations. While traditional ETL tools continue to serve specific use cases (particularly in highly regulated industries requiring strict data governance) the shift toward ELT architectures has opened new possibilities for transformation approaches. Cloud-native data warehouses provide increasingly sophisticated transformation capabilities built directly into their platforms. These capabilities, combined with modern transformation tools like dbt, enable organizations to build robust, scalable transformation workflows without relying on traditional ETL infrastructure. The choice between different transformation approaches often depends on organizational factors including technical expertise, compliance requirements, data volumes, and existing infrastructure investments. Some organizations adopt hybrid approaches, using traditional ETL for sensitive data processing while leveraging ELT patterns for more flexible analytics workflows. ## Conclusion Data transformation is not only possible without traditional ETL tools: it has become the preferred approach for many modern data organizations. The evolution toward ELT architectures, combined with the computational power of cloud data warehouses, has created new opportunities for flexible, scalable data transformation. However, the absence of traditional ETL tools doesn't eliminate the need for proper transformation management. Organizations must still address challenges around consistency, documentation, collaboration, and quality assurance. Modern transformation tools like dbt have emerged to fill this gap, providing the governance and management capabilities needed for reliable, scalable transformation workflows. The key insight for data engineering leaders is that data transformation represents a fundamental capability that can be implemented through various technological approaches. The choice of tools and architecture should align with organizational needs, technical capabilities, and business requirements rather than being constrained by traditional ETL paradigms. Success depends not on the specific tools chosen, but on implementing proper practices for managing transformation logic, ensuring data quality, and enabling effective collaboration across data teams. ## Data transformation without ETL tools FAQs **What is the difference between ETL and ELT?** ETL (Extract, Transform, Load) is a sequential pipeline where data is extracted from source systems, transformed according to business requirements, and then loaded into a target system like a data warehouse. In contrast, ELT (Extract, Load, Transform) first extracts data from source systems and loads it into a data warehouse in its raw form, where transformations occur afterward. ELT leverages the scalability and computational power of cloud-native data platforms, offering greater flexibility since raw data becomes immediately available for analysis while transformations can be applied iteratively as business needs evolve. **** **What are the different types of data transformation approaches available?** **What challenges can arise when performing data transformation without traditional ETL tools?** Organizations often encounter several significant challenges including consistency issues when different teams develop their own transformation approaches, leading to repeated work and inefficient resource utilization. Documentation and lineage tracking become difficult when transformations are scattered across various scripts and systems, making debugging and maintenance challenging. Collaboration suffers when transformation logic exists in isolated scripts, leading to conflicting metrics across business units. Quality assurance becomes more difficult without systematic testing frameworks, increasing the risk of data quality issues that can undermine business decision-making. --- --- title: "How dbt enhances your Redshift data stack" description: "Redshift is powerful. dbt adds structure, testing, and governance to help your team scale analytics with speed and confidence." url: "https://www.getdbt.com/blog/redshift-dbt" date: "2025-05-13" authors: ["Daniel Poppy"] categories: ["Learn"] --- # How dbt enhances your Redshift data stack Redshift is a powerful data warehouse—but as data complexity grows, teams need more structure to manage transformations, collaboration, and quality. That’s where dbt comes in. ## The challenge of data transformation in Redshift [Redshift](https://aws.amazon.com/redshift/) excels at storing large volumes of data and executing complex queries efficiently. However, as your data ecosystem expands, several challenges emerge. Managing SQL complexity becomes a major hurdle. For example, when your marketing team needs to analyze customer behavior across touchpoints, you may accumulate dozens of disconnected SQL scripts. Calculating customer lifetime value means joining data from orders, customer profiles, and marketing campaigns—quickly leading to fragile and unwieldy logic. Data quality presents another challenge. Redshift lacks native tools for ensuring your transformations produce expected results. Without systematic testing, errors can silently affect critical reports for months before discovery—potentially leading to poor decisions or missed revenue opportunities. Collaboration becomes harder as your team grows. Without a structured workflow, knowledge stays siloed, documentation becomes stale, and changes lack peer review. This creates risk when team members move on or when business requirements shift. ## How dbt transforms your Redshift experience [dbt](https://www.getdbt.com/product/dbt) brings structure and consistency to SQL-based workflows—turning transformation logic into an engineering discipline. In Redshift, dbt helps teams organize models logically, define clear dependencies, and document every step. This makes it easier to understand how data flows through your system and reduces duplication of effort. All transformations in dbt are stored as code in a Git repository, enabling modern software development practices like branching, code review, and version control. Teammates can review changes before they go live, track the history of edits, and roll back if something breaks. This is especially powerful when multiple analysts work on shared datasets. [dbt’s testing framework](https://docs.getdbt.com/docs/build/tests) ensures your data meets business and technical expectations. You can write tests to enforce primary key uniqueness, validate relationships between models, check value ranges, and flag violations of business logic—catching issues early, before they affect dashboards or decision-making. dbt also auto-generates documentation, creating a searchable catalog of your Redshift models and their dependencies. Business users and data team members alike can trace how each model is built, what logic it contains, and how it’s used downstream. ## Real-world benefits of using dbt with Redshift Redshift’s tightly coupled views and tables can make iterative development difficult—especially when replacing tables triggers dependency errors. dbt solves this with bind=False, enabling late-binding views that decouple table updates from dependent objects. This lets teams iterate without disruption. Beyond that, [dbt helps you get the most from Redshift’s architecture](https://www.getdbt.com/data-platforms/redshift). You can configure sort and distribution keys to match query patterns, choose the right [materialization strategy](https://docs.getdbt.com/docs/build/materializations) (view, table, or incremental), and build efficient incremental models that only update new records—cutting down on compute time and cost. Team productivity improves dramatically. dbt promotes a shared, modular transformation layer where analysts reuse each other’s logic instead of reinventing the wheel. Standardized model structure, testing, and documentation reduce ramp-up time for new team members and help them contribute faster. The [dbt DAG](https://docs.getdbt.com/docs/dbt-versions/2022-release-notes#dag-updates-and-performance-improvements) (directed acyclic graph) makes your project more navigable and easier to reason about. You can see exactly how models relate, assess the blast radius of a change, and identify bottlenecks before they become issues. Together, these capabilities make your Redshift workflows more maintainable, scalable, and resilient—enabling data teams to move faster with fewer errors. ## Implementation approach [Getting started with dbt and Redshift](https://docs.getdbt.com/guides/redshift?step=7) is straightforward—but like any transformation initiative, it benefits from a phased, intentional rollout. Most teams begin by creating a dbt project that includes models, tests, and macros. For production use, many build a Docker image of their project and store it in a container registry (like Amazon ECR). This image is then executed on a schedule using an orchestrator such as Airflow, AWS Step Functions, or Dagster. Start small. Choose a business-critical area with clear goals but manageable complexity. Build a few foundational models, add tests, and document as you go. This pilot will help you prove value quickly and establish best practices before scaling across teams or domains. Your team structure will shape your approach. Some organizations centralize dbt ownership in a core data team that manages shared, enterprise-wide models. Others take a federated approach, empowering domain teams to own their own transformations. Both can work—what matters is clear ownership and consistent standards. Training is key. While dbt uses SQL, its concepts—modularity, testing, documentation, CI/CD—require a mindset shift. Invest in onboarding and upskilling early. Most teams find the learning curve shallow and the returns immediate: cleaner data, faster development, and fewer fire drills. ## Conclusion Amazon Redshift is a powerful cloud data warehouse—but it wasn’t built with modular development, testing, or version control in mind. That’s where dbt comes in. dbt brings structure, collaboration, and quality control to your transformation layer. By layering dbt on top of Redshift, you get [a framework that applies software engineering best practices to your analytics workflows](https://www.getdbt.com/resources/the-analytics-development-lifecycle)—resulting in more reliable data, faster development cycles, and better business decisions. The combination of [Redshift’s performance and dbt’s transformation framework](https://www.getdbt.com/data-platforms/redshift) empowers data teams to scale with confidence. Whether you’re building complex transformations, modeling slowly changing dimensions, or ensuring data quality, dbt makes it easier to move fast—without breaking things. **Ready to level up your Redshift workflows?** [Try dbt for free](https://www.getdbt.com/signup) or [explore the docs](https://docs.getdbt.com/guides/redshift?step=7) to get started. ## dbt and Redshift FAQs **Can you use dbt with Amazon Redshift?** Yes. [dbt integrates seamlessly with Amazon Redshift](https://www.getdbt.com/data-platforms/redshift), adding structure and governance to your transformation workflows. It helps manage SQL complexity, automate testing, standardize development, and solve Redshift-specific challenges—like late-binding views and incremental loading. Explore the [dbt + Redshift integration guide](https://docs.getdbt.com/guides/redshift?step=7) to learn more. **What is dbt in AWS?** dbt is a transformation framework that brings software engineering best practices—like modular development, version control, and automated testing—to your data workflows on AWS. With Redshift, you can deploy dbt using Docker, schedule it with Step Functions or other orchestration tools, and manage credentials securely with AWS services like Secrets Manager. **Is Amazon Redshift a relational database?** Yes. Amazon Redshift is a relational database designed for analytical workloads. It supports SQL and uses a columnar storage architecture optimized for complex queries on large datasets—making it ideal for business intelligence and reporting, not transactional workloads. **Is Amazon Redshift an ETL tool?** No. Amazon Redshift is a data warehouse—not an ETL tool. While you can write SQL transformations in Redshift, it doesn’t handle extraction or loading from source systems. For full ETL workflows, teams typically pair Redshift with tools like AWS Glue, Fivetran, or dbt to manage and transform data efficiently. --- --- title: "The next era of analytics: How AI is changing the game" description: "How will AI change analytics and data transformation? dbt Labs CTO Mark Porter offers his insights." url: "https://www.getdbt.com/blog/analytics-next-era" date: "2025-05-12" authors: ["Mark Porter"] categories: ["Insights"] --- # The next era of analytics: How AI is changing the game AI is having a tremendous impact on every company. Make no mistake—the fundamentals of producing good, high-quality data remain unchanged. At the same time, AI is changing how everyone—engineers, analytics, decision-makers, and end users—works with, understands, and uses data. In this article, I’ll talk a lot about the Analytics Development Lifecycle (ADLC), which we here at dbt Labs view as an important way of thinking about analytics. I’ll speak a little about how dbt has evolved into a data control plane that supports every phase of the ADLC. Finally, I’ll look at what AI means for data and analytics and how I think businesses can adapt—and how dbt Labs is adapting to meet the new challenges and opportunities that generative AI presents. **** ## The importance of data quality in analytics Our founder and CEO, Tristan, has talked about [how dbt Labs views the Analytics Development Lifecycle](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) (ADLC) as the best path to building a mature analytics practice within an organization of any size. Our mission here at dbt Labs has always been to empower data practitioners to provide great data and information that companies can utilize to alter their direction. Using the ADLC, teams and companies can do just that by talking to each other using the same language. Let’s see why that’s important. Every year at dbt Labs, we run the State of Analytics Engineering survey to understand the pains, gains, and areas of investment for global data teams . Here are [last year’s results](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024). ![State of AE Report results](https://cdn.sanity.io/images/wl0ndo6t/main/6d027d952e3d2198e83966e3467675c7eb338893-1640x882.png) **** When we founded dbt Labs in 2016, building data transformations and discovering already available data products were huge problems. We’re proud to see that that’s gone down. Constraints on computing resources also went down thanks to products such as [Snowflake](https://www.snowflake.com/), [Databricks](https://www.databricks.com/), [BigQuery](https://cloud.google.com/bigquery), [Athena](https://aws.amazon.com/athena/), and [Redshift](https://aws.amazon.com/redshift/). So let’s look at the big problems _now_. - One is **ambiguous data ownership**. People can get at the data, but they don't know who owns it—and presumably where it came from and how high-quality it is. - **Data literacy** is a big problem—it’s challenging to be an effective analytics engineer or analyst when your stakeholders don’t understand what the data means or how it can affect business decisions. - Sadly, the largest problem is **data quality**. When you think about that for a moment, that’s a big problem. If we can’t trust the data, then how useful can our work be? How confident should we and our stakeholders be making decisions on that data? I’ve spent most of my career building and operating OLTP databases, and data quality in those systems is largely around accuracy, i.e. “Is the bank balance right?”. That’s one kind of data quality. But when you're running analytics, there are lots of different dimensions of quality. To list a few: completeness, consistency, timeliness, traceability, lineage, and uniqueness. You can't, or at least shouldn’t, present a dashboard to your CEO at 9am on Monday that doesn’t have clean data on every one of these dimensions—decisions will be made based on that data. Business is moving faster than ever before, is more complex than ever before, and those poor decisions have real consequences. This means that data quality has risen to a whole new level of importance. And that’s where the ADLC and dbt both come into play. ## The initial problem that dbt solved I used to do analytics for a very small company. We had about 350,000 users, and we had eight [MySQL](https://www.mysql.com/) databases on which I ran all the analytics. I had one file in Vim that was 38,000 lines long. Every day, I cut and pasted parts of that SQL script into the various MySQL databases to do our analytics. You can guess how well that worked. To fix this, dbt Labs brought the concept of software development to analytics. Using dbt, you can run all of your analytics code through a process modeled after the [Software Development Lifecycle (SDLC)](https://aws.amazon.com/what-is/sdlc/). ![ADLC loop](https://cdn.sanity.io/images/wl0ndo6t/main/1fc981ff485ca62fb80c5c9d4bde3789904514f8-2400x1260.jpg) dbt enables creating data transformations in the form of vendor-agnostic [dbt models](https://docs.getdbt.com/docs/build/models). You can then [test your transformations](https://docs.getdbt.com/docs/build/data-tests) instead of shipping them to production prematurely and inadvertently trashing someone’s dashboard. Once tested, you can deploy your changes [using Continuous Integration (CI)](https://docs.getdbt.com/docs/deploy/continuous-integration) and observe them. You can also use dbt to find and discover data and unearth facts. Of course, once you discover facts, you're never actually done because your questions lead to more questions. And so you analyze those facts, and then you plan and develop more analytics code to provide new answers. And that's why the ADLC is a figure-eight. ### The evolution of dbt and the ADLC dbt has grown a lot over the years. At first, we ran SQL and Python data transformations, and that was it. But that was big! It changed analytics from “Mark owns the SQL script” to “everyone in the company who deals with analytics can see, share, test, and check in changes to the SQL script.” You could also roll back to the last known good state if you had an “oh crap, I broke the dashboard the CEO uses to keep the company running” moment. Next, we added orchestration, so that you could automatically run jobs and see if things were taking too long to run. Then, to make a very long story currently involving around 580 people short, we added all these other capabilities. The [Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl), for example, means you can have one definition of what a term (such as “fiscal year”) represents, rather than having every team define its own, divergent version. We have [a visual editor](https://docs.getdbt.com/docs/cloud/visual-editor) for writing SQL and YAML code. We have a number of new features to help manage data and speed deployment of new data-driven projects: - Most recently, and most exciting, we have [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot), which can fetch the data from your data warehouse and build a dbt model around it. - We added [dbt Mesh](https://docs.getdbt.com/best-practices/how-we-mesh/mesh-1-intro), so that teams in large companies could own their own data and make data contracts with other parts of the company. - We’re also focused on optimizing costs, providing support for the Iceberg open table format, and launching a new Advanced CI feature where you can learn more about what’s going on with your data. ## dbt grows into a data control plane The product that dbt Labs owns today has expanded beyond transformation. No, we haven’t expanded into ingestion. There are lots of tools that do that well. And we've not expanded into AI and BI analytics; instead, we partner with AI and BI analytics tools that do what they do well. What we _have_ done via [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) is expand into a complete orchestration and observability system. This way, when your data warehouses are burning all their compute overnight, you know what they're doing and whether it worked. We call this [the data control plane](https://www.getdbt.com/blog/data-control-plane-introduction): an abstraction layer that sits across your data stack, unifying capabilities for orchestration, observability, cataloging, semantics, and more. Features such as the Semantic Layer and our [data catalog](https://docs.getdbt.com/docs/collaborate/explore-projects) round out our support of the ADLC and create a single place to observe and orchestrate data across your company. We did something else pretty wild: [we acquired SDF](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs). SDF gives us complete SQL understanding, unlike dbt’s previous simple text parsing. With SDF, we'll catch all syntax errors before sending them to the warehouse and help you fix them in the editor. We'll also quickly query your warehouse's data dictionary to identify non-existent tables or columns. This fundamentally changes how developers work with dbt, providing: - Much faster developer experience - Reduced warehouse costs - Better DAG lineage tracing, which is especially important for [GDPR](https://gdpr.eu/what-is-gdpr/) compliance This integration work is still ongoing, but will provide faster development cycles and additional cost savings for data teams. ## How businesses can adapt to the AI era That raises the question of what’s next for analytics. Particularly, it raises the question of how AI impacts analytics and the work of data engineers. If we think about previous disruptive shifts—e.g., the shift from on-premise to cloud, or the shift from databases to data warehouses—the trends were much slower. They took five, 10, or 15 years and were less disruptive. By contrast, AI is an immediately disruptive shift. Take [agents](https://aws.amazon.com/what-is/ai-agents/), for example. Did anyone actually think that computers were going to formulate English questions to each other to communicate? That was silly. It wasn't like one computer was going to say, “Hey Joe, can you tell me about revenue last week?” And yet, here we are. With agents and new innovations like [Model Context Protocol (MCP)](https://www.anthropic.com/news/model-context-protocol), we’re on the verge of seeing an absolute explosion—data talking to data, computers talking to computers, and analysts being able to ask different questions. So AI _is_ different. At the same time, part of our job as a technology company is not to buy into hype. When it comes to AI, that means being realistic and setting expectations. When I was at Grab, I had 1,700 people reporting to me in engineering. 150 of them were AI engineers. And they were developing completely traditional models—not using GenAI [Large Language Models (LLMs)](https://www.ibm.com/think/topics/large-language-models). So when it comes to adapting your business to AI, I have three recommendations: - **Lean into every traditional model**. They’re incredibly valuable. You don't need to use LLMs. You don't need to have a huge bill with your favorite LLM company. Build models, fine-tune models, work on models. - **Start simple**. Most companies are having trouble getting started with AI. Start simple - summarize support tickets, use Notion summaries. Start getting familiar.To be clear, I don’t think there’s an AI chatbot I’ve used yet on a company website that didn’t require human intervention. We’re still in the early days here.  - **Data quality is key**. AI doesn’t know its data better than any expert or random social media user does. The quality of data—particularly, the quality of its source and whether we can trust that source—will be key. I think we’ll see AI systems start to do two things. One, they’ll be able to trace data. Two, and more importantly, they’ll start to grade data, tell us how well-governed it is, and where it came from. And then that happens, it’s going to be game-changing. ### AI and the developer experience Another angle to consider isn’t just how AI is disrupting business, but how it’s disrupting data engineering. [Tristan recently wrote a piece on this](https://www.getdbt.com/blog/how-ai-will-disrupt-data-engineering) in which he talked about AI’s disruptive role in how we develop and manage analytics code. I manage a group of software engineers. For them, AI is changing how they code, document, and even debug. It’s also changing how my product team and designers are doing research. They use a deep research site or [ChatPRD](https://www.chatprd.ai/) and get a great PRD done in minutes, not days. In data, too, tools like dbt Copilot are changing the way data engineers create models as well as how business users interact with data. I think you’ll see a lot of fast development here, particularly with [agentic AI](https://blogs.nvidia.com/blog/what-is-agentic-ai/). There’s all this hype in the world about everyone losing their jobs to AI. My son, who’s in the industry, sees it differently. He said that AI is going to let his team do the jobs the rest of the business always wanted to do at the speed, quality, and intensity that they always wanted. Back in the 1960s, computers could do X amount of work and people wanted us to do Y amount. It was the same in the 1980s. It’s the same today. We’re still, even with AI, not even close to computers being able to do the things that people want them to do. As long as that band gap exists of human ingenuity, of human cleverness, of human directed creativity, I don't think anyone has to worry about their jobs. ## How dbt is changing to serve the next era of analytics and AI When we launched dbt, it was for a very specific persona: the data nerd. The data nerd loves to sit and write some SQL statements, chain them all together, and feel proud. And that’s great. Nothing wrong with that. dbt and dbt Cloud remain powerful tools for data engineers and [analytics engineers](https://www.getdbt.com/blog/what-is-analytics-engineering). But over time, more and more analysts started using dbt to drive business insights. So one of the things that you're going to see us doing more is really catering to a set of workflows that help analysts have a lower-friction workflow. For example, today with [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot), you can ask it to build you a model based on your data and a natural language description of the output you need. Alternatively, you can take an existing model developed by someone else that’s extremely complex and say, “Tell me what this does.” This is a critical shift, as there are anywhere from five to 15 times more analysts at a company than there are data engineers and analytics engineers. Giving them more power to transform and manage data means greater data democratization and faster data project velocity. There’s this third persona we’re targeting, and that’s the data executive. They need their data to be accurate and to be assured that data fed to AI systems is governed properly. And they also care about tracking costs. In the past, the people pushing for dbt and the ADLC at companies were data and analytics engineers. Today, we find analysts and data executives joining the call. As dbt continues to evolve, you’ll see more of this evolution: using the power of AI and other technologies to get high-quality data quickly into the hands of everyone who participates in the data lifecycle. --- --- title: "The essential skills for data engineers in 2025" description: "Discover the core skills that define successful data engineers in 2025—from SQL to AI-powered workflows." url: "https://www.getdbt.com/blog/data-engineer-skills-2025" date: "2025-05-12" authors: ["Joey Gault"] categories: ["Pulse"] --- # The essential skills for data engineers in 2025 Despite technological advances, fundamental technical skills continue to form the backbone of effective data engineering. SQL proficiency remains non-negotiable, as it serves as the primary language for data manipulation across virtually all modern data platforms. Data engineers must demonstrate advanced SQL capabilities, including complex query optimization, window functions, and the ability to write efficient queries that perform well at scale. Programming languages, particularly [Python](https://www.python.org/), have become increasingly important as data pipelines grow more sophisticated. Python's rich ecosystem of libraries for data manipulation, API integration, and automation makes it indispensable for modern data engineering workflows. Java also maintains relevance, especially in big data environments and enterprise systems where performance and scalability are paramount. Cloud platform expertise has shifted from advantageous to essential. Data engineers must be proficient with at least one major cloud provider (AWS, Google Cloud Platform, or Microsoft Azure) and understand how to leverage cloud-native services for data storage, processing, and orchestration. This includes familiarity with managed services like Amazon Redshift, Google BigQuery, or Snowflake, which have become the standard for modern data warehousing. Data modeling skills have evolved to encompass both traditional dimensional modeling and modern approaches suited for cloud data warehouses. Engineers need to understand when to apply different modeling techniques, how to design schemas that balance query performance with maintainability, and how to structure data for both analytical and operational use cases. ## The transformation layer: modern data transformation practices The rise of the [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) paradigm has fundamentally changed how data engineers approach transformation work. Rather than transforming data before loading it into storage systems, modern practices emphasize loading raw data first and performing transformations within the data warehouse itself. This shift has made tools like [dbt](https://www.getdbt.com/product/what-is-dbt) central to the data engineer's toolkit. dbt has become particularly important because it [brings software engineering best practices to data transformation work](https://www.getdbt.com/resources/the-analytics-development-lifecycle). Data engineers using dbt can create modular, version-controlled SQL transformations that are testable, documented, and maintainable. The tool's approach to building reusable models and managing dependencies has made it possible to treat data transformations with the same rigor as application code. dbt is also evolving into a broader data control plane with the [dbt Fusion engine](https://www.getdbt.com/product/fusion) and state-aware orchestration for faster, more efficient builds, richer metadata, and portability across platforms. For AI-aligned workflows, [dbt Agents](https://www.getdbt.com/product/dbt-agents) and the [dbt MCP server](https://www.getdbt.com/blog/build-reliable-ai-agents-with-the-dbt-mcp-server) provide governed project context to AI systems, enabling reliable automation and assistance over trusted, versioned code [Understanding ETL and ELT frameworks](https://www.getdbt.com/blog/etl-vs-elt) more broadly remains crucial, as different use cases may require different approaches. Data engineers need to know when to apply each pattern and how to implement both effectively using modern tooling and cloud infrastructure. ## Infrastructure and orchestration: managing complexity at scale As data systems become more complex, orchestration and automation capabilities have become essential. Tools like [Apache Airflow](https://airflow.apache.org/) enable data engineers to manage complex workflows with multiple dependencies, error handling, and retry logic. Or you can default to [dbt jobs](https://docs.getdbt.com/guides/airflow-and-dbt-cloud?step=1) for transformation orchestration and, where eligible, leverage Fusion’s state-aware orchestration to rerun only what’s needed and cut compute, and use Airflow or similar for cross-system workflows that coordinate beyond transformations. Big data technologies continue to play important roles, particularly for organizations dealing with massive datasets or real-time processing requirements. [Apache Spark](https://spark.apache.org/) remains relevant for large-scale data processing, while [Apache Kafka](https://kafka.apache.org/) has become the standard for streaming data architectures. Data engineers should understand when these technologies are necessary and how to implement them effectively. API integration skills have grown in importance as organizations increasingly rely on third-party services and need to extract data from various SaaS platforms. Understanding RESTful APIs, authentication mechanisms, and rate limiting is essential for building robust data ingestion pipelines. ## Governance, security, and quality: the operational imperatives Data governance and security have evolved from compliance requirements to business imperatives. Data engineers must implement comprehensive access controls, understand data lineage tracking, and ensure compliance with regulations like [GDPR](https://gdpr.eu/) and [CCPA](https://oag.ca.gov/privacy/ccpa). This includes implementing data masking, encryption, and audit trails throughout the data pipeline. Data quality management has become more sophisticated, requiring engineers to implement automated testing, monitoring, and alerting systems. The ability to build data quality checks directly into transformation pipelines (rather than treating quality as an afterthought) is now expected. This includes understanding statistical methods for detecting anomalies and implementing business rule validation. [Monitoring and observability capabilities](https://docs.getdbt.com/docs/deploy/job-scheduler) have expanded beyond simple pipeline success/failure alerts. Modern data engineers need to implement comprehensive monitoring that tracks data freshness, volume changes, schema evolution, and performance metrics. This operational awareness enables proactive problem resolution and builds trust in data systems. ## The AI transformation: adapting to an AI-enabled future Artificial intelligence is reshaping data engineering in profound ways. While AI won't replace data engineers, it will significantly change how they work. Many routine tasks (writing basic transformation code, generating documentation, and even debugging pipeline failures) are becoming AI-assisted or fully automated. In dbt, [Copilot](https://www.getdbt.com/product/dbt-copilot) accelerates model, test, and docs creation; dbt Agents and the dbt MCP server enable governed agentic workflows over your semantic and lineage context. Data engineers need to understand how to work effectively with AI tools while maintaining the judgment to know when human oversight is required. This includes understanding the limitations of AI-generated code and maintaining the ability to review, test, and validate automated solutions. The rise of AI also creates new requirements for data engineers. Machine learning workloads have different data requirements than traditional analytics, often requiring real-time feature stores, model versioning, and specialized data formats. Understanding these requirements and how to build infrastructure that supports both traditional analytics and ML use cases is becoming increasingly valuable. ## Collaboration and communication: the human element As data teams become more specialized, collaboration skills have become more critical. Data engineers increasingly work alongside analytics engineers, data scientists, and business stakeholders, requiring strong communication abilities to translate technical concepts for non-technical audiences and understand business requirements. The ability to work in cross-functional teams and participate in agile development processes has become standard. Data engineers need to understand how their work fits into broader business objectives and be able to prioritize tasks based on business impact rather than purely technical considerations. Documentation and knowledge sharing skills have grown in importance as data systems become more complex and teams become more distributed. The ability to create clear, maintainable documentation and share knowledge effectively across teams is now a core competency. dbt’s unified workflow, lineage, and docs support [governed collaboration across personas](https://www.getdbt.com/blog/why-governed-collaboration-is-the-key-to-modern-analytics-workflows), enabling analysts and analytics engineers to contribute safely while data engineers focus on higher-value work. ## Adaptability and continuous learning: staying current The rapid pace of change in data technology makes adaptability one of the most important skills for data engineers. New tools, frameworks, and best practices emerge regularly, and successful engineers must be comfortable with continuous learning and experimentation. Problem-solving abilities remain crucial, but the nature of problems is evolving. Modern data engineers need to think systematically about complex, distributed systems and be able to debug issues that span multiple technologies and platforms. Project management skills have become more important as data engineers often lead initiatives that span multiple teams and systems. Understanding how to break down complex projects, manage dependencies, and communicate progress to stakeholders is increasingly valuable. ## Looking ahead: preparing for continued evolution The data engineering field will continue to evolve rapidly, driven by advances in AI, changes in data architecture patterns, and growing business demands for real-time insights. The most successful data engineers will be those who combine strong technical fundamentals with the ability to adapt to new tools and approaches. Organizations should focus on building teams with diverse skill sets that complement each other, rather than expecting every individual to master every technology. The combination of strong technical skills, business acumen, and adaptability will continue to define successful data engineering careers. The integration of AI into data engineering workflows will accelerate, making it essential for data engineers to understand how to leverage these tools effectively while maintaining the critical thinking and domain expertise that AI cannot replace. Those who can successfully combine human judgment with AI capabilities will be best positioned for success in this evolving landscape. As the field continues to mature, the most valuable data engineers will be those who can bridge the gap between technical implementation and business value, building systems that are not just technically sound but also aligned with organizational objectives and capable of evolving with changing requirements. Expect continued consolidation toward [open data infrastructure](https://www.getdbt.com/blog/what-is-open-data-infrastructure) that unifies data movement and transformation while preserving choice of compute. The [dbt Labs–Fivetran merger](https://www.getdbt.com/blog/dbt-labs-and-fivetran-merge-announcement) underscores this trajectory and our commitment to open standards for analytics and AI. ## Data engineering FAQs **What does a data engineer do?** Data engineers build and maintain the infrastructure and systems that enable organizations to collect, store, process, and analyze data at scale. They design and implement data pipelines that extract data from various sources, transform it into usable formats, and load it into data warehouses or other storage systems. Modern data engineers work with cloud platforms, orchestrate transformation with tools like dbt and Apache Airflow, and implement data quality monitoring and governance measures. They also collaborate closely with data scientists, analytics engineers, and business stakeholders to ensure data systems meet organizational needs and support both traditional analytics and machine learning use cases. **Which programming languages are data engineers commonly proficient in?** Data engineers are commonly proficient in SQL and Python as their core programming languages. SQL remains non-negotiable as the primary language for data manipulation across virtually all modern data platforms, requiring advanced capabilities including complex query optimization and window functions. Python has become increasingly important due to its rich ecosystem of libraries for data manipulation, API integration, and automation, making it indispensable for modern data engineering workflows. Java also maintains relevance, particularly in big data environments and enterprise systems where performance and scalability are paramount. **What does a data engineer do when creating big data ETL pipelines, and what aspects of production readiness do they focus on?** When creating big data ETL pipelines, data engineers focus on implementing the modern ELT (Extract, Load, Transform) paradigm, where raw data is loaded first and transformations are performed within the data warehouse. They use tools like dbt to create modular, version-controlled SQL transformations that are testable and maintainable. For production readiness, they implement comprehensive monitoring and observability systems that track data freshness, volume changes, and performance metrics. They also focus on data governance and security by implementing access controls, data lineage tracking, and compliance measures. Additionally, they build automated data quality checks directly into transformation pipelines and design orchestrated workflows with proper error handling and retry logic using tools like Apache Airflow. --- --- title: "Why compilers matter" description: "We continue our season on developer experience by looking at compilers with the SDF Labs cofounder, Lukas Schulte." url: "https://www.getdbt.com/blog/why-compilers-matter" date: "2025-05-11" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Why compilers matter _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/why-compilers-matter-w-lukas-schulte). _ Tristan Handy dives deep into the world of compilers in this episode of The Analytics Engineering Podcast with Lukas Schulte, cofounder of SDF Labs (not to be confused with [last episode’s guest—Lukas’ dad and fellow SDF cofounder Wolfram Schulte](https://roundup.getdbt.com/p/the-evolution-of-databases-w-wolfram)). Tristan and Lukas discuss what compilers are, how they work, and what they mean for the data ecosystem. SDF, which was [recently acquired by dbt Labs](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs), builds a world-class SQL compiler aimed at abstracting away the complexity of warehouse-specific SQL. The conversation covers the evolution of compiler technology, what software engineering has gotten right over the past several decades, and w[hy the data ecosystem is poised for similar transformation](https://www.getdbt.com/blog/how-ai-will-disrupt-data-engineering). Lucas and Tristan explore why SQL has lagged behind other programming ecosystems, and how new compiler infrastructure could lead to package management, interoperability, and greater innovation across data platforms. It’s a fascinating (and timely) episode: [Get ready for the new dbt engine](https://www.getdbt.com/blog/how-to-get-ready-for-the-new-dbt-engine). _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ### Chapters - 02:40 The vision behind SDF Labs - 04:00 What is a compiler? - 05:00 Components of a compiler: frontend, IR, backend - 08:00 Syntax vs. semantics and the role of parsing - 10:00 Logical vs. physical plans in SQL compilers - 13:00 Historical context: mainframes to LLVM - 16:00 Cross-architecture portability in Rust & other compilers - 18:00 What is LLVM and why it matters - 20:00 Bootstrapping and the self-recursive nature of compilers - 21:00 Compilers in Java, TypeScript, and dbt - 23:00 Why compilers are foundational to software ecosystems - 26:00 The SQL dialect problem in data warehouses - 29:00 Can SQL get its own LLVM? - 31:00 How Substrate and DataFusion aim to standardize SQL - 35:00 Package management and the path toward SQL abstractions - 38:00 The future of the data ecosystem with a common SQL compiler ## Key takeaways from this episode ### What is a compiler? **Tristan Handy:** What is a compiler? **Lukas Schulte:** It's something that takes higher-level human-readable code and translates, compiles, rewrites it into lower-level machine code that is much harder for humans to understand and much easier for machines to understand. Compilers typically have phases. They have a frontend that deals with the language you're working with, a middle component—usually called an IR or intermediate representation—and a backend that takes that IR and compiles it into machine code. ### Compiler phases: frontend, IR, backend **Tristan Handy:** How does it all come together? **Lukas Schulte:** There’s a preprocessor that handles macros, removes comments, and prepares the text. Then a lexer converts it into tokens. These tokens get assembled into a tree that the compiler can understand. That’s where syntax validation and semantic analysis happen. From there, we build a logical representation of the operations we want to perform. That transitions to a physical plan, which starts considering the hardware: how many cores, how much memory, which files we’re accessing. After that, optimizations are applied and it compiles to actual machine code using a toolchain like LLVM. ### Syntax vs. semantics **Lukas Schulte:** Let’s break down syntax vs. semantics. Imagine the code `x = x + 1`. That has valid syntax. Its meaning—its semantics—is that we’re incrementing `x` by 1. Now, you could also write `x += 1`. Different syntax, same semantics. So syntax defines structure, and semantics define meaning. That distinction is important when you’re analyzing or transforming code. ### LLVM and portability **Tristan Handy:** Have we been building abstraction layers like this for decades? **Lukas Schulte:** Absolutely. That’s what LLVM does. It provides a consistent intermediate representation that compilers can use to target multiple backends—Intel, ARM, different OSes. Apple invested early in LLVM to support custom chips. With Rust, for example, LLVM is what lets us build binaries that behave the same on macOS, Windows, and Linux with relatively little effort. ### Bootstrapping compilers **Tristan Handy:** So there’s this recursive loop—compilers being built with other compilers? **Lukas Schulte:** Exactly. Rust wasn’t always written in Rust—it started in C++. Eventually, the compiler was rewritten in Rust itself. Now, Rust compiles Rust. It’s fully self-hosted. That’s common with mature languages—it shows the compiler ecosystem is stable and powerful enough to sustain itself. ### Why compilers matter **Tristan Handy:** You said once that compilers are the foundation of every software ecosystem. What did you mean? **Lukas Schulte:** There are two big drivers in software: abstractions and standards. You want one way to interface with a USB device—not ten. Same for software. You want one standard way to express a Python program, a JavaScript app, etc. Compilers enforce those standards and make sure the same code works across platforms. That consistency powers things like package managers, shared libraries, and open ecosystems. ### SQL dialects and fragmentation **Tristan Handy:** Are there ecosystems that are doing worse than others? **Lukas Schulte:** SQL does a particularly bad job. Anyone who's used more than one data warehouse knows you can't take the same SQL statement and expect it to work the same way. Casting, case sensitivity, functions—every engine handles these things differently. ### Toward a universal SQL compiler **Tristan Handy:** Can you convince me this problem is solvable? **Lukas Schulte:** Yes. That's what we're working on with SDF—creating a shared intermediate representation for SQL. If we can express SQL logic in a unified form, we can compile it to any dialect—BigQuery, Snowflake, Redshift, and so on. That allows developers to build reusable libraries, just like in other languages. It also makes governance, validation, and testing easier. ### Future of data ecosystems **Tristan Handy:** What would that future look like for practitioners? **Lukas Schulte:** One major change would be the emergence of robust SQL libraries. Today, there’s no `import` system for SQL. Everyone writes similar logic over and over. A shared compiler abstraction would let us reuse components, collaborate across companies, and build an ecosystem of packages for transformations, metrics, and validations—similar to how we use NPM or PyPI. --- --- title: "Docusign’s path to 40% cost savings and 60% increased productivity" description: "Discover how DocuSign modernized its data infrastructure with dbt." url: "https://www.getdbt.com/blog/modernizing-data-at-scale-docusign-s-path-to-40-cost-savings-and-60-increased-productivity" date: "2025-05-09" authors: ["Chakshu Mehta", "Hrishi Kulkarni"] categories: ["Product"] --- # Docusign’s path to 40% cost savings and 60% increased productivity At Docusign, the goals of the data team are threefold: help leaders make informed decisions, get ahead of market trends and competition, and enable the business to achieve strategic goals. But none of this is possible without trusted data. “Every day, I start with one question: have I set things up so my customers can trust the data?” says Bishal Gupta, Analytics Engineering Leader at DocuSign. “If the answer is no, then I still have work to do.” The Docusign data team knew their legacy infrastructure couldn’t keep up with their data needs, so they decided to modernize their data stack using dbt. In this blog post, we’ll take a look at how that shift has transformed the business, increased cost savings—and even recovered $3 million in lost revenue. ### From legacy stored procedures to modular dbt models It’s a scenario you likely know all too well: complex architecture, legacy stored procedures, and siloed SQL logic. That’s where Docusign started. These procedures, while functional, were difficult to trace, hard to debug, and introduced inconsistencies across reports and dashboards. This led to a tangled data architecture that stifled scalability and slowed down insights. “As the business grows, the data grows,” says Gupta. “dbt gave us the agility to adapt, scale, and stay ahead as the business evolved.” ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/b544dd99ec9b4a582ee38f7c408f24f1854877d2-512x291.png) By migrating these procedures to modular, reusable dbt models, Docusign could easily trace and address problems in the data. The result was an organized and flexible data model that enabled faster issue resolution and a streamlined experience for the data team. “Right off the bat, adopting dbt’s modular approach increased our team’s productivity by 60%,” says Gupta. “Meanwhile, our cost savings improved by 40%.” Gupta explains that these cost savings came from three factors: first, dbt increased developer productivity by enabling faster model creation. Second, it reduced time spent troubleshooting and addressing data-quality issues. Lastly, dbt's integration with Snowflake optimized data materialization, cutting down on storage and processing costs. ### Building a modern data stack by migrating to dbt Cloud Impactful as the move has been, it was just the first step of adopting a modular, scalable approach to data. To bring software engineering rigor and a true Analytics Development Lifecycle (ADLC) to their workflows, the data team migrated dbt Core models to [dbt Cloud](https://www.getdbt.com/product/dbt-cloud). This transition generated numerous cascading benefits like: - **Improved collaboration.** With Git integration and CI/CD pipelines built into dbt Cloud, collaboration across teams became much smoother. Developers can easily manage code versions more effectively and automate deployments with confidence, just like a modern software engineering team. - **Centralized data models. **Docusign uses [dbt Mesh](https://www.getdbt.com/product/dbt-mesh) to create standardized shared models across different teams. Now each team accesses and uses the same trusted base models—ensuring consistency and reducing duplication. - **Standardized macros.** By centralizing business logic in standardized sources and reusable macros, Docusign ensures that the same logic is applied consistently throughout the data pipeline. This has helped reduce errors and rework that previously occurred when multiple teams implemented their own versions of the same transformations. “Stakeholders want as many metrics as possible, as fast as possible—and they want it yesterday,” says Gupta. “Because of dbt, we don’t cut corners to get our stakeholders the accurate metrics they need.” ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/21c85c0f6b93f24708fda21d33e5292f59aa9773-512x270.png) ### Deep dive: modularity, reusability, and testability To illustrate dbt’s impact, let’s walk through how the Docusign data team solved a costly product issue. When a Docusign customer signs an agreement on the Docusign platform, they often need to submit a payment. DocuSign simplifies this process by allowing them to make payments natively on the platform. But sometimes these payments wouldn’t go through. This was a huge problem: it created a poor user experience, led to customer churn, and resulted in an estimated $3 million in lost revenue annually. To tackle this issue, the data, finance, and payments teams decided to work together. Using dbt, the data team: - Refactored legacy logic into reusable, testable models - Standardized payment-data pipelines across systems - Built lineage-aware, production-ready datasets - Implemented unit and data testing to ensure integrity at every layer This collaborative, dbt-driven approach helped identify the root cause—and gave teams the tools and confidence to fix the issue quickly and prevent it from happening again. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/5e1757af39b14c5110370d58ec0b126a895a2275-512x292.png) “Before, we weren’t following the [DRY principle of ‘Don’t Repeat Yourself](https://www.getdbt.com/blog/dry-principles),’” says Guptal. “The result was a weekly headache involving multiple versions of the same data, where the numbers didn’t match. “So we put the DRY principle in action: we put our code into a macro, which holds all common operations and logics,” Gupta continues. “By using dbt macros and centralized models, we eliminated duplicative logic and reduced errors. Most importantly, we increased productivity and trust.” The results were striking: the data team improved reporting while also driving significant business impact. The insights they uncovered helped reduce passive churn, identified $3 million in lost revenue, and enabled the business to make better financial decisions. ### Lessons for data teams Docusign’s journey from tangled, complex architecture to a governed, scalable dbt ecosystem illustrates what’s possible when teams align technical vision with business outcomes. Whether you’re modernizing a legacy stack or migrating to dbt Cloud for the first time, Gupta recommends the following: - **Start small with a proof-of-concept.** Prioritize a high-impact use case and use metrics to validate. - **Use the transition to reevaluate business logic.** Think through how it will help you scale for new use cases. - **Prioritize modularity, documentation, and test coverage. **Focus on the ways you can improve coding efficiency, flexibility, and team productivity. - **Invest in training.** Tools like the [dbt Fundamentals Certification](https://learn.getdbt.com/courses/dbt-fundamentals) can help your team quickly learn the foundational steps of transforming data in dbt Cloud. Watch Docusign's Coalesce 2024 session here [Watch video](https://www.youtube.com/watch?v=N-7Q0WuZ_Oo) If you’re looking to scale a modern data platform, investing in modular analytics engineering with dbt pays off. With dbt, you can build a stack that’s not just smarter—but also more trusted, agile, and impactful. Contact us to [book a demo](https://www.getdbt.com/contact), or [sign-up for dbt Cloud](https://www.getdbt.com/signup) to connect your data warehouse and start building. --- --- title: "How to get ready for the new dbt engine" description: "The next-generation dbt engine is coming. Here's what you need to know." url: "https://www.getdbt.com/blog/how-to-get-ready-for-the-new-dbt-engine" date: "2025-05-09" authors: ["Joel Labes", "Azzam Aijazi"] categories: ["Learn"] --- # How to get ready for the new dbt engine Big things are coming to dbt. On May 28, we’re releasing a major upgrade: an all-new, next-generation dbt engine, built in Rust using the technology from [our recent acquisition of SDF Labs](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs). It’s designed to bring unprecedented speed and cost-savings to your data workflows, while serving as the foundation for the [next era of data work](https://www.getdbt.com/blog/how-ai-will-disrupt-data-engineering). **We’re excited to share more** about the new engine on May 28th. To ensure the smoothest possible upgrade path when it’s released, there are a couple of simple steps you can take right now. Let’s walk through what’s coming, and how to get prepared. ## What’s coming The next-generation dbt engine will bring with it major improvements to how you develop, deploy, and manage your data products. Here’s what you can expect: #### ⚡ Lightning-fast development The new engine will bring deep SQL comprehension to dbt for the first time. dbt won’t just _pass along_ your SQL to a data warehouse, it will actually be able to _understand_ it. This unlocks powerful capabilities like live error detection as you write code, smart autocomplete suggestions, compiled code previews, and more. All that, coupled with model parse times 30x faster than before, amounts to a best-in-class developer experience for data practitioners. #### 💸 Built-in cost efficiency By emulating your data platform locally, the new engine will be able to validate code and catch errors as it’s written... and **before any warehouse compute is used.** Plus, with capabilities like state-aware orchestration, you’ll be able to automatically run only what's changed in your DAG, helping you avoid unnecessary runs and reduce warehouse spend. #### 🔍 Better visibility and trust SQL comprehension brings even more rich metadata to dbt. Because the engine understands your SQL ahead of execution, it will be able to generate rich model- and column-level lineage out of the box, allowing for faster debugging, better visibility, and impact analysis. In the near future, this will also unlock features like built-in PII tracing and policy enforcement. ## How to prepare your projects today During development of the new engine, we've taken the opportunity to reconsider and resolve some [long-standing quirks of dbt](https://github.com/dbt-labs/dbt-core/discussions/11493). These quirks made it harder to guarantee consistent behavior in dbt and can cause problems for developers of third-party tools that integrate with dbt, maintainers of dbt itself, and especially end-users. To take advantage of the new engine, you will likely need to give your project(s) a quick spring clean. We’re building tools and guides to make this upgrade simple, but there are two key things you can do _right now_ to be ready when the new engine arrives. ### 1. Upgrade to the latest version of dbt dbt Core v1.10 contains new [deprecation warnings](https://docs.getdbt.com/reference/deprecations) for soon-to-be-unsupported behaviors. In 1.10 they are informative only—your project will still run as normal—but you'll need to resolve them to successfully use the new engine once it's available. That means you'll want to upgrade to the new engine from a project on the [`Latest` release track](https://docs.getdbt.com/docs/dbt-versions/cloud-release-tracks) (for dbt Cloud users) or dbt Core v1.10 (for self-hosted users). This will ensure the simplest, most predictable experience: you can pre-validate that your project doesn't rely on deprecated behaviors. If you're one of the vast majority of customers already using release tracks, sit tight! You will automatically see relevant deprecation warnings (along with all other vetted new dbt capabilities and fixes) ahead of May 28th. 📘 Learn how to [switch to the Latest release track →](https://docs.getdbt.com/docs/dbt-versions/cloud-release-tracks) ### 2. **Resolve deprecation warnings** Many deprecations (such as moving arbitrary configs into the `meta` dictionary) can be resolved automatically. Users with a local filesystem (using the Cloud CLI or dbt Core) will be able to use [this auto-fix script](https://github.com/dbt-labs/dbt-cleanup) developed by dbt Labs and then review any other deprecations requiring manual attention. The dbt Cloud IDE will surface the same list of configuration or syntax issues that need updating. It will soon provide an interface to the same auto-fix script as well. Resolving these warnings helps ensure your project will be compatible with the new engine when it’s released. 📘 See how to [fix deprecation warnings →](https://docs.getdbt.com/reference/deprecations) ## What happens on May 28? We’ll provide migration guidance on [docs.getdbt.com](http://docs.getdbt.com/) as part of the launch of the new engine. In the meanwhile, [**join us live on May 28**](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase) for the dbt Launch Showcase, where our product team will go into more depth on the new engine, and share everything else that’s new in dbt. We’ll be showing: - The powerful new **VS Code extension** for dbt development powered by the new engine - A **reimagined visual editing experience** for analyst- and governance-friendly modeling - **Cost management features** to help you understand and optimize resource use … and much more. You’ll get a front-row seat to the future of data development, and we’ll show you exactly how to take advantage of it. We’ll catch you at the dbt Launch Showcase. ![Promo for the dbt Launch Showcase on May 28th. ](https://cdn.sanity.io/images/wl0ndo6t/main/dbfc61e26586b02d3c2a4108ebc05fcb128b94de-1800x945.png) We’ll catch you at the dbt Launch Showcase. ![Promo for the dbt Launch Showcase on May 28th. ](https://cdn.sanity.io/images/wl0ndo6t/main/dbfc61e26586b02d3c2a4108ebc05fcb128b94de-1800x945.png) --- --- title: "Moving data teams from cost center to profit driver" description: "Data is cheaper than ever, but it’s gotten harder to control costs. Here’s how data teams can grow without breaking the bank." url: "https://www.getdbt.com/blog/cost-center-profit-driver-data-teams" date: "2025-05-08" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Moving data teams from cost center to profit driver Data teams are having their [Jevons Paradox](https://en.wikipedia.org/wiki/Jevons_paradox) moment. The historical price of computer memory and storage has plummeted. In 1990, a terabyte of storage would cost you $10 million. Today, that same terabyte costs a mere $50. Similarly, cloud providers such as AWS continue to slash prices on services; AWS has famously lowered prices 107 times since its inception. That doesn’t mean, however, that data teams can stop worrying about costs. Because, as costs have gone down, demand for data has gone up. That has data teams consuming more resources than ever. That’s the Jevons Paradox: Increased efficiency creates increased demand. It’s like adding new lanes to a highway. An added lane will alleviate traffic for a while. Eventually, more people start driving because they realize they can get around faster. Pretty soon, the road’s backed up again. To address this challenge, more data teams are transitioning from cost centers into strategic drivers of value. We talked with some experts—Ben Kramer, senior director of data analytics at Bilt Rewards, and Colin Lennon, an analytics engineer at ClickUp—about how they’re navigating this transition through deliberate strategies and cultural shifts. **** ## The mindset shift: From reactive to proactive Moving from a cost center to a profit driver means going from a reactive to a proactive mindset when it comes to data. That means going out and seeking out value, rather than waiting for stakeholders to approach with questions. At ClickUp, a project management platform, the analytics team implemented this mindset shift, positioning themselves as value creators rather than service providers. Similarly, at neighborhood loyalty program Bilt, data engineering leaders looked for opportunities beyond reporting and analytics, extending their reach into operational use cases. That’s moved the data team from the background to being a product displayed directly to end customers. This transition from a background support function to a front-and-center role in product delivery fundamentally changes how data teams are perceived. With greater visibility comes increased responsibility and more stringent Service Level Agreements (SLAs). The upside? Additional resources and recognition of the data team's critical function within the organization. ## Quality and standardization as cornerstones When data becomes customer-facing, quality standards naturally increase. That requires doubling down on data quality practices and instituting gates and processes to ensure all new data projects meet a high bar for release. For Bilt, the shift to profit driver led them to lean more heavily into good [analytics development lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle) practices. These include: - Increased code reviews - Using tools such as the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) to create singular definitions of metrics used across teams - An increased focus on data [documentation](https://docs.getdbt.com/docs/build/documentation), so that engineers, analysts, and business decision-makers better understand the data they’re using Engineers at Bilt feel more bought into creating high-quality data because they know they’re more directly embedded in the business. “The engineers understand the importance of the data that they create so that it can be displayed to the customer directly,” Kramer said. ClickUp found that aligning with the business on standardized key metrics was critical for success. By establishing common definitions for both themselves and downstream functions, they created efficiencies that ultimately reduced costs while improving value delivery. ## Measuring ROI: The perennial challenge For many data teams, demonstrating return on investment remains challenging, particularly for platform work. Changing this requires a shift in thinking—both on the part of data teams as well as on the part of the business. Bilt found that their operational integration - embedding data teams deeply within product teams—provided natural Return on Investment (ROI) justification. When data directly supports customer-facing features, the business value becomes self-evident. In this situation, it’s less likely that every new query will come under stringent quarterly cost scrutiny. ClickUp takes a direct approach to ROI measurement. When they take in new work—whether work that comes to them reactively or that they find proactively—they associate it with Annual Recurring Revenue (ARR) or OKRs (Objectives and Key Results). That enables them to show how their work ties back directly to the business. This focus on measurable outcomes helps data leaders make strategic decisions about where to invest their limited resources. If a project can't be tied to business growth or operational improvement, it may be deprioritized in favor of initiatives with more straightforward value propositions. ## Balancing cost control with innovation One concern with over-focusing on cost control is that it might stifle innovative thinking. Maintaining innovation is critical, especially as organizations grapple with how best to incorporate AI into their businesses. The key is to keep innovation at the forefront, but always have costs in the back of your mind. Tying back experiments to business OKRs alleviates much of this concern. “Especially with AI,” Lennon said, “there needs to be a little bit of room to kind of experiment and say hey, I'm willing to take this bet and this is going to provide value for us.” If the bet doesn’t pan out, he said, it can be refined and tried again. ClickUp isn’t shy about experimenting with different AI-based approaches—whether that’s leveraging [Snowflake’s Cortex](https://www.snowflake.com/en/product/features/cortex/) technology or feeding product review data to a [Large Language Model (LLM)](https://aws.amazon.com/what-is/large-language-model/) to extract sentiment. The data team is primarily using AI now to facilitate existing workflows built upon [dbt models](https://docs.getdbt.com/docs/build/models) to make them more efficient, with a focus on growth marketing. Bilt follows much the same tack, experimenting with building a natural language interface for their business users. Instead of users asking data questions via Slack, which requires an engineer to stop their work, go into the data warehouse, make a query, etc., they can self-service answers using a natural language UI that connects [Claude](https://www.anthropic.com/claude) to [Google BigQuery](https://cloud.google.com/bigquery). That experimentation, however, can lead to budget pressure. “Self-service is awesome until you look at your BigQuery bill or Snowflake bill at the end of the quarter,” said Kramer. Bilt handles this by building out alerts for high query or per-user costs. It’s also committed to obsoleting high-cost tests or dbt model builds. ## Creating a culture of cost consciousness Part of reducing costs while innovating is fostering a cost-consciousness culture. As the keepers of the data keys, data teams are uniquely positioned to drive this shift. Bilt’s data team partners with finance to review contracts, analyze tool usage, and identify opportunities for efficiency. For example, the data team led an initiative to reduce overhead within both their eventing tools and the BigQuery engine. While very much a behind-the-scenes workflow, it’s an important one that supports the company’s growth-oriented teams, such as platform or engineering. At Clickup, the team asks their stakeholders to help them identify how to associate a given initiative with business goals or OKRs. That drives judgments about whether the team can cut costs. ClickUp’s Lennon also said that concerns about costs tend to ebb and flow and that the data team remains flexible in response to that. Before the COVID-19 pandemic, for example, the focus was on shipping fast and making products “good enough.” During the pandemic, the focus shifted back to focusing on costs. With things settling back down now, Lennon says the team is finding “a happy medium” that balances rapid delivery with cost consciousness. ## How a $2 test led to thousands in savings Another area of focus for both teams is hidden costs. Tool sprawl and unused data models represent significant hidden expenses for many organizations. Small inefficiencies, left unmonitored and unaddressed, can compound into significant expenses over time. ClickUp experienced unexpected cost spikes that necessitated an investigation into historical usage patterns. By implementing cost monitoring dashboards, they gained visibility into spending trends and could identify anomalies for further investigation. Bilt implemented a two-layer approach to cost visibility: internal dashboards that track hourly costs by user and automated alerting for high-cost operations. This visibility enabled them to identify a downstream test costing approximately $2 each time it ran (hourly). On further inspection, they realized the test had been rendered redundant by upstream testing. Removing this single test translated to thousands of dollars in monthly savings. **** ## Balancing access with control As Bilt realized, self-service is great until the organization finds itself saddled with a big bill. As self-service analytics capabilities expand, especially with AI-driven natural language processing, organizations need to balance [democratized access](https://www.getdbt.com/blog/managing-data-democratization) with cost and security implications. Bilt has felt the effects of the Jevons Paradox on its own data team. A year ago, the team had two people. The team is now six people. And those six are just as busy, if not more so, than the two were last year. As the excitement around data and ease of access grows, so does the demand. Using Claude and BigQuery, Bilt can do far more with its six people than it could otherwise. However, that requires constant monitoring to ensure costs don’t spiral out of control. ClickUp said it had to shift its mindset around new data projects. Previously, the goal was to build things that were useful to the business. Now, as their work has gone global, the team has to consider other factors, including: - How sensitive is the data? - How valuable is the use case we’re supporting? Within this framework, says ClickUp’s Lennon, the company takes a stance of granting access to data, on the assumption that greater access will drive faster project completion and get the business to its OKRs. ## Looking to the future Both Bilt and ClickUp see themselves facing challenges with scaling. As Bilt grows as a company quarter over quarter, it’s challenging the data team, both in terms of servicing the business as well as retaining innovation without increasing costs. Bilt’s focus for the near term is two-fold: - Deeper product analysis to better understand customers - Leverage AI in both internal and external use cases—e.g., enabling customers to interact directly with an ecosystem-connected LLM Both of these initiatives, while valuable, could lead to cost spikes that threaten business ROI. That makes monitoring and reining in costs an ongoing priority. Meanwhile, ClickUp is asking similar questions of its burgeoning AI initiatives: Is this given project helping us be more effective? Or is it just incurring costs? The company is looking at how it can build AI functionality into its existing infrastructure, as well as leveraging AI tools to boost productivity within existing workflows. There may be no escape from the Jevons Paradox. However, as Bilt and ClickUp have shown, fostering a culture of cost consciousness and making smart, strategic bets on technologies like AI can enable data teams to do more with less. By working closely with business stakeholders and tying data projects back to business objectives, data teams can shift from being perceived as a cost center to being regarded across the organization as a key driver of value. --- --- title: "How to build trust in your data products" description: "Build confidence in your data products with metadata, governance, and clear product ownership that scales trust across your org." url: "https://www.getdbt.com/blog/build-trust-in-data-products" date: "2025-05-06" authors: ["Joey Gault"] categories: ["Pulse"] --- # How to build trust in your data products A [data product](https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/scenarios/cloud-scale-analytics/architectures/what-is-data-product) represents more than just a dataset or report—it's a curated asset designed to solve specific business problems while maintaining the attributes that foster trust. These assets include database tables, dashboards, machine learning models, or any deliverable that serves as the output of a data producer for consumption by others. The distinction between regular data deliverables and true data products lies in the additional properties that make data assets trustworthy. Data products provide three key advantages: discoverability, access control, and backward compatibility. These advantages directly address the trust deficit that often exists between data producers and consumers. To achieve these trust-building benefits, data products must be discoverable through searchable catalogs, preventing valuable data from becoming "dark data" that consumes resources while generating no business value. They need unique identifiers that enable consistent access across teams, whether through database connections, S3 URIs, or HTTP URLs. Most importantly, they must be trustworthy and observable, allowing consumers to inspect data origins, update frequencies, and transformation logic. Self-describing metadata ensures that business context travels with the data, while interoperability mechanisms enable seamless integration across different tools and workflows. Security and governance controls provide the foundation for regulatory compliance and appropriate access management. When these elements work together, they create data assets that teams can confidently build upon. ## Implementing 'data as a product' thinking The mindset shift toward treating data as a product fundamentally changes how teams approach data development and maintenance. This approach applies product management principles to datasets, ensuring they possess the qualities that build user trust and drive adoption. Product-like management means data producers actively engage with stakeholders to understand requirements and create development backlogs that address real business needs across multiple releases. This proactive approach replaces the reactive cycle of one-off requests that often leads to inconsistent or contradictory outputs. Interfaces and contracts become critical trust-building mechanisms. Data producers create explicit specifications that define the structure, fields, and types for each version of their data products. These contracts serve as agreements with consumers, providing the predictability necessary for building dependent systems and analyses. Versioning enables controlled evolution of data products without breaking existing consumers. When breaking changes become necessary—such as removing fields or changing data types—new versions are created while maintaining support for previous versions during defined transition periods. This approach prevents the sudden disruptions that erode trust in data systems. Access rules become integral to every release, ensuring that security and compliance considerations are built into the product rather than added as an afterthought. This systematic approach to access control helps maintain trust while enabling appropriate data sharing across the organization. ## Organizational benefits of the product approach The product-oriented approach to data management delivers benefits that extend beyond individual datasets to transform organizational data capabilities. For data development teams, this approach organizes operations around standardized practices for documentation, version control, and troubleshooting. Teams gain a systematic way to track their work and support streamlined dataflows at scale. Development priorities become aligned with business needs rather than driven by the loudest or most recent request. With proper standards in place, data development becomes more strategic and proactive. New versions can address diverse stakeholder needs systematically rather than through ad hoc responses that may conflict with each other. The ability to discover and reuse existing work accelerates new data product development while reducing the waste and inaccuracies that come from duplicative efforts. Teams can build more quickly when they can confidently leverage the work done by others, knowing that proper documentation, testing, and contracts ensure reliability. For the broader organization, data products enable self-service capabilities that remove traditional barriers to data usage. Instead of requiring custom pipeline development through centralized data engineering teams, consumers can leverage existing data products to address their specific needs. This shift reduces bottlenecks while maintaining quality and governance standards. Data silos break down when organizations establish consistent criteria for exposing data across teams. The standardized approach to discoverability, addressability, and security streamlines data access while ensuring appropriate controls remain in place. This balance between accessibility and governance builds trust at the organizational level. ## Technical implementation with modern tools Building trustworthy data products requires tools that support the full lifecycle of product development, from discovery through deployment and maintenance. A robust data platform architecture provides the foundation, with data catalogs serving as the single source of truth for discovering assets regardless of their location within the organization. Data transformation tools like [dbt](https://www.getdbt.com/product/what-is-dbt) play a crucial role in enabling teams to create trustworthy data products. Through dbt, teams can create data models that import from various sources while maintaining clear lineage and dependencies. The ability to create comprehensive tests verifies data quality automatically, while auto-generated documentation provides the metadata and context that consumers need to use data products confidently. Model contracts guarantee a model’s columns and data types before it builds. When you need breaking changes, model versions provide a migration window and smoother upgrades for downstream consumers. The combination of testing, documentation, and contracts creates a framework for building and maintaining trust systematically. The discovery capabilities built into modern data platforms enable teams to find and understand existing data products across the organization. This discoverability reduces duplication while encouraging reuse of proven, tested assets. When teams can easily find relevant data products and understand their provenance, quality measures, and business context, they're more likely to build upon existing work rather than creating redundant solutions. ## Establishing governance and quality frameworks Trust in data assets requires consistent approaches to governance and quality that scale across the organization. Rather than relying on manual processes or ad hoc quality checks, successful organizations embed governance into their data product development workflows. Automated testing becomes a cornerstone of trustworthy data products. By defining tests that verify data quality, completeness, and consistency, teams can catch issues before they propagate to consumers. These tests should cover not just technical correctness but also business logic validation, ensuring that data products meet their intended purposes. Data lineage tracking provides transparency that builds confidence in data products. When consumers can see exactly where data originates, what transformations have been applied, and when updates occurred, they can make informed decisions about how to use the data. This visibility also enables faster troubleshooting when issues arise. Access controls and security measures must be designed into data products rather than layered on top. Role-based access control, encryption at rest and in transit, and audit logging provide the security foundation that enables broader data sharing while maintaining appropriate protections. When security is built into the product development process, it becomes a trust enabler rather than a barrier. ## Measuring and maintaining trust Building trust in data assets is not a one-time effort but an ongoing process that requires measurement and continuous improvement. Organizations need metrics that help them understand whether their data products are meeting trust and quality objectives. Usage metrics provide insights into which data products are gaining adoption and which may have trust issues that prevent broader use. Low adoption rates often signal problems with discoverability, documentation, or quality that need to be addressed. High adoption rates, conversely, indicate successful trust-building that can be replicated in other data products. Quality metrics should track both technical accuracy and business relevance. Data freshness, completeness, and consistency provide technical quality indicators, while business metrics might include user satisfaction scores or the frequency of data-driven decisions based on specific products. Feedback mechanisms enable continuous improvement of data products based on consumer experiences. Regular surveys, usage analytics, and direct feedback channels help data producers understand how their products are being used and where improvements are needed. This feedback loop ensures that data products continue to meet evolving business needs while maintaining the trust that enables their adoption. The path to building trust in data assets requires commitment to systematic approaches that prioritize user needs, quality, and transparency. By treating data as products with proper governance, documentation, and lifecycle management, organizations can create data assets that teams confidently build upon. The investment in trust-building mechanisms pays dividends through increased data adoption, reduced duplication, and more effective data-driven decision making across the organization. Success comes not from implementing any single tool or process, but from creating a comprehensive approach that makes trustworthiness a fundamental characteristic of every data asset. **What is a data product? ** A data product represents more than just a dataset or report—it's a curated asset designed to solve specific business problems while maintaining attributes that foster trust. These assets include database tables, dashboards, machine learning models, or any deliverable that serves as the output of a data producer for consumption by others. The distinction between regular data deliverables and true data products lies in the additional properties that make data assets trustworthy, including discoverability, access control, and backward compatibility. **How do you manage data assets effectively to reduce costs, increase value, and maintain inventory and security?** Organizations can manage data assets effectively by implementing a product-oriented approach that includes several key elements. This involves creating searchable catalogs to prevent valuable data from becoming "dark data," establishing unique identifiers for consistent access across teams, and implementing automated testing to verify data quality. Security and governance controls should be built into the product development process rather than added as an afterthought. Additionally, organizations should establish data lineage tracking for transparency, implement role-based access controls, and create feedback mechanisms for continuous improvement based on consumer experiences. **What are the key components of data assets, and how do they add value?** The key components of trustworthy data assets include discoverability through searchable catalogs, unique identifiers for consistent access, self-describing metadata that provides business context, and interoperability mechanisms for seamless integration across tools. Security and governance controls provide the foundation for regulatory compliance and access management. Quality management is achieved through automated testing that verifies data completeness and consistency, while data lineage tracking provides transparency about data origins and transformations. These components work together to create data assets that teams can confidently build upon, reducing duplication and enabling self-service capabilities. --- --- title: "DevOps meets DataOps: Streamlining analytics with dbt and Snowflake" description: "Transform analytics workflows with DevOps for data. Automate testing, CI/CD, and deployments using dbt + Snowflake." url: "https://www.getdbt.com/blog/devops-dataops-dbt-snowflake" date: "2025-05-05" authors: ["Luis Leon"] categories: ["Learn"] --- # DevOps meets DataOps: Streamlining analytics with dbt and Snowflake For years, deploying data changes to production meant running manual processes based on imperative data pipelines. These processes were often error-prone and lacked even basic quality checks. In the software world, [DevOps](https://aws.amazon.com/devops/what-is-devops) changed the way that software engineers deploy applications. The use of version control, testing, and automation brought a new level of consistency and quality to deployments, resulting in fewer production bugs and less downtime. A similar transformation is happening in the data world. DataOps adopts DevOps to data workflows, using techniques such as: - Version control to track changes - Automated deployment pipelines driven by code review approvals - Automatically run tests to verify transformation logic in pre-production environments I’ll dive into how DevOps principles are transforming data teams, changing the way the industry does analytics. I’ll also look at how you can use dbt and Snowflake to achieve faster time to value with declarative data transformations and automated deployments. **** ## DevOps + DataOps: A new era of data management Traditional data workflows are typically written in imperative languages, often using different languages and workflow tools for different workflows. Code often isn’t stored in a central location or deployed in a regular manner. This scattershot approach hampers the scalability of data workflows in a number of ways: - No one knows where the code for a given workflow resides - There’s little code reuse across projects - Analytics code changes aren’t adequately reviewed or tested prior to deployment, resulting in report downtime and additional dev work - Changes can’t be easily reverted if a problem is detected in production A DataOps approach to analytics code addresses these issues by adopting [a mature analytics workflow](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle), based on DevOps principles, that involves all data stakeholders. It treats every analytics code change as a software deliverable that involves planning, development, testing, deployment, and operationalization. ![ADLC loop](https://cdn.sanity.io/images/wl0ndo6t/main/948eb5cda47eacf1bf78a2268c0666be43706ce5-4581x2126.png) DevOps practices integrate once-siloed teams across the software development life cycle, from Dev to QA to Ops, resulting in both faster innovation and improved product quality. Used together, dbt and Snowflake support a DevOps approach to analytics workflows that results in higher quality and faster time to value. [dbt](https://www.getdbt.com/) is the industry standard for data transformation at scale. It acts as a [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) for your company’s data estate, providing a single, cross-vendor solution for [modeling data transformations](https://docs.getdbt.com/docs/build/models) declaratively using just SQL and YAML. Snowflake is a powerful AI Data Cloud that allows you to build data-intensive applications without the operational overhead. Snowflake [supports a DevOps approach](https://docs.snowflake.com/en/developer-guide/builders/devops) to accelerating development lifecycles through a number of features, such as API automation support to manage schema changes, dynamic tables, and more. You could opt to use other, more imperative approaches in Snowflake to transform your data— e.g., Snowflake stored procedures. But these bring in more technical complexity and are less easy to test and deploy. Because it uses only SQL and YAML, dbt provides a significant productivity boost, offering a declarative and easy-to-understand model for data transformation. In speaking with Snowflake and dbt users, we have heard universally that dbt’s declarative approach complements Snowflake’s DevOps support, providing a higher time-to-value than building imperative pipelines. ## Hands-on: Deploying a DevOps-driven data pipeline Let me show how this works in practice. In my [last article](https://www.getdbt.com/blog/data-pipelines-snowflake-dbt), I showed how to model a Profit & Loss data pipeline using dbt and Snowflake. In this article, I’ll narrow in on how dbt and Snowflake features support a DevOps approach to analytics workflow deployments. You can [access the source code](https://github.com/Snowflake-Labs/sfguide-deploying-pipelines-with-snowflake-and-dbt-labs?_fsi=b8uZIrMV&_fsi=b8uZIrMV) for this project on Github. We also have [an end-to-end quickstart](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html) you can follow if you’re new to either dbt or Snowflake. ### Prerequisites - A [Snowflake account](https://app.snowflake.com/), plus a user with the ACCOUNTADMIN permissions - A [GitHub account](https://github.com/signup) into which you can import (fork) the example code. You can join GitHub for free if you don’t have an account you can use for this walkthrough. ### Setting up a development environment A hallmark of the DevOps process is committing all code to a source control repository such as GitHub. This means committing any SQL Data Definition Language (DDL) commands for creating data warehouse schemas as well as all data transformation code. This enables anyone to find data transformation code and contribute changes. That accelerates data pipeline development while also providing detailed tracking for all proposed and approved code changes. Another DevOps feature is developing and testing in pre-production environments. This prevents changes from going live to production, not just until development is complete, but until: - Another engineer on the team can perform a code review; and - The changes are thoroughly tested against sample data In Snowflake, you can use the EXECUTE IMMEDIATE command to run a SQL script directly from your GitHub repository. This enables you or any developer to create your own isolated environment in which you can load data, develop, and test changes: ```sql USE ROLE ACCOUNTADMIN; USE DATABASE SANDBOX; ALTER GIT REPOSITORY DEMO_GIT_REPO FETCH; EXECUTE IMMEDIATE FROM @DEMO_GIT_REPO/branches/main/scripts/deploy_environment.sql USING (env => 'DEV'); ``` Once you’ve created an environment, you can create a dbt project and begin creating data transformation models. The GitHub project already has a sample dbt project, which you can run [after configuring your project to access your Snowflake instance](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html#1). Besides simplifying data transformation development, dbt also supports DevOps deployments via support for [testing](https://docs.getdbt.com/docs/build/data-tests) and [documentation](https://docs.getdbt.com/docs/build/documentation). Developers can write data tests alongside their models and run them before, during, and after deployment. Documentation generated from model metadata gives data consumers details on what data means, where it comes from, and how to use it. ### Automating deployments with dbt and Snowflake Deploying your development pipeline with dbt is simple—all you need is to issue a dbt run command: ```sql cd dbt_project dbt run ``` Similarly, you can deploy your dbt documentation to end users with a few simple commands: ```sql dbt docs generate dbt docs serve ``` You can create as many stages as you need in your DevOps pipeline. For example, you can use the same script above with a target of PROD to create a production environment: ```sql USE ROLE ACCOUNTADMIN; USE DATABASE SANDBOX; ALTER GIT REPOSITORY DEMO_GIT_REPO FETCH; EXECUTE IMMEDIATE FROM @DEMO_GIT_REPO/branches/main/scripts/deploy_environment.sql USING (env => 'PROD'); ``` You can also control [how data is **materialized** by dbt](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html?index=..%2F..index#6) based on environment. By default, dbt will use a CREATE TABLE AS SELECT query to re-create an entire table with each deployment. This probably isn’t what you want in a production deployment. Instead, for prod, you can switch to [incremental materialization](https://docs.getdbt.com/docs/build/incremental-models), which only creates new rows according to filter criteria that you control. I already discussed how to deploy a new database schema using EXECUTE IMMEDIATE. But what if you need to make more complicated changes and migrations to existing data as part of a deployment? For that, you can write Python scripts that use the [Snowflake API](https://docs.snowflake.com/en/developer-guide/snowflake-python-api/snowflake-python-managing-databases) to instrument creating or altering tables, performing table operations, swapping table names, and other operations. ### Monitoring and quality control The job isn’t over once your changes are deployed. You need to monitor your pipeline for performance, as well as check incoming data to ensure that your pipeline anticipates data edge cases and transforms fields correctly. The easiest way to do this is to run your dbt tests on incoming production data. You can automate this [using CI/CD features in dbt Cloud](https://docs.getdbt.com/docs/deploy/continuous-integration), the commercial version of dbt built to support DevOps workflows. You can start with basic tests, [expanding your test suite over time](https://www.getdbt.com/blog/data-testing) to become more proactive, avoid “alert fatigue” with meaningful and well-crafted alerts, and decrease time to issue detection and resolution. ## Dynamic tables for near-real-time analytics Snowflake is regularly adding new features to streamline the DevOps deployment process. One of my favorites is [dynamic tables](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html#7). With dynamic tables, you can tell Snowflake to use a query to transform data from one or more base objects. You can do this by making a simple dbt project configuration change: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f114159011cf60dddb91d46995588228d4bf84df-458x216.png) You can fine-tune dynamic tables to specify a target freshness, or lag. This ensures that your data remains up to date within a specific freshness threshold. For example, you could specify a lag of five minutes to ensure your destination table is never more than five minutes behind your base table. You can lower this value further to deliver near-real-time analytics using a few simple project configuration changes as opposed to writing a bunch of imperative logic. ## Conclusion: The shift to DevOps-style automation As AI increases the demand for high-quality data, data engineering teams need to move faster than ever. That’s why dbt and Snowflake are focused on streamlining analytics workflows with easy, declarative tooling and support for DevOps best practices. Going forward, we’re looking at ways to make the analytics lifecycle even shorter and faster without sacrificing quality. [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot), for example, is leveraging AI to reduce the time required to generate models, documentation, and [metrics](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl). And Snowflake is working hand-in-hand with dbt on deeper integrations that will reduce cycles even further. See for yourself how dbt and Snowflake can accelerate your analytics workflows by [signing up for a free dbt Cloud account](https://www.getdbt.com/signup) and [exploring Snowflake’s DevOps capabilities](https://docs.snowflake.com/en/developer-guide/builders/devops). --- --- title: "How analytics engineers can use AI in their everyday workflow" description: "Five powerful ways to bring AI into your dbt workflow without sacrificing quality." url: "https://www.getdbt.com/blog/how-analytics-engineers-can-use-ai-in-their-everyday-workflow" date: "2025-05-02" authors: ["Carolina Moura"] categories: ["Learn"] --- # How analytics engineers can use AI in their everyday workflow _This guest post comes from Carolina Moura, an analytics engineer at [Indicium](https://indicium.ai/)._ If you’re an analytics engineer, you’ve likely spent hours documenting dbt models, debugging cryptic errors, or chasing down bugs caused by a stray apostrophe. Complex SQL calculations, cross-platform migrations, and repetitive documentation tasks slow you down. But AI can help. Used strategically, AI boosts efficiency, reduces errors, and frees up time for high-impact work. In this article, I’ll break down five powerful ways to bring AI into your dbt workflow, without sacrificing quality. ## 1. Generate dbt documentation Documenting dbt models is essential for maintaining data governance, but writing descriptions manually is time-consuming and often inconsistent (who hasn’t missed an indentation in a Jinja block?). Large language models (LLMs) makes this process faster and more standardized. Generative models can analyze SQL code to create descriptions for tables and columns, and even suggest data quality tests. Let’s use the _stg_orders.sql_ model from the Northwind database as an example: ![Code snippet showing the stg_orders.sql dbt model from the Northwind database](https://cdn.sanity.io/images/wl0ndo6t/main/715906282543fb88b2d29e7c63c24f2078523bfc-414x235.jpg) We can use an LLM to automatically generate documentation for the model by prompting something like: _“Based on the dbt model code stg_orders.sql, generate the corresponding stg_orders.yml documentation file. Include descriptions for the table and columns. Also, suggest relevant generic tests to ensure data quality, such as null value validation, uniqueness checks, and date range validation.”_ Here’s how AI might structure the YAML, with descriptions and tests inferred from the schema: ![YAML file with table and column descriptions generated by AI for the stg_orders model](https://cdn.sanity.io/images/wl0ndo6t/main/aa4d969c26d4e094c383fde3f7544d0ec9f73bc2-895x613.jpg) That said, for columns with business-specific logic, AI may miss important nuances. It’s essential to review and adjust accordingly. Keep an eye out for some common pitfalls, which I’ll cover at the end of the article. dbt also offers [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot), built directly into dbt Clouds IDE. With dbt Copilot teams can autogenerate documentation by leveraging deep context from your models capturing metadata, relationship, and structure so you get more accurate, reliable outputs with less manual effort. ## 2. Interpret error logs Who hasn’t run into an indecipherable error and spent far too long trying to understand the root cause? Logs can be lengthy and hard to interpret, especially when they involve references between models, custom macros, or dbt_project.yml configurations. Instead of analyzing each detail manually, you can use an LLM to interpret the error and suggest a fix. When running a dbt model, you might encounter this error: ![Screenshot of a dbt error message encountered during model execution](https://cdn.sanity.io/images/wl0ndo6t/main/c7de07075e9db27c0b2d28614f39107670ee82e8-948x54.jpg) After submitting the error to the LLM with the prompt _“Explain this dbt error and suggest a fix”_, you might receive a response like this: ![Text generated by AI providing a diagnosis and suggested fix for a dbt model error](https://cdn.sanity.io/images/wl0ndo6t/main/c863f9a55906e4445232aa53eee66536c0555386-1024x379.jpg) In seconds, you’ll get a potential diagnosis and solution, saving time and reducing frustration. And if you don’t want to context switch to an external LLM, dbt Copilot makes it even easier. Built into the dbt Cloud IDE, Copilot lets you highlight an error directly in your SQL, hit Command+B, and automatically generate a prompt to ask for an explanation—so you can diagnose and fix issues without ever leaving your workflow. ## 3. Convert functions between data warehouses In data migration projects, adapting functions from one data warehouse to another (e.g. Redshift to Snowflake or BigQuery to Databricks) can be time-consuming. Each platform has its own syntax, operators, and native functions, which require manual adjustments to ensure compatibility. AI can automate much of this process. Suppose you're migrating this model from Redshift to Snowflake: ![SQL code example of a dbt model written for Redshift](https://cdn.sanity.io/images/wl0ndo6t/main/c36c533b46f7fa96a0dabed7b263cb3ce5cacda3-541x199.jpg) You can use a prompt like this to generate a version compatible with Snowflake: _“Convert the following dbt model from Redshift to Snowflake, keeping the logic intact.”_ ![SQL code converted by AI from Redshift to Snowflake syntax](https://cdn.sanity.io/images/wl0ndo6t/main/83dd9b76484eb698a579cc61132480464b2b6074-566x199.jpg) The LLM adapts the syntax to the new platform while preserving the original logic, which reduces manual work and human error. This speeds up migrations and helps maintain consistency across environments.** **And you can do this right in dbt: just highlight a block of SQL from a dbt model, use dbt Copilot, and prompt it to convert the code, without ever leaving your IDE. ## 4. Optimize dbt models Beyond syntax conversion, AI can suggest improvements for performance and readability. It can help refactor subqueries into Common Table Expressions (CTEs), reduce duplicated logic, and simplify joins. These optimizations can reduce execution time, improve maintainability, and make your models easier to understand, all without changing the underlying logic. Here’s an example of a dbt model that could benefit from refactoring: ![SQL code of a dbt model with potential for performance and readability improvements](https://cdn.sanity.io/images/wl0ndo6t/main/f2e457c62bca5bcbc14c3df8b4aecb144392613a-979x343.jpg) You can use a prompt like this: _“Optimize the following dbt model for performance and readability without changing the logic. Follow SQL code style best practices.”_ The AI-generated version might look like this: ![Refactored dbt model generated by AI with improved performance and readability](https://cdn.sanity.io/images/wl0ndo6t/main/855aa1ec03d0d341eb1833434e9cd1eb356a71f3-659x757.jpg) For larger projects with complex transformations, small optimizations like these can scale into meaningful gains, not just in performance, but also in clarity and team collaboration. And you can do this right in dbt Cloud’s IDE: just highlight a block of SQL in-line, use dbt Copilot, and prompt it to optimize for performance, readability, or even apply your custom SQL style guide—all without leaving your workflow. ## 5. Automate dimensional modeling Designing a scalable data model is a challenge, especially when working with multiple tables. AI can suggest optimized structures for fact and dimension tables, helping you build more efficient and scalable models. Using the Northwind database diagram as a reference, you can prompt: _“Based on the diagram provided, design a dimensional model.”_ ![Entity-relationship diagram of the Northwind database](https://cdn.sanity.io/images/wl0ndo6t/main/582edc554720799f9a6fec174d8a8f27941eb7b3-974x747.jpg) You might receive an output like this: ![Dimensional model generated by AI using the Northwind database structure, including suggested fact and dimension tables](https://cdn.sanity.io/images/wl0ndo6t/main/e1c315d398fad9f9a70cb46129dccbc7eab89eaa-1024x397.jpg) Once the structure is defined, you can also ask the LLM to: - Create a mapping table between the transactional and dimensional models. - Generate the corresponding dbt models using the transactional tables as sources. - Suggest naming conventions and folder structures to organize your dbt project. ## Best practices when working with AI AI can be a powerful ally, but only when used thoughtfully. Here are some key practices for safe, effective usage: - **Always validate AI output:** Review any AI-generated code or documentation carefully. It may contain syntax errors or misinterpretations. - **Test before deploying:** Run tests on any generated or modified code before pushing to production. - **Keep sensitive data private:** Never input confidential data, credentials, or proprietary business logic into AI tools. - **Learn from AI, don't just copy:** Use the output as a learning opportunity to understand new techniques and patterns. - **Iterate your prompts:** The quality of AI output often depends on how clearly and specifically you ask. Don’t be afraid to refine and experiment. The AI landscape is evolving fast, and tools are becoming more customizable and integrated into data platforms. If you haven’t started using AI in your daily workflow, now’s a great time to begin. Start small: documentation and error debugging are low-risk areas where mistakes are easy to catch. As you explore more use cases, consider building a prompt library tailored to your common tasks. And don’t keep it all to yourself. Share prompts and insights with your team helps to foster a culture of learning and experimentation. Over time, you’ll gain a clearer sense of where AI delivers the most value, and how to make the most of it in your work. --- --- title: "The evolution of databases" description: "Wolfram Schulte, the cofounder and CTO of SDF Labs, now a part of dbt Labs, discusses databases, compilers, and dev tools." url: "https://www.getdbt.com/blog/evolution-of-databases" date: "2025-04-28" authors: ["Daniel Poppy"] categories: ["Insights"] --- # The evolution of databases _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/the-evolution-of-databases-w-wolfram). _ Welcome to our new season of The Analytics Engineering Podcast. This season, we’re focusing on developer experience. We’ll explore the developer experience by tracing the lineage of foundational software tools, platforms, and frameworks. From compilers to modern cloud infrastructure and data systems, we’ll unpack how each layer of the stack shapes the way developers build, collaborate, and innovate today. It’s a theme that lends itself to a lot of great conversations on where we’ve come from and where we’re headed. In our first episode of the season, Tristan talks with Wolfram Schulte. Wolfram is a distinguished engineer at dbt Labs. He joined the company via the [acquisition of SDF Labs](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs) [Labs](https://docs.getdbt.com/blog/sql-comprehension-technologies), where he was [co-founder and CTO](https://www.getdbt.com/blog/building-the-next-gen-dbt-engine). He spent close to two decades in Microsoft Research and several years at Meta building their data platform. One of the amazing things about Wolfram is his love of teaching others the things that he's passionate about. In this episode, he discusses the internal workings of data systems. He and Tristan talk about [SQL parsers](https://docs.getdbt.com/blog/the-levels-of-sql-comprehension), [compilers](https://roundup.getdbt.com/p/the-power-of-a-plan-how-logical-plans), [execution engines](https://docs.getdbt.com/blog/sql-comprehension-technologies), [composability](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle), and the world of heterogeneous compute that we're all headed towards. While some of this might seem a little sci-fi, it’s likely right around the corner. And Wolfram is inventing some of the tech that's going to get us there. Join Tristan May 28 at the [**2025 dbt Launch Showcase**](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase) for the latest features landing in dbt to empower the next era of analytics. We'll see you there. [Register now](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase) _Please reach out at podcast@dbtlabs.com for questions, comments, and guest suggestions._ **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ### Chapters - 01:35 Introduction to dbt Labs and SDF Labs collaboration - 04:42 Wolfram's journey from monastery to tech innovator - 07:55 The role of compilers in database technology - 11:05 Building efficient engineering systems at Microsoft - 14:13 Navigating data complexity at Facebook - 18:51 Understanding database components and their importance - 24:44 The shift from row-based to column-based Storage - 27:40 Emergence of modular databases - 28:44 The rise of multimodal databases - 30:45 The role of standards in data management - 35:04 Balancing optimization and interoperability - 36:38 Conceptual buckets for database engines - 38:46 DataFusion compared to DuckDB - 40:44 ClickHouse - 44:20 Bridging the gap between SQL and new technologies - 50:55 The future of developer experience ## Key takeaways from this episode ### From monastery to Microsoft: Wolfram’s journey **Tristan Handy: Can you walk us through the Wolfram Schulte origin story?** **Wolfram Schulte:** I was born in rural Germany—Sauerland—and ended up in a monastery boarding school after my father passed away. Their goal was to train monks and priests, but that didn’t stick for me. Later I went to Berlin—back then you had to cross East Germany to get there—and began studying physics. But I realized everyone else understood physics better than I did! One day I walked past a lecture on data structures and algorithms, and I was hooked. I hadn’t written a line of code at that point, but I switched to computer science immediately. After my PhD in compiler construction, I joined a startup, then landed at Microsoft Research in 1999 thanks to a chance encounter with the logician Yuri Gurevich. ### Inside Microsoft Research and Cloud Build At Microsoft Research, we were like Switzerland—neutral across teams like Office, Windows, and Bing. We’d invent tools and ideas, but often the business units didn’t trust them. That changed when I was asked to build an engineering org. We created **Cloud Build**, a distributed build system like Google’s Bazel. It reduced build times from hours to minutes and had a huge impact on iteration speed, productivity, and even morale. People stayed in flow. Builds were faster, cheaper, and smarter—running mostly on spare capacity. ### Janitorial work at Meta: cleaning up big data **You later joined Facebook (Meta). What was that like?** A different world. No titles for engineers. Egalitarian, fast-moving. I joined to clean up the data warehouse—what they called “janitorial work.” At Meta, each type of workload had its own engine: time-series, batch, streaming, etc. This made understanding lineage and dependencies across systems extremely hard. We responded by building UPM, a SQL pre-processor that stitched metadata across engines. It became part of Meta’s privacy infrastructure and compliance tooling, especially after the fallout from Cambridge Analytica. ### Databases as compilers **Let’s shift gears. Can you walk us through how analytical databases actually work—like a professor at a whiteboard?** Sure. Think of a database like a compiler: 1. **Parsing & analysis:** Is the SQL valid? Are the types correct? 2. **Optimization:** SQL is declarative, so you can reorder joins, push down filters—based on algebraic laws like associativity. 3. **Execution:** Often done in parallel, especially in modern warehouses. 4. **Storage:** Columnar vs. row-based; optimized formats like Parquet or ClickHouse’s custom format. Historically, storage and compute were bundled. Now they’re decoupled. But when the engine understands the format deeply, performance is much better. ### The rise of modular and composable data platforms **How did we get from monolithic systems to the composable database architectures we have today?** It started with the rise of big data—Hadoop, HDFS, MapReduce. That decoupled compute from storage. Columnar formats like Parquet enabled analytical workloads. Then came Iceberg, Delta Lake, and similar standards that enabled multiple engines to share data. Modern databases are modular. For example, Postgres is transactional, but you can bolt on an OLAP engine for analytical queries. You can mix and match based on your workload. The result is a data ecosystem that’s far more flexible—but also more complex. ### Engine families: Snowflake, DuckDB, ClickHouse **Can you help us bucket the different kinds of engines out there?** Totally. Here are three buckets: - **Cloud-native engines:** Snowflake, BigQuery. They’re optimized for massive scale, often with their own proprietary storage. - **Embedded/single-node engines:** DuckDB, DataFusion. Great for local dev or embedded analytics. DuckDB is for users; DataFusion is for database builders. - **Real-time/high-throughput engines:** ClickHouse, Druid. Tuned for streaming and extremely fast aggregations. Each has its trade-offs. Increasingly, projects are combining these. For example, you can plug DuckDB or DataFusion into Spark to speed up leaf-node execution. The whole engine space is getting more composable—and more interchangeable. ### The role of SDF in dbt’s future **If you think about the future where SDF is fully integrated into dbt Cloud, what does that enable?** Initially, it might feel the same—but faster, smarter. Longer-term, we can give developers superpowers. Imagine your dev environment proactively surfaces: - “This data looks different than yesterday—want to investigate?” - “You’re missing a metric that’s often used alongside this one.” - “This join will behave differently on engine X—here’s what to change.” That’s the kind of intelligent, predictive developer experience we’re building. We’re catching SQL up to what IDEs have done for code. And if we can make logical plans portable across engines, dbt becomes the consistent interface across heterogeneous compute. --- --- title: "ELT best practices for Databricks workflows" description: "Explore ELT best practices on Databricks: layering architecture, SQL warehouses, delta optimizations, and dbt integration." url: "https://www.getdbt.com/blog/elt-best-practices-databricks" date: "2025-04-25" authors: ["Joey Gault"] categories: ["Pulse"] --- # ELT best practices for Databricks workflows When architecting ELT workflows on [Databricks](https://www.databricks.com/), your choice of compute infrastructure directly impacts both performance and cost efficiency. Databricks offers several compute options—SQL warehouses, All-Purpose Compute, and Jobs Compute—but for ELT workloads, SQL warehouses represent the optimal choice. SQL warehouses are specifically optimized for SQL workloads and provide built-in features like query history for auditing and optimization. They scale both vertically to handle larger datasets and horizontally to support concurrent operations. Among the available warehouse types—Serverless, Pro, and Classic—serverless warehouses offer the most compelling advantages for ELT workflows. [Serverless warehouses](https://docs.databricks.com/aws/en/admin/sql/serverless) dramatically reduce spin-up times and scale quickly when workloads demand additional resources. This eliminates the need to keep clusters idle, as serverless warehouses activate rapidly when work begins and shut down when complete. They also leverage [Databricks' Photon engine](https://www.databricks.com/product/photon) automatically, providing optimal performance for both transformation and serving workloads. Sizing your SQL warehouses requires balancing data volume, complexity, and latency requirements. A practical approach is starting with a Medium warehouse and adjusting based on performance observations. Larger warehouses often prove more cost-effective for complex workloads because they complete tasks faster—a Small warehouse taking an hour to complete a pipeline might finish the same work in thirty minutes on a Medium warehouse. Consider provisioning separate warehouses for different workload types. ELT pipelines have different compute patterns than ad-hoc analysis, so dedicated "pipeline" warehouses sized for data volumes and SLAs work more efficiently than shared resources. Factor in dbt's thread count when sizing—higher parallelism requires more compute capacity to maintain performance. Configure auto-stop settings aggressively for serverless warehouses. Since they spin up in seconds, setting auto-stop to five minutes (or as low as one minute via API) won't impact user experience while minimizing costs. ## Architectural patterns: staging, intermediate, and marts Successful ELT implementations on Databricks follow a layered approach that mirrors the [medallion architecture](https://www.databricks.com/glossary/medallion-architecture): bronze (staging), silver (intermediate), and gold (marts). This structure provides clear data lineage, enables incremental development, and supports different consumption patterns across your organization. The bronze layer handles raw data ingestion and initial storage. For datasets in cloud storage, leverage Databricks' `COPY INTO` functionality rather than staging external tables. `COPY INTO` operates incrementally and ensures data is written in Delta format, providing the performance, reliability, and governance advantages that Delta tables offer. You can implement `COPY INTO` as a pre-hook before building downstream models or invoke it using dbt's run-operation command. While staging external tables remains an option for teams migrating from other cloud warehouses, this approach prevents you from leveraging Delta's benefits and requires additional maintenance like running repair operations for new partition metadata. The silver layer focuses on cleaned, modeled data optimized for performance and cost. Many organizations implement incremental processing at this stage, processing only new or updated records rather than recreating entire tables. dbt's incremental model materialization facilitates this approach by creating temporary views with data snapshots and merging them into target tables. Configure the temporal range of your snapshots using conditional logic in your `is_incremental` blocks—the most common pattern merges data with timestamps later than the current maximum in the target table. For teams with tight SLAs, several advanced optimization techniques can improve merge performance significantly. Enable auto compaction to maintain optimal file sizes between 32MB and 256MB. Databricks handles this automatically with optimized writes enabled by default in SQL warehouses, but opting into auto compaction provides additional benefits. Leverage data skipping through Z-ordering for high-cardinality columns frequently used in joins or filters. The syntax is straightforward: `OPTIMIZE table_name ZORDER BY (col1, col2, col3)`. Limit Z-ordering to three columns maximum and run it either as a post-hook after model builds or as a scheduled job on a regular cadence. Maintain current statistics with the `ANALYZE TABLE` command to ensure optimal join plans. Run this for columns frequently used in joins, either as a post-hook or scheduled job: `ANALYZE TABLE mytable COMPUTE STATISTICS FOR COLUMNS col1, col2, col3`. Consider implementing `VACUUM` operations to remove unused files from Delta tables. While deleted records are soft-deleted from the transaction log, underlying files remain in storage. `VACUUM` removes these files to reduce storage costs and improve merge performance, but be cautious about retention periods—vacuuming files older than seven days prevents restoring table versions that depend on those files. The gold layer delivers business-ready marts that stakeholders access through BI tools. Apply the same optimization techniques as the silver layer, with particular attention to Z-ordering since these tables typically have stricter SLA requirements. Gold tables are ideal candidates for defining metrics using dbt's semantic layer capabilities, ensuring consistency across key business KPIs. ## Performance optimization and troubleshooting Effective performance management requires both proactive optimization and reactive troubleshooting capabilities. Databricks provides several tools to identify and resolve performance bottlenecks in ELT workflows. The SQL warehouse query profile serves as your primary troubleshooting tool. It provides detailed information about query execution, including time spent in tasks, rows processed, and memory consumption. The profile offers both tree and graph views to identify slow operations and understand data transformation flows. Common performance issues revealed by query profiles include inefficient file pruning, full table scans, and exploding joins. Address file pruning issues by reordering columns in your bronze-to-silver transformations. Databricks collects statistics on the first 32 columns by default, so position numerical keys and high-cardinality query predicates before the 32nd column, with strings and complex data types after. Full table scans indicate queries scanning entire tables rather than leveraging file-level statistics. File compaction and Z-ordering techniques described earlier help alleviate this problem by improving data layout and enabling better file skipping. Exploding joins produce result sets much larger than input tables, often creating [Cartesian products](https://www.datacamp.com/tutorial/cartesian-product). Prevent these by making join conditions more specific and preprocessing data through aggregation, filtering, or sampling before join operations. For incremental models, rely on merge strategies as the recommended approach for most use cases. Databricks has significantly improved merge performance with low-shuffle merge and Photon optimizations. Optimize merge operations by reading only relevant partitions through filters and `incremental_predicates`, updating only necessary rows and columns, and defining single materialized keys for efficient lookups. ## Monitoring and observability Understanding your ELT performance requires comprehensive monitoring of both [dbt](https://docs.getdbt.com/docs/deploy/monitor-jobs) operations and Databricks resource utilization. dbt generates metadata on timing, configuration, and freshness with each job run, accessible through the Discovery API—a GraphQL service that supports queries on this metadata. Teams can analyze this data like any other business intelligence source by piping it into their data warehouse. The [Model Timing tab in dbt](https://docs.getdbt.com/blog/how-we-shaved-90-minutes-off-model#your-new-best-friend-the-model-timing-tab) provides visual identification of models requiring optimization or refactoring. The [Admin API](https://docs.getdbt.com/guides/optimize-dbt-models-on-databricks?step=1) enables pulling dbt artifacts from runs, allowing you to stage manifest.json files and model the data using the dbt artifacts package. This package helps identify inefficiencies and optimization opportunities across your dbt models. Combine [dbt monitoring with Databricks' native observability features](https://docs.getdbt.com/guides/productionize-your-dbt-databricks-project?step=1) to gain complete visibility into your ELT workflows. Monitor warehouse utilization, query performance, and cost trends to identify optimization opportunities and ensure efficient resource usage. ## Implementation considerations Successfully implementing ELT on Databricks requires careful attention to both technical and organizational factors. Start with clear data contracts that define schema expectations, freshness requirements, and ownership responsibilities between data producers and consumers. These contracts reduce friction and prevent silent failures as your pipelines evolve. Implement version control and CI/CD practices from the beginning. Treat your analytics code like application code, with peer reviews, automated testing, and safe deployment practices. dbt's integration with Git-based workflows makes this straightforward while providing the collaboration benefits that modern data teams require. Consider incremental adoption when migrating from existing ETL processes. Begin with less critical datasets to validate your approach and build team expertise before tackling mission-critical workflows. This approach reduces risk while allowing your team to develop best practices specific to your organization's needs. Plan for schema evolution and data governance from the start. ELT's flexibility in handling schema changes is a significant advantage, but it requires thoughtful approaches to managing those changes across your pipeline. Establish conventions for handling new columns, data type changes, and structural modifications to source systems. ## Conclusion ELT on Databricks represents more than a technical architecture choice—it's a foundation for building scalable, maintainable, and collaborative data workflows. By leveraging serverless SQL warehouses, implementing layered data architectures, and following performance optimization best practices, data engineering leaders can deliver faster insights while reducing operational overhead. The combination of Databricks' cloud-native capabilities with dbt's transformation framework provides a powerful foundation for modern analytics engineering. Teams that embrace these patterns find themselves spending less time on infrastructure management and more time delivering business value through trusted, well-governed data products. Success with ELT requires commitment to software engineering best practices, proactive performance monitoring, and clear organizational processes around data ownership and quality. When implemented thoughtfully, these approaches enable data teams to scale efficiently while maintaining the reliability and trust that business stakeholders demand. ## ELT for Databricks FAQs **What are the similarities and differences between ETL and ELT? ** While the article focuses on ELT implementation, the key difference lies in when transformation occurs. ELT (Extract, Load, Transform) loads raw data first into the target system and then transforms it, leveraging the compute power of modern cloud data platforms like Databricks. This approach provides more flexibility in handling schema changes and allows for faster initial data ingestion, as demonstrated by the medallion architecture pattern with bronze (raw), silver (cleaned), and gold (business-ready) layers. **** **When should teams choose Delta Live Tables over SQL notebooks or PySpark jobs for ELT in Databricks?** For most ELT workloads, SQL warehouses represent the optimal choice over other compute options. SQL warehouses are specifically optimized for SQL workloads, provide built-in features like query history for auditing, and scale both vertically and horizontally. Serverless warehouses offer the most compelling advantages with dramatically reduced spin-up times, quick scaling, and automatic leverage of Databricks' Photon engine. Teams should consider provisioning separate warehouses for different workload types, as ELT pipelines have different compute patterns than ad-hoc analysis. **How do Auto Loader and Delta Lake with MERGE enable scalable, incremental ELT pipelines on Databricks?** For incremental processing, leverage Databricks' `COPY INTO` functionality rather than staging external tables, as it operates incrementally and ensures data is written in Delta format. Delta Lake's merge strategies are recommended for most incremental use cases, with Databricks having significantly improved merge performance through low-shuffle merge and Photon optimizations. Optimize merge operations by reading only relevant partitions through filters, updating only necessary rows and columns, and defining single materialized keys for efficient lookups. Additional optimizations include enabling auto compaction, using Z-ordering for high-cardinality columns, and implementing `VACUUM` operations to remove unused files. --- --- title: "Data engineering tools: How do they fit together?" description: "Explore how modern data engineering tools work together to power reliable pipelines—with dbt at the core of transformation." url: "https://www.getdbt.com/blog/data-engineering-tools" date: "2025-04-25" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data engineering tools: How do they fit together? Data infrastructure today centers on a modular modern data stack rather than a single [Extract, Transform, and Load (ETL)](https://www.getdbt.com/blog/data-transformation-vs-etl) suite. Businesses assemble dedicated tools for each stage of the lifecycle for ingestion, transformation, and storage. This decomposition makes it simple to build, update, and scale each layer independently without rebuilding the entire pipeline. The outcome is not just a pile of tools. It’s a living system for creating dependable data products that’s greater than the sum of its parts. But how do these components interact to convert raw, fragmented inputs into clean, production-grade outputs? In this article, we’ll deconstruct how these elements combine to create the cohesive, automated pipeline of the modern [data engineering](https://www.getdbt.com/blog/dataops-vs-data-engineering) toolchain. We’ll also discuss the benefits and challenges teams encounter in building resilient systems. ## Data ingestion: Getting raw data into the warehouse A robust data ingestion layer decouples data extraction from transformation. Teams can add or modify data sources with little to no downstream effect. Collecting all raw data in a common storage tier like a data warehouse allows easy reprocessing or manipulation of data without re-extracting it. Data ingestion platforms like [Fivetran](https://fivetran.com/docs/getting-started) connect to a variety of sources, including Software as a Service (SaaS) applications, relational, and NoSQL databases. They can be used in batch and streaming mode, including [micro-batch](https://docs.getdbt.com/docs/build/incremental-microbatch) to support near-real-time requirements. The majority of modern pipelines follow an [ELT (Extract, Load, and Transform)](https://www.getdbt.com/blog/etl-vs-elt) pattern, where the tool extracts data sources and loads them in raw form. Transformation engines perform relevant operations to bring the [data into a usable format](https://www.getdbt.com/blog/analytics-engineering-transformation). This is opposed to traditional ETL, which performs all complex transformations before centralizing the data. ## Transformation: Modeling data with software engineering principles A modern data transformation process treats transformations as code, with peer reviews, testing, and automated deployments. dbt is the standard tool for this, enabling modular SQL models to transform data no matter where it lives in your enterprise. The raw tables ingested by ingestion tools are cleaned and transformed into business-ready fact and dimension tables, all expressed as SQL or Python code. Using version control driven by [Git](https://docs.getdbt.com/docs/cloud/git/git-configuration-in-dbt-cloud) and automated deployment workflows, dbt supports an [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) that mirrors a software development lifecycle (SDLC), with changes developed, tested, and released iteratively. Teams can run their data pipelines on schedules or trigger them through orchestrators to keep data up-to-date. Tests can be run on incoming data to ensure continuous data quality and prevent drift. **** ## Storage and compute: Cloud data warehouses In the modern data stack, cloud warehouses have eliminated the need for special processing engines for analytics. Under the ELT paradigm, raw data is loaded first, then transformed directly within the warehouse for a specific business use case. This positions the cloud warehouse as the central, governed layer connecting upstream ingestion and collection with downstream analytics and data products. The stack is built on modern warehouses such as [Snowflake](https://docs.snowflake.com/?_ga=2.131702730.419740291.1754826711-485818585.1754826711), [Google BigQuery](https://cloud.google.com/bigquery), and [Amazon Redshift](https://docs.aws.amazon.com/redshift/) that offer decouple storage and compute into separately scalable services. Cloud data warehouses use a fully managed [massively parallel processing (MPP) architecture](https://www.techtarget.com/searchdatamanagement/definition/MPP-database-massively-parallel-processing-database) to scale SQL workloads across compute clusters, efficiently processing petabyte-scale data and complex joins. The architecture decouples storage and compute, allowing storage and query processing to scale independently, and billing at a granular level. ## Observability: Monitoring pipeline health and data quality Observability tools monitor data health and detect anomalies, maintaining quality in the complex pipeline. Data observability uses DevOps-like monitoring to data pipelines, giving you end-to-end visibility into data quality, freshness, lineage, and performance. A good observability system reports standard metrics, run statuses, record counts, schema versions, and latency into the centralized backend with SLA thresholds and statistical alerts. Once signals are in place, teams can use an [observability platform](https://www.getdbt.com/blog/observability-within-dbt) to automate anomaly detection, root-cause analysis, and incident workflows. ## Analytics and BI: Delivering insights to users Data engineering’s goal is to provide insights to the business. At the pipeline’s end, BI and analytics tools allow users to explore and visualize the transformed data. Tools such as [Looker](https://cloud.google.com/looker/docs), [Tableau](https://www.tableau.com/blog), and [Mode](https://mode.com/blog/five-ways-data-analysis-improves-product-development) can access the cloud warehouse and allow analysts to create reports and dashboards on top of the cleaned, transformed datasets. BI tools in a modern stack have a single source of information since the transformation logic is centralized in dbt and the warehouse. In addition to direct querying, you can use the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/setup-sl) as part of BI semantic systems. The dbt Semantic Layer centralizes the definition and naming of critical cross-team business metrics so that they are accessible and consistent for everyone. For example, the metrics layer and exported metadata in dbt can feed Looker and produce LookML definitions automatically. This integration synchronizes KPIs and measures coded in dbt with dashboards and reports. ## Orchestration: Scheduling and dependency management Orchestration tools provide reliable, repeatable processes for ingestion, transformation, testing, and BI refreshes. While dbt provides basic orchestration out of the box, [tools like Airflow and Prefect](https://docs.getdbt.com/docs/deploy/deployment-tools) can expand dbt’s capabilities by building custom integrations with other third-party systems. ## Example end-to-end workflow Let’s take a closer look at a production-grade data pipeline in action. ### Ingestion Fivetran utilizes CDC and batch connectors to replicate [Salesforce](https://www.salesforce.com/) objects and Google Analytics logs into a raw schema in Snowflake. It auto-maps fields, auto-adds columns, and coerces data types, landing as Snowflake micro-partitioned files. A webhook executes custom Airflow processing. ### Transformation dbt transforms the data using SQL, creating a dependency-aware [directed acyclic graph (DAG)](https://www.getdbt.com/blog/dag-use-cases-and-best-practices) that ensures the order of execution. Staging models filter and cast raw data, parsing and deduplicating it, and then materialize as views or temporary tables. Mart models insert or rebuild aggregated tables in-place using [incremental](https://docs.getdbt.com/docs/build/incremental-models-overview) or full-table [materializations](https://docs.getdbt.com/docs/build/materializations) and allow partitioning and clustering setups in dbt. The warehouse persists these tables in an optimized compressed columnar form on managed object storage, micro-partitions, and clusters to achieve efficient pruning and cost savings. ### Testing dbt runs schema and data tests (uniqueness, not-null, relationships) on freshly constructed models. Monte Carlo consumes metadata events provided by Fivetran, dbt, and the data warehouse to track SLA-based freshness and row-count anomalies, using statistical detection on metric distributions to raise alerts on unexpected schema changes. When it detects an anomaly, Monte Carlo will automatically analyze the impact by examining downstream assets. ### BI refresh A dbt post-hook or Airflow task calls the Tableau Server REST API to refresh Hyper extracts. The refresh is supported by incremental extract logic, which updates only partitions that change to optimize compute and I/O. The orchestrator queries the queryTask endpoint and retries on failure, as well as logs execution metrics in the DAG UI. When it is done, the metadata catalog is refreshed through the API, and dashboards are pointed to the most recent tables. ### Iteration An analyst requests a new KPI through the issue tracker or chat. The [analytics engineer](https://www.getdbt.com/blog/analytics-engineer-vs-data-analyst-vs-data-engineer) forks the dbt repo, updates the SQL in the relevant model, and creates new tests. The pre-merge diff engine that [Datafold](https://docs.datafold.com/integrations/orchestrators/dbt-cloud) uses runs in the CI pipeline and compares the dev and prod model outputs to expose drift. CI runs dbt run and [dbt test](https://docs.getdbt.com/reference/commands/test), and blocks merges on any test or diff failures. ## Benefits of data engineering tools The success of a modern data platform strongly relies on the selection of tools at every step of the data pipeline. Choosing the right tools brings a number of benefits, including: - **Agility and speed:** Modeling workflows can be run in parallel with APIs, files, or event streams, lessening bottlenecks. The cloud-native infrastructure of a modern data pipeline allows scaling the compute, storage, and pipeline concurrency on demand with limited manual work. - **Scalability: **Increasing data quantities are addressed through compute-setting (or schedule) tuning rather than pipeline redesign. Containerized stateless components can be horizontally scaled and can run parallel processes in high-throughput environments. - **Resilience: **Modern data engineering tools integrate observability and orchestration to provide pipeline resilience. Observability monitors the performance and identifies anomalies in real time. Orchestration automates recovery using SLA sensors, retries, and conditional execution. ## Trade-offs and challenges Modern data engineering tools are more flexible and powerful, but also add architectural and operational complexity. A multi-tool stack with modules needs to be well-designed, integrated, and aligned by the team to avoid a brittle or fragmented stack. - **Tool sprawl:** Supporting multiple platforms for ingestion, transformation, orchestration, and observability increases the number of integrations. Minimize tool sprawl through platform consolidation, integration standardization, and automation of provisioning to avoid environment drift. - **Debugging:** The cause of a failure can be anywhere, such as upstream data issues, pipeline problems, model logic issues, or BI configuration issues. Without centralized observability and alerting, the resolution time rises, and incident response becomes reactive. - **Cost and complexity:** Licensing, cloud services, and overhead costs of running various services can become costly. Teams can control costs and complexity with transparent tracking, strong governance, and a streamlined architecture. - **Learning curve: **The new stacks require the ability to work with tools like orchestration frameworks, transformation layers, Git workflows, and cloud warehouses. Onboard new engineers through well-organized onboarding, project-based work, clear documentation, and practical exercises in all tools of the stack and workflows. ## Conclusion Modern data stacks utilize a layered architecture to transform raw data into trusted insights, enabling parallel processing, scaling, and rapid response to data requirements. Transformation is the backbone of this stack, turning fragmented inputs into clean, validated, and well-documented models. dbt’s modular design and support for testing, automated deployment, and monitoring simplify data workflows and strengthen their resilience. To learn more, [schedule a demo of dbt today](https://www.getdbt.com/contact). --- --- title: "ELT best practices for Snowflake workflows" description: "Learn ELT best practices for Snowflake: staging patterns, warehouse strategy, dbt integration, and performance tuning." url: "https://www.getdbt.com/blog/elt-best-practices-snowflake" date: "2025-04-25" authors: ["Joey Gault"] categories: ["Pulse"] --- # ELT best practices for Snowflake workflows [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) flips the traditional data processing sequence by loading raw data into your data warehouse first, then performing transformations using the warehouse's computational power. This approach aligns perfectly with [Snowflake's architecture](https://www.snowflake.com/en/), which separates storage and compute to provide virtually unlimited scalability. The fundamental [difference between ETL and ELT](https://www.getdbt.com/blog/etl-vs-elt) lies in when and where transformations occur. With ETL, data undergoes transformation before reaching the warehouse, often creating bottlenecks and limiting flexibility. ELT loads raw data immediately, making it available for analysis while transformations happen within Snowflake's powerful compute environment. This shift enables data teams to work more iteratively, applying transformations as business requirements evolve rather than being locked into predetermined data structures. Snowflake's cloud-native architecture makes it particularly well-suited for ELT workflows. The platform's ability to scale compute resources on-demand means transformation workloads can leverage massive processing power when needed, then scale back down to control costs. This elasticity, combined with Snowflake's columnar storage and advanced query optimization, creates an ideal environment for the ELT approach. ## Architectural foundations for ELT success Implementing ELT effectively on Snowflake requires thoughtful architectural decisions that support both current needs and future growth. The foundation starts with a well-structured database design that separates raw data from transformed analytics-ready datasets. A proven approach involves establishing two primary databases: a raw database for untransformed source data and an analytics database for business-ready datasets. The raw database serves as the landing zone for all extracted data, maintaining source system fidelity while providing a stable foundation for transformations. This separation ensures that raw data remains unchanged and auditable while transformed data can evolve with business requirements. The analytics database becomes the domain of your transformation layer, where [dbt](https://www.getdbt.com/product/what-is-dbt) and other tools create the models, views, and tables that power business intelligence and reporting. This clear separation provides several advantages: it maintains data lineage visibility, supports regulatory compliance requirements, and enables different access patterns for different user types. Compute resource allocation represents another critical architectural decision. Snowflake's virtual warehouses should be sized and configured to match specific workload patterns. A typical ELT setup benefits from dedicated warehouses for different functions: a loading warehouse for data ingestion tools, a transforming warehouse for dbt runs, and a reporting warehouse for BI tools and analyst queries. This separation prevents resource contention and enables fine-tuned cost optimization. ## Optimizing Snowflake configuration for ELT Snowflake's unique architecture offers numerous configuration options that can significantly impact ELT performance and cost efficiency. Understanding these options and implementing them strategically creates the foundation for successful ELT operations. Warehouse sizing and auto-scaling policies require careful consideration based on workload characteristics. For transformation workloads, larger warehouses often provide better price-performance ratios for complex [dbt runs](https://docs.getdbt.com/reference/commands/run), as they complete faster and reduce overall compute time. However, reporting workloads may benefit from smaller, auto-scaling warehouses that can handle variable query patterns efficiently. Auto-suspend settings play a crucial role in cost management. Transformation warehouses used by dbt can typically use longer auto-suspend times (5-10 minutes) to avoid frequent startup costs during model runs, while reporting warehouses benefit from shorter suspend times (1-2 minutes) to minimize idle costs between user queries. Storage optimization becomes increasingly important as data volumes grow. Snowflake's automatic clustering can improve query performance on large tables, but it comes with additional costs. Implementing clustering keys strategically on frequently queried columns in your most important analytics tables can provide significant performance benefits. Similarly, leveraging [Snowflake's time travel](https://docs.snowflake.com/en/user-guide/data-time-travel) and fail-safe features appropriately balances data protection needs with storage costs. Security configuration deserves special attention in ELT environments where raw data flows through the system. Network policies should restrict access to approved IP addresses, and role-based access control should follow the principle of least privilege. A well-designed role hierarchy typically includes loader roles for ingestion tools, transformer roles for dbt and data engineers, and reporter roles for analysts and BI tools. ## Data modeling strategies for ELT workflows ELT workflows enable more flexible and iterative data modeling approaches compared to traditional ETL. This flexibility requires establishing clear patterns and conventions to maintain consistency and quality across your data models. The staging layer becomes particularly important in ELT workflows, serving as the bridge between raw source data and business logic. Staging models should focus on light transformations: casting data types correctly, standardizing naming conventions, and handling basic data quality issues. This approach creates a clean foundation for downstream models while maintaining clear lineage back to source systems. [Dimensional modeling principles](https://www.getdbt.com/blog/guide-to-dimensional-modeling) remain relevant in ELT environments, but the implementation becomes more flexible. Rather than requiring upfront decisions about grain and dimensionality, ELT allows teams to build multiple views of the same data for different use cases. Fact tables can be modeled at different grains, and dimension tables can be created with varying levels of detail based on specific analytical needs. Incremental modeling strategies become crucial for managing large datasets efficiently. [dbt's incremental models](https://docs.getdbt.com/docs/build/incremental-models-overview) allow processing only new or changed records, significantly reducing transformation time and compute costs. Implementing effective incremental strategies requires careful consideration of update patterns, merge logic, and data freshness requirements. ## Implementing robust data quality practices ELT workflows require strong data quality practices since raw data enters the warehouse before extensive validation. Building comprehensive testing and monitoring into your transformation pipeline ensures data reliability while maintaining the speed advantages of ELT. [dbt's built-in testing capabilities](https://docs.getdbt.com/docs/build/data-tests) provide an excellent foundation for data quality assurance. Generic tests for uniqueness, not-null constraints, referential integrity, and accepted values should be implemented systematically across your models. Custom tests can address business-specific logic and complex validation rules that generic tests cannot cover. Data freshness monitoring becomes critical in ELT environments where multiple systems depend on timely data availability. Implementing automated checks for data arrival times, record counts, and key metric variations helps identify issues before they impact downstream consumers. These checks should trigger alerts that enable rapid response to data quality issues. Schema evolution handling requires particular attention in ELT workflows. Since raw data structures can change as source systems evolve, your transformation layer must be resilient to these changes. Implementing schema tests and using dbt's source freshness checks helps identify schema changes early, while flexible model designs can accommodate minor variations without breaking. ## Performance optimization techniques Snowflake's performance characteristics differ significantly from traditional databases, requiring specific optimization approaches for ELT workloads. Understanding these differences and implementing appropriate techniques can dramatically improve transformation performance and reduce costs. Query optimization in Snowflake benefits from understanding the platform's columnar storage and automatic query optimization features. Writing SQL that takes advantage of Snowflake's strengths (such as leveraging column pruning, predicate pushdown, and join optimization) can significantly improve performance. Avoiding unnecessary data movement and minimizing result set sizes through effective filtering and aggregation strategies reduces both execution time and costs. [Materialization strategies in dbt](https://docs.getdbt.com/guides/create-new-materializations?step=1) should align with usage patterns and performance requirements. Tables provide the fastest query performance but consume storage and require refresh cycles. Views offer storage efficiency and always-current data but may have slower query performance for complex logic. Incremental models balance performance and efficiency for large, frequently updated datasets. Concurrency management becomes important as ELT workflows scale. Snowflake's multi-cluster warehouses can automatically scale to handle concurrent workloads, but this scaling comes with costs. Designing transformation schedules to minimize peak concurrency while meeting business requirements helps optimize resource utilization. ## Cost management and monitoring ELT workflows on Snowflake can provide excellent cost efficiency when managed properly, but they require active monitoring and optimization to prevent unexpected expenses. Implementing comprehensive cost management practices ensures that the flexibility benefits of ELT don't come at the expense of budget control. Compute cost optimization starts with right-sizing warehouses for specific workloads. Transformation workloads often benefit from larger warehouses that complete faster, while interactive workloads may be more cost-effective on smaller warehouses. Regular analysis of warehouse utilization patterns helps identify optimization opportunities. Storage cost management requires understanding Snowflake's storage pricing model and implementing appropriate data lifecycle policies. Time travel and fail-safe features provide valuable data protection but consume storage. Setting appropriate retention periods based on business requirements and regulatory needs balances protection with costs. Query monitoring and optimization should be ongoing practices. Snowflake's query history and performance monitoring tools help identify expensive queries and optimization opportunities. Implementing query tags and monitoring dashboards provides visibility into cost drivers and usage patterns across different teams and use cases. ## Governance and collaboration frameworks ELT workflows often involve more team members in data transformation activities, requiring robust governance frameworks to maintain quality and consistency. Establishing clear processes and standards enables productive collaboration while preventing chaos. Development workflow standards should leverage software engineering best practices adapted for analytics work. Git-based version control, branch-based development, and code review processes help maintain code quality and enable collaboration. Establishing clear environments for development, testing, and production ensures changes are properly validated before affecting business users. Documentation and metadata management become increasingly important as ELT workflows democratize data transformation. dbt's automatic documentation generation provides a strong foundation, but teams should establish standards for model descriptions, column documentation, and business logic explanation. This documentation serves as crucial institutional knowledge and enables self-service analytics. Data lineage and impact analysis capabilities help teams understand dependencies and assess change impacts. dbt's lineage graphs provide technical dependency information, while business glossaries and metric definitions help bridge the gap between technical implementation and business understanding. ## Future-proofing your ELT implementation As data volumes and complexity continue to grow, ELT implementations must be designed for scalability and evolution. Building flexibility into your architecture and processes ensures your data platform can adapt to changing requirements and new technologies. Modular design principles help create maintainable and scalable ELT workflows. Breaking complex transformations into smaller, focused models improves maintainability and enables parallel processing. Establishing clear interfaces between different layers of your data models creates flexibility for future changes. Technology evolution considerations should influence architectural decisions. While Snowflake provides excellent capabilities today, maintaining some level of abstraction through tools like dbt helps protect against future platform changes. Similarly, designing transformation logic that can adapt to new source systems and data types provides resilience against business evolution. Monitoring and alerting infrastructure should be designed to scale with your data operations. As ELT workflows become more complex and critical to business operations, comprehensive monitoring becomes essential for maintaining reliability and performance. The ELT approach on Snowflake represents a fundamental shift in how organizations can approach data transformation. By leveraging Snowflake's unique architecture and implementing these best practices, data engineering leaders can build more flexible, scalable, and cost-effective data platforms. The key lies in understanding both the technical capabilities and the organizational changes required to fully realize ELT's potential. Success requires not just technical implementation but also cultural adaptation to more collaborative, iterative approaches to data work. ## ELT for Snowflake FAQs **Why is Snowflake well-suited for ELT compared to traditional ETL, and how does its architecture affect the workflow?** Snowflake's cloud-native architecture makes it particularly well-suited for ELT workflows because it separates storage and compute to provide virtually unlimited scalability. Unlike traditional ETL where transformations create bottlenecks before data reaches the warehouse, ELT loads raw data immediately into Snowflake and performs transformations using the warehouse's computational power. Snowflake's ability to scale compute resources on-demand means transformation workloads can leverage massive processing power when needed, then scale back down to control costs. This elasticity, combined with columnar storage and advanced query optimization, creates an ideal environment where data teams can work more iteratively and apply transformations as business requirements evolve. **How do you configure dbt and Airflow (via Astronomer Cosmos) to run ELT transformations in Snowflake, including setting up the Snowflake connection?** To configure dbt and Airflow for ELT transformations in Snowflake, you need to establish dedicated virtual warehouses for different functions: a loading warehouse for data ingestion, a transforming warehouse for dbt runs, and a reporting warehouse for BI tools. The setup requires proper role-based access control with loader roles for ingestion tools, transformer roles for dbt and data engineers, and reporter roles for analysts. Warehouse sizing should match workload patterns, with larger warehouses often providing better price-performance ratios for complex dbt runs. Auto-suspend settings should be configured appropriately: transformation warehouses can use longer auto-suspend times (5-10 minutes) to avoid startup costs during model runs, while reporting warehouses benefit from shorter suspend times (1-2 minutes) to minimize idle costs. **What Snowflake objects and permissions (warehouse, database, role, schema) are required to enable an ELT pipeline?** A well-structured ELT pipeline on Snowflake requires two primary databases: a raw database for untransformed source data and an analytics database for business-ready datasets. You need dedicated virtual warehouses for different functions: loading, transforming, and reporting warehouses to prevent resource contention. The role hierarchy should include loader roles for ingestion tools, transformer roles for dbt and data engineers, and reporter roles for analysts and BI tools, following the principle of least privilege. Security configuration should include network policies restricting access to approved IP addresses and proper role-based access control. This separation maintains data lineage visibility, supports compliance requirements, and enables different access patterns for different user types. --- --- title: "AI data pipelines: Critical components and best practices" description: "Explore how to design AI-ready pipelines — and how dbt helps deliver high-quality data for scalable, intelligent AI systems." url: "https://www.getdbt.com/blog/ai-data-pipelines" date: "2025-04-23" authors: ["Daniel Poppy"] categories: ["Learn"] --- # AI data pipelines: Critical components and best practices As businesses integrate AI across more functions, the need for robust, scalable data pipelines has never been greater. Traditional pipelines often rely on batch processing and static dashboards. But AI demands something different: real-time ingestion, dynamic processing, and automated model retraining. Without the right infrastructure, AI systems risk learning from outdated, biased, or low-quality data — leading to poor predictions and costly mistakes. In this article, we’ll break down the core components and best practices for building AI-ready data pipelines — and explore how dbt helps teams deliver the high-quality, trustworthy data that modern AI demands. ## Why are AI data pipelines important? As companies seek to innovate by incorporating AI into their operations, AI-driven data pipelines ensure they can harness the full potential of their data. An AI data pipeline is a structured workflow that moves raw data through ingestion, transformation, and processing to feed machine learning models. It ensures that AI systems receive clean, optimized, and continuously updated data to generate accurate insights. Today’s GenAI applications and agents go beyond “traditional” AI by generating new content, automating creative processes, and enabling autonomous agentic activity. AI pipelines support a wide range of transformative AI business use cases, including an increasing number of [agentic AI applications](https://hbr.org/2024/12/what-is-agentic-ai-and-how-will-it-change-work): - **Customer engagement **- AI powers chatbots and virtual assistants that provide human-like interactions, handle inquiries and provide personalized responses, improving customer service while reducing human workloads. - **HR & recruiting** - AI agents automate resume screening, generate personalized interview questions, and assist in onboarding, streamlining HR processes. - **Fraud detection and risk management** – Businesses use AI to identify suspicious financial transactions, detect cybersecurity threats, and automate risk assessments. - **Code generation and software development **- AI accelerates development by automating code writing, debugging, and documentation, helping engineers build applications faster. - **Ad personalization and marketing automation** - Businesses use AI to create dynamic ad copy, optimize and build campaign strategies, and personalize ad content based on audience preferences, improving customer engagement. - **Data analysis and AI-driven insights** – AI quickly summarizes reports, detects trends, and generates insights from complex datasets, making analytics more accessible and actionable. - **Autonomous systems and robotics** – Self-driving vehicles, robotic process automation (RPA), and industrial automation all rely on AI to continuously monitor, update, and manage their operations, usually without human intervention. In general, AI applications cannot run on traditional data pipelines, which rely on manual input, predictable workloads, and are more suited for structured reporting. To effectively launch AI-driven applications, you need the advantages that an AI data pipeline provides: - **Real-time decision making** - AI applications need continuous, trustworthy data that traditional pipelines can't provide. With a robust AI pipeline, businesses can respond to market changes, customer behaviors, and operational challenges in real time. - **Scalability and automation** - As data volumes grow, manual data handling becomes more onerous and error-prone. AI pipelines require automated data ingestion, transformation, and delivery, ensuring seamless scaling without human intervention. - **Faster model training and iteration **- With automated data flows, new data is continuously fed into AI models, allowing them to retrain and improve over time, which is essential for adapting to shifting trends. - **Data quality and consistency** - AI models are only as good as the data they learn from. A robust AI pipeline ensures data is clean, structured, and free from inconsistencies, preventing biased or inaccurate predictions. - **Competitive advantage **- Businesses that integrate AI pipelines can identify patterns, predict outcomes, and optimize processes faster than competitors using manual data processing, leading to smarter, data-driven decisions. - **Cost efficiency** - A well-designed AI pipeline minimizes manual work, which improves processing efficiency and reduces costs by ensuring that AI systems run optimally with the most relevant, high-quality data. To bridge the gap between traditional pipelines and AI-driven applications, businesses must adopt AI data pipelines that enable automation, scalability, and continuous data flow. These pipelines provide the foundation for real-time insights and timely decision-making, which are essential for staying competitive in an increasingly data-driven landscape. ## Critical components of AI data pipelines Traditional pipelines were built for static dashboards and structured reporting. But AI applications require something more: pipelines that are dynamic, scalable, and automated by design. Because AI models depend on the most timely and reliable data, the pipelines supporting them must incorporate the following critical components: ### Data ingestion AI pipelines ingest data from diverse sources — databases, APIs, data lakes, event streams, and unstructured formats. Unlike batch-oriented pipelines, AI systems require real-time or near-real-time data flow to stay accurate. Efficient ingestion ensures models are trained on the most current and relevant inputs. ### Data transformation Raw data isn’t AI-ready. It needs to be cleaned, structured, and tested before it can power models. dbt is the industry standard for [data transformation](https://www.getdbt.com/blog/dbt-explained) in the modern enterprise. By using a [modular approach](https://www.getdbt.com/blog/modular-data-modeling-techniques), dbt defines reproducible data transformations as SQL models instead of lengthy queries or scripts. This ensures consistency across AI workflows, allowing pipelines to continuously deliver high-quality, structured datasets. ### Feature engineering and updates Unlike standard data pipelines that focus on predefined metrics, AI pipelines rely on uncovering meaningful patterns within data for predictive accuracy. [Feature engineering](https://www.ibm.com/think/topics/feature-engineering) transforms raw data into the features used by AI models to make predictions by extracting the most relevant and valuable aspects of a dataset. With dbt, teams can build [reusable, version-controlled feature sets using SQL](https://www.hopsworks.ai/post/feature-engineering-with-dbt-for-data-warehouses), and keep them updated [incrementally](https://docs.getdbt.com/docs/build/incremental-models) as new data arrives. This reduces engineering overhead and speeds up iteration. ### Model training, fine-tuning, and RAG Once transformed, data is fed into machine learning models for training. Fine-tuning improves model performance with fresh domain data. [Retrieval-Augmented Generation (RAG)](https://mindsdb.com/blog/whats-the-difference-between-fine-tuning-retraining-and-rag) enhances [Large Language Models (LLMs)](https://aws.amazon.com/what-is/large-language-model/) output by injecting real-time context from internal datasets. dbt supports training and model optimization by automating data transformation and providing incremental data updates, ensuring models can be updated efficiently without full reprocessing of the data. ### Monitoring and feedback In addition to continuous fine-tuning, AI model performance must be monitored in real time to prevent degradation. AI pipelines need data drift detection and feedback loops to detect shifts in data quality. By providing [column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage), dbt can improve the auditability and reliability of AI pipelines. ### Security and governance AI systems must operate on secure, governed data. dbt embeds governance into the transformation layer with version control, role-based access, and automated testing. It integrates with platforms like [Alation](https://www.alation.com/partners/dbt-labs/) and [Atlan](https://atlan.com/partners/dbt/) to expose lineage, ownership, and compliance metadata — helping you meet AI audit and regulatory requirements. Together, these components ensure your AI pipelines are fast, flexible, and production-grade — ready to support everything from GenAI applications to real-time autonomous agents. ## Best practices for building AI data pipelines Building AI pipelines isn’t just about infrastructure. It’s about building the right habits — scalable, efficient, and trustworthy ones. High-quality, structured data is essential for AI models to generate accurate insights and adapt to changing conditions. dbt is a powerful data transformation tool that automates data workflows, enforces quality checks, and ensures the reproducibility of data transformations. With a modular approach and built-in testing, dbt is ideal for building resilient AI data pipelines — from ingestion to model monitoring. You can easily incorporate dbt into the following best practices for AI pipeline development. ### Eliminate "Garbage in, Garbage out" AI models are only as good as the data they receive. Ensuring high-quality, well-structured data is essential for accurate predictions and insights. - **Implement data quality tools** – Automated validation tools detect anomalies, missing values, and inconsistencies before data reaches AI models. dbt’s built-in [testing and observability](https://www.getdbt.com/product/test-and-observe) ensure data integrity throughout the pipeline. - **Data transformation** – Raw data must be cleaned, structured, and optimized for AI workflows. dbt automates transformations, ensuring datasets are consistent, reproducible, and analytics-ready. - **Document data sources and transformations** – Transparency is key. dbt’s data lineage helps teams trace transformations, ensuring accountability and auditability. ### Automate wherever possible Manual processes slow down AI pipelines and introduce errors. Automation ensures efficiency, scalability, and reliability. dbt enables [pipeline automation](https://www.getdbt.com/product/deploy) from development to production. - **Data transformation** – dbt automates SQL-based transformations, reducing manual intervention and ensuring consistent, repeatable workflows. The [dbt Fusion](https://www.getdbt.com/product/fusion) engine dramatically accelerates AI pipeline deployment—and feature engineering—with 30X faster SQL parsing speeds. - **Feature engineering **- dbt enables teams to define reusable SQL models, making it easier to create and refine features without relying on lengthy queries. It works alongside [Snowflake’s feature store](https://docs.getdbt.com/blog/snowflake-feature-store), ensuring consistent, governed datasets for machine learning workflows. Finally, incremental transformations ensure that models receive fresh, optimized features. - **Data quality checks and validation **– dbt’s [testing ](https://www.getdbt.com/product/test-and-observe)frameworks catch errors before they impact AI models, ensuring high-quality datasets. - **Monitoring and maintenance **– AI pipelines require continuous monitoring to detect data drift and maintain accuracy. Built-in [observability ](https://www.getdbt.com/product/test-and-observe)in dbt helps you maintain data quality and reliability at scale. ### Use cloud services for scalability Cloud-native platforms offer elastic scale, cost efficiency, and reliability — making them ideal foundations for AI pipelines. - **Scale with dbt** – dbt offers cloud-native scalability with cloud services like [Snowflake](https://www.getdbt.com/blog/data-pipelines-snowflake-dbt), [BigQuery](https://www.getdbt.com/data-platforms/bigquery), and [Databricks ](https://www.getdbt.com/data-platforms/databricks)to enable dynamic scaling, ensuring AI models receive optimized, high-performance data. - **Partner integrations** – dbt collaborates with [consulting partners](https://www.getdbt.com/partner-directory) to help businesses implement scalable, automated data pipelines. - **Use dbt platform **– The cloud-hosted dbt platform manages deployments, schedules jobs, and [integrates with CI/CD](https://docs.getdbt.com/docs/deploy/continuous-integration) tools — no manual maintenance required. ### Leverage AI AI-powered applications excel at enhancing efficiency by automating routine or previously-time consuming manual tasks. Teams can benefit from AI agents to monitor model performance, adjust hyper-parameters, and retrain models as new data arrives, ensuring continuous improvement. In the pipeline creation process, AI tools in dbt streamline workflows by automating repetitive tasks, accelerating and optimizing workflows. - **dbt Copilot** – An [AI assistant](https://docs.getdbt.com/docs/cloud/dbt-copilot) that accelerates the analytics workflow by automating code generation, enabling users to generate SQL queries, documentation, tests, metrics, and semantic models using natural language prompts. - **dbt Canvas** – An AI-powered [visual editing tool](https://www.getdbt.com/blog/dbt-canvas-is-ga) that empowers teams to accelerate dbt model development in a governed environment, ensuring consistency and improving collaboration. - **dbt Insights** – An AI-powered [query tool ](https://www.getdbt.com/product/analyst)that enables analysts to explore, validate, and analyze data efficiently. It integrates with dbt’s metadata and governance framework, ensuring queries align with structured data models. ### Don't forget security AI pipelines must be secure, compliant, and auditable to protect sensitive data. - **Access controls** – dbt supports [role-based access permissions](https://www.getdbt.com/security), ensuring only authorized users can modify transformations. - **Audit and compliance checks** – dbt’s lineage tracking, and planned built-in data governance with support for [tagging of PII/PHI and enforcement](https://www.getdbt.com/blog/get-to-know-the-new-dbt-fusion-engine-and-vs-code-extension) of data policies help businesses maintain transparency and regulatory compliance. - **Automated security checks** – Security should be built into AI pipelines. dbt integrates with [cloud security frameworks](https://www.getdbt.com/security) to ensure compliance without manual intervention. By following these best practices, businesses can build resilient, scalable, and secure AI data pipelines, ensuring models operate efficiently and reliably. Using dbt’s latest advancements — dbt Fusion and dbt MCP Server — will further turbocharge your AI data pipeline creation. Fusion’s lightning-fast parse times and [state-aware orchestration](https://www.getdbt.com/product/fusion) enable organizations to further streamline data transformations. The[ MCP Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) uses [Model Context Protoco](https://modelcontextprotocol.io/introduction)l to enable seamless, universal connectivity between AI-powered applications and the governed, structured data in dbt. MCP Server connects AI applications to dbt data, allowing AI agents to autonomously interpret data models, relationships, and structures without human intervention— leading to faster insights and optimized decision-making. These latest innovations enhance dbt’s already vital role in AI data workflow creation, ensuring businesses can build the scalable, efficient, and secure AI data pipelines required for today’s advanced AI applications. To learn more, [ask us for a demo](https://www.getdbt.com/contact). --- --- title: "Tableau Conference: Not even an earthquake could stop the fun" description: "Tableau was a blast and we're glad you came. See you next year!" url: "https://www.getdbt.com/blog/tableau-conference-recap-2025" date: "2025-04-18" authors: ["Jeff Mills"] categories: ["Partnerships"] --- # Tableau Conference: Not even an earthquake could stop the fun This week team dbt joined over 9000 attendees at Tableau Conference in sunny San Diego to share all the innovation we’ve been working on. We will soon be supporting all three deployment options of Tableau (Desktop, Sever and Cloud) with our semantic layer. The week started with an earthquake, [literally](https://www.sandiegouniontribune.com/2025/04/14/magnitude-5-2-earthquake-near-julian-jolts-san-diego-county-and-beyond/), and ended with a continued commitment to bring dbt, and our trust and rigor, to all Tableau Data Rockstars out there. dbt was featured in sessions from The Information Lab, a feature during Devs on Stage, and a packed booth full of great questions and fun conversations. We can’t wait to see what both communities (Tableau’s and dbt’s) build in the coming year. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/51ef7bf9a2ab7807f7992ea6aa32b72d70dff358-4284x5712.jpg) ## Thank you to the Information Lab [Harriet Owen](https://www.linkedin.com/in/harriet-owen-4862b5145/) and [Jack Parry ](https://www.linkedin.com/in/jaackparry/)from The Information Lab, a premier Tableau and dbt partner, introduced and demoed how to use dbt models within Tableau to bring trust and rigor to the analytic process. It started with a pretty hilarious definition of what dbt is (that dbt and DBT are in fact not the same thing) and the went into a live demo of leveraging the dbt Semantic Layer inside of Tableau desktop. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/194fe5d295a867469e170b61c0dbd2edd7b87ef8-1512x2016.jpg) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/53c6b531917d43e388890237d8d4e8652b6a1b66-5712x4284.jpg) ## Devs on Stage featured dbt - Thank you Patrick! Devs on Stage is arguably the most popular session a Tableau Conferences. It dates back to the very first event in 2008. This is when the actual product engineers get on stage and share what they’ve been building. [Patrick Green](https://www.linkedin.com/in/patrickgreen2014/) shared how Tableau Admins can use dbt to ‌better understand what’s going on in the Tableau Server and better serve their analytics deployments. Nothing better than seeing dbt and Tableau on the big stage getting loud applause. You can [watch the whole session here](https://www.salesforce.com/plus/experience/tableau_conference_2025/series/tableau_conference_2025_highlights/episode/episode-s1e5?t=1444)—it’s a must-see. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/9ded75cb41b3315ecf72b5f262f5f77a1edd0372-4032x3024.jpg) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/39c164816a4b810c601f48f5c54f5483fe1b6a97-3024x4032.jpg) ## It wasn’t all work—we had some fun too Many thanks to Interworks and The Information Lab for hosting social gatherings. We had fun meeting with customers and partners at the Padres baseball game on Tuesday night and the Wild Hare on Wednesday night. It was so nice to connect in such beautiful settings. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/3ad1da5a9e62f14192cdf8538d34fcd7a54061d2-3024x4032.jpg) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/439d7b317834c617986232949239076d8844fa72-4032x3024.jpg) ## The momentum continues If you want to learn more about the dbt-Tableau integration—please take a look at our [docs](https://docs.getdbt.com/docs/cloud-integrations/semantic-layer/tableau). Or you can request a demo [here](https://www.getdbt.com/contact). And please be sure to catch our [dbt Launch Showcase](https://www.getdbt.com/) at the end of May. We can’t wait to see what you all build with dbt and Tableau. --- --- title: "M1 Finance powers AI self-service with Claude and dbt structured data" description: "See how M1 Finance uses dbt’s structured data and Semantic Layer to fuel a self‑service chatbot delivering accurate AI insights." url: "https://www.getdbt.com/blog/m1-finance-ai-self-service-claude-dbt" date: "2025-04-17" authors: ["Hrishi Kulkarni", "Chakshu Mehta"] categories: ["Product"] --- # M1 Finance powers AI self-service with Claude and dbt structured data You can imagine that as a fintech company, M1 deals with a lot of complex, highly regulated data. For its analysts and business users, fast access to accurate data is critical, both for strategic decision making and strict regulatory compliance. Relying on the wrong figures could directly affect customer portfolios and misinform executive planning. To meet this need, the data team rolled out a self‑service AI layer powered by Anthropic’s Claude LLM, hoping to let anyone pose questions in natural language But going from a simple business question to reliable insight wasn’t so simple. ### **Bottlenecks to accessing data** To provide self-service analytics for its business users, M1 uses Superset. It’s a powerful tool for querying the data warehouse directly—but it required a level of SQL proficiency that most business stakeholders weren’t comfortable with. As a result, users relied heavily on the data team to help them write SQL queries and find answers. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/5c71a4bb868d6e57cd63626dd3bf7e3f563da6dd-512x285.png) “The data team was becoming a bottleneck to the business,” says Brady Dauzat, Machine Learning Engineer at M1. “Instead of enabling people to make data-driven decisions, we were actually slowing them down because they needed to wait for us to do their jobs.” To truly unlock self-serve for business users, the data team decided to build an LLM-powered conversational interface (ie chatbot) to enable business users to directly access and analyze data. Let’s walk through how they built a hallucination-proof SQL AI using [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer). ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/20bfa6e9959ad5e1769b05130e43524418c72e1f-512x368.png) ### The problem: frequent LLM hallucinations For their first iteration, the data team built an LLM chatbot using a “direct catalog” approach. When the user asked a natural-language query, the LLM (Claude by Anthropic) generated a SQL query for them to run. To create the SQL query, the LLM accessed the entire M1 data catalog, like schema, column names, and descriptions. Unfortunately, the LLM frequently responded with hallucinations. While the SQL queries could run in the data warehouse and return data, they often retrieved the wrong data. This is a well-known challenge with LLMs when given overly broad contexts. For example, the LLM might invent a column name that sounded correct but in fact didn’t exist. Unless the user deeply understood SQL and could identify errors in the query, they would never know there was an issue in the first place. “To understand why this was challenging for the LLM, think of it this way. Say you provided your full data catalog to a smart person who had never seen it before,” explains Dauzat. “If you asked them to produce a valid query for key business metrics, would you trust them to do it? “We were essentially asking the LLM to sort through all kinds of complex contexts around our entire data warehouse just to answer a question,” he continues. “That's very difficult to do and still produce a correct answer.” Essentially, providing the LLM with the entire data catalog without constraints meant it had to "understand" a vast amount of context. This method was prone to errors, leading to incorrect queries or invented column names. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f806403161e7f9376b7e781f4df3f7cc35e2edc4-454x512.png) ### The solution: a “Mad Libs” approach to ensure high-quality AI outputs using dbt Semantic Layer Clearly, this was a big problem. To overcome hallucinations, the data team tried a different approach—making the LLM serve as an interface between the user and dbt Semantic Layer. “Going back to our earlier example, let’s say you gave a smart person a list of key business metrics in a fill-in-the-blank template for how to answer your question,” says Kelly Wolinetz, Senior Data Engineer at M1. “My trust in them generating the correct answer goes way up.” In other words, the LLM no longer has to understand the full data warehouse. It just needs to understand how to fill out a form. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/983aa1e40b53b67bcde1e503789a8a985467f6b9-512x300.png) With this approach, the LLM no longer generates the SQL statements on its own. Instead, the data team asked the LMM to write a dbt Semantic Layer command, where key business metrics are predefined. To do that, the data team provided it with a system prompt and the user question—basically a [“Mad Libs” approach](https://benn.substack.com/p/llms-shouldnt-write-sql) to finding and providing the answers. **** Now when the user asks a natural-language question, the dbt Semantic Layer command runs in the background to produce the actual SQL, and the LLM returns the SQL query to the user, ensuring it conforms to vetted metrics and standard query templates. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/286c05ef3b204d98709b4cac4844d060a8b7262d-512x78.png) ### Impact: increased user engagement and reduced hallucinations For the data team, dbt Semantic Layer has been pivotal for ensuring data accuracy, quality, and consistency. Hallucinations are almost completely eliminated, with SQL syntax errors and misinterpretations virtually nonexistent. The SQL output always conforms to how the data team wants it queried, so that business users only ever submit queries for predefined, vetted metrics. It provides a robust check against the LLM’s tendency to invent or misinterpret data points. As a result, the process of pulling data is dramatically simpler for business users. They’re engaging with the tool, and the data team has received positive feedback about their experience so far. Anecdotally, ad-hoc Slack requests to the data team have decreased. Meanwhile, the team is receiving new requests for metrics—showing that users are engaging with the tool and learning which questions they need answered. “When we’ve held information sessions for our engineering and non-engineering teams, the feedback has been really encouraging,” says Wolinetz. “We love it when people use our tools, and it’s our non-technical stakeholders who are using it the most.” ### Advancing AI with dbt-based architecture As for what’s next, the data team is exploring the following: - Adding more metrics to the tool. - Integrating the LLM directly into Superset, so that users can interact with the chatbot right where they work. - Experimenting with different error-handling approaches. - Restructuring the LLM to answer general questions about the data warehouse. (e.g., “What dimensions are available for metric X?”) - Leveraging dbt Model Context Protocol server to directly integrate their dbt Semantic Layer and structured data to their LLM. “For our business needs, our investment in AI and dbt Semantic Layer has been worth it,” concludes Wolinetz. Watch M1 Finance's Coalesce session here: [Watch video](https://www.youtube.com/watch?v=ovs3-TdxqtI) If you’re thinking about how to add AI to your data workflows, we’d love to chat. Reach out to [book a demo](https://www.getdbt.com/contact), or [sign-up for dbt Cloud](https://www.getdbt.com/signup) to connect your data warehouse and start building. --- --- title: "AI is Driving a Surge in Data Budgets, According to New Report from dbt Labs" description: "2025 State of Analytics Engineering Report reveals how the demand for quality data to feed AI is fueling data team growth" url: "https://www.getdbt.com/blog/ai-driving-surge-in-data-budgets-2025-state-of-analytics-engineering" date: "2025-04-16" authors: ["Elaine Green"] categories: ["Press"] --- # AI is Driving a Surge in Data Budgets, According to New Report from dbt Labs **PHILADELPHIA**, Apr. 16, 2025 -- [dbt Labs](https://www.getdbt.com/), the pioneer in analytics engineering, today published its third annual [State of Analytics Engineering Report](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2025) sharing new insights into the evolution of the data industry in the artificial intelligence (AI) era. The report reveals that AI is the catalyst for significant investment in data teams as enterprises require higher-quality data to power their AI applications. **Data Budgets Spike** According to the report, data budgets are growing significantly this year with 30% of participants reporting budget growth compared to just 9% last year. Additionally, AI tooling was recognized as the largest area of investment for the year ahead with 45% of respondents citing it as a key priority. Data team sizes are increasing too (40% reported growth, compared to 14% last year), clearly demonstrating how AI is driving demand for larger teams that can better ensure high quality, governed data. Additionally, as data budgets increase, so have salaries in North America, with 80% of individual contributors making over $100,000 (compared to 69% last year) and 49% of managers making over $200,000 (compared to 32% last year). **Data Teams Increase AI Usage** AI isn’t just fueling investment, it’s also disrupting the way data professionals operate. The report found that 80% of respondents are using AI in their daily workflow, compared to just 30% last year. Of those using AI daily, 70% said they use AI for code development and 50% use it for documentation. Organizations are clearly increasing investments in tools to accelerate – not replace – their data teams, and leveraging AI in the data workflow is improving developer productivity while bolstering data quality in the process. As a result, company perception of data teams is positive, with respondents overwhelmingly agreeing (75%) that their organization values the data team. “AI is disrupting the way that teams work with organizational data,” said Mark Porter, CTO of dbt Labs. “As companies increase AI investments, leaders are prioritizing the teams responsible for data quality and governance—the essential foundation for AI effectiveness. At the same time, data engineers are turning to AI to automate routine tasks, completely changing how data is delivered to the business. Because of this, the strategic role of the data team continues to grow, with AI as the catalyst. It’s a symbiotic relationship – data professionals make AI better, and AI makes data teams better.” **Looking to the Future** Effective AI requires high quality inputs, and poor data quality continues to be the challenge most frequently reported (56% of respondents). That’s why building trust in data is cited as the top priority for growing data teams, accentuating the importance of data governance and observability – and data professionals are hopeful that AI can help. Data teams are optimistic about the potential impact of AI in the analytics workflow, citing its ability to help bridge the data quality gap with features like proactive data monitoring and pipeline debugging. “This past year has shown us that investing in data is critical for AI success,” said Piyush Bhargava, Sr. Director Global Data & Analytics, Docusign. “Establishing trust in Artificial Intelligence starts with trust in enterprise data, which is why we’ve invested in a modern technology stack with dbt as a key pillar. By leveraging dbt to build reusable data assets, we’ve built a scalable data foundation and are now looking to boost productivity through new AI tools in dbt Cloud. This report confirms what we’ve seen firsthand: good data is the bedrock of strong AI, and with dbt as the cornerstone of our strategy, we’re well-positioned to continue driving innovation.” On April 30, dbt Labs will host a panel of industry experts for the [2025 State of Analytics Engineering Virtual Event](https://www.getdbt.com/resources/webinars/2025-state-of-analytics-engineering-virtual-event). The conversation will focus on strategies for building effective data organizations in a rapidly shifting landscape and address how data teams are integrating generative AI, adapting to economic shifts, and addressing persistent industry challenges. To view the 2025 State of Analytics Engineering report, visit [https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2025](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2025). **Methodology** dbt Labs surveyed 459 data practitioners and leaders from October 8, 2024 through December 27, 2024. 70% of survey respondents are individual contributors (ICs) and 30% are managers. Analytics engineers made up 48% of IC respondents, 36% of IC respondents are data engineers, and 16% of IC respondents are data analysts. **About dbt Labs** Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 50,000 teams using dbt every week. To learn more about dbt Labs, visit [getdbt.com](http://getdbt.com/) and follow us on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "Do you need a data orchestration platform?" description: "Understand what orchestration platforms do and how to tell when your data pipelines are ready for one." url: "https://www.getdbt.com/blog/data-orchestration-platform-necessary" date: "2025-04-14" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Do you need a data orchestration platform? Most companies don’t have a clear, centralized view of what’s happening in their data pipelines. That’s because their data pipelines are scattered across multiple technologies and systems, making them difficult to manage, troubleshoot, and scale. A data orchestration platform provides a single location for building, deploying, and monitoring data pipelines, no matter where the data itself lives. In this article, we’ll break down what a data orchestration platform does, how it works, and how to know when it’s time to invest in one. ## What is a data orchestration platform? [Data orchestration is the process of moving data from multiple source](https://www.getdbt.com/blog/data-orchestration-vs-etl)s into a single source. It involves collecting, [transforming](https://www.getdbt.com/blog/data-transformation), and storing data so that it can be used for a specific purpose, such as business analytics or AI workflows. A data orchestration platform is the tool that makes this coordination possible at scale. It enables teams to: - Pull data from multiple sources (APIs, SaaS apps, event streams, databases, etc.) - [Run data transformation workflows](https://www.getdbt.com/blog/data-transformation-best-practices) automatically at scheduled intervals or in response to events (e.g., a file uploads new data files into an [Amazon S3](https://aws.amazon.com/s3/) bucket) - Detect failures and alert engineers when something breaks - Monitor pipeline health and status across teams - Scale pipelines as data volume grows In early-stage or low-volume environments, manual scripting might get the job done. You can write SQL or Python scripts to trigger transformations and refresh data as needed. But as use cases grow and data dependencies stack up, this approach becomes brittle, time-consuming, and tough to scale. This is especially true if your data pipeline processing requires connecting multiple components of your data architecture together. If a component fails and the failure goes undetected, your stakeholders won’t have the timely data they need for driving critical business decisions. A data orchestration platform brings reliability, [scalability](https://www.getdbt.com/blog/common-challenges-to-scale-data-operations), and observability to your data stack — especially once your pipeline ecosystem becomes too large or business-critical to manage manually. ## How data orchestration works A data orchestration platform brings structure to the chaos of modern data systems. It coordinates every step of the pipeline — so data flows from source to insight with reliability and scale. Here’s what that typically includes: **Ingestion. **Pulls data from a variety of sources, including: - **Relational databases** (e.g., PostgreSQL, BigQuery, Snowflake) - **Nonrelational stores** (like MongoDB) - **Structured files** (such as CSV or JSON) **Workflow management. ** Orchestration platforms allow teams to define and run multi-step workflows. These workflows can include data extraction, transformation, quality checks, and activation. Most platforms support Python and offer SDKs or decorators that simplify building and managing DAGs (directed acyclic graphs). Tools like [Apache Airflow](https://airflow.apache.org/) and [Prefect](https://www.prefect.io/) are common examples. **Activation.** Delivers trusted data where it’s needed. Whether that’s dashboards for the sales team, feature sets for machine learning, or structured inputs for AI — workflows tailor data to each downstream need. **Observability.** [Detects and issues alerts on errors.](https://www.getdbt.com/blog/data-observability) If a pipeline fails or a task takes longer than expected, teams are notified instantly. Logs and metadata help engineers quickly trace the issue and resolve it before it impacts stakeholders. ### Benefits of a data orchestration platform Manual scripts and cron jobs can get you part of the way. But they fall short when it comes to scale, reliability, and cross-team collaboration. A data orchestration platform unlocks more than automation—it gives you structure and confidence in your data operations. ### Ensure data freshness Data orchestrations ensure that data pipelines are run according to the needs of the business. This can be done on a schedule (e.g., every hour) or in response to a signal that new data is available. The result? Stakeholders always get up-to-date data, without having to ask for it. That trust builds confidence—and enables faster, more reliable decision-making. ### Break down data silos Siloed data lives in spreadsheets, team-specific databases, or buried somewhere in marketing’s SaaS tools. When teams can’t access each other’s data, you get: - Duplication - Inconsistencies - Governance gaps - Missed opportunities Orchestration helps unify your data. It brings sources together into one platform, standardizes the format, and makes it easier for teams to discover and collaborate on trusted data assets. ### Gain visibility into your pipelines Without orchestration, pipelines often run on different servers, in different languages, managed by different teams. There’s no easy way to answer: “What ran? Did it succeed? How much did it cost?” A data orchestration platform gives you a centralized view of your data workflows—what’s running, where, how often, and how reliably. Many tools also let you track cost, duration, and performance over time. ### Improve reliability and recover faster Data breaks. But when you’re flying blind, you don’t know what’s broken — or where to look. Orchestration platforms let you detect failures fast and trace them to specific tasks in your workflow. Instead of guessing where the issue is, you get alerts and logs that point directly to the problem. That means fewer disruptions, faster fixes, and fewer messages asking “why is my dashboard blank?” ## Determining whether you need data orchestration If you’re running more than a handful of data pipelines across multiple teams, it’s time to consider orchestration. [It’s not just about automation](https://www.getdbt.com/blog/data-pipeline-automation)—it’s about trust, collaboration, and visibility. Common signals that you’ve outgrown manual processes include: - Frequent pipeline failures with long resolution times - Stakeholders complaining about stale or incorrect data - Scrambling to locate a pipeline when something breaks - No centralized access to pipeline logs for debugging - A mysteriously ballooning cloud bill from rogue jobs In short: If your data environment feels like it’s held together with duct tape and dashboards, orchestration is no longer a nice-to-have. ### What orchestration won't fix But orchestration alone isn’t enough. You also need a way to build, test, and scale pipelines consistently—across teams, tools, and clouds. Without that, you’ll still struggle with collaboration, trust, and governance. To build truly trustworthy pipelines, you also need: **A uniform framework for building and testing data transformations**. It’s hard for teams to collaborate on data pipelines if they’re defined in multiple systems and languages. **A way to see data lineage**. [Data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) traces the flow of data across your data estate. It’s an invaluable tool to see the origin of data and perform root cause analysis on sticky data transformation issues. **A single place to discover and use datasets**. A great dataset isn’t valuable unless others can find it. Besides orchestration, you also need to make data available to users so they can find it, learn about it, and bring it into their reporting dashboards, data-driven apps, and AI data pipelines. ### dbt is more than just orchestration dbt is a modern data control plane that combines transformation, testing, documentation, and orchestration in a single, unified platform. With dbt, you get: ✅ A version-controlled repo for all transformation logic, with CI/CD workflows and pull request reviews ✅ [Automated lineage](https://docs.getdbt.com/docs/explore/column-level-lineage) showing upstream and downstream dependencies ✅ Built-in testing and validation—so you catch issues _before_ they hit production ✅ [Auto-generated documentation](https://docs.getdbt.com/docs/explore/build-and-view-your-docs) with every run ✅ A searchable [dbt Catalog](https://docs.getdbt.com/docs/explore/explore-projects) for discoverability and context And for teams that need to scale orchestration further, dbt’s Fusion engine introduces [state-aware orchestration](https://docs.getdbt.com/docs/deploy/state-aware-about). That means: - Only models with updated inputs get re-run — saving time and compute. - You can control refresh logic with source freshness checks or custom intervals. Want more flexibility? You can define models in SQL or Python, depending on what your transformations require. And if you’re already using tools like Apache Airflow or Dagster, [dbt integrates seamlessly](https://www.getdbt.com/product/integrations) — letting you standardize transformation while using the orchestration tool that fits your broader stack. ## Conclusion Once you’re managing more than a few data pipelines, orchestration stops being optional—it becomes essential. A data orchestration platform helps you ensure data freshness, catch issues before they hit stakeholders, and break down data silos across your entire organization. dbt provides data orchestration on top of a best-in-industry data transformation framework. With dbt, you can monitor, test, and document your entire data estate—while keeping your workflows fast, visible, and reliable. Best of all, dbt fits into your existing stack. It’s cross-platform, SQL-native, and built for scale. No matter where it’s stored, with dbt, everyone who works with data in your company can speak the same language. **Ready to simplify and scale your data workflows?** [Start your free dbt trial](https://www.getdbt.com/signup) and experience smarter orchestration today. ## Data orchestration platforms FAQs **What is the relationship between data transformation and orchestration?** [Data transformation ensures the data is prepared and validated.](https://www.getdbt.com/blog/data-transformation) Data transformation focuses on cleaning, enriching, and structuring raw data for analysis or AI workflows. [Data orchestration manages the flow and execution of data tasks](https://www.getdbt.com/blog/data-orchestration-vs-etl). A comprehensive data strategy often requires both: orchestration to reliably coordinate pipelines and transformations to ensure that the data in those pipelines is trustworthy and usable. **How does a data orchestration platform improve data quality?** Data orchestration platforms coordinate every step of the pipeline, ensuring that data flows reliably from source to insight. With robust error detection, alerting, and logging capabilities, these platforms flag and capture issues before they impact production. **When should a company consider a data orchestration platform?** If you’re managing more than a few data pipelines across multiple teams then it’s probably time to consider a data orchestration platform. Keep an eye out for indicators like these: - Frequent pipeline failures with long resolution times - Stakeholders complaining about outdated or incorrect data - Difficulty locating pipelines when issues arise It also becomes essential if there's no centralized access to pipeline logs for debugging or if cloud bills are unexpectedly increasing due to unmanaged jobs. **How does dbt enhance traditional data orchestration capabilities?** [dbt combines transformation, testing, documentation, and coordination](https://www.getdbt.com/product/dbt) into a unified platform. It offers version-controlled logic, automated lineage for root cause analysis, and built-in testing to prevent issues from reaching production. [dbt's Fusion engine](https://www.getdbt.com/product/fusion) introduces state-aware orchestration, intelligently re-running only necessary models based on updated inputs, which optimizes compute resources and processing time. --- --- title: "What’s new in dbt Cloud - April 2025" description: "Get the scoop on all the latest features landing in dbt." url: "https://www.getdbt.com/blog/whats-new-in-dbt-cloud-april-2025" date: "2025-04-10" authors: ["Alexis Jones", "Sara Gawlinski"] categories: ["Product"] --- # What’s new in dbt Cloud - April 2025 Spring is springing 🌷 and it’s time for another dbt product update! Lots of exciting new features to share below, so keep reading for all the details. Also, ICYMI, we’re fresh off of our most recent product launch event, dbt Developer Day, where we showcased new features coming to dbt that will dramatically improve the developer experience so that teams can move faster, without sacrificing quality. Start your engines 🏁 and catch the replay [here](https://www.getdbt.com/resources/webinars/dbt-developer-day). While you’re at it, be sure to save your seat for [our next big launch event](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase), the dbt Launch Showcase, happening on May 28. Now onto what’s new in dbt Cloud this month. ## Cross-team collaboration 🕵️‍♀️ **Model query history is GA.** Quickly track model usage and popularity with this helpful feature, embedded right into dbt Explorer. By quickly grasping the relative popularity of your models, you have a data-driven “to do” list of where you should allocate your engineering resources, as well as built-in empathy for the consumers behind the models you build. Generally available to dbt Enterprise customers on Snowflake and BigQuery, with Redshift and Databricks coming later this year. [Read the docs](https://docs.getdbt.com/docs/collaborate/model-query-history) to learn more. ![Visual of dbt dag with model consumption lenses](https://cdn.sanity.io/images/wl0ndo6t/main/c3542b35d6bbdc678003e26f1a2b65908c992420-3164x1710.png) 📊** The Power BI integration for the dbt Semantic Layer is now in beta**. This one has been a long time coming and we’re excited to finally announce that our customers can now query metrics defined in the dbt Semantic Layer directly from Power BI. This makes it possible for data teams to maintain metrics in dbt while enabling business users to consume those metrics in Power BI. [Read the docs](https://docs.getdbt.com/docs/cloud-integrations/semantic-layer/power-bi) to learn more, and contact your dbt Labs account representative to request access to the `.msi` installer. ## Developer productivity 🪄**dbt Copilot is GA.** Our AI-powered data assistant here! dbt Copilot is redefining data engineering with context-aware AI that captures rich context from your dbt projects—metadata, model structures, and semantics—to auto-generate tailored SQL, documentation, tests, and semantic models that align perfectly with your team's conventions. dbt Copilot transforms your analytics workflows from manual, routine work into an AI-enhanced discipline, making data engineering smarter, faster, and more accessible than ever before. [Read the launch blog](https://www.getdbt.com/blog/dbt-copilot-is-ga) to learn more. ⚡ **The next-generation dbt engine is now in private beta.** Made possible by our recent acquisition of SDF Labs, [this new engine brings dramatic improvements to developer experience and efficiency](https://www.getdbt.com/blog/dbt-developer-day-2025). Think: sub-second parse times, intelligent SQL autocompletion, and error detection without hitting the warehouse. The new engine will soon power both the dbt Cloud IDE and a brand-new, dbt-built VS Code Extension, also in private beta. Broader availability is planned for late May, and you can [express interest in joining the beta now](https://docs.google.com/forms/d/1JElfCGT_fU1HlI-XUPs5SV9_nwNALaxTf0T7a8-RSJA/edit). 🧪 **dbt Core 1.10 is now in beta,** which includes sample mode and (soon) stricter validation. [Sample mode](https://docs.getdbt.com/docs/build/sample-flag) lets you build just a subset of your data in dev or CI—perfect for large time-based datasets—so you can iterate quickly and cut down on warehouse costs. You can define trailing or historical time windows using the `--sample` flag, or set a default window at the environment level. [Learn more](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.10) in the dbt 1.10 docs. ## Platform 😎 **Dark mode is** **GA:** Rejoice! You can now build your pipelines in the dbt Cloud IDE from the comfort of dark mode. Just navigate to your account name and set your theme to Dark. **🔒 Redshift External OAuth support via Identity Center for Okta and Microsoft Entra is GA**. It’s now easier and more secure than ever to set up development credentials with Redshift —no more managing long-lived passwords or access keys. [Read the docs](https://docs.getdbt.com/docs/cloud/manage-access/external-oauth) to learn more. **** **⏫ The oldest dbt versions are being deprecated.** To ensure the best possible usage experience in dbt Cloud, we are ending support for [dbt versions 1.0-1.2](https://docs.getdbt.com/docs/dbt-versions/core#latest-releases). Any workloads you may still have running on these versions will be automatically upgraded to dbt v1.3 by April 30th. Soon all updates for dbt Cloud customers will occur via [release tracks,](https://docs.getdbt.com/docs/dbt-versions/cloud-release-tracks) a configuration that _automatically_ upgrades your workloads to either the very latest or a recent version of dbt. As part of this move, within the coming months the dbt Cloud IDE will no longer support development for dbt versions <1.6. Please take note and proactively update your jobs and environments accordingly. **🔐 Azure DevOps Service Principal authentication is now live**. This brings a big security upgrade to replace the legacy Service User approach for customers using ADO. The enhancement introduces short-lived, 1-hour access tokens in place of long-lived PATs, aligning with Microsoft’s upcoming 2FA enforcement for both user and service accounts. If you're still using the old method, [now’s the time to migrate](https://docs.getdbt.com/docs/cloud/git/setup-service-principal#migrate-to-service-principal)—your security team will thank you. ## Partnerships ☁️ **GoogleNext Recap:** dbt went big at [Google Next](https://www.getdbt.com/blog/dbt-labs-launches-on-google-cloud-and-google-cloud-marketplace) this week. We know many of you use BigQuery and have been asking for support for Google Cloud. We’ve got great news for you: - dbt Cloud is now available in Preview for deployment on Google Cloud. If you’re interested, reach out to us at [sdr@dbtlabs.com](mailto:sales@getdbtlabs.com). - You will soon be able to find us for purchase on the Google Cloud Marketplace so you can procure dbt from the same place you get all your Google Cloud technologies. - dbt now supports DataFrames on BigQuery. Now you can execute dbt python models directly on BigQuery instead of Dataproc. - We also now support Workload Identity Federation so users can avoid having to use service keys to authenticate to BigQuery and use Oauth instead. 🔌 **Support for more data platforms**: We introduced our [Teradata adapter](https://docs.getdbt.com/docs/core/connect-data-platform/teradata-setup) in late 2024. After great customer feedback and a rigorous beta program, we’re thrilled to announce that this adapter is now generally available. Also, IBM watsonx.data and IBM Netezza are now supported on dbt Core, maintained by the IBM team. ### Lots more to come on May 28 We hope you’re as jazzed about these platform updates as we are! We look forward to seeing what you build and to hearing your feedback. [Be sure to save the date](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase) for our next major launch event, coming to an internet browser near you on May 28—the dbt Launch Showcase. Lots of exciting updates to share there, including the latest on the new SDF-powered dbt engine, our VS Code extension, bringing the power of dbt to more data collaborators, and much much more. See you there. --- --- title: "dbt Cloud 🤝 Google Cloud" description: "dbt and Google Cloud expand both their product integrations and their stratgic partnership at Google Next." url: "https://www.getdbt.com/blog/dbt-cloud-google-cloud" date: "2025-04-09" authors: ["Sean McIntyre", "David Tishgart"] categories: ["Partnerships"] --- # dbt Cloud 🤝 Google Cloud ## dbt Cloud now available on Google Cloud If you’re running dbt Cloud on BigQuery – and we know many of you are – today’s announcement that dbt Cloud is now available on Google Cloud is super exciting. It means your data control plane and your data platform are now both hosted natively on the Google ecosystem. Today our customers have even more flexibility in how they build and scale their data and AI environments. Whether you’re centralizing spend, focused on compliance and data residency, or building data pipelines, you now have all the power of dbt Cloud—on Google’s trusted infrastructure. If you run your data architecture on Google Cloud, but are not yet a dbt customer, you will soon be able to find and purchase dbt Cloud directly from the Google Marketplace. That means you can get up and running faster, simplify procurement, and align spend with your existing cloud commitments. For those unfamiliar, dbt Cloud is a [Data Control Plane](https://www.getdbt.com/dbt-cloud/get-started) that helps organizations build and manage their pipelines, so teams can deliver high-quality, trusted data to the business faster and at a reduced cost. dbt takes a code-first approach, with a variety of governed development environments that cater to a range of technical aptitudes. This includes a cloud-connected command-line interface, a graphical IDE, an AI-powered Copilot, or a visual drag-and-drop experience – it’s all interoperable. Data teams regularly build up data warehouses with dbt: dimensional modeling, data vault, data marts, or one big table. Whether the end application is AI, Looker for BI, or operational analytics, dbt is the new standard for data transformation and provides a solid foundation on which organizations to build revenue-generating data products. Check out the Dev Day blog to learn about the [new features of dbt Cloud](https://www.getdbt.com/blog/dbt-developer-day-2025) that are turbocharging data development. ## Why dbt Cloud on Google Cloud? For many teams, data architectures are getting more complex. There's a growing need to deploy workloads flexibly, meet region-specific data residency requirements, and manage cost—without sacrificing performance. By bringing dbt Cloud to Google Cloud, we’re giving customers the ability to deploy where it makes the most sense for their business. And since we already support deep integrations with [BigQuery](https://www.getdbt.com/data-platforms/bigquery) you’ll be able to run fast, efficient transformation workflows on one of the most scalable engines out there. This includes new enhancements like full support for BigQuery DataFrames, so Python workflows and machine learning pipelines can live right alongside your SQL models. BigQuery customers that are recognizing business value from dbt Cloud include [RocketMoney](https://www.getdbt.com/case-studies/rocket-money), [Bilt Rewards](https://www.getdbt.com/case-studies/bilt-rewards-professional-services), and [Virgin Media O2.](https://www.youtube.com/watch?v=iY7rbutVwm0&t=3s) There are a number of advantages to building your data pipelines in BigQuery with dbt Cloud: **Faster Time to Value - **Build data pipelines using modular SQL-focused code, CI/CD, versioning, and scheduling to speed development and reduce errors. dbt now supports BigQuery DataFrames, so you can execute dbt python models directly on BigQuery and Dataproc. **Increased Governance - **Built-in data testing, automated documentation, and detailed data lineage ensure that transformed data in BigQuery is accurate, reliable, and easy to audit. **User empowerment - **With governed development environments that appeal to any level of technical acuity - Cloud CLI, graphical IDE, or drag & drop visual editor - dbt Cloud makes it easier for analysts to collaborate and bring critical context to data development. **Scalability + management - **Out of the box** **Incremental models, partitioning, and clustering help reduce data scanned and improve efficiency. **** ## Platform flexibility dbt Cloud now works on all the major hyperscalers and most popular cloud data platforms. With out-of-the-box support for the Iceberg open table format (coming soon), dbt gives you greater choice in where you store your data and where you work with it, eliminating the need for you to extract and load data into the platform directly. For example, transformed data stored in Google Cloud Storage in Iceberg format can be read by multitudes of applications and cloud data platforms (Snowflake, Databricks, Redshift) – using the tools or platform that best suits your team needs and requirements. For companies with multi-cloud or multi-platform environments, or who empower domain teams to manage their own pipelines, dbt and Iceberg provide a common way of working across projects enforced by data contracts. This enables teams to move faster while staying within governance and compliance boundaries. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/70f396647d9433c8b9ef7179e992b1e774c70ae0-872x516.png) Caption: _cross-project and cross-platform references in dbt Cloud_ ## AI runs on data. Data runs on dbt. You’re probably familiar with the old proverb in computer science, “garbage in, garbage out.” Organizations use dbt Cloud to produce higher quality datasets with data quality checks at every step of the data pipeline: from source testing and freshness checks on the raw data, all the way to data contracts and tests at the consumption layer. Inspired by how software developers work, dbt adds testing, documentation, and versioning to ensure datasets are high quality and reusable for years to come. dbt Copilot, the AI experience embedded within the data control plane, helps data teams leverage the power of GenAI with their rich metadata context to speed up analytics development across dbt Cloud. This includes AI to auto-generate documentation, data tests, semantic models, metric definitions, and inline SQL. With dbt Copilot, data teams have a context-aware, AI assistant directly in the dbt Cloud IDE—and soon, the visual editor – that leverages dbt model structures, relationships, metadata, and lineage to deliver smart, actionable recommendations. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a888c629af75c3f011d7b94ace329a8da1e9275c-2143x1293.gif) Caption:_ dbt Copilot autogenerating documentation_ When data is well-governed and created with confidence, it improves the quality and veracity of AI outputs and gets us closer to finally realizing value from AI. With dbt’s python models, running data science workloads and machine learning during the process, provides a reliable way to bring data warehouse datasets downstream for training and running models using an MLOps approach. dbt Cloud also integrates with Vertex, allowing dbt to query BigQuery with Vertex AI functions, to apply key insights to data, and store them back into BigQuery, for other applications to pick up on and easily build on. We're co-hosting a webinar with the Google Cloud team on **May 7** to walk through everything that’s new, complete with demos and tips for getting started. **** --- --- title: "dbt Labs Launches on Google Cloud and Google Cloud Marketplace to Power Data and AI Workflows" description: "Expanded Google partnership delivers new flexibility while enhanced BigQuery platform integration supports machine learning." url: "https://www.getdbt.com/blog/dbt-labs-launches-on-google-cloud-and-google-cloud-marketplace" date: "2025-04-09" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Launches on Google Cloud and Google Cloud Marketplace to Power Data and AI Workflows **PHILADELPHIA, April 9, 2025 – **dbt Labs, the pioneer in analytics engineering, today announced dbt Cloud is now deployable on Google Cloud. This empowers organizations with flexibility to choose how they architect their data and AI environments, while centralizing spend. In addition, dbt Cloud will soon be available via Google Cloud Marketplace, giving customers the freedom to access the secure, scalable data platform via a new channel. This new multi-tenant instance allows organizations to leverage dbt Cloud and meet often complex data residency and compliance requirements. Simultaneously, dbt will continue to enable data teams to build and manage faster, more efficient data and AI workflows on Google Cloud’s [BigQuery](https://www.getdbt.com/data-platforms/bigquery), a unified, AI-ready data analytics platform. "We are excited to launch our partnership with Google Cloud to expand customer access to dbt Cloud, enable our customers to deploy how and where they want to manage their data estates, and continue to deliver a world class experience for current and future joint customers,” said Brandon Sweeney, President and Chief Operating Officer at dbt Labs. “As the standard for AI on structured data, dbt is helping teams boost data quality while shipping trusted data faster. Together, we can empower more teams to embrace the Analytics Development Lifecycle (ADLC).” dbt Labs is also strengthening its integration with the BigQuery platform by launching full support for BigQuery DataFrames, a Python API for data analysis and building machine learning (ML) workflows. In conjunction with new authentication enhancements, users can combine dbt and BigQuery to improve data quality and trust, key requirements for AI on structured data. As BigQuery ranks in the top three data connections used by the 100,000-plus members of the dbt Community, this more robust integration will better support data teams as they analyze data and execute ML tasks in the platform. More efficient data teams supported by dbt Copilot, an AI-powered data assistant in dbt Cloud, will enable BigQuery users to build their pipelines more accurately and more efficiently. “Bringing dbt Cloud to Google Cloud Marketplace will help customers quickly deploy, manage, and grow their data and AI environments on Google Cloud’s trusted, global infrastructure,” said Dai Vu, Managing Director, Marketplace & ISV GTM Programs at Google Cloud. “dbt Labs can now securely scale and support customers on their digital transformation journeys.” This expanded relationship with Google Cloud is the latest innovation milestone for dbt Labs, which recently crossed the [$100M ARR threshold](https://www.getdbt.com/blog/dbt-labs-100m-arr-milestone). Earlier this year, [dbt Labs acquired SDF Labs](https://www.getdbt.com/blog/dbt-labs-announces-sdf-labs-acquisition) to bring SQL comprehension into dbt and usher in a new era of ‘what’s possible’ for analytics: supercharging developer productivity and heightening data quality, all while optimizing data platform costs. In addition, dbt Copilot is now [generally available](https://www.getdbt.com/blog/dbt-copilot-is-ga). With dbt Copilot, users can leverage the full context of their data—its relationships, metadata, and lineage—to automate routine tasks and consistently uphold key ADLC best practices like documentation, data testing, semantic modeling, and SQL formatting. The result is a refined, governed dataset that serves as a solid foundation for organizations to build high-quality analytics and advanced AI systems, leading to better business decisions. This week, the dbt Labs team will be at Google Cloud Next (booth #2895) to demo the latest innovations in dbt, including dbt Copilot with Google Cloud’s BigQuery platform. On May 7, dbt Labs will host a webinar with Google Cloud teams to discuss how to leverage all of the new product integrations, the power of AI-driven analytics, and real demos to get you started. Learn more and register to attend at [https://www.getdbt.com/resources/webinars/dbt-bigquery-whats-new-from-google-next](https://www.getdbt.com/resources/webinars/dbt-bigquery-whats-new-from-google-next). **About dbt Labs** Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 50,000 teams using dbt every week. To learn more about dbt Labs, visit [getdbt.com](https://www.getdbt.com/) and follow us on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/). --- --- title: "How to build scalable data pipelines with Snowflake and dbt" description: "Build fast, scalable data pipelines with Snowflake and dbt—backed by AI, DevOps, and dynamic cost optimization." url: "https://www.getdbt.com/blog/data-pipelines-snowflake-dbt" date: "2025-04-07" authors: ["Luis Leon"] categories: ["Learn"] --- # How to build scalable data pipelines with Snowflake and dbt The demand for high-quality data is accelerating. That’s especially true with the emergence of AI and machine learning use cases. This means analytics engineers need tools that enable them to create, test, document, and deploy data pipelines faster than ever before. Over the past several years, dbt Labs and Snowflake have both spent a lot of person-power figuring out how to remove friction from the data pipeline development experience. Using the two systems together, analytics engineers can bring new data products from conception to reality quickly—without sacrificing data quality. I’ll dive into the benefits of using dbt together with Snowflake, then show the two in action by implementing a data pipeline that analyzes trading profit and loss. I’ll also dig into some advanced use cases, such as leveraging a Large Language Model (LLM) to extract patterns and sentiments from trader execution notes, and discuss other ways that dbt and Snowflake are leveraging AI to accelerate data pipeline production. **** ## Why dbt + Snowflake for data pipelines? dbt is the industry standard for data transformation at scale. It’s a framework that sits on top of your data warehouse and lets anybody who can write SQL statements deploy production-grade pipelines on top of Snowflake. One thing that makes dbt such a strong framework is that [it implements DevOps-style software engineering best practices](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) in the data world. This means that dbt brings ideas like version control and modularity to your data engineering pipelines. Snowflake is a powerful AI Data Cloud that allows you to build data-intensive applications without the operational overhead. Its unique architecture and years of innovation make it the best platform for mobilizing data in your organization. Snowflake has also released a number of features that support a [DevOps](https://www.snowflake.com/guides/conquering-devops-data/) (or “DataOps”) approach to data pipelines. These include source control repository integration via Git, building Snowflake data primitives from your Git repo code, and running scripts and deploying complex objects—data models, Machine Learning (ML) models, native apps, etc.—with the Snowflake CLI. You can also use the Snowflake API from your Python scripts to work with Snowflake in a more Pythonic way. With dbt plus DevOps support from Snowflake, you can now mix and match features to leverage advanced change management and build flexible Continuous Integration/Continuous Deployment (CI/CD) pipelines. This enables you to elevate your data engineering practices to the next level without resorting to third-party tools such as Terraform. dbt and Snowflake both work well together here, in part, because both take a **declarative** approach to DevOps for data. Both systems work together to remove a lot of the lower-level complexity involved in working with data. That automation increases the number of people who can meaningfully transform data, leading to increased data democratization. ## Building your dbt-Snowflake pipeline Here’s how these features come together to make it easier to build data pipelines. In this walkthrough, you’ll use Snowflake and dbt to analyze trading profit and loss (P&L). ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/c673b9f57548692db8a50313e7dd9229a53db3fc-512x309.png) First, you’ll calculate P&L by normalizing trade data across multiple currencies by using current foreign exchange (FX) rates and compare actual performance against the portfolio targets. To do this, you’ll leverage datasets from Snowflake Marketplace on FX rates and US equity price history. You’ll blend these with a couple of manual sources coming from CSV files that contain history from the trading desk, as well as the target allocation ratios we’ll use as a configuration. Additionally, you’ll leverage the power of an LLM to extract patterns and sentiment from trader execution notes, giving analysts more insight into the decision-making behind their activity. This article gives a lot of technical background on each step. It’ll also discuss options for various steps in case your scenario differs slightly. If you want to focus on just setting up and running the code, [you can use the step-by-step guide here](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html?index=..%2F..index#0). You can also see a detailed walkthrough of this demo [in our webinar replay](https://www.snowflake.com/webinars/virtual-hands-on-labs/deploying-data-pipelines-with-snowflake-and-dbt-labs-2025-03-12). ### Prerequisites - A [Snowflake account](https://app.snowflake.com/), plus a user with the ACCOUNTADMIN permissions - A [GitHub account](https://github.com/signup) into which you can import (fork) the example code. You can join GitHub for free if you don’t have an account you can use for this walkthrough. ### Access data To get started, [import the two data products](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html?index=..%2F..index#2) from Snowflake Marketplace: the US Stock Prices table and the Currency Exchange Rates data. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f207f813b84aa58225a574310565d35571989eec-512x212.png) Snowflake's unique value proposition started with a completely new way of sharing and collaboration both within and between organizations. In this scenario, you’re going to leverage this power to connect to two free public resources available to all Snowflake users. This data is immediately available in your account as a source. There's no ETL. There’s zero latency and no additional storage cost, because the data still sits on the provider side. You’ve just become a consumer of it. After this, you’ll have two data sets available as if they’re a part of your system: the Forex currency exchange and stock trading stock prices data. Next, you run a script from Snowflake to connect to GitHub and [set up your development environment](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html?index=..%2F..index#3). Once that's done, your Snowflake environment will have a pointer to [your forked Git repository](https://github.com/Snowflake-Labs/sfguide-deploying-pipelines-with-snowflake-and-dbt-labs?_fsi=b8uZIrMV&_fsi=b8uZIrMV). If you click on it, you can see exactly the same files you have in GitHub. In the GitHub repo’s deploy_environment.sql file, you’ll see a file that has two unique attributes: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/68c1b7c8ee5b8b8b26c947847acbd1f39446a34a-512x83.png) First, it has these curly brackets that represent the [Jinja templating language](https://jinja.palletsprojects.com/en/stable/). This is exactly the same syntax you see everywhere in dbt, which is also Jinja template-based. In this case, the script has templatized environments with the {{env}} variable so you can re-run this script to create your development, staging, testing, or production environments. Second, there are the [CREATE OR ALTER object commands](https://docs.snowflake.com/en/sql-reference/sql/create-or-alter) - DDL language for creating tables and views. With this script, you don't need to care about the state of your environment. You can rerun that script repeatedly and recreate new environments for dev, staging, or production that have everything you need. You can set up your environment right from your Git repo by using the [EXECUTE IMMEDIATE FROM](https://docs.snowflake.com/en/sql-reference/sql/execute-immediate-from) command, which can execute a file written with Jinja: ```sql USE ROLE ACCOUNTADMIN; USE DATABASE SANDBOX; ALTER GIT REPOSITORY DEMO_GIT_REPO FETCH; EXECUTE IMMEDIATE FROM @DEMO_GIT_REPO/branches/main/scripts/deploy_environment.sql USING (env => 'DEV'); ``` ### Transform data with dbt Now, let’s see how things work on the dbt side. First, use dbt to initialize your empty environment with the [dbt seed](https://docs.getdbt.com/docs/build/seeds) command. This will upload your CSV files containing your trading desk history and target allocation ratio data. ```sql cd dbt_project dbt seed ``` Next, run your [dbt models](https://docs.getdbt.com/docs/build/models), which specify how data will be transformed from the source and stored in the target tables. These can be different data systems; in this case, you’ll use Snowflake as both the source and the target. The project divides its models into three logical groupings—staging, intermediate, and marts—which is one effective way to model your data for different purposes by segmenting them into different schemas. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f601d6ee33719a803df2acc45cede2775374631b-512x359.png) I won’t go too deep into the dbt models here. Suffice it to say, I generated a bulk of this code using [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot), which leverages LLMs and information stored in a project to generate code based on natural-language descriptions. I’m excited about how well this works and the potential to boost developer productivity that this brings. You can run these models with the dbt run command. dbt provides flexibility in how the new data is represented, or [materialized](https://docs.getdbt.com/docs/build/materializations), in Snowflake. Initially, the models are configured to materialize as views, which speeds up development and testing. For production, you’d likely use table materialization for better query performance. You can change the materializations at any time—and set them differently in different environments—to suit your use case and optimize your pipelines. For example, if you’re materializing as tables (which will make querying in production more efficient), every time you run your model, it will copy all of the data over. If you want to be more efficient, you can change the [materialization option](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html?index=..%2F..index#6) for a given model to switch from table to [incremental materialization](https://docs.getdbt.com/docs/build/incremental-models?_fsi=b8uZIrMV&_fsi=b8uZIrMV&_fsi=b8uZIrMV), which only surfaces the rows you tell dbt to filter for. For example, you can tell dbt to only import those rows that were created or modified after a specific date. ### Scale with Snowflake DevOps features Finally, you can leverage a combination of Snowflake and dbt features to [deploy your pipeline to production](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html?index=..%2F..index#8). Here, you use the same approach you used to set up your dev environment by leveraging the EXECUTE IMMEDIATE FROM command but with an environment value of prod supplied as an argument: ```sql USE ROLE ACCOUNTADMIN; USE DATABASE SANDBOX; ALTER GIT REPOSITORY DEMO_GIT_REPO FETCH; EXECUTE IMMEDIATE FROM @DEMO_GIT_REPO/branches/main/scripts/deploy_environment.sql USING (env => 'PROD'); Then you can use dbt to deploy the seed data and run your data transformations: dbt seed --target=prod dbt run --target=prod ``` You _could_ have used Snowflake stored procedures for all of this work. However, in my view, dbt represents a huge productivity boost in a number of ways: - Because dbt uses a declarative vs. an imperative model, you save a lot of time in building declarative models versus imperative workflows. - The dbt logic is also portable. Since all of your data transformations are code, you can put them under version control, where they can be change managed, and other analytics engineers can find them. These devs can easily run it on other Snowflake instances, leveraging the code to recreate this pipeline themselves or build more sophisticated ones. - dbt’s support for testing improves data quality, while its support for shipping rich documentation with its models makes it easier for stakeholders to understand and use the resulting data. There will still, however, be custom components specific to Snowflake—tasks, streams, stored procedures, etc.— which might not be supported directly in dbt. In the past, you’d have to set these up with something like Terraform or DIY your own solution. New Snowflake DevOps features—such as the [Snowflake CLI](https://docs.snowflake.com/en/developer-guide/snowflake-cli/index) and the [Python API](https://docs.snowflake.com/en/developer-guide/snowflake-python-api/snowflake-python-overview)—mean you can manage your data transformation models and all Snowflake objects via automated CI/CD workflows. ### Optimize for performance and cost Using dbt and Snowflake features, you can optimize this even further. For example, you can leverage [Dynamic Tables](https://docs.snowflake.com/en/user-guide/dynamic-tables-intro) in Snowflake to [deploy your intermediate and mart models](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html?index=..%2F..index#7). With Dynamic Tables, Snowflake takes care of the table creation, so you don’t even need to create a separate data pipeline. Using Dynamic Tables is [as simple as turning on the relevant option](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html?index=..%2F..index#7) for the relevant model group in your dbt_project.yml: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/9591c947d9c75de6ec0157ec65d22f7030a2f127-1600x1183.png) Note that if you’ve defined [data quality tests](https://docs.getdbt.com/docs/build/data-tests) in dbt, these will still be run even when using Dynamic Tables. So you get the benefit of fast, low-overhead creation of data while also checking for and issuing alerts on any data quality issues with every run. Another great advantage of using dbt with Snowflake from a cost optimization standpoint is that you can dynamically control what size Snowflake instances you use for your runs at a granular level. You can, for example, say for one model you want to resize up from a small warehouse instance (the default) to a medium for the model run — and then immediately size it back down. You can control this for each model, for a subset of models, and even based on the environment. I’ve seen customers resize dynamically based on how much data was in the input. It’s a powerful feature that grants customers a lot of flexibility. ## The future of data pipelines: AI and automation You can take your dbt and Snowflake pipelines even further by leveraging new functionality from each platform. For example, the [intermediate/int_extracted_entities.sql model](https://github.com/Snowflake-Labs/sfguide-deploying-pipelines-with-snowflake-and-dbt-labs/blob/main/dbt_project/models/intermediate/int_extracted_entities.sql) utilizes [Snowflake Cortex](https://www.snowflake.com/en/product/features/cortex/) to extract the trader’s notes and uses an LLM to detect what market signal is driving the trade. It also classifies the note based on what execution strategy the trader seems to be taking—e.g., taking profit or cutting losses. ```sql with trading_books as ( select * from {{ ref('stg_trading_books') }} ), -- Extract sentiment using SNOWFLAKE.CORTEX.SENTIMENT cst as ( select trade_id, trade_date, trader_name, desk, ticker, quantity, price, trade_type, notes, SNOWFLAKE.CORTEX.SENTIMENT(notes) as sentiment, SNOWFLAKE.CORTEX.EXTRACT_ANSWER(notes, 'What is the signal driving the following trade?') as signal, SNOWFLAKE.CORTEX.CLASSIFY_TEXT(notes||': '|| signal[0]:"answer"::string,['Market Signal','Execution Strategy']):"label"::string as trade_driver from trading_books where notes is not null ) select * from cst ``` On the dbt side, [the company recently acquired SDF](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs), a high-performance toolchain that acts as a multi-dialect SQL compiler, linter, and language server, among other things. SDF can faithfully emulate cloud data warehouses such as Snowflake, catching breaking changes during development before analytics engineers even check in a single line of code. This speed means SDF can catch errors as you’re typing—well before they even first type in dbt run. Finally, I already mentioned how I leveraged [dbt Copilot ](https://docs.getdbt.com/docs/cloud/dbt-copilot)to generate most of my models. dbt Copilot integrates with the dbt Cloud Integrated Development Environment (IDE) to assist in the generation of [code](https://docs.getdbt.com/docs/cloud/use-dbt-copilot), [documentation](https://docs.getdbt.com/docs/build/documentation), [tests](https://docs.getdbt.com/docs/build/data-tests), [metrics](https://docs.getdbt.com/docs/build/metrics-overview), and [semantic models](https://docs.getdbt.com/docs/build/semantic-models). This greatly reduces the time spent writing models, accelerating the deployment of new data pipelines. ## Bringing AI-powered pipelines to life with dbt + Snowflake dbt and Snowflake both implement a number of declarative features that unlock faster and more efficient end-to-end data workflows. Used together, teams can initialize, model, test, document, and deploy data pipelines powered by the latest advances in AI faster than ever. See it first-hand for yourself. Sign up for [a free Snowflake account](https://signup.snowflake.com/?utm_cta=trial-en-www-homepage-top-right-nav-ss-evg&_ga=2.77451413.1258348163.1743559550-689731207.1741746584) and [a free dbt Cloud account](https://www.getdbt.com/signup) and try it out yourself [by following our walkthrough](https://quickstarts.snowflake.com/guide/data_engineering_deploying_pipelines_with_snowflake_and_dbt_labs/index.html?index=..%2F..index#0). --- --- title: "How AI will disrupt BI as we know it" description: "A continuation-in-spirit from my recent post, “How AI will disrupt data engineering as we know it.”" url: "https://www.getdbt.com/blog/how-ai-will-disrupt-bi" date: "2025-04-06" authors: ["Tristan Handy"] categories: ["Insights"] --- # How AI will disrupt BI as we know it _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/how-ai-will-disrupt-bi-as-we-know). _ Business intelligence is on a collision course with AI. The collision itself hasn’t happened yet, but it’s clearly coming. The inevitability of this has been clear roughly since the launch of ChatGPT, but no one knew exactly what shape that would take. Today I want to propose how that collision is going to happen and what will happen in its aftermath. I think it will be a very good thing for data practitioners of all stripes—those who officially have the word ‘data’ in their title but also everyone else who simply uses data in the service of their larger job. So: I’m all for it. Before getting into AI part of the story, I need to introduce two specific mental models. Let’s go. ## BI is a portfolio of stuff We all use the term “BI” but have become inured to what an Orwellian term it is. “Business intelligence” isn’t descriptive, it is industry-speak for a bunch of stuff glued together in order to achieve a desired user outcome: know facts about a business using tabular data. For a long time, BI included a bunch of stuff that it no longer does. Like: data processing. Pre-cloud, BI tools processed data locally and often had proprietary processing engines. They competed on being fast. With the cloud, that evaporated. Local data processing was anathema. BI tools got easier to build but gave up a part of their value proposition. In today’s post-cloud world, I would suggest that BI tools have three jobs: 1. **Modeling:** Define the semantic concepts behind your structured data: metrics, dimensions, joins, etc. Think: LookML. 2. **Exploratory data analysis (EDA):** The iterative process of exploring data in search of useful insights. Highly iterative, flow-state, and unpredictable. Think: Looker explore window. 3. **Presentation:** The aggregation of multiple data artifacts together to present a single cohesive narrative that can be shared out to potentially many others within an organization, all governed by a permission model. Think: Looker dashboard. ![The 3 jobs of a BI tool in 2025](https://cdn.sanity.io/images/wl0ndo6t/main/5c38d953cb21b81c1d663759ca41845fa72ca8e9-1456x549.jpg) Some tools skip modeling and just allow users to do EDA without a model. EDA and presentation are the most core jobs of any BI tool and every BI tool I’m familiar with does both. And it is the fact that BI tools facilitate the EDA process that enables them to govern and share the presentation of that analysis. ## Scaling the criticality of an analysis _All credit to my collaborator Dave Connors for this mental model 🦞_ Generally speaking, artifacts pass through a few lifecycle stages as they mature into data products supporting production use cases. Think about these stages as the ‘production line’ of BI. ### Phase 0: Exploratory analysis The first thing data practitioners do when faced with a business question is to start developing low fidelity sketches to try to answer it. The vast majority of the work generated here will be thrown away, so there are low expectations code quality and governance. The primary goals of the best EDA experiences are iteration speed, flow state, and flexibility. ### Phase 1: Personal reporting At a certain point, some exploratory analysis will cross over into a true insight; your question is answered, your curiosity sated. The question is important enough that you want to make sure you can return to it later. But it is not yet “ready for prime time”—you’re not ready to share it with others and have it be a part of someone’s operating cadence. Some BI tools have a separate section for your “personal space”—think about your personal folder in Looker. ### Phase 2: Shared reporting The moment that a report gets shared with another person, the required governance characteristics of a data artifact increase significantly. When you create a report you understand its context; when someone else starts using it they just expect it to be correct. In phases 0 and 1, there may not be any governance applied—all governance may be applied at the compute layer with grants. But once you share an artifact, it is the governance at the BI layer that determines who gets to see what. This is simply because _most data consumers don’t have accounts within the data platform_ and so the BI tool takes over as the arbiter. In phases 0 and 1, there is also no auditability requirement. Auditability, change tracking, and general data ops best practices are introduced when artifacts are shared with others in Phase 2. ### Phase 3: Production artifact When shared reporting reaches a very high level of criticality (frequent access by a large number of end users, agreed upon SLAs, supports a critical business process, dynamic features), it’s officially “in production” and needs to be owned and operated like any other production data asset. === If you think about these stages as the ‘production line’ of BI, the most important job of a BI tool is to be the conveyor belt through all of these stages. Start with raw materials, end with a production data product. At each phase of maturity, it’s easy to extend the product to support the next set of capabilities: governance, dynamic filters, SSO, etc. You never think about those things during Phase 0, but as your work progresses, the BI tool makes it straightforward to progressively add those capabilities. But for this all to work, you gotta start the process inside the BI tool all the way back from Phase 0. You can’t do your EDA in Jupyter & Pandas and expect to ship it to users in Tableau…that’s not how that works. So: you gotta do your EDA in a BI tool to take advantage of the “production line”. But…are BI tools typically the best way to do EDA? We’ll return to that later. ## MCP and AI-as-aggregator The final thing we need to understand is the impacts of a _context protocol_. I wrote about this [a few weeks ago](https://roundup.getdbt.com/p/how-ai-will-disrupt-data-engineering): > The easiest thing to do for any technology vendor at the very onset of the AI era was to take all of the domain-specific context that you had and surface it to users in a chat interface. And we did the same thing. It was (and is) quite good—it does a great job of allowing users to ask business questions and answering them with semantic-layer-governed responses. > > The problem with this approach is that users don’t actually want to interact with dozens of chat interfaces. They don’t want to remember to go to a given tool to get one type of answer and another tool for another type of answer. There will not be 30 chat experiences all with different context. There will be one…or maybe just a few. But likely a single dominant one. > > This is how [aggregators](https://stratechery.com/aggregation-theory/) work. You likely don’t use a bunch of different search engines—you probably just use one, and it is probably Google. This is how chat will go as well. > > The problem is, Google could scrape the web and respond to all queries based on that knowledge. But ChatGPT cannot know all of the information you want to ask it questions about (at least, yet). That lack of business context is the problem. > > That’s where a _context protocol_ comes in. A context protocol—a somewhat new topic in the public AI conversation—is a standardized way for services to provide additional context to models via an open protocol. The most promising one today is called [MCP](https://modelcontextprotocol.io/introduction), but whether or not MCP wins, the awareness/excitement/support for this idea has developed a ton of momentum and I am fairly convicted that _something like this_ will become real and widely-supported. > > There will be a large number of context providers (every source of valuable enterprise context) and a large number of context consumers (different products with AI capabilities). There is no way to create point-to-point integrations to facilitate this. A protocol will be needed if we are going to see the right type of advancements, and I think it will happen. > > Imagine that your license to ChatGPT enterprise or Claude Desktop or whatever _already came with_ a connection to all of the metadata about every piece of structured data you had access to. What was there, how trustworthy it was, how suitable it was for the analysis you were describing, etc. Well, in the intervening weeks since I wrote this, a couple of things have happened. First, this: ![Sam Altman tweet](https://cdn.sanity.io/images/wl0ndo6t/main/ba42b81ac5081b43c387972cb925744b9f92a25c-1690x1376.png) …then this: ![Sundar Pichai tweet](https://cdn.sanity.io/images/wl0ndo6t/main/8d0715323fcdb7c6bc2caa1689f460caeb479367-1686x996.png) Clearly this thing is going somewhere. Just as momentously, I have gotten access to an internal-only dbt/MCP-powered experience in Claude Desktop. In it, I can ask every type of metadata question I might want (powered by our Metadata API) and I can also ask questions about all of our business metrics (powered by our Semantic Layer API). It is incredible. I don’t want to share too much right now, but … having your data and metadata available in the context of a modern reasoning model is incredible. ## BI in an AI-first world Ok, we now understand: the jobs of a BI tool, BI conveyor belt, and how to get structured data context into your AI-of-choice. We’re finally in position to tackle the coming collision. Here it is; plain and simple: 1. AI is going to be meaningfully better at exploratory data analysis than any BI tool. 2. If you take away EDA from BI, the ‘conveyor belt’ model breaks down. And the conveyor belt model is the primary reason you use your current BI tool. 3. It is not yet clear how the BI ecosystem will adapt to this new reality. That’s it. That’s my entire argument. Let’s see if it holds up. ## Artificial intelligence will far outstrip business intelligence for exploratory data analysis There are a lot of data tasks that AI is good at. I’ve talked about a lot of these in the context of data engineering [here](https://roundup.getdbt.com/p/how-ai-will-disrupt-data-engineering). But the area of data analysis that is going to be _most_ benefited from AI is EDA. I am confident about that for two reasons. First, I have empirically validated this first-hand. The dbt + MCP + Claude 3.7 combo that I outlined earlier is just dramatically better at EDA than anything I’ve experienced in my life, and it’s getting better fast. But I am not ready to show you that (it’s single-digit weeks away from a public demo!), so you may not believe me. Fair. The second reason I’m confident about this is the fact that most time spent in EDA is writing code (whether done by hand or via a GUI). And we now know how good leading-edge models are at writing code when supplied with the right context. Whether you want to reference [individual developer testimonials](https://www.techsistence.com/p/up-to-90-of-my-code-is-now-generated) or [the head of YC](https://www.cnbc.com/2025/03/15/y-combinator-startups-are-fastest-growing-in-fund-history-because-of-ai.html?utm_source=tldrnewsletter) or [Andrej Karpathy](https://x.com/karpathy/status/1886192184808149383) or [Google](https://www.forbes.com/sites/jackkelly/2024/11/01/ai-code-and-the-future-of-software-engineers/), it all lines up. And it just so happens that the two software engineers whose opinions I trust most in the world—my cofounders Drew and Connor—have gone all in on Cursor over the last 3 months and are not-quite-but-almost religious about the experience. If you find yourself skeptical of this, here are a few things to keep in mind. 1. You don’t need the LLM to answer ‘why’ questions, or generate hypotheses, to have it be far superior than your current workflow. Rather—it just makes you _a lot faster_ because it can write EDA code a whole lot faster than you can (whether you’re writing Excel formulas or dataframe operations). 2. Accuracy is a non-issue as long as you ask a question that can be governed by a semantic layer. The code written tends to be: get data from the SL, manipulate it in Python, generate a chart using some dynamic javascript library. If you can’t get a dataset governed by an SL query, text-to-SQL does continue to improve with sufficient context. Just imagine: an interface that allows you to just have your questions answered far faster. You remain the objective function and the creative drive behind the process, AI is simply better and faster than you are at writing analytical code. IMO that shouldn’t feel threatening, **that should feel empowering**. I seriously lost it the first time I interacted with our internal data in this type of experience. The primary value prop of a data analyst shouldn’t be writing code, it should be analytical problem solving and generating action. ### The conveyor belt model breaks down You typically don’t tend to use your BI tool because it is the fastest or most delightful EDA experience. You use it because, when you have something to publish to your coworkers, you know exactly how to do that. But what if another tool were _so much better_ at EDA that _you would be handicapping yourself if you didn’t use it?_ What would you do? There are likely three answers. First, you could go back to publishing one-off assets. If you ask any AI experience to “give me that in an Excel file” most of them have no problem doing that. So maybe you just go back to shipping attachments. But that doesn’t feel like progress. Second, having iterated and found the insight you were looking for, you now have to reconstitute that analysis inside of your BI tool of choice. In practice this will likely only happen rarely; it is not a stable equilibrium because every human hates double work. Third, and hopefully preferable, is that we find some way to pull back in the results of an exploration into the governed framework of the BI tool. Imagine asking “make a PowerBI worksheet out of this analysis.” We will need to get deeper into the MCP era to see exactly how this will play out, but I’m optimistic that it will be possible. The third option still sees the BI tool as an important governance and presentation layer but pulls out the most strategic responsibility (EDA) from its portfolio. ## A very different BI tool BI tools used to ship with compute engines. Today they do not. What if BI tools were no longer the primary way EDA was done? What if their primary job was to render data artifacts a governed, interactive environment? That is still an incredibly valuable thing, and needed as long as humans are going to continue to interact with structured data (IMO: a long time). But it’s not what BI tools look like today. Most BI tool vendors want to pull this new EDA experience _inside their Chrome_—exposing AI-powered interfaces inside their products. I don’t believe this will be how most users do EDA, for three reasons: 1. **User behavior: **Aggregation theory will dominate, every knowledge worker inside of a company needs access to this functionality and they’re not all going to think to go to a specific tool first, they’re going to prefer to simply ask data questions in the same place they ask all of their other questions. Claude, ChatGPT Enterprise, whatever. 2. **Tool combinations: **MCP is not only powerful because it lets you use a single tool, it is powerful because it is a pluggable framework to pull in all kinds of tools for the model to use. You’ll be able to ask a BI question (”Show me our most important renewals for the coming quarter”) and then immediately act on it in another tool (”Email the main point of contact on the account to set up a check-in meeting”). Having all of these tools interact together inside of a single interface is combinatorially powerful. There is already a large ecosystem of tooling available and community-driven innovation is happening _fast_. 3. **Tech: **Except for MSFT, current BI vendors are not AI research labs. They are just not going to create better models or be the primary destination for all AI interactions within a company. ## My predictions I think that the BI workflow that has dominated for the past ~15 years is going to change significantly over the next 2. EDA will significantly migrate over to AI interfaces, enabled by MCP. I think this will be incredibly positive for all knowledge workers throughout a company. It will enable more users to create sophisticated analytics and will enable existing data practitioners to move significantly faster. I think this will be a headwind to many current BI vendors. BI is extremely sticky and this change isn’t going to happen overnight, but it will be a headwind. I think there is likely space for new players to innovate: to be the best place to aggregate and govern all of the artifacts built in this new workflow. I’ll return to this post after six months and see how my predictions are faring! --- --- title: "How to reduce BigQuery costs without compromising performance" description: "Learn how teams use dbt with BigQuery to reduce costs, improve performance, and scale analytics faster." url: "https://www.getdbt.com/blog/reduce-bigquery-costs" date: "2025-04-03" authors: ["Daniel Poppy"] categories: ["Learn"] --- # How to reduce BigQuery costs without compromising performance BigQuery’s fully managed, serverless model makes it a favorite for fast, scalable analytics—but its ease of use can come at a price. Without intentional design, teams often face unpredictable [BigQuery](https://cloud.google.com/bigquery?hl=en) bills driven by inefficient queries, poorly organized tables, and unchecked data growth. To keep BigQuery costs under control, you need to understand where spend comes from—and how to design your data workflows for efficiency. ## Understanding BigQuery's pricing model BigQuery’s costs break down into two core components: - **Compute**: You pay for the queries you run, based on the number of bytes processed. You can choose between **on-demand pricing** (pay-per-query) or **capacity-based pricing** (pre-purchased slots, like virtual CPUs). - **Storage**: Charged based on the amount of data stored in BigQuery tables. Active storage is priced per GB per month, with lower rates for long-term storage. There are also add-on costs for advanced features like BigQuery ML, BI Engine, and streaming inserts. 💡 **Tip**: You can monitor your project’s spend in real time using [Google Cloud billing reports](https://cloud.google.com/billing/docs/how-to/reports). Understanding these cost drivers is the foundation for strategic optimization. In this post, we’ll walk through practical ways to reduce BigQuery costs—highlighting real examples from [dbt](https://www.getdbt.com/product/what-is-dbt) users who’ve dramatically lowered spend using incremental models, better table design, and smarter query logic. **** ## Using incremental models to reduce processing costs One of the biggest cost-saving moves in BigQuery is adopting [incremental models with dbt](https://docs.getdbt.com/docs/build/incremental-models-overview). Rather than rebuilding entire tables every time a model runs, incremental models only process new or updated data—saving time and compute. This approach is ideal for large datasets that change frequently but only in parts (e.g., daily transactions, event logs, or application activity). ### Real results from teams using dbt and BigQuery - [**Bilt Rewards**](https://www.getdbt.com/case-studies/bilt-rewards-professional-services) cut $20,000 in monthly BigQuery costs by switching to efficient incremental models with dbt. - [**Enpal**](https://www.getdbt.com/case-studies/enpal) slashed their monthly data spend by **70%** after implementing dbt across their modern data stack. - [**Symend**](https://www.getdbt.com/case-studies/symend) decreased their daily warehouse usage by 70% while cutting data latency from 12 hours to 2 hours—all with incremental modeling. > > > — Ben Kramer ### Best practices for incremental models in BigQuery - **Use filters on both source and target tables** Apply `WHERE` clauses to narrow the scope of new data—and avoid scanning the entire table unnecessarily. - **Add clustering keys** to improve block pruning Columns like `user_id`, `event_date`, or `order_id` help BigQuery skip irrelevant data when processing. - **Use incremental predicates to reduce scan volume** Example: Instead of scanning the full target table, filter for just the last 2–3 hours. Yes, there’s a tradeoff between accuracy and cost—but it’s often worth it for lower-sensitivity datasets. - **Use [dbt’s merge strategy](https://docs.getdbt.com/guides/custom-cicd-pipelines?step=1)** when updates matter If your target table needs to track late-arriving records or changes, the merge strategy ensures updates are applied efficiently—without full refreshes. - **Bonus: Combine with CI/CD and testing **When you treat your models like software—with CI pipelines, version control, and tests—you unlock even more efficiencies. dbt makes this easy with built-in [testing](https://docs.getdbt.com/docs/build/tests), documentation, and scheduling. ## Design tables the minimize scan costs BigQuery charges based on how much data your queries scan—so the way you design tables has a direct impact on cost. Partitioning and clustering help reduce the amount of data BigQuery processes, making queries both cheaper and faster. ### Use partitioning to scan only what you need Partitioning splits your table into logical segments—typically by a `DATE` or `TIMESTAMP` column like created_at or event_date. This helps BigQuery scan only the relevant slice of data for each query. **** 💡 **Best practice**: If your data doesn’t include a time column, consider adding one upstream in your dbt models. ### Cluster by frequently filtered columns Clustering organizes data within each partition based on one or more columns (like user_id, region, or campaign_id). When you filter on these columns, BigQuery can prune irrelevant data blocks—reducing scan size and improving performance. **** 💡 **When to cluster**: Use it for high-cardinality columns that are frequently filtered—like `account_id`, `campaign_id`, or `region`. ### Combine partitions and clusters for compound savings For best results, use both: partition by time and cluster by filter-heavy fields. A marketing events table partitioned by `event_date` and clustered by `campaign_id` enables fast, targeted lookups—without scanning full partitions. ### Design with your queries in mind Use BigQuery’s [INFORMATION_SCHEMA](https://cloud.google.com/bigquery/docs/information-schema-intro) or query plan tools to analyze which columns are commonly used in filters, sorts, or joins. Then design partitions and clusters around those patterns. ### Use dbt to build efficient, scalable models The way you build and maintain data models in BigQuery has a major impact on performance and spend. [dbt makes it easier to manage transformations as modular, testable code](https://docs.getdbt.com/reference/resource-configs/bigquery-configs)—enabling cost-efficient practices like incremental processing, filtering, and reusability. - **Use incremental models to limit data scanned.** Instead of rebuilding entire tables, dbt can process only new or changed records—cutting compute and speeding up workflows. For large datasets, this alone can save thousands in query costs. - **Filter source and target data.** Narrow your model’s scope by applying filters on both sides of the transformation. This minimizes data scanned during incremental runs and keeps tables lean. - **Cluster by unique keys.** Apply clustering to columns used in filters (like user_id or event_date) to enable block pruning. This lets BigQuery skip over irrelevant data during processing. - **Use incremental predicates strategically.** There’s a tradeoff between cost and completeness. A rolling time window (e.g. “last 24 hours”) might miss late-arriving data but dramatically reduces bytes processed. ## Write cost-efficient queries In BigQuery, query design directly impacts your bottom line. Efficient SQL can dramatically reduce the amount of data scanned—saving both time and money. Here’s how to trim unnecessary compute from your queries: - **Avoid `SELECT *`.** Always specify the columns you need. Pulling all columns—even when unnecessary—forces BigQuery to scan more data. - **Filter early and often.** Apply `WHERE` clauses as soon as possible in your query logic. The sooner irrelevant rows are excluded, the less data gets processed. - **Use approximate functions** like `APPROX_COUNT_DISTINCT()` when perfect precision isn’t necessary. These reduce data scanned while delivering fast, directional insights. - **Pre-aggregate for reuse.** If your team repeatedly runs the same metrics, materialize them in summary tables or dbt models. This avoids recalculating expensive aggregations on raw data. - **Review high-cost queries regularly.** Use the BigQuery UI or [`INFORMATION_SCHEMA` ](https://cloud.google.com/bigquery/docs/information-schema-intro)to identify the queries consuming the most resources. A small refactor can often lead to outsized savings. 💡 **Tip: **Pair query optimization with [dbt’s model-level control](https://docs.getdbt.com/docs/build/groups#adding-a-model-to-a-group). Breaking complex transformations into modular steps makes it easier to spot inefficiencies and reuse logic across your project. ## Cut costs with caching and temp tables in BigQuery BigQuery offers built-in caching and flexible table options that can cut costs without changing your query logic. ### Take advantage of automatic caching BigQuery caches query results for 24 hours by default—at no additional cost. If an identical query runs again within that window (and the underlying data hasn’t changed), BigQuery serves the cached results instantly. To make the most of this: - Standardize common queries across teams - Avoid unnecessary query variation (e.g., column order, whitespace) - Use views to enforce consistent query patterns ### Use temporary or materialized tables For expensive queries that power multiple downstream analyses, consider writing results to: - **Temporary tables** – These exist only during the session and can support iterative development without long-term storage costs. - [**Materialized tables**](https://docs.getdbt.com/guides/create-new-materializations?step=1) – For repeated use, materializing results (e.g., with dbt) saves on reprocessing. You’ll pay once for the query, then reuse the results cheaply. This approach is especially helpful when: - Multiple teams access the same outputs - Queries require multiple joins or transformations - You’re powering dashboards or reports with tight SLAs ## Real-world success stories These aren’t just theoretical strategies—organizations using BigQuery and dbt together are seeing real, measurable cost savings and performance gains. ### Bilt Rewards By implementing efficient incremental models in dbt, Bilt reduced its BigQuery costs by $20,000/month. They also cut analytics spend by 80% after adopting the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), moving away from embedded BI tools and eliminating duplicated logic across the stack. [Read the case study --> ](https://www.getdbt.com/case-studies/bilt-rewards-professional-services) ### Symend Symend decreased daily warehouse credit usage by 70% with an incremental model strategy powered by dbt. This also slashed data latency from 12 hours to just 2 hours, improving timeliness for downstream teams. [Read the case study --> ](https://www.getdbt.com/case-studies/symend) ### Siemens With better table design and smarter data engineering practices, Siemens reduced dashboard maintenance costs by 90%, and dropped daily load time from 6 hours to just 25 minutes. [Read the case study --> ](https://www.getdbt.com/case-studies/siemens) ### AXS AXS used incremental models and modular dbt pipelines to speed up deployments by 50% and reduce maintenance effort by 40%. This let their data team ship faster without increasing BigQuery costs. [Read the case study -->](https://www.getdbt.com/case-studies/axs) **** ## Build fast, spend smart with BigQuery + dbt Managing BigQuery costs isn’t about cutting corners—it’s about building a smarter analytics foundation. With thoughtful strategies like incremental modeling, partitioned table design, query optimization, and smart caching, teams can significantly reduce spend without sacrificing performance. And when you combine [BigQuery’s scalable processing with dbt’s modular, testable transformations](https://www.getdbt.com/data-platforms/bigquery), cost-efficiency becomes a built-in part of your workflow: - **Reduce data scanned** by processing only new or changed records with [incremental models](https://docs.getdbt.com/guides/bigquery?step=1). - **Prevent rework and errors** with automated testing and CI/CD for SQL. - **Improve team velocity** by building reusable logic in a shared, governed layer. Whether you’re running mission-critical pipelines or scaling analytics org-wide, dbt helps you get the most from BigQuery—on your timeline and your budget. 🚀 Ready to see how dbt can help your team control BigQuery costs? [Try dbt](https://www.getdbt.com/signup/) or [book a demo](https://www.getdbt.com/contact/) today. ## BigQuery Cost Optimization FAQs **Is BigQuery free or paid?** BigQuery is a paid service, but Google offers a generous free tier: 10GB of storage and 1TB of query processing per month. Beyond that, you pay for two things: - **Compute**: Based on data scanned per query (on-demand) or via pre-purchased slots (capacity pricing). - **Storage**: Charged per GB, with discounts for long-term storage. **Why is BigQuery so expensive?** BigQuery’s serverless model is powerful, but cost control requires intention. Expenses climb due to: - Inefficient queries (e.g., `SELECT * `or poor filtering) - Lack of table partitioning and clustering - Rebuilding full tables instead of using incremental updates - Not leveraging caching or materialized results With strategies like [incremental models](https://docs.getdbt.com/docs/build/incremental-models), better table design, and dbt’s modular transformations, teams can dramatically reduce BigQuery spend. **Is BigQuery cheaper than Snowflake?** It depends on your workload. - [**BigQuery**](https://cloud.google.com/bigquery?hl=en) is often more cost-effective for intermittent, bursty workloads due to its pay-per-query model. - [**Snowflake**](https://www.snowflake.com/) may be better for consistent, high-volume processing thanks to its independently scalable compute. Both platforms can be efficient—with the right modeling, optimization, and governance. We’ve seen dbt help teams reduce costs across both. **Can I learn BigQuery for free?** Yes. Google Cloud offers free learning resources and usage credits: - 1TB of query and 10GB of storage monthly (free tier) - [Cloud Skills Boost](https://www.cloudskillsboost.google/) courses and hands-on labs - BigQuery [documentation and tutorials](https://cloud.google.com/bigquery/docs) - YouTube, blogs, and community forums for practical examples 💡**Pro tip:** Pair BigQuery practice with dbt to learn how real teams build and optimize pipelines in production. --- --- title: "Creating reliable data products with analytics engineering" description: "Explore how companies can build dependable data products through analytics engineering with the Analytics Development Lifecycle." url: "https://www.getdbt.com/blog/creating-reliable-data-products-with-analytics-engineering" date: "2025-04-02" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Creating reliable data products with analytics engineering ## The need for a mature analytics workflow Even with cloud scale, technological advances, and growing use of tools, many teams remain stuck with point solutions or fragmented workflows. Consider how, at most organizations: - Version control for ingestion pipelines is rare. - Testing and Service Level Agreements (SLAs) for dashboards are nearly non-existent. - Collaboration often depends on informal, manual processes. - Incident handling is ad hoc rather than structured. The result is "data products"—dashboards, models, reports—whose reliability and business fitness is difficult to prove or maintain. For analytics to create value at scale, mature workflows, not just point technologies, are required. #### Example Suppose a marketing manager needs customer conversion metrics updated daily. The data engineering team relies on ad hoc SQL scripts, run manually, and emailed as spreadsheets to the analytics team, who then create dashboards. If a schema changes or data quality issue arises upstream, errors propagate—conversion rates may be reported incorrectly for days before anyone realizes. There is no auditable process to roll back, test, or track changes, and the business loses trust in the analytics system. ## Requirements of a mature analytics workflow A truly reliable data product arises from a workflow embodying the following characteristics: - **Data and collaboration scale:** Can handle growing data volumes and more contributors without process breakdown. - **Accessibility:** Enables all relevant personas (engineers, analysts, decision-makers) to contribute. - **Velocity and agility:** Supports fast, iterative analysis with minimal process overhead. - **Correctness and validation:** Embeds automated tests to ensure data accuracy. - **Auditability:** Every change and result is reproducible and traceable. - **Governance:** Access, compliance, and usage policies are integrated from the outset. - **Criticality, reliability, resilience:** Data products can scale from experiments to business-critical use with confidence; error detection and recovery are systematic, not ad hoc. Imagine a retailer launching a flash sale campaign. The product team needs rapid insights on which channels drive real-time sales. With a mature workflow, campaign data ingestion can scale up quickly, analysis can be pushed to multiple analysts, findings are validated by automated tests, and any errors or schema changes are automatically detected, triaged, and communicated. Confidence in insights remains high, facilitating decisive action during the high-stakes campaign. ## The Analytics Development Lifecycle (ADLC) [The ADLC is a structured workflow modeled on proven software engineering practices](https://www.getdbt.com/resources/the-analytics-development-lifecycle). It emphasizes the continuous, collaborative, and iterative nature of analytics work and applies equally to every "artifact"—data pipelines, models, dashboards, or derived datasets. The ADLC consists of the following stages: 1. **Plan** 2. **Develop** 3. **Test** 4. **Deploy** 5. **Operate** 6. **Observe** 7. **Discover** 8. **Analyze** Each stage interacts in a loop, creating a cycle of improvement and adaptation, not a one-way flow. ### 1. Plan Planning sets the foundation for reliability and value. It includes: - Clarifying the business need or hypothesis driving the change. - Involving relevant stakeholders early. - Assessing downstream impacts of any changes (e.g., which dashboards depend on a specific model). - Designing for maintainability, data security, and stakeholder access. - Chunking large projects into small, manageable iterations. #### Example A financial analyst wants to introduce a "customer lifetime value" metric. The plan involves identifying required data sources, reviewing existing models for possible reuse, evaluating privacy implications, and setting up a process so marketing and product teams can provide feedback before the metric is finalized. ### 2. Develop Development is not just about technical coding; it instills collaboration and best practices: - All business logic is captured as code, regardless of interface (SQL, Python, visual tools). - Development environments are flexible; contributors can use tools suited to their workflow. - A style guide enhances consistency, making code maintainable by others. - Functionality and clarity are prioritized over premature optimization. - Code review by peers ensures robustness and knowledge sharing. - Vendor lock-in is minimized by favoring open standards. #### Example Two data modelers, using different interfaces, collaborate on the same data model representing "active users." Thanks to standardized code in source control, reviews, and a shared style guide, the team can efficiently merge improvements without confusion. ### 3. Test Testing is the ADLC's backbone for reliability: - **Unit tests:** Ensure each function or model behaves as intended. - **Data tests:** Validate that actual data conforms to logic and assumptions. - **Integration tests:** Catch issues arising from the combination of upstream and downstream components. Testing is mandatory before promotion to production and is run automatically as part of continuous integration. #### Example An update to the sales pipeline model triggers automated tests that check for referential integrity, realistic sales values, and non-breaking schema changes in downstream dashboards. Any failures block deployment, avoiding disruptions. ### 4. Deploy Deployment is automated, transparent, and safe: - Triggered by merging code into a main branch. - Handles environment promotion (dev → staging → production). - Does not create user-facing downtime. - Supports automated, safe rollback in case of unseen errors. #### Example When a new product categorization model is approved, deployment happens automatically upon merge. If an error is detected in production (e.g., a missing category), it is quickly reverted, minimizing business disruption. ### 5. Operate and 6. Observe Once live, the system is actively operated and observed: - Production systems are always-on, or have clear and minimal planned downtimes. - Resilience is designed in: error handling, monitoring, and automated remediation. - Incidents—data load failures, slow queries, stale reports—are detected and investigated before users notice. - Key metrics (such as uptime, freshness, latency) are tracked and guide process improvements. #### Example The analytics team for an e-commerce site is notified by monitors that yesterday's transactions failed to load due to an upstream API change. Automated incident tracking and on-call procedures ensure rapid diagnosis and resolution, with zero impact on downstream sales reporting. ### 7. Discover and 8. Analyze The final (and looping) stage is where business value is extracted. Discovery involves making all data assets (datasets, dashboards, metrics) easily searchable, accessible, and understandable—removing friction for both analysts and decision-makers. Analysis builds on these assets, using governed, trusted data to conduct investigations, answer ad hoc questions, iterate on hypotheses, and produce shareable, maintainable outputs. Key requirements: - Search and access to all governed data artifacts without bottlenecks. - Direct feedback, annotation, and improvement loops embedded in the tools. - Analysis outputs can themselves cycle back into the Plan stage for continued improvement. - Environments (development, staging, production) are transparent and selectable by users based on their needs. #### Example A supply chain analyst discovers a new anomaly in warehouse returns using a standardized, documented dataset. Her exploratory notebook, once validated and reviewed, becomes a maintained dashboard, with lineage and reproducibility guaranteed by the ADLC workflow. ## Stakeholders: collaboration across personas A mature analytics workflow recognizes that roles are flexible—individuals may put on different "hats" depending on the need: - **Engineer:** Builds reusable data pipelines and models. - **Analyst:** Explores data, validates hypotheses, produces recommendations. - **Decision-maker:** Consumes insights and acts upon them. The ADLC's greatest value emerges when these roles collaborate seamlessly within the same workflow and tooling. Organizational agility is maximized when hand-offs disappear, and individuals can transition between hats as projects demand. #### Example A small SaaS startup's product leader creates a quick usage metric, validates initial results, and, after peer review, productionizes it for quarterly board reporting—all within the ADLC framework, skipping no quality gates. ## Instituting the ADLC: principles and long-term value To create reliable data products, organizations must treat analytics systems as software systems—inherently collaborative, modular, testable, and auditable. This means: - Every artifact is versioned, tested, and documented. - Feedback loops are explicit, encouraging continuous improvement. - Errors are anticipated; processes for detection, mitigation, and communication are built-in. - SLAs for data products (availability, correctness, freshness) are defined, measured, and met. - Governance is not an afterthought, but intrinsic. Long-term, this creates data products with high trust, low maintenance overhead, and scalability as business stakes rise. ## Conclusion The analytics development lifecycle (ADLC) offers a definitive, end-to-end workflow for building mature, reliable, and value-generating data products. By adopting its principles—drawn from decades of software engineering experience—organizations can align people, processes, and tools, achieving both agility and governance. The path to data maturity is ongoing—a shared endeavor among practitioners, leaders, and technology providers. By consistently applying the ADLC, companies can close the gap between the promise of analytics and practical, dependable delivery of insights, driving better outcomes for every stakeholder. Learn more about analytics engineering best practices at [getdbt.com/blog](https://getdbt.com/blog) ## Analytics engineering FAQ ### What is analytics engineering? Analytics engineering is a discipline that combines software engineering principles with data analytics. It focuses on creating reliable, maintainable data products through structured workflows. Analytics engineering implements processes like version control, testing, automation, and governance to ensure data products (dashboards, models, reports) are accurate and trustworthy. It bridges the gap between raw data and actionable insights by applying software development best practices to analytics workflows. ### What do analytics engineers do? Analytics engineers build and maintain data pipelines, models, and other analytics infrastructure while ensuring reliability and scalability. Their key responsibilities include: - Creating reusable data pipelines and models using code - Implementing testing frameworks to validate data accuracy - Setting up version control for analytics assets - Automating deployment processes - Establishing monitoring systems to detect issues - Collaborating with analysts and decision-makers - Building discoverable, well-documented data assets - Ensuring governance and compliance requirements are met - Supporting the entire Analytics Development Life Cycle (ADLC) --- --- title: "How to reduce Snowflake costs without sacrificing performance" description: "Learn how teams use dbt to optimize Snowflake costs across compute, storage, and code—without losing performance." url: "https://www.getdbt.com/blog/reduce-snowflake-costs" date: "2025-04-02" authors: ["Daniel Poppy"] categories: ["Learn"] --- # How to reduce Snowflake costs without sacrificing performance [Snowflake’s usage-based pricing](https://www.snowflake.com/en/pricing-options/) gives teams flexibility—but without intentional management, it’s easy for costs to spiral. Most organizations overspend not because [Snowflake](https://www.snowflake.com/) is expensive by default, but because of inefficient patterns in how compute, storage, and data movement are managed. - **Compute** typically drives the largest share of spend. Each second a warehouse runs (even idle), it consumes credits. - **Storage** costs build up from retained tables, Time Travel history, and unmonitored staging areas. - **Data transfer** fees can surprise teams when data moves across clouds, regions, or external integrations. These factors don’t exist in silos. Poor data modeling or bloated storage can increase query complexity, which drives up compute. A lack of warehouse segmentation can lead to resource contention and slow performance—pushing teams to overprovision as a quick fix. In this post, we’ll break down practical strategies for reducing Snowflake costs across your stack—from warehouse and storage optimizations to query tuning, architecture decisions, and team practices. These aren’t theoretical tips; they’re based on real customer outcomes and best practices from teams using [**dbt and Snowflake together**](https://www.getdbt.com/data-platforms/snowflake) to build cost-efficient analytics workflows. ## Optimize warehouse configuration and usage Warehouse optimization is often the fastest—and most overlooked—path to Snowflake savings. By tailoring warehouse size, schedule, and scope to actual workload needs, teams can significantly reduce compute costs without degrading performance. ### Match warehouse size to workload Overprovisioned warehouses are one of the most common sources of overspend. Start by right-sizing based on query volume and complexity. **💡 Tip:** Use warehouse monitoring tools or Snowflake's query history to identify peak usage vs. idle time. ### Use auto-suspend and auto-resume settings Idle warehouses still burn credits. Configure auto-suspend after a short inactivity window—typically 2 to 5 minutes—and enable auto-resume to ensure processes don’t get blocked. ### Align warehouse schedules with actual usage Not all workloads run 24/7. Use orchestration tools (like [dbt](https://www.getdbt.com/product/what-is-dbt) or Airflow) to start and stop warehouses around business hours, especially for regional or time-zoned teams. ### Segment by workload type Sharing a single warehouse across dev, reporting, and data science creates performance bottlenecks—and wastes compute. Create purpose-specific warehouses to optimize sizing and concurrency for each use case. ## Optimize data storage Snowflake makes it easy to store massive volumes of data—but without a retention and organization strategy, storage costs can quietly compound. While storage is typically cheaper than compute, poor data hygiene can indirectly increase compute costs by slowing queries or bloating transformations. Here’s how to make storage work harder (and cheaper) for your team: ### Tier data based on value Not all data needs to stay in Snowflake forever. By adopting a tiered storage model, teams can preserve what matters and archive the rest. Consider moving cold data to object storage or a data lake, while retaining hot data—like metrics and model outputs—in Snowflake for speed. ### Tune Time Travel for the real world [Snowflake’s Time Travel](https://docs.snowflake.com/en/user-guide/data-time-travel) is powerful for recovery and auditing—but the default 90-day window might be overkill. Shortening retention on less critical tables can yield quick wins. **💡 Tip: **Set different defaults at the schema level to reflect each domain’s needs. ### Organize tables for performance and compression Storage and compute are connected. Better-organized data compresses more efficiently and supports faster queries—reducing both costs. Using clustering keys that align with your most common filter conditions pays dividends across your stack. ### Automate cleanup of temp and dev data Staging tables, temp models, and old test data often go unnoticed—but Snowflake charges for it all. Put automated cleanup jobs in place to track and remove these regularly. Tools like [dbt’s run-operation](https://docs.getdbt.com/reference/commands/run-operation) can help you automate cleanup scripts directly within your project. ## Improve query and code efficiency Poorly written queries are among the fastest ways to inflate Snowflake costs. By improving how data is transformed and accessed, organizations can significantly reduce compute usage—without compromising performance or insights. ### Start with better SQL hygiene A few targeted improvements can make a big dent in warehouse usage: - **Filter early and often** – Narrow scans with `WHERE` clauses before joining. - **Avoid Cartesian joins** – Explicit join conditions prevent runaway costs. - **Use SELECT with intention** – Only pull what’s needed; skip `SELECT *` - **Summarize when possible** – Aggregated views are faster and cheaper than raw detail. ### Modularize transformations with dbt Large, monolithic SQL scripts are fragile and hard to optimize. With dbt, you can break down complex pipelines into modular, testable models—each designed for a single purpose. Need to go faster? Use [incremental models](https://docs.getdbt.com/docs/build/incremental-models) to transform only new or changed records, reducing warehouse load. ### Test early to reduce reprocessing Silent errors in transformations can lead to expensive re-runs. Adding [automated testing](https://docs.getdbt.com/category/tests) catches issues before they escalate. Start with these built-in tests: - `not_null` - `unique` - `relationships` - `accepted_values` Then expand into custom assertions as your models mature. ### Materialize and cache repeat logic If a query (or subquery) is run frequently, [materialize](https://docs.getdbt.com/docs/build/materializations) the results. Use table, incremental, or ephemeral models depending on how often the data changes. This is especially valuable when paired with a [semantic layer](https://www.getdbt.com/product/semantic-layer)—one version of the logic, reused across your stack. ## Make architecture work for your budget You can optimize individual queries and warehouses all day—but unless your architecture supports those efforts, costs will keep creeping up. Structural choices like how you route workloads, manage scaling, and enforce limits play a critical role in long-term Snowflake spend. ### Balance performance with dynamic scaling For spiky workloads (think: end-of-day reports or campaign launches), consider multi-cluster warehouses with auto-scaling. This lets you serve concurrent queries efficiently—without paying for idle capacity the rest of the day. ### Enforce guardrails with resource monitors No team wants to be surprised by an end-of-month bill. [Snowflake’s resource monitors](https://docs.snowflake.com/en/user-guide/resource-monitors) let you: - Set spend thresholds (e.g., alert at 80%, suspend at 95%) - Automate alerts and take corrective action early - Keep non-critical workloads from consuming critical capacity ### Route workloads intentionally Different query types deserve different environments. Splitting out BI dashboards, ad hoc exploration, and data science into dedicated warehouses: - Minimizes contention - Optimizes performance per workload - Makes cost attribution and forecasting easier ### Use Snowflake's built-in accelerators (selectively) Snowflake offers features like [search optimization](https://docs.snowflake.com/en/user-guide/search-optimization-service) and [automatic clustering](https://docs.snowflake.com/en/user-guide/tables-auto-reclustering) to accelerate query performance. These can increase storage costs but often pay off when used for the right tables. Consider enabling these on: - High-volume tables with frequent filters (e.g., `WHERE user_id =`) - Slowly changing dimension tables - Models powering operational dashboards ## Build a culture of cost ownership You can optimize compute and storage—but sustainable Snowflake cost management comes down to people. The most successful teams treat cost awareness as a shared responsibility across engineering, analytics, and business stakeholders. ### Make costs visible—and attributable Visibility drives accountability. Use Snowflake’s [tagging](https://docs.snowflake.com/en/user-guide/object-tagging/introduction) and [query history](https://docs.snowflake.com/en/user-guide/ui-snowsight-activity) features to track usage by team, project, or environment. Then surface that data in internal dashboards to enable cost ownership. ### Foster a culture of shared learning Documentation and education go further than mandates. Host recurring “cost review” sessions where teams share wins, inefficiencies, and lessons learned. Encourage experimentation with techniques like: - Warehouse resizing - Query optimization - Staging table cleanup ### Make optimization continuous Treat cost reviews the way you treat product retros. A retail data team holds quarterly usage audits to: - Flag cost regressions - Revisit underperforming models - Reallocate resources as business needs evolve This process helped them maintain stable Snowflake costs even as data volume tripled over two years. ### Focus on cost-to-value, not just cutting spend Cost optimization isn’t always about spending less—it’s about spending smarter. Use metrics like: - Cost per insight delivered - Query cost per stakeholder group - Revenue influenced by data products This shifts the focus from blanket cuts to intelligent trade-offs. ## Cut costs, not capabilities, with Snowflake and dbt Managing Snowflake spend isn’t about sacrificing performance or putting limits on innovation. It’s about building smarter systems that scale efficiently—and that’s where [dbt](https://www.getdbt.com/product/dbt) comes in. Organizations that pair **Snowflake’s scalable compute** with **dbt’s modular, testable transformations** unlock a powerful model for cost control: - **Fewer redundant queries.** dbt models are reusable and version-controlled, so teams write once and use everywhere—reducing warehouse load and developer rework. - **More efficient pipelines.** Incremental models and optimized materializations ensure Snowflake only processes what’s new, cutting compute costs over time. - **Cleaner, more trustworthy data.** Built-in tests and documentation stop bad data from flowing downstream, preventing costly reprocessing and boosting confidence in outputs. - **Greater visibility.** [Column-level lineage](https://docs.getdbt.com/docs/explore/dbt-explorer-faqs#column-level-lineage), semantic consistency, and environment-aware execution help teams understand—and justify—how resources are used. The result? Teams can scale analytics workloads and meet SLAs without running up costs behind the scenes. But the biggest differentiator is **discipline**: organizations that treat data like code—with [CI/CD](https://docs.getdbt.com/docs/deploy/about-ci), testing, and review—are the ones who control their spend while accelerating delivery. With dbt and Snowflake working in tandem, cost efficiency becomes part of your architecture—not an afterthought. That’s how you deliver reliable, business-ready insights without overspending. ## Snowflake Cost Optimization FAQs **Can I use Snowflake for free?** Snowflake offers a free trial for new users to test the platform’s capabilities. After the trial period, Snowflake operates on a consumption-based pricing model, where you pay for compute, storage, and data transfers. While there’s no permanent free tier for production use, you can manage costs effectively by enabling auto-suspend, right-sizing warehouses, and applying cost optimization best practices. **How much does it cost to get Snowflake certified? ** Snowflake certification exams typically range from $175 to $375 USD, depending on the certification type and level. The SnowPro Core certification (foundational level) costs around $175, while specialty or advanced certifications are more expensive. Pricing may vary by region and promotions. **What is the minimum billing for Snowflake?** Snowflake bills compute by the second (with a 60-second minimum) when a warehouse is running. You’re only charged for the storage you use, making it easy to start with minimal cost. For organizations that need predictable billing, Snowflake also offers capacity-based pricing models with agreed-upon minimum commitments. **Why is Snowflake worth so much? ** Snowflake’s high valuation is due to its cloud-native architecture that separates compute from storage, delivering unmatched scalability and concurrency. Its ability to support multi-cloud deployments, handle massive datasets, and maintain consistent performance has driven rapid enterprise adoption. The consumption-based pricing model aligns with customer growth, creating predictable revenue streams—making Snowflake a leader in the modern data platform market. --- --- title: "What should a data analytics workflow look like?" description: "A modern analytics workflow is built like software: scalable, governed, and ready for business-critical decisions." url: "https://www.getdbt.com/blog/what-should-a-data-analytics-workflow-look-like" date: "2025-04-02" authors: ["Daniel Poppy"] categories: ["Insights"] --- # What should a data analytics workflow look like? A value-creating analytics workflow goes beyond ad hoc spreadsheet analysis. It must meet several core requirements: - Data scale: Analytical systems must handle large, variable volumes of data without manual adjustments. - Collaboration scale: The workflow should support small teams or thousands of contributors equally well—without coordination bottlenecks. - Accessibility: Different user personas (engineers, analysts, decision-makers) should participate as peers, not in silos. - Velocity: The workflow should facilitate rapid analysis without introducing excessive overhead or governance friction. - Correctness: Results must be trustworthy, with built-in mechanisms for validation and error detection. - Auditability: All transformations, analyses, and outputs must be reproducible, with tracked changes. - Governance: Data access and usage must be controlled to comply with internal rules and external regulations. - Criticality: The system should support both explorative work and mission-critical production needs, without requiring refactoring. - Reliability an resilience: The system should remain robust in the face of failures, minimizing both downtime and the impact of errors. **Example:** A retail company has seasonal sales spikes. During holiday campaigns, their data platform must seamlessly scale data pipelines and support concurrent collaboration between category managers, financial analysts, and data engineers—all while ensuring that promotional reports are correct, auditable, and accessible only to authorized personnel. ## The stakeholders: hats, not badges Every organization leveraging analytics has individuals playing three roles: ### 1. The engineer Designs and maintains reusable data assets (pipelines, models, metrics) that enable the broader analytics function. ### 2. The analyst Explores data, derives insights, and develops recommendations that inform business choices. ### 3. The decision-maker Acts on the data-driven recommendations and is responsible for directing business action based on insights. Note: Individuals can and should shift between these 'hats' as needed. In a mature workflow, this flexibility allows for reduced friction and fosters innovation. For example, an analyst should be empowered to adjust model transformations if needed, without waiting in a ticket queue for engineering. **Example: **At a SaaS startup, a product manager (decisionmaker) collaborates with a growth analyst (analyst) to evaluate a new feature's impact. When they discover that the relevant event data isn't yet available, the analyst applies their engineering skills to adjust the event pipeline, unblocking the analysis in hours instead of weeks. ## The Analytics Development Lifecycle (ADLC) [The Analytics Development Lifecycle (ADLC) draws directly from the software development lifecycle, recognizing that analytics systems are ultimately software systems](https://www.getdbt.com/resources/the-analytics-development-lifecycle). The ADLC applies to all analytical assets—pipelines, models, dashboards, and more—by guiding them through eight cyclical stages: ### 1. Plan Changes to analytics systems start with planning. This includes identifying business requirements, consulting stakeholders, anticipating downstream impacts, and planning for long-term maintenance. The goal is to clarify why a change is needed and how it will be safely and sustainably implemented. **Example:** A marketing team wants to understand conversion rates by new customer acquisition channels. The analyst gathers requirements, checks if existing models capture the necessary data, and consults channel owners on business definitions. ### 2. Develop Actual creation or modification of analytical assets happens here. This should be anchored in code—whether SQL, Python, or another open language—to maximize reproducibility, reviewability, and independence from any single tool. Best practices in this stage include following style guides, documenting changes, prioritizing readability, and ensuring code is generic and reusable where possible. **Example:** The analyst writes new transformations to classify historical leads by updated channel definitions. They use existing macros and test helpers to avoid duplicating code present elsewhere. ### 3. Test All changes must be thoroughly tested before release. This involves three broad types: - Unit Tests: Confirm logical correctness of new code. - Data Tests: Ensure the code behaves properly with actual data. - Integration Tests: Verify that changes don't break dependencies or complementary systems. Automated testing is mandatory—manual checks are not scalable or reliable. **Example:** Before merging, the analyst writes unit tests that check customer channel categorization logic, and data tests to validate that all records have a valid channel. Integration tests confirm that existing dashboard queries referencing this field continue to function. ### 4. Deploy Moving changes into production should be automated and pain-free. Deployments must be triggered from source control, happen without causing user interruptions, and include automated rollback in case of failure. **Example:** Once all tests pass and code is peer-reviewed, merging to the main branch triggers an automated deployment. Smoke tests post-deployment quickly catch unexpected edge cases, and any failures auto-revert to the previous version. ### 5. Operate and observe Once in production, the system must be continuously monitored. This includes system health, data quality, and user-impacting issues. Mature systems detect and resolve issues quickly—ideally before users notice. Key metrics (uptime, throughput, error rates) must be defined and measured. Incident response processes and on-call rotations ensure rapid remediation. **Example:** Post-launch, data completeness metrics are monitored; if a channel is found with missing assignment, an alert triggers so engineering and analytics teams can resolve it quickly. ### 6. Discover Users must be able to easily find existing data assets—tables, dashboards, reports—via simple interfaces, without informal bottlenecks or reliance on internal networks. **Example:** A new analyst joins the team and, using the platform's universal search bar, locates all assets related to "customer acquisition," including source models, transformation code, and trusted dashboards. ### 7. Analyze The real business value of analytics is created here, through exploration and interpretation of data. Mature systems allow users to conduct exploratory analysis in-place, and as insights mature, fold findings back into the broader system via the Plan and Develop stages. Feedback and questions from users should spark continuous improvement cycles. **Example: **The analyst uses the discovered data assets to identify a surprising drop in conversion for a newly launched channel. They share interim results, gather feedback from marketing, and formalize findings into a new dashboard, which then undergoes its own cycle of testing and deployment. ## The loop: continuous improvement The ADLC is intentionally cyclical. Any insight or change at the Discover & Analyze phase should motivate a new Plan phase to improve or modify existing analytical assets. This feedback flow is crucial—mature analytics systems improve over time, informed by both system monitoring and user experience. ## Hypothetical end-to-end scenario Imagine a financial services company seeking to reduce fraudulent account signups: 1. Plan: Compliance, risk, and data teams define requirements for a new fraud detection model. They involve customer support (a key stakeholder) to align on business needs and downstream notification requirements. 2. Develop: Data engineers extend existing user event pipelines, while analysts prototype new feature engineering logic within reusable models and scripts. 3. Test: All components are unit tested, data validated on a staging environment, and integration tests confirm that fraud alerts don't erroneously trigger notifications to legitimate users. 4. Deploy: Following code review and merge, the full set of changes is deployed automatically, with no downtime. 5. Operate & Observe: Engineers set up real-time monitoring for false positives/negatives. An incident is quickly detected when a legitimate signup is flagged, leading to a rollback and targeted fix. 6. Discover & Analyze: Analysts explore fraud patterns in new data assets, sharing findings with compliance and operations, and suggesting further refinements. 7. The loop restarts as these ideas go back to planning for the next iteration. ## Requirements for the discover and analyze phase A mature workflow must guarantee: - Frictionless discovery of assets and data lineage - In-place operations—no downloading or emailing files - Built-in feedback mechanisms for every analytical artifact - Transparent processes for requesting/approving access - Provenance and change tracking for all data and assets - Flexible environment selection (dev, staging, prod) - Minimal exposure to technical implementation details for end users ## Conclusion A mature analytics workflow is not a luxury—it's a necessity for any organization that seeks to build trust in data, act quickly on insights, and scale analytics impact across the business. The Analytics Development Lifecycle (ADLC) provides a practical, software-inspired framework that integrates best practices at every layer of the stack, among all participants. Fully realizing this vision will require both iterations in process and continued evolution of analytical tooling. By embracing clear workflow stages—Plan, Develop, Test, Deploy, Operate, Observe, Discover, and Analyze—and breaking down organizational silos, companies can achieve analytics that are not only fast and flexible, but also governed, reliable, and value-generating. ADLC is the new benchmark for modern analytics. Organizations that embrace this approach will have a decisive edge—not just in technological capability, but in business adaptability and innovation. Frequently asked questions about analytics workflows ### What is an analytics workflow? An analytics workflow is a structured approach to processing and analyzing data that goes beyond ad hoc spreadsheet analysis. A mature analytics workflow must handle large volumes of data, support collaboration at scale, remain accessible to different user types (engineers, analysts, decision-makers), deliver insights quickly, ensure correctness and auditability, maintain proper governance, support both exploratory and mission-critical needs, and provide reliability even when facing system failures. ### What is an analysis workflow? An analysis workflow is the structured process that guides how data is transformed into actionable insights. In mature organizations, this follows the Analytics Development Lifecycle (ADLC), which applies software development principles to analytics. The workflow involves planning based on business requirements, developing analytical assets using code, testing for correctness, deploying changes safely, operating and monitoring systems, discovering existing assets, and analyzing data to create business value. ### What are the steps of workflow? While traditional workflows may have 5 steps, the Analytics Development Lifecycle (ADLC) presents a more comprehensive 8-stage approach: 1. Plan - Identify requirements and consult stakeholders 2. Develop - Create or modify analytical assets using code 3. Test - Verify correctness through unit, data, and integration testing 4. Deploy - Move changes to production through automated processes 5. Operate & Observe - Monitor system health and data quality 6. Discover - Find existing data assets through simple interfaces 7. Analyze - Explore and interpret data to create business value 8. Loop back to planning for continuous improvement ### What is the actual workflow of data analysis? The actual workflow of data analysis follows a cyclical pattern where insights drive improvements. Starting with planning, analysts define requirements with stakeholders. They then develop solutions using code, thoroughly test for correctness, and deploy changes to production environments. Once live, systems are monitored for performance and data quality. Users discover relevant data assets through search interfaces, then analyze this data to extract business insights. These insights often identify new opportunities or improvements, beginning the cycle again with planning for the next iteration. The most effective workflows support three key stakeholder roles—engineers who build data assets, analysts who derive insights, and decision-makers who act on recommendations—while allowing individuals to shift between these roles as needed to reduce friction and foster innovation. --- --- title: "Improving analytics quality with the analytics development lifecycle" description: "Treat analytics like software. The Analytics Development Lifecycle boosts data quality, trust, and speed—at scale." url: "https://www.getdbt.com/blog/improving-analytics-quality-with-the-analytics-development-lifecycle" date: "2025-04-02" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Improving analytics quality with the analytics development lifecycle For years, analytics teams have borrowed superficially from software engineering: implementing source control for data pipelines, adding basic tests, or requiring code reviews. While these practices add value, they typically remain confined to the most technical segments of the analytics stack, such as data transformation. Critical processes like ingestion, analysis, and dashboarding often lack the rigor required to deliver business-critical insights with confidence. The consequences are familiar to most data leaders: - Eroded trust in outputs: Business users question dashboards when unexplained changes or errors appear - Decision-making bottlenecks: Teams waste time investigating metric discrepancies rather than driving action - Scaling challenges: As data teams grow, inconsistent processes breed duplicated efforts, conflicting logic, and longer development cycles These are fundamentally process problems, not technology problems. The solution requires a lifecycle approach that treats the entire analytics workflow as a production software system—one subject to continuous improvement, quality assurance, and collaborative development practices. ## The Analytics Development Lifecycle (ADLC): A strategic framework [The Analytics Development Lifecycle (ADLC) provides a structured approach for developing, deploying, and managing analytics assets with consistent quality and operational efficiency](https://www.getdbt.com/resources/the-analytics-development-lifecycle). Adapted from the Software Development Lifecycle (SDLC), this framework addresses the unique requirements of analytics workflows while maintaining enterprise-grade standards. The ADLC encompasses six core phases: 1. Plan 2. Develop 3. Test 4. Deploy 5. Operate & Observe 6. Discover & Analyze This lifecycle operates iteratively: insights from analysis inform future planning, and improvements flow continuously through production deployment. Each phase establishes clear expectations and proven practices that support cross-functional collaboration, complete traceability, and production reliability. ## Requirements for enterprise analytics maturity A well-implemented ADLC addresses several critical operational dimensions: - Data scale: Maintains performance and reliability across varying data volumes - Team scale: Supports individual contributors through hundreds of team members with robust collaboration workflows - Accessibility: Enables meaningful participation from engineers, analysts, and business decision-makers - Development velocity: Minimizes time from concept to insight without introducing costly handoffs - Data correctness: Ensures outputs are accurate, testable, and dependable for business decisions - Auditability: Provides complete change tracking and reproducibility for stakeholder confidence - Governance: Integrates security and compliance requirements throughout the development process - Use-case flexibility: Supports both exploratory analysis and mission-critical reporting - Operational resilience: Enables rapid recovery from failures while minimizing business impact While analytics organizations often structure teams around "data engineer," "analyst," or "business user" roles, effective ADLC implementation recognizes that these functions are fluid. The framework defines three core personas that may be distributed across job titles or embodied within individual contributors: - Engineer: Develops reusable data pipelines, models, and metrics infrastructure - Analyst: Leverages data assets to generate insights, reports, and recommendations - Decisionmaker: Consumes analytical outputs to drive business strategy and operations The ADLC explicitly encourages role flexibility, enabling team members to transition seamlessly between analyzing data patterns, proposing infrastructure changes, and interpreting insights for business impact. Tooling and workflow design should facilitate this adaptability rather than enforce artificial boundaries. ## ADLC implementation: Phase-by-phase quality improvement ### 1. Plan The planning phase establishes the foundation for all downstream work through clear business alignment, stakeholder identification, impact analysis, and technical design. **Implementation example:** A marketing analyst identifies requirements for enhanced customer segmentation to improve campaign targeting effectiveness. Rather than developing an isolated solution, the analyst documents the business case, collaborates with data engineering on required schema changes, and evaluates impacts on existing dashboards and reporting workflows. #### Critical Success Factors: - Document comprehensive business rationale and success metrics - Engage relevant stakeholders during requirements definition - Design validation approaches and edge case handling upfront - Define migration paths for changes affecting existing assets - Plan for ongoing maintenance, access control, and documentation ### 2. Develop The development phase implements planned changes using version-controlled code, whether SQL, Python, or declarative configuration languages. **Implementation example:** A data engineer implements customer segmentation logic within a version-controlled repository, following established coding standards, writing maintainable and documented code, and extending shared libraries to minimize duplication. #### Critical success factors: - Maintain human-readable code as the authoritative source of truth - Implement tool-agnostic solutions where possible to minimize vendor lock-in - Provide flexible development environments while enforcing consistent standards - Require comprehensive peer review for all production changes ### 3. Test Production analytics assets require comprehensive automated testing that mirrors software engineering standards. The ADLC distinguishes between three essential testing categories: - Unit tests: Validate business logic independent of underlying data - Data tests: Assert expectations against production data samples - Integration tests: Verify system-wide functionality, including cross-asset dependencies **Implementation example: **Before deploying updated revenue calculation logic, the team implements data tests validating that revenue values remain positive and integration tests confirming that dependent dashboards render correctly with updated calculations. #### Critical success factors: - Enforce testing requirements for all production deployments - Implement continuous integration (CI) pipelines that execute comprehensive test suites - Establish and monitor test coverage metrics to ensure meaningful protection ### 4. Deploy Deployment processes must be automated, reliable, and reversible. Manual deployment steps introduce risk and bottlenecks, while automated rollback capabilities ensure rapid recovery from production issues. **Implementation example: **Following successful peer review and testing, an engineer merges changes to the main branch, automatically triggering deployment pipelines that promote changes to production. If regression testing identifies issues, automated processes execute immediate rollbacks to previous stable states. #### Critical success factors: - Implement branch-based environment management (development, staging, production) - Automate all deployment and rollback procedures - Design zero-downtime deployment capabilities for critical systems ### 5. Operate and observe Production environments require continuous monitoring and proactive incident response. Data quality, system uptime, and processing performance must be tracked with automated alerting for rapid issue resolution. **Implementation example:** When monitoring systems detect a failed nightly sales data load, automated alerts notify the responsible team immediately. Investigation traces the failure to an upstream schema change, enabling rapid correction and redeployment. Automated communications update business stakeholders on both the issue and resolution. #### Critical success factors: - Maintain high-availability standards for business-critical data systems - Implement comprehensive monitoring with automated alerting for rapid issue detection - Establish clear incident response procedures and regularly review operational metrics - Continuously improve documentation and runbook procedures ### 6. Discover and analyze This phase encompasses both data asset discovery and analytical work performed using those assets. Many organizations lack maturity in this area compared to their engineering-focused infrastructure. **Implementation example:** A product manager investigating declining user engagement can efficiently search for certified datasets on user activity, understand data lineage and quality, perform exploratory analysis, and—upon developing actionable insights—create reusable dashboards subject to full ADLC governance. #### Critical success factors: - Provide intuitive search and discovery interfaces for curated, documented data assets - Enable collaborative analysis with integrated feedback and review mechanisms - Implement transparent access controls that maximize appropriate self-service capabilities - Apply version control, testing, and maintenance standards to key analytical artifacts ## Real-world implementation scenarios ADLC adoption delivers measurable improvements in quality, velocity, and organizational trust regardless of company size or industry. **Scenario 1: Accelerating critical reporting issues** An enterprise discovers data discrepancies in quarterly revenue reports during board preparation. Traditional processes require tickets to data engineering with unclear timelines and lengthy email investigations. With ADLC implementation, business analysts can trace data lineage, propose corrections within core models, and execute automated testing and deployment workflows. Issues resolve in hours rather than weeks, maintaining stakeholder confidence and decision-making velocity. **Scenario 2: Enabling rapid innovation with compliance** A growth-stage company requires rapid iteration on customer segmentation while maintaining strict privacy compliance. ADLC enables development in isolated branches with code-enforced access controls. All changes undergo peer review and automated testing for both business logic and regulatory compliance before deployment. If production issues arise, automated rollbacks minimize exposure and stakeholder impact. **Scenario 3: Scaling through organizational growth** An enterprise managing rapid acquisition growth faces the challenge of integrating multiple data teams. ADLC provides common workflows for building, testing, deploying, and documenting assets across teams. Centralized, reusable logic prevents duplication while distributed teams contribute specialized enhancements. Success is measured through increased usage and trust in certified datasets across business units. ## Implementation strategy for data leaders - **Incremental adoption:** Begin with rigorous version control and testing for core pipelines, then extend mature practices to dashboards, notebooks, and broader workflows - **Organizational alignment:** Foster culture change that empowers team members to contribute across traditional role boundaries while removing unnecessary process bottlenecks - **Strategic tooling decisions:** Select platforms supporting reproducible workflows, automated testing, cross-functional collaboration, and asset discoverability - **Continuous improvement:** Build feedback collection and operational review processes into every ADLC phase Learn more about analytics workflows and best practices at [getdbt.com/blog](https://getdbt.com/blog). --- --- title: "Data integration: The 2025 guide for modern analytics teams" description: "A practical guide to modern data integration: techniques, architectures, tools, and why dbt is key to building scalable pipelines." url: "https://www.getdbt.com/blog/data-integration" date: "2025-04-02" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data integration: The 2025 guide for modern analytics teams Analytics teams continue to spend valuable time reconciling data that should already be accessible. Siloed systems, inconsistent logic, and fragile pipelines introduce delays at every stage of analysis. As software as a service (SaaS) adoption accelerates and real-time demands increase, the analytics environment has become fragmented across cloud platforms, tools, and teams. Some approaches to data integration aren’t keeping pace with today’s requirements. Rigid, monolithic pipelines struggle with schema drift, slow iteration, and limited observability. Modern teams need data integration workflows that are modular, testable, and built to adapt. Analytics engineers are responsible for delivering timely, accurate insights from data that is constantly changing in structure, volume, and source. To meet this demand, they need integration systems that can adapt quickly, support collaboration, and maintain trust in every downstream output. In this article, we’ll look at the right way to build reliable and scalable data integration pipelines to produce analysis-ready outputs - and the tools you can use to implement them quickly and securely. ## What is data integration? At its core, data integration is the process of combining data from multiple sources into a single, unified view. The goal? Make data accurate, fresh, and query-ready, no matter where it comes from or how often it changes. A modern data integration pipeline flows through five core stages: - Source - Ingest - Store - Transform - Consume Teams start by ingesting raw data via APIs, logs, or [prebuilt connectors](https://www.fivetran.com/blog/database-connector-setup-a-step-by-step-guide). That raw data lands in a centralized platform, such as a cloud warehouse or [lakehouse](https://www.databricks.com/blog/2020/01/30/what-is-a-data-lakehouse.html), where it is transformed into structured, business-ready models. These models power everything from dashboards to AI/ML pipelines. While this process appears static, it rarely stays that way. Organizations frequently update their source systems, onboard new tools, and adapt to evolving data structures. These shifts introduce complexity that static pipelines can’t handle alone. To keep pace, [analytics engineers](https://www.getdbt.com/blog/what-is-analytics-engineering) build repeatable, testable, and modular integration workflows that are resilient to change. When designed well, these pipelines reduce manual cleanup, maintain high [data quality](https://www.getdbt.com/blog/building-a-data-quality-framework-with-dbt-and-dbt-cloud), and provide the foundation for fast, reliable, and AI-ready insights. ## Core data integration techniques: Practical insights for 2025 Modern data integration isn’t just about data movement—it’s about **choosing the right architecture** for your latency, scale, and governance needs. Each technique has trade-offs in complexity, performance, and flexibility. There’s no one-size-fits-all solution. The key is selecting the method that aligns with your **data shape**, **update frequency**, and **use case**. Here are the most common techniques used in today’s modern data stacks: ### ETL vs. ELT Both [ETL (Extract-Transform-Load) and ELT (Extract-Load-Transform)](https://www.getdbt.com/blog/etl-vs-elt) move data from disparate sources into a centralized analytics environment. The difference lies in _when_ the data transformations happen, influencing scalability, cost, and agility. - In **ETL**, data is transformed _before_ it’s loaded—typically on a standalone ETL server. This made sense when compute was expensive and storage was limited. - In **ELT**, raw data is loaded _first_, then transformed inside the warehouse using native SQL and cloud compute. This shift gives teams more control, flexibility, and speed—especially for iterative modeling and analytics. ETL still has a place in regulated industries or legacy systems. But in most modern stacks, ELT is the standard. It’s cheaper to run, easier to scale, and better aligned with tools like [dbt](https://www.getdbt.com/product/dbt) that bring software engineering best practices to analytics workflows. ### Change data capture (CDC) and streaming [Change data capture (CDC)](https://docs.getdbt.com/blog/change-data-capture) supports near-real-time pipelines by detecting and synchronizing source system changes as they occur. - **Log-based CDC** reads directly from the database’s transaction logs, offering a low-latency, low-overhead approach. - **Trigger-based CDC** emits change events from within the application when direct log access is unavailable. This added complexity is justified when data freshness for near-real-time scenarios is essential. CDC is especially useful in personalization, fraud detection, or operational reporting, where a delayed view can lead to poor decisions or missed opportunities. To make CDC pipelines reliable, design for exactly-once delivery. Handle schema changes up front. Transform events into [incremental models](https://docs.getdbt.com/docs/build/incremental-models) within the warehouse to maintain clean and easy-to-test downstream logic. ### Data virtualization and federation Data virtualization allows teams to query data across systems without actually moving it. This is useful when building quick proofs of concept or when working with data that cannot be moved due to compliance or privacy requirements. That said, data virtualization poses considerable challenges: - Live queries can introduce latency and performance risk - Joins across systems are often fragile and slow - Permissions don’t always flow downstream, complicating governance If the data needs to be queried regularly, joined with other sources, or analyzed deeply—it’s almost always better to load it into a centralized platform like a lakehouse or warehouse. Save virtualization for **edge cases**, not your core pipeline. ## Common architectures and patterns Once you’ve selected your integration techniques, the architecture you build around them determines how well your data pipelines scale, how your team collaborates, and how durable your models are in the face of change. Here are three of the most common data integration architectures in practice today: ### Batch hub-and-spoke The batch hub-and-spoke architecture uses a central database as a “hub” that receives scheduled extracts from various source systems ("spokes"). The hub processes the data and pushes curated outputs downstream to business intelligence (BI) tools or reporting marts. It remains common in environments where source systems are on-prem or where compliance requires tight control over scheduling. However, the pattern has its limits. The rigid nature of batch workflows often results in lengthy refresh windows, making real-time analytics challenging or impossible to achieve. As jobs grow, monolithic ETL scripts become harder to debug or extend. Without modularity, parallel development is constrained, and even small changes can trigger full pipeline reruns. As modern workloads demand agility, batch architectures are increasingly giving way to real-time or ELT-based designs. ### Cloud warehouse or lakehouse as the central hub In this model, raw data is ingested directly into scalable cloud platforms like [BigQuery](https://cloud.google.com/bigquery), [Snowflake](https://www.snowflake.com/en/), or [Databricks](https://www.databricks.com/). All transformations happen inside the platform, using **native compute **— which is ideal for ELT workflows. This setup aligns well with ELT workflows and supports diverse use cases across analytics, machine learning, and real-time reporting. That said, teams that don’t manage compute carefully may face ballooning costs. Without clear conventions, ad-hoc transformations can diverge across teams, creating inconsistencies in how key metrics are defined. Access control and metadata must be explicitly configured, as these platforms don’t inherit the role-based security models of traditional enterprise systems. To manage complexity, many teams pair their warehouse with a **semantic layer**—a single source of truth for metrics, dimensions, and business logic that powers consistent insights across tools. **** ### Semantic layer: the new standard As self-service BI grows, and more teams work across multiple tools, the semantic layer becomes essential. It lets you define metrics once — like “customer churn” or “monthly recurring revenue” — and reuse them across dashboards, notebooks, or AI applications. But like any standardization effort, it requires: - **Cross-team agreement** on definitions before modeling - **Structured governance** to prevent version drift - **Upfront investment** to ensure performance at scale If implemented poorly, the semantic layer can introduce performance lags and user confusion. But when done right? It unlocks faster decision-making, better data governance, and a truly scalable data foundation. ## Benefits of effective data integration Modern analytics teams don’t just move data — they scale insights, streamline operations, and reduce risk. Effective data integration unlocks value across five core dimensions: - **Get insights faster:** Clean, query-ready data enables analysts to answer questions quickly and iterate faster. - **Improved data quality:** Well-modeled pipelines, backed by automated [data testing](https://docs.getdbt.com/docs/build/tests) and lineage tracking, help teams catch issues early and build trust in the results. - **Scalability and elasticity:** Modern ELT platforms and cloud-native tools scale dynamically as your data grows. No rebuilds. No rewrites. No bottlenecks. - **Cost optimization:** Efficient integration workflows reduce redundant logic and cut down on compute waste. A [Forrester study](https://aws.amazon.com/blogs/publicsector/forrester-study-examines-data-integration-roi-for-public-sector-organizations/) found that public-sector organizations achieved a 33% return on investment (ROI) over five years after adopting modern integration systems. - **Stronger collaboration and governance:** Version-controlled models and shared logic create consistency across teams, enabling faster delivery and more auditable workflows. It’s analytics, built like software — because it should be. ## Best practices for successful data integration Even the best tools and architectures fall short without the right operational discipline. Successful data integration depends as much on process as it does on platforms. The practices below help teams build pipelines that are reliable from the start and resilient over time. ### Define ownership and data contracts early Every model in your pipeline should have a clear owner—someone accountable for maintaining logic, reviewing updates, and communicating upstream changes. Enter: [**data contracts**](https://atlan.com/data-contracts). These are explicit agreements between producers and consumers about schema, freshness, and reliability. Contracts reduce friction, prevent silent breakages, and make debugging faster when things go sideways. ### Build incrementally with tests and alerts Start small. Scale smart. [**Incremental models**](https://docs.getdbt.com/docs/build/incremental-models) only process new or updated records, which: - Cuts down runtime and compute costs - Reduces strain on source systems - Speeds up feedback loops Pair this with automated testing and alerting to catch bad data before it hits dashboards. Need a primer? [See the dbt docs on incremental models →](https://docs.getdbt.com/docs/build/incremental-models) ### Enforce version control and CI/CD Treat your analytics code like application code. [Version-controlled models](https://docs.getdbt.com/docs/cloud/git/version-control-basics) allow teams to experiment safely, review changes, and roll back when needed. With [Continuous Integration and Continuous Deployment (CI/CD) pipelines](https://docs.getdbt.com/guides/custom-cicd-pipelines?step=1) in place, every commit runs tests, checks dependencies, and confirms that outputs meet expectations. ### Monitor lineage, freshness, and spend Integration doesn’t stop at deployment. Use observability tools (like [dbt Catalog](https://www.getdbt.com/product/dbt-catalog), [Monte Carlo](https://www.montecarlodata.com/), or [Datafold](https://www.datafold.com/)) to: - Monitor lineage across sources, models, and outputs - Track data freshness and update frequency - Audit resource usage and identify cost leaks You can’t improve what you can’t see — so bake visibility into your integration workflows from day one. ## Evaluating data integration tools in 2025 When evaluating data integration tools in 2025, look beyond UI and licensing costs. Your tools should support scalable, observable, and trustworthy data workflows. Here’s what to prioritize when evaluating integration platforms: - **Prebuilt connectors** Support for common databases, cloud warehouses, SaaS platforms, and APIs to reduce custom code and accelerate onboarding. - **Orchestration and scheduling** Built-in tools (or integrations with tools like [Kestra](https://kestra.io/) or [Airflow](https://airflow.apache.org/)) to handle dependency management, trigger runs, and monitor pipeline flow. - **Testing and observability hooks** Native support or easy integration with testing frameworks and data observability platforms like [dbt tests](https://docs.getdbt.com/docs/build/tests), [Monte Carlo](https://www.montecarlodata.com/), or [Datafold](https://www.datafold.com/). - **Metadata and lineage tracking** Integration with your [data catalog](https://atlan.com/what-is-a-data-catalog/) or [semantic layer](https://www.getdbt.com/product/semantic-layer/) to trace transformations and understand downstream impact. - **Schema enforcement and data contracts** Support for [data contracts](https://atlan.com/data-contracts/), schema validation, and policy enforcement to ensure reliable, governed data flows. No tool will check _every_ box—but focusing on capabilities that support your architecture, governance model, and team workflows will pay dividends. ## dbt: The transformation layer for modern data integration Let’s cut to it: dbt is the standard for transforming raw data into trustworthy, analytics-ready models. It’s the glue between ingestion and insights—and the reason many data teams are finally sleeping better at night. dbt brings [software engineering best practices to SQL](https://www.getdbt.com/resources/the-analytics-development-lifecycle), letting teams define, test, and deploy transformations as modular, version-controlled code. Out of the box, dbt supports the key capabilities analytics engineers need to maintain high data quality at scale: - [Version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) and [CI/CD](https://docs.getdbt.com/docs/deploy/continuous-integration) for safe, peer-reviewed deployments - Modular [SQL modeling](https://docs.getdbt.com/docs/build/sql-models) that encourages reusable, testable transformations - [Built-in tests](https://docs.getdbt.com/docs/build/data-tests) and assertions to catch issues before they reach dashboards - [Lineage documentation](https://docs.getdbt.com/docs/collaborate/explore-projects) to trace data from source to insight - [Integrated development workflows](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/ide-user-interface) that reduce time to production and simplify debugging If you’re investing in scalable, governed analytics, dbt isn’t optional. It’s the transformation backbone your team needs to build clean, trustworthy pipelines—without slowing down. Ready to get started? [See dbt in action today.](https://www.getdbt.com/contact) ## Data integration FAQs **What is data integration?** Data integration is the process of combining data from multiple sources into a single, governed view that is accurate, fresh, and analytics-ready. Modern pipelines typically move through five stages: Source → Ingest → Store → Transform → Consume. **How does the data integration process work from source identification and extraction through transformation, loading, and ongoing synchronization?** - Identify sources and owners; define data contracts for schemas, freshness, and SLAs. - Ingest via connectors, APIs, files, or logs (batch or CDC/streaming). - Land raw data in a centralized platform (warehouse or lakehouse) to preserve lineage. - Transform into business-ready models with modular code, tests, and incremental processing. - Orchestrate and deploy with CI/CD; monitor lineage, freshness, and costs. - Keep data synchronized with CDC/streaming for low-latency use cases; use alerts to detect breakages. **When should an organization choose ELT over ETL for data integration, and what are the trade-offs between the two approaches?** - Choose ELT when using cloud warehouses/lakehouses, you need fast iteration, scalable compute, and flexible modeling. Pros: agility, scalability, simpler handling of schema drift, lower ops overhead. Cons: risk of cost sprawl and inconsistent logic without governance. - Choose ETL when regulations or legacy constraints require transforming before data lands in the target, or when targets can’t handle heavy compute. Pros: tighter control on what enters the target. Cons: slower iteration, less scalable, harder to maintain. - Many teams standardize on ELT and add governance (semantic layer, tests, contracts) to manage trade-offs. **How does using a virtual mediated schema with wrappers/adapters enable real-time data integration compared to traditional ETL into a data warehouse?** - Virtualization (mediated schema + wrappers) provides a unified logical view across sources and pushes queries to them, avoiding data movement and enabling near-real-time access—useful for quick prototypes or data that cannot be moved. - Trade-offs: variable latency, fragile cross-source joins, and governance/permissions gaps. - Traditional ETL/ELT loads and transforms data into a central store, adding latency but delivering better performance, reliable joins, consistent governance, and reusability for recurring analytics. Use virtualization for edge cases, not core analytics. **How can data integration tools handle schema evolution and DDL changes during CDC without disrupting pipelines?** - Use contract-first schemas with versioning; prefer additive changes and deprecation windows for breaking changes. - Stage raw events separately, then map to canonical models; allow flexible staging (e.g., capturing unexpected columns) and apply strict schemas downstream. - Automate DDL propagation and schema reconciliation; maintain a schema registry or metadata layer. - Ensure idempotent, exactly-once processing with ordering keys and robust retry/rollback. - Build incremental, testable transformations; add tests/alerts for missing columns, type mismatches, and null spikes; monitor lineage and freshness. - For major changes, run dual-writes/dual-reads during cutovers and backfill to align history. --- --- title: "Iceberg?? Give it a REST" description: "The new abstraction that changes nothing... and everything." url: "https://www.getdbt.com/blog/iceberg-give-it-a-rest" date: "2025-03-30" authors: ["Anders Swanson"] categories: ["Learn"] --- # Iceberg?? Give it a REST _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/iceberg-give-it-a-rest). _ The analytics engineering landscape is shifting beneath our feet as our familiar data warehouse coalesces into the data engineer's lake house—all thanks to a powerful new abstraction. For us SQL lovers, the future paradoxically resembles both the present and past, yet the opportunity ahead is simply too compelling to ignore. Today, I’m going to sketch out for you: 1. What exactly is this abstraction of abstractions at the heart of this sea change 2. The lay of the land today: how far things have come, what’s still holding us back, and open questions 3. [EXTRA CREDIT] Technical weeds: table format convergence, S3 tables, vended credentials, and more To not bury the lede any further, I’ll be talking about [Apache Iceberg™️](https://iceberg.apache.org/) , and a further abstraction: the [Iceberg REST Catalog Specification](https://iceberg.apache.org/terms/#catalog) (IRC). The current state of Iceberg isn't easy to navigate. Despite all the buzz, the technology is still young. The ecosystem changes quickly—each day brings something new, from proposals to private previews to updates in `pyiceberg`. So what's really going on here? What matters most? Why should you care? And if even Iceberg's creator says we shouldn't have to think about it (more below), why is everyone talking about it? Over the past year, I’ve been working with many data teams to learn and implement Iceberg in production. I'm convinced of the Iceberg’s potential to impact many more analytics engineering teams. True Iceberg adoption will happen once robustly integrated with all major data platforms but even where it has been integrated there’s last-mile user experience missing that’s dampening the adoption curve. But it’s improving everyday! So let’s get into it. ## Iceberg: A tough nut to crack ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0f6fef22acf23639af69a09c64234a92fc2a0c6f-500x560.jpg) Understanding Apache Iceberg is a “tough nut to crack” because it’s easy to get lost in the technical weeds and miss the big picture. Ironically, Iceberg exists so that most people don’t need to think about it at all! In the [first five minutes of his talk](https://www.youtube.com/watch?v=_GW3GYZK66U) at Data Council last year, Ryan Blue, a creator of Apache Iceberg says almost exactly that: _Iceberg should be invisible [in that it should aim to]:_ - Avoid_ unpleasant surprises_ - Don’t_ steal attention and reduce context switching_ This sounds a lot to me like a powerful abstraction that lets you focus on the task at hand without getting bogged down in details. “Bogged down in details” is an apt description for data engineering until recently. MapReduce, Hadoop, Hive, and Spark were all powerful tools that got the job done, but no one will claim that these were easy to use. You could never just write SQL — a portion of your brain was always reserved for reasoning about where and how the data was written and avoiding unpleasant surprises and edge cases. Your resulting pipeline could process petabytes of data, and you had the sweat to show for it. “Bigger data → more work” is a reasonable heuristic, but the impetus for Iceberg was an attempt to minimize the cognitive burden with a new abstraction of a table that just works like a database’s table (e.g. Postgres or SQL Server). Iceberg isn’t a silver bullet that solves all problems with large analytic data, but it’s a stronger, empowering abstraction. ## The IRC (no, not that IRC) ### The summer of ~~love~~ Iceberg catalogs Ten months ago now, in June 2024, during what we colloquially refer to as “Summit Season”, two hallmark announcements were made within 24 hours of each other. “Iceberg steals the Summits spotlight” and “Iceberg wins the table format war!” comprise the the gist of many folks’ reactions. I largely agree, with a small tweak: the real winner was the Iceberg REST Catalog. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/df3475f3e3771adb4e1e0059dee71037435cd0f5-1342x448.jpg) ### What does an IRC do for me? If you want to know what an IRC is and does, I’ve put that section at the bottom of this post in hopes of avoiding the technical weeds and staying high-level The IRC reminds me of when I first started using Dropbox in college. Countless times, I’d often stay up all night before a deadline writing in Microsoft Word on my Macbook Pro. In the morning, I’d run across campus to a desktop PC at the computer lab. In a minute I could get my paper off the internet opened in Word again so I could print it. It’s easy now to take for granted the power of the abstraction that Dropbox represented even at a time when the Internet already existed. There was complexity behinds the scenes, but the core UX was magic in it’s simplicity: It was a folder of files that was where you wanted it to be and it just worked like a folder should. This is what I feel the IRC represents for us in data. Coming back to data warehousing, we can replace a “folder of a dozen files” with a “schema of a dozen tables”. So imagine that you have this schema of tables and access to many query engines like: Databricks, Snowflake, Redshift, DuckDB, Trino, and Spark. How powerful would it be if you connect any of the engines to the schema above to read those tables, modify their data, and make new tables. Also, that others using other query engines could see those changes. On top of that, this system does not include expensive copying of data, or working with FTP servers, Google Drive, or directly interacting with Azure Blob storage. You just connect your SQL engine to an Iceberg catalog to read and write your data. This is the promise of the IRC in conjunction with your data platform. How might your data team operate differently in this world? This is why [we’re launching cross-platform Mesh](https://www.getdbt.com/blog/introducing-cross-platform-dbt-mesh) to support these exact multi-stack engine scenarios that more than 50% of our Cloud customers already find themselves in today. ## A year of progress So where are we at today with Iceberg and where are we going as we enter summit season 2025 and beyond? The threads I want to pull on are about end-user adoption, data platform vendor integrations, and open source catalogs. I don’t have a crystal ball, but I’ll prognosticate a smidge. ### Curiosity? High! Adoption? Lukewarm (but growing!) At my Iceberg breakout session at Coalesce in Las Vegas last October, I asked the analytics engineers in attendance to raise their hand if they’d heard of Iceberg — all of the hands went up. When I asked those with their hands up to keep them up if they felt they could explain Iceberg to the person sitting next to them, nearly all of the hands went down. Tellingly, this included the folks who said their teams were already using Iceberg in production. This isn’t a problem: it’s Ryan Blue’s vision in action! More-so, this is the opportunity of Iceberg via IRCs: understanding the technology isn’t necessarily a prerequisite for adoption. Maybe one person on the team sets it up. For everyone else it’s business (analytics) as usual. ### Data Platforms are showing up in a big way for the IRC So what have the data platforms and other independent software vendors (ISVs) been up to in the past year? HOLY COW — SO MUCH! It's remarkable to see the entire ecosystem embrace an open-source Apache project as the foundation for their products. The vendors that have integrated deserve a huge round of applause. Yes, they’re just responding to customer demand, and yes, a real reason to invest in Iceberg is you can reallocate engineers away from maintaining proprietary table formats and work that drives more revenue. Still, the industry's investment deserves praise, especially since taking a more self-interested and cynical approach would have been easier, at least in the short-term. Six months ago, we predicted internally that most vendors would support the IRC spec within 6–12 months. Today, after evaluating more private previews than I could possible count, what progress can we observe? If we can interpret “Iceberg support” as being compliant with the spec as of six months ago, then our prediction is looking good. The only major outstanding work is something known as “external writes”. However, as I’ve mentioned above, Iceberg itself is still evolving, so our prediction was poorly framed in the first place. Maybe the right question to ask is _When will IRCs be a stable abstraction such that:_ - _End users have a stable, fully-featured interface_ - _The Iceberg spec can continue to evolve under-the-hood without heavily burdening data teams using Iceberg?_ Perhaps this moment comes when data platform catalogs support external writes, and this will be true in six months. Time will tell! ### OSS catalogs: Important but not for end users Databricks and Snowflake also deserve credit for also open sourcing their catalogs: [Unity Catalog](https://github.com/unitycatalog/unitycatalog) and [Polaris](https://github.com/polaris-catalog/polaris), respectively. [Lakekeeper](https://github.com/lakekeeper/lakekeeper) is another worth calling out for being written in Rust and improving quickly. When data teams ask if I recommend self-hosting a catalog, my answer is largely “No!”. The exception here are teams that have either or both of - Enterprise security requirements (think: on-prem, self-managed data centers) - A dedicated data platform team with the know-how to deploy critical data infrastructure The challenge here is that of uptime and availability. If the IRC is unresponsive, you can’t query the tables any more. A minority of teams will sign themselves up for this challenge. For most, I think your time is better invested elsewhere. Beyond this small minority of data teams, the real value of these projects are for: - Data SaaS vendors: who need some catalog functionality - Prospective data platform customers. who need help committing to use a proprietary catalog (”worst case we migrate away and run the OSS catalog ourselves!”) I don’t say this to cast doubt on the technology, in fact quite the opposite. All of these projects are being used today in production and are “battle-tested”. This usage serves to further refine the IRC as a standard. Everyone benefits from this, even users of proprietary catalogs. ## What might data platforms do differently? IRCs are the clearest option for making Iceberg truly an implementation detail, but adoption is hindered when data platforms don’t truly integrate the concept into their products. Some examples of this include requiring users to: - Create a second catalog within the data platform to make data available elsewhere - Choose a unique object store path for the data when creating an Iceberg table - Mount tables individually and manage their refresh Some data platforms are taking a cautious approach to Iceberg and REST catalogs, worrying that these might create a disjointed experience alongside their native, proprietary table formats. These platforms are instead focused on streamlining their lake house experience within their own product suite. While this concern is understandable, this becomes a game of chicken. Customers want interoperability so they risk losing customers by having a walled garden. Iceberg has fundamentally changed how data teams evaluate tools—any platform without a clear Iceberg strategy now receives a "lock-in" red flag during vendor evaluations, even if said team has yet to start using Iceberg. ## What questions are on my mind for this summer’s Summits and beyond? [Iceberg Summit](https://www.icebergsummit2025.com/) is happening next week, both IRL in SF as well as virtual. You should check it out! As far as what Iceberg announcements I’m hoping for and expecting come June, here’s a list of things that, if announced, would be leading indicators for accelerated Iceberg adoption: - Support query engines to write directly to external Iceberg REST catalogs - Support mounting of a schema’s worth of Iceberg tables - Full support for catalog vended credentials - Any differentiated features that go beyond the scope of the actual Iceberg spec and are focused on UX and developer productivity If we get all of this and more, I still have some open questions - **What’s the multi-region and/or multi-cloud story of Iceberg catalogs?** Right now everything presumes the same cloud and same region or suffer painful egress and latency costs - **How to federate RBAC across query engines?** We still heavily rely upon data bases to `GRANT` access to data. If the data its RBAC is managed in the IRC catalog, how is the query engine configured? - **Wha**t** are best practices for working with multiple catalogs?** More on that in a future post 😉 Thanks so much for reading — as always the comments and my DMs are open. Should you be left wanting more, there’s four more sections that shy less away from the technical weeds. ## Technical weeds ### What about Delta Lake? Some of you will be frustrated that I didn’t bring up Delta Lake. At the time of the Tabular acquisition I remember some people speculating things like this > Databricks acquired Tabular to squash Iceberg in favor of their open table format Delta Lake. It was refreshing to see that cynical take be put to rest so soon when [this interview](https://vimeo.com/1012543474) was posted between Michael Armbrurst and Ryan Blue (creators of Delta and Iceberg, respectively). I love this quote so much: > It was never our intention to start a "format war" and have people spend so much time thinking about storage. It should just work and very few people should have to think about it. You should be able to focus on doing analytics. To achieve this north star of "you don't have to think about it," they aim to standardize the two projects as much as possible. This isn't just lip service! One example touched on was their plan to standardize the `VARIANT` type implementation by [pushing it upstream into parquet itself](https://github.com/apache/parquet-format/blob/master/VariantEncoding.md). Another great example came through [Deletion Vectors](https://docs.delta.io/latest/deletion-vectors.html) (DVs)—a feature that Delta tables had but Iceberg lacked. While Iceberg had a comparable feature called "equality deletes," it wasn't nearly as performant. Now this work has been merged into the spec, slated for release with the Iceberg V3 table spec. This work represents a true data industry team effort with contributions from engineers at Databricks, Snowflake, Netflix, Google and more. If you're feeling brave, curious, and reading “roaring bitmap” doesn’t send you running for the hills, check out [the PR](https://github.com/apache/iceberg/pull/11238) and click around! There's been much discussion about technology that converts between table formats, like Databricks' UniForm and Apache XTable. While these tools are essential in the short term, they'll ultimately become redundant. I'm seeing strong signals that the Delta and Iceberg teams agree not only on what the most important problems are, but also on how they should be solved. But maybe I’m being overly optimistic! ### What about S3 tables? I’ve long been bullish on the IRC, but the [announcement of S3 Tables Buckets](https://aws.amazon.com/blogs/aws/new-amazon-s3-tables-storage-optimized-for-analytics-workloads/) and [Nikhil Benesch’s analysis](https://meltware.com/2024/12/04/s3-tables.html) made me question the assumption. I had been thinking of IRC as an abstraction over object storage, i.e. the REST Catalog would deal with creating, naming, finding iceberg files without you having to think about it. With Table Buckets it’s the converse. you think about S3, but don’t have to create/manage reason about an IRC. This isn’t necessarily a bad thing for both query engine developers nor end users. For a query engine developer, you could argue that it’s easier for query engines to integrate with S3 than it is to integrate with a still evolving OpenAPI spec. They’re all already familiar with object storage! For end users analytics engineers like us, IRCs can be a hurdle to initial Iceberg adoption because you have to set one up before you can create a single table. S3 Table buckets radically simplifies this, in that they have their own catalog behind the scenes. Not only is this catalog wildly performant like many products out of AWS, it also automatically handles maintenance tasks like file compaction. This approach has already borne fruit imho given there’s a plethora of Iceberg quickstart tutorials out there now ([Snowflake](https://aws.amazon.com/blogs/storage/connect-snowflake-to-s3-tables-using-the-sagemaker-lakehouse-iceberg-rest-endpoint/), [DuckDB](https://duckdb.org/2025/03/14/preview-amazon-s3-tables.html)). ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/4eb67ca7eda3e9fd218383d38b8deecbb10c30bd-1422x296.jpg) I’m still very wary about asking analytics engineers to think about S3 paths when writing SQL to the extent that I still think it is an anti-pattern. This is why [dbt by default will manage the path for you when materializing a Snowflake-managed Iceberg table](https://docs.getdbt.com/reference/resource-configs/snowflake-configs#base-location). With S3 Table buckets there’s not a clear notion of a namespace to hierarchically organize tables (think: `datababe.schema.table` ). However, just a few months later AWS S3 announced that Table buckets are available via the IRC API, so S3 Table Buckets has proven to not be opinionated on API. Perhaps their approach is the correct one in providing both UXes. However, there’s an opportunity to simplify. While it is only natural that the S3 team would collaborate with the AWS Data Catalog team, the result is a rather disjoint end user experience. It should not surprise us that the AWS S3 team wants to bring their expertise to making data lakes management easier and cheaper. I’d count on the team to continue evolving this product over time, so you should keep your eyes peeled as well. ### IRCs: What specifically do they do? At the risk of oversimplifying, what the IRC does is closes some remaining gaps that kept SQL on data lakes from feeling like the SQL you’d expect. ### ~~Attention~~ Naming is all you need One powerful abstraction of traditional SQL databases: all you need to query a table is its name, and you never have to think about where the table’s data is stored. You’ve likely never even thought about how much easier this makes your life until you don’t have it anymore. But, in data lakes, often you need to know the table’s path in the object store (e.g. S3) for it’s data. refers to example normal SQL three-part name `my_db.my_schema.my_table` data lake object store path `S3://my-data-lake/some/folders/my-table/metadata.json` I believe that asking analytics engineers to think about S3 paths when writing SQL is an anti-pattern. This is why [dbt by default will manage the path for you when materializing a Snowflake-managed Iceberg table](https://docs.getdbt.com/reference/resource-configs/snowflake-configs#base-location) ### Sir, were you aware that was a red light you just drove through? The other problem that the IRC solves is more behind-the-scenes. When I run a query in Postgres, I never think: - I hope this file lands on disk successfully - I hope no one else is trying to write to this table right now - What if someone else deletes the files I'm writing? We SQL users take this all for granted, but this isn’t possible with a data lake unless you have a catalog! Postgres and many other DBs play “traffic cop” for you so you don’t have to. The IRC fills this role for you on the lake ### One API to rule them all The last problem relates to simplifying how data platforms and query engines integrate with Iceberg. Spark has never had a problem integrating with Iceberg because Iceberg is implemented in Java. But how do you - integrate the Iceberg Java library if your database is written in Python? - Read from an Iceberg catalog written in Go with a query engine written in Rust? The IRC solves this problem by proposing a language agnostic API and a spec for a backend service that does some work that a query engine developer previously would have had to build. This is great because it lowers the barrier to adoption by reducing the required engineering effort to integrate. ## What about IRC’s vended credentials? Once you already have an IRC set up and configured (non-trivial work in it’s own right), the next step is to give a query engine access to it. To do so, by default the query engine must authenticate to two things in order to be able to read and write to the IRC: - The IRC itself (typically with a personal access token) - The object store that has the files associated with the Iceberg table Not only is this a high-friction set-up, the experience isn’t very intuitive. For example, in this setup, when you ask the IRC for a particular table that you’d like to read, it will return to you an object store path for a file that has more info. If you don’t have access to this file in (e.g. S3), you’re SoL. That’s why this pattern also requires that the query engine also has direct access to the object store. However, it doesn’t have to be this hard! With Vended Credentials, you only need to authenticate to the IRC, and the IRC will provision you access to the files in object store. This is a much simpler workflow than what I experienced my first time using IRCs over a year ago. Vended credentials have been in the Iceberg spec since last June, but only recently has it been supported in platforms like Snowflake, Databricks, and SageMaker Lakehouse after a number of preview periods. One query engine writing directly to an external IRC is also vastly simplified by vended creds. You just connect to their IRC and write the table directly without ever knowing where the data is stored. How great to live in a world where when another team needs data from you, you never have to connect to their FTP server, Google Drive, Azure blog storage account to put the data, you just write to their IRC. A consequence of vended credentials is that the IRC becomes critical path for accessing data, it means that you’ll have to refactor your connection later should you decide to stop using an IRC or select a different one. However the abstraction is more simple because you only need to tell your query engine about the IRC and not about object storage anymore. The bear case here for vended credentials here is that it introduces a third access model distinct from the native RBAC of storage (i.e. IAM Policy) and the query engine (think database roles and privileges). However, you can’t have a catalog without RBAC, and the closer that RBAC lives to the data the better. It doesn’t make sense that a query engine should have roles for accessing the data, especially in a world where multiple query engines will access it! --- --- title: "Data transformation: Overcoming common challenges" description: "Learn how to scale data transformation with modern platforms, consistent metrics, and engineering best practices." url: "https://www.getdbt.com/blog/data-transformation-overcoming-common-challenges" date: "2025-03-27" authors: ["Joey Gault"] categories: ["Pulse"] --- # Data transformation: Overcoming common challenges [Data transformation](https://www.getdbt.com/blog/data-transformation) sits at the heart of modern analytics workflows, serving as the bridge between raw source data and meaningful business insights. This process involves converting one materialized data asset—such as a table or view—into another purpose-built for analytics through a series of SQL or Python commands. The transformation journey typically encompasses four critical stages: discovery and profiling to assess data structure and quality, cleansing to correct inaccuracies and remove duplicates, mapping to align data with target system requirements, and storage in centralized repositories like data warehouses. The evolution from traditional [ETL (Extract, Transform, Load)](https://www.getdbt.com/blog/extract-transform-load) to modern [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) approaches has fundamentally changed how organizations handle transformation workflows. While ETL transforms data before loading it into storage systems, ELT leverages the scalability and flexibility of cloud infrastructure to transform data after it's been loaded into warehouses. This shift enables more agile development cycles and allows multiple teams to transform the same raw data according to their specific analytical needs. However, this flexibility comes with complexity. As organizations scale their ELT implementations, they often encounter challenges around data consistency, pipeline management, and maintaining visibility across increasingly complex transformation networks. Without proper governance and tooling, teams can find themselves managing hundreds of disconnected transformation processes, each with its own logic, testing standards, and deployment procedures. ## The consistency challenge One of the most persistent challenges in data transformation involves maintaining consistency across multiple datasets and transformation processes. As organizations grow, different teams often develop their own approaches to common transformation tasks, leading to divergent naming conventions, conflicting business logic implementations, and inconsistent data quality standards. This fragmentation creates several downstream problems that can undermine the reliability of analytics systems. When datasets don't follow standardized conventions, analysts risk duplicating work as they recreate similar transformations across different projects. More critically, inconsistent timezone handling, metric definitions, and data relationships can lead to conflicting reports that erode stakeholder confidence in data-driven insights. For example, if the marketing team calculates customer lifetime value differently than the finance team, executive dashboards may present contradictory views of business performance. The challenge extends beyond technical implementation to organizational alignment. Without clear data modeling conventions established before transformation work begins, teams may create incompatible data structures that become increasingly difficult to reconcile as systems mature. This technical debt accumulates over time, making it harder to scale transformation processes and reducing the overall readability and maintainability of data pipelines. Addressing consistency challenges requires establishing enterprise-wide standards for data modeling, transformation logic, and quality testing. Organizations need to define clear guidelines for naming conventions, data types, and business rule implementation while providing tools and processes that make it easier for teams to adhere to these standards than to work around them. The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) helps solve this by allowing teams to define metrics and dimensions once, in code, and reuse them everywhere — in dashboards, AI copilots, or downstream tools. This shared, governed layer creates a single source of truth that scales with the organization and reduces the risk of conflicting definitions or duplicated effort. ## Standardizing core business metrics Perhaps no challenge is more critical, or more complex, than ensuring consistent definitions and calculations of key performance indicators across an organization. When different teams generate conflicting reports about the same business metrics, decision-makers lose confidence in data-driven insights, and strategic initiatives can be derailed by disagreements about fundamental business performance measures. The root of this challenge often lies in the distributed nature of modern data teams. As organizations scale, multiple groups may independently develop their own calculations for seemingly straightforward metrics like monthly recurring revenue, customer acquisition cost, or inventory turnover. These calculations may differ in subtle but important ways: perhaps one team includes trial customers in their user counts while another excludes them, or different teams apply different time windows for calculating retention rates. Without [version-controlled, centrally-defined metric calculations](https://www.getdbt.com/blog/build-centralize-and-deliver-consistent-metrics-with-the-dbt-semantic-layer), these discrepancies compound over time. Business intelligence tools may pull from different transformation outputs, creating dashboards that tell different stories about the same underlying business performance. This fragmentation not only wastes time as teams debate which numbers are "correct," but it can also lead to poor strategic decisions based on inconsistent or inaccurate data. Successful organizations address this challenge by treating metric definitions as critical business assets that require the same rigor as software code. This means implementing version control for business logic, establishing clear ownership and approval processes for metric changes, and ensuring that standardized calculations are accessible across all downstream analytics tools. The goal is to create a single source of truth for each key business metric while maintaining the flexibility to evolve these definitions as business requirements change. ## Scaling transformation architecture As data volumes grow and use cases multiply, organizations face the challenge of scaling their transformation architecture without sacrificing performance, reliability, or maintainability. What begins as a manageable set of transformation scripts can quickly evolve into a complex web of interdependent processes that become increasingly difficult to monitor, debug, and optimize. The scalability challenge manifests in several ways. First, as the number of data sources increases, transformation logic must handle more diverse input formats and data quality issues. Second, as more teams rely on transformed data, the performance requirements for transformation processes become more stringent: delays in data processing can cascade through multiple downstream systems and impact business operations. Third, as transformation logic becomes more sophisticated, the computational resources required to execute these processes can grow exponentially. Traditional approaches to managing transformation complexity often fall short at scale. Ad hoc scripts scattered across different systems become impossible to maintain and optimize. Manual deployment processes create bottlenecks that slow down development cycles and increase the risk of errors in production systems. Without proper monitoring and alerting, transformation failures may go undetected until they impact critical business processes. [Modern transformation architectures](https://www.cloudera.com/resources/faqs/faqs-resources-modern-data-architecture.html#:~:text=data%20architecture%20FAQs-,What%20is%20modern%20data%20architecture?,making%20and%20drive%20business%20growth.) address these challenges through several key principles. Modular design allows transformation logic to be broken into reusable components that can be independently developed, tested, and optimized. Automated testing and deployment processes reduce the risk of errors while enabling faster iteration cycles. Comprehensive monitoring and observability tools provide visibility into transformation performance and data quality metrics, enabling proactive identification and resolution of issues. ## Tool selection and implementation challenges Choosing the right data transformation tools represents a critical decision point that can significantly impact an organization's ability to scale its analytics capabilities. The modern data landscape offers numerous options, each with different strengths, limitations, and implementation requirements. Data engineering leaders must navigate complex trade-offs between functionality, cost, technical complexity, and organizational fit. The build-versus-buy decision presents particular challenges for transformation tooling. Building custom transformation solutions offers maximum flexibility and control but requires significant engineering resources and ongoing maintenance. Organizations that choose this path must account for the total cost of ownership, including the need to hire and retain specialized talent, maintain infrastructure, and continuously evolve the platform to meet changing requirements. Commercial and open-source solutions offer different value propositions. [Open-source tools provide flexibility and cost advantages](https://www.getdbt.com/blog/dbt-labs-and-fivetran-product-vision) but require technical expertise to implement and maintain. Software-as-a-Service solutions offer managed infrastructure and support but introduce vendor dependencies and recurring costs. The choice between these approaches often depends on organizational factors such as available technical resources, budget constraints, and risk tolerance. Beyond the technical considerations, tool selection must account for the human factors that determine adoption success. The learning curve associated with new transformation tools can significantly impact productivity during transition periods. Tools that require specialized programming skills may limit participation in transformation development to a small subset of team members, creating bottlenecks and reducing the overall agility of data operations. ## Engineering best practices and governance Implementing software engineering best practices in data transformation workflows addresses many scalability and reliability challenges, but it requires significant organizational change and technical investment. Many data teams operate with less rigorous development practices than their software engineering counterparts, leading to transformation code that is difficult to test, debug, and maintain at scale. Version control represents a foundational requirement for mature transformation workflows. Without proper version control, teams cannot track changes to transformation logic, making it difficult to identify the root cause of data quality issues or roll back problematic deployments. However, implementing [version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) for data transformation requires more than just storing SQL files in a repository: it requires establishing branching strategies, code review processes, and deployment workflows that account for the unique characteristics of data pipelines. Automated testing presents particular challenges in data transformation contexts. Unlike traditional software applications, data transformations operate on datasets that change over time, making it difficult to establish stable test conditions. Effective data testing strategies must account for both the logic of transformation code and the quality of input data, requiring sophisticated approaches to test data management and assertion design. Documentation and collaboration tools become increasingly important as transformation systems grow in complexity. Transformation logic that makes sense to its original author may be incomprehensible to other team members months later. Automated documentation generation can help maintain visibility into transformation processes, but it must be supplemented with human-authored explanations of business logic and design decisions. ## The path forward with modern transformation platforms Modern data transformation platforms like [dbt](https://www.getdbt.com/product/what-is-dbt) address many of these challenges by providing integrated solutions that embed engineering best practices into the development workflow. Rather than requiring teams to build their own infrastructure for version control, testing, and deployment, these platforms provide opinionated frameworks that guide teams toward scalable, maintainable transformation architectures. dbt's approach to transformation development exemplifies how modern platforms address common challenges. By representing all transformation logic as modular SQL or Python models, dbt enables teams to build reusable transformation components that can be independently developed and tested. [Built-in testing frameworks](https://docs.getdbt.com/docs/build/data-tests) make it easier to implement data quality checks, while automatic documentation generation provides visibility into transformation logic and data lineage. The platform's integration with version control systems and [CI/CD pipelines](https://docs.getdbt.com/guides/custom-cicd-pipelines?step=1) enables teams to apply software engineering practices to data transformation without building custom infrastructure. This integration supports collaborative development workflows where multiple team members can contribute to transformation logic while maintaining code quality through automated testing and peer review processes. Perhaps most importantly, modern transformation platforms provide centralized governance capabilities that help organizations maintain consistency across distributed teams. Features like centralized metric definitions, standardized testing frameworks, and unified documentation help ensure that transformation logic remains aligned with business requirements even as systems scale. ## Real-world implementation success Organizations across industries have successfully addressed transformation challenges by adopting comprehensive approaches that combine modern tooling with organizational best practices. These implementations demonstrate that while transformation challenges are complex, they are not insurmountable with the right combination of technology, processes, and organizational commitment. [Large media companies](https://www.getdbt.com/industry/telecommunications) have used transformation platforms to simplify complex data architectures while reducing the burden on engineering teams. By standardizing transformation processes and enabling self-service analytics capabilities, these organizations have freed up technical resources for higher-value projects while improving the speed and reliability of data delivery to business stakeholders. [Financial services organizations](https://www.getdbt.com/industry/financial-services) have leveraged transformation platforms to overcome engineering bottlenecks that were slowing down critical reporting processes. By centralizing and automating transformation workflows, these companies have significantly reduced the time required to produce business-critical reports while improving data quality and consistency across different business units. Technology companies have used transformation platforms to enhance their analytics capabilities while maintaining operational efficiency. By automating data transformation processes and implementing comprehensive testing frameworks, these organizations have been able to scale their data operations without proportionally increasing their engineering headcount. The common thread across these successful implementations is a recognition that transformation challenges require both technological and organizational solutions. Modern platforms provide the technical foundation for scalable transformation workflows, but success ultimately depends on establishing clear governance processes, training team members on best practices, and maintaining organizational commitment to data quality and consistency standards. The future of data transformation lies in platforms that continue to abstract away technical complexity while providing powerful capabilities for managing transformation logic at scale. As these platforms mature, they will enable organizations to focus more on deriving business value from their data and less on managing the technical infrastructure required to make that data usable. ## Data Transformation FAQs **What is data transformation?** Data transformation is the process of converting one materialized data asset—such as a table or view—into another purpose-built for analytics through a series of SQL or Python commands. It serves as the bridge between raw source data and meaningful business insights, typically encompassing four critical stages: discovery and profiling to assess data structure and quality, cleansing to correct inaccuracies and remove duplicates, mapping to align data with target system requirements, and storage in centralized repositories like data warehouses. **How does the placement of data transformation differ between ETL and ELT pipelines?** The key difference lies in when transformation occurs relative to data loading. ETL (Extract, Transform, Load) transforms data before loading it into storage systems, while ELT (Extract, Load, Transform) leverages the scalability and flexibility of cloud infrastructure to transform data after it's been loaded into warehouses. This shift to ELT enables more agile development cycles and allows multiple teams to transform the same raw data according to their specific analytical needs, though it can introduce complexity around data consistency and pipeline management. **Why do we need to use transformation on our data?** Data transformation is essential because raw source data is rarely in the format needed for meaningful business analysis. Without transformation, organizations face challenges including inconsistent data quality, conflicting business logic implementations across teams, and incompatible data structures that become difficult to reconcile over time. Transformation ensures data consistency, enables standardized metric calculations across the organization, and converts diverse input formats into purpose-built datasets that support reliable analytics and decision-making. --- --- title: "Data pipelines: Critical components and best practices" description: "Discover how to build modern, scalable data pipelines with dbt—from ingestion to analysis—with trust, governance, and efficiency." url: "https://www.getdbt.com/blog/data-pipelines" date: "2025-03-27" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data pipelines: Critical components and best practices According to [David Menninger](https://isg-one.com/about-us/people/david-menninger), Executive Director of software research at ISG, data integration and data engineering continue to be the biggest challenges in today’s AI and analytics initiatives. Without robust data pipelines, organizations will struggle to ingest, transform, and govern data at scale, limiting their ability to extract meaningful insights from their data. AI-driven analytics and automation are increasing the demand for real-time, high-quality data pipelines. Modern data pipelines must be scalable and efficient, especially as businesses rely on analytics—and increasingly AI—to process vast amounts of structured and unstructured data. ## What is a data pipeline? In simplest terms, a data pipeline is a series of steps that move data from one or more data sources to a destination system for storage, analysis, or operational use. The [three elements](https://www.databricks.com/glossary/data-pipelines) of a pipeline are sources, processing, and destination. Data pipelines convert raw, messy data into something trustworthy, structured, and ready for analysis. While [batch pipelines](https://www.databricks.com/glossary/data-pipelines), which process data on a fixed schedule, can still be useful, today’s data pipelines must be powerful enough to handle huge volumes of data being processed at speed. Thanks to cloud platforms that decouple storage and compute, batch pipelines have largely given way to [real-time (streaming) and cloud-native pipelines](https://www.ibm.com/think/topics/data-pipeline). ### ETL vs ELT pipelines The shift toward data processing at scale and speed required a new way of transforming data. Traditional ETL (Extract, Transform, Load) architectures transform data before loading it into a warehouse. In contrast, ELT (Extract, Load, Transform) pipelines load raw data into the warehouse, transforming it in [parallel ](https://aws.amazon.com/compare/the-difference-between-etl-and-elt/)using affordable on-demand cloud computing. The comparison table highlights the key differences between ETL and ELT. Today, many practitioners consider ETL pipelines as a [subcategory](https://www.ibm.com/think/topics/data-pipeline) of data pipelines, used mainly to load [legacy databases](https://aws.amazon.com/compare/the-difference-between-etl-and-elt/) in the data warehouse. dbt is purpose-built for [ELT workflows](https://www.getdbt.com/blog/etl-vs-elt) where data is already loaded into the data warehouse for processing. The remainder of this article focuses primarily on ELT pipelines. ## Components of a data pipeline Data flows through a pipeline from source to destination, passing through multiple stages along the way. While some frameworks simplify this into three steps—ingestion, transformation, and storage—these can be broken down further to reveal the complexity and value of modern data pipelines. The following components reflect that expanded view: - **Ingestion **- Ingestion involves selecting and pulling raw data from source systems into a target system for further processing. Data engineers [evaluate](https://www.databricks.com/glossary/data-pipelines)—either manually or using automation—data variety, volume, and velocity to ensure that only valuable data is ingested. - **Loading **- In this step, the raw data lands in the cloud data warehouse or lakehouse platform (e.g. [Snowflake](https://www.snowflake.com/), [BigQuery](https://cloud.google.com/bigquery), [Redshift](https://aws.amazon.com/redshift/), or [Databricks](https://www.databricks.com/)). We’ve broken out this step from ingestion to emphasize that this the “L” in ELT that allows the data to be transformed within the data repository. - **Transformation **- This is where raw data is cleaned, modeled, and tested inside the warehouse. This includes filtering out irrelevant data, normalizing data to a standard format, and aggregating data for broader insights.‍ With [dbt](https://www.getdbt.com/product/dbt), these transformations become modular, version-controlled code. This makes data workflows more scalable, testable, and collaborative. - **Orchestration **- Sometimes subsumed into the “transformation” component, automated [orchestration ](https://www.getdbt.com/blog/data-orchestration-vs-etl)schedules and manages the execution of pipeline steps. This ensures transformations run in the right order at the right time. - **Observability and Testing** - The [observability](https://www.montecarlodata.com/product/data-observability-platform/) and [testing ](https://www.getdbt.com/product/test-and-observe)component includes data quality checks, lineage tracking, and freshness monitoring. These are critical for building trust, ensuring governance, and catching issues before they impact downstream analytics. - **Storage **- The transformed data is then [stored](https://www.ibm.com/topics/data-storage) within a data repository, typically a centralized data warehouse or data lake, where it can be retrieved for analysis, business intelligence, and reporting. - **Analysis **- In this last step, data teams ensure that the stored data is ready for analysis, documenting final models, aligning them with a [semantic layer](https://www.getdbt.com/product/semantic-layer), and making them easily discoverable. The semantic layer turns complex data structures into familiar business terms, making it easier for analysts and data scientists to explore data using SQL, machine learning, and BI tools. ## Common data pipeline challenges The point of a data pipeline is to move and transform raw data into reliable, structured insights that drive business decisions. Data scientists and engineers face several common challenges when building and deploying data pipelines. ### Lack of observability across the data estate Today’s modern data stack is complex and fragmented, making it difficult for businesses to gain visibility across their data estate. Without [observability](https://www.getdbt.com/blog/observability-within-dbt), data engineers are flying blind, unable to detect anomalies, trace root causes, assess schema changes, or ensure reliable data for analytics and AI. ### Ingestion of low-quality data When data enters the pipeline with missing values, schema drift, or inconsistent formats, it can silently corrupt insights and models. Without a clear understanding of your data sources—including their reliability, quality, and [governance](https://www.dataversity.net/data-governance-trends-in-2025/) policies—teams risk producing inaccurate outputs and drawing flawed conclusions. ### Pipeline scalability The explosion of data volume and sources presents an ongoing challenge to building fast, reliable data pipelines. Large, monolithic SQL scripts are hard to debug and maintain, while pipeline bottlenecks slow down data processing, delaying insights. The continued use of manual rather than automated processes also negatively impacts scalability and speed. ### Untested transformations When data transformations are deployed without automated testing, even minor logic errors or schema mismatches can later propagate through the pipeline. The result is broken downstream models, inaccurate dashboards and reports, or flawed machine learning outputs. ### Lack of trust in data outputs In complex data pipelines, trust breaks down when consumers struggle to identify reliable models or metrics. [Data trust](https://www.dataversity.net/what-is-data-trust-and-why-does-it-matter/) depends not only on accuracy, completeness, and relevance but also on transparency. Without clear lineage, semantic definitions, and documented assumptions, consumers may question a metric’s validity or ignore it altogether. ## Best practices for data pipelines Today’s pipelines are largely built on cloud-native architectures that decouple storage and compute, enabling scalable, cost-efficient data processing. But even with cloud-native infrastructure, challenges in observability, data quality, scalability, and trust persist. The following best practices help teams overcome these obstacles to ensure reliable, high-performance data workflows. ### Adopt a data product mindset Data pipelines aren’t just about moving data—they’re about creating trustworthy, usable outputs that drive business value. ELT data pipelines accelerate this transformation, but today’s organizations must go further by adopting a [data product mindset](https://www.ascend.io/blog/data-pipeline-best-practices). This approach focuses on delivering well-documented, version-controlled, and consumption-ready data products to the end user, reinforcing data trust and empowering self-service analytics. #### How dbt can help As organizations adopt a data product mindset, ensuring trust, transparency, and usability in pipeline outputs becomes critical. dbt provides key capabilities to reinforce data trust and self-service analytics: - **Gain visibility across the data estate with column-level lineage** - dbt Catalog’s [column-level lineage](https://www.getdbt.com/blog/proactively-improve-your-dbt-projects-with-new-dbt-explorer-features) helps consumers understand the journey of individual columns from raw input to final analytical models. This transparency builds trust by allowing users to trace data origins and transformations. - **Ensure metric consistency with semantic alignment** - dbt’s [Semantic Layer](https://www.getdbt.com/product/semantic-layer) centralizes metric definitions, ensuring consistency across all pipelines and datasets. This prevents metric drift and accelerates the process of creating reliable, reusable data products. - **Strengthen governance with centralized workflow management** - dbt’s [workflow governance](https://www.getdbt.com/product/governance) enables analysts and engineers to standardize on a single platform, ensuring version control, lineage tracking, and access management. This reinforces trust by making data transformations auditable and reliable. ### Ensure end-to-end data quality Poor-quality data entering a pipeline leads to silent errors, broken dashboards, and flawed insights, eroding trust in analytics. Early data quality checks prevent errors from spreading, while automated validation, anomaly detection, and schema enforcement ensure reliable downstream analytics and applications. #### How dbt can help - **Integrate with industry-leading data quality, observability, and governance tools** - dbt [integrates ](https://www.getdbt.com/product/integrations)with leading data quality and data governance tools to ensure that the data entering your pipelines won’t cause downstream errors later. - **Identify pipeline errors early** - Without guardrails, teams risk unreliable outputs. dbt [tests ](https://www.getdbt.com/blog/adlc-test)validate uniqueness, non-null values, and referential integrity. And by treating data like code with [automated ](https://www.metaplane.dev/blog/dbt-test-examples-best-practices)tests and reviews, you can prevent failures from reaching production. - **Continuously monitor pipelines** - With dbt, teams can embrace [continuous integration (CI)](https://docs.getdbt.com/docs/deploy/continuous-integration) to test code before deployment and [monitor](https://www.getdbt.com/product/test-and-observe) pipeline health. dbt tracks the state of production environments, allowing CI jobs to validate modified models and their dependencies before merging. This reduces the risk of breaking changes and ensures smooth pipeline execution. ### Optimize for scalability Modern data pipelines must adapt to changing data structures, optimize query performance to ensure fast analytics, and support real-time ingestion and integration. This demands balancing scalability with cost. [Linear scalability](https://www.ascend.io/blog/data-pipeline-best-practices), where teams grow as pipelines proliferate, is unsustainable. Technologies like Snowflake [Snowpipe Streaming](https://docs.snowflake.com/en/user-guide/data-load-snowpipe-streaming-overview) and Databricks [Lakeflow](https://www.databricks.com/product/data-engineering) enable efficient ingestion and transformation by leveraging high-throughput, low-latency processing, ensuring pipelines scale effectively without excessive overhead. #### How dbt can help - **Implement modular data transformations** - dbt enables [modular](https://www.getdbt.com/blog/modular-data-modeling-techniques), version-controlled transformations, making pipelines easier to debug and maintain. Each model is self-contained, making it easier to isolate and fix errors without affecting the entire pipeline. - **Use version control and enable collaboration **- dbt’s [Git-based version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) tracks data changes as code, enabling teams to collaborate, audit, revert updates, and maintain a single source of truth for scalable pipeline management. - **Optimize with incremental model processing **- dbt’s [incremental model processing](https://docs.getdbt.com/docs/build/incremental-models) transforms only new or updated data, reducing costs, minimizing reprocessing, and improving efficiency. This approach enhances query performance, lowers warehouse load, and accelerates transformations. - **Improve performance with parallel execution** - dbt’s [parallel microbatch execution](https://docs.getdbt.com/docs/build/parallel-batch-execution) processes data in smaller, concurrent batches, reducing processing time and improving efficiency compared to sequential runs. ### Automate orchestration for scalable pipelines Automated orchestration ensures data pipelines run efficiently by synchronizing processes from ingestion to analysis. Event-driven triggers and CI/CD automation detect failures early, reducing downtime and improving reliability. By automating execution, teams can scale pipelines without manual intervention, ensuring consistent and optimized workflows. #### How dbt can help - **Automate pipeline execution with state-aware orchestration** – dbt [Fusion ](https://www.getdbt.com/product/fusion)optimizes workflows by running models only when upstream data changes, reducing redundant executions and improving efficiency. - **Build event-driven execution with hooks and macros** – dbt’s [hooks](https://docs.getdbt.com/docs/build/hooks-operations) automate operational tasks like managing permissions, optimizing tables and executing cleanup operations. Macros bundle logic into reusable functions, enabling parameterized workflows and event-driven execution. - **Enable continuous integration (CI) for reliable deployments** – dbt integrates with [CI/CD workflows](https://docs.getdbt.com/docs/deploy/continuous-integration), automatically testing modified models and their dependencies before merging to production. Whether you're building or refining data pipelines, streamlining workflows, and ensuring data trust are key. By implementing dbt as your pipeline backbone, you gain automation, orchestration, and version control, helping you create scalable, reliable pipelines that drive business value. To learn more about using dbt to implement effective data pipelines, [request a demo today](https://www.getdbt.com/contact). ## Data pipeline FAQs **What is a data pipeline, and how does it ingest, transform, and store data from multiple sources?** A data pipeline is a set of automated processes that move data from sources (databases, applications, events, files) to a destination (data warehouse/lake or operational system) for analysis and use. It ingests by extracting and selecting valuable data, loads it into a centralized platform, transforms it by cleaning, standardizing, modeling, aggregating, and testing inside that platform, and then stores curated outputs for consumption by BI tools, machine learning, or APIs—supported by governance, lineage, and monitoring. **When should an organization use a batch processing pipeline versus a streaming data pipeline?** Use batch when data freshness can be measured in minutes to hours, workloads are predictable (e.g., nightly reports), costs must be minimized, and source systems change infrequently. Choose streaming when low-latency insights are needed (e.g., fraud detection, personalization, IoT telemetry), events arrive continuously at high velocity, or operational decisions depend on up-to-the-second data. Microbatching/hybrid designs can balance latency with cost and complexity. **What components make up a data pipeline architecture, and how do ingestion, transformation, storage, and delivery stages work together?** Core components typically include: - Ingestion: Select and extract raw data from source systems. - Loading: Land raw data in a cloud warehouse or lakehouse. - Transformation: Clean, normalize, model, aggregate, and test data within the platform. - Orchestration: Schedule and sequence tasks so dependencies run in the right order. - Observability and testing: Validate quality, monitor freshness, track lineage, and detect anomalies. - Storage: Persist curated datasets and models for reliable access. - Analysis and delivery: Expose standardized, discoverable models (often via a semantic layer) to BI, ML, and operational tools. These stages form a loop of continuous improvement where monitoring informs fixes, transformations evolve, and delivery remains consistent and trustworthy **What are some common challenges in building data pipelines?** Frequent issues include limited observability across tools and platforms, low-quality or changing source data (missing values, schema drift, inconsistent formats), scalability bottlenecks from monolithic code or manual processes, inadequate automated testing that lets logic errors propagate, and low trust in outputs due to unclear lineage, inconsistent metric definitions, or poor documentation. **What are some best practices for building scalable data pipelines?** Treat datasets as products with clear ownership, documentation, and SLAs; use modular, version-controlled transformations; enforce data quality with automated tests, anomaly detection, and schema validation; adopt CI/CD to catch breaking changes before production; leverage incremental processing and parallel execution to reduce costs and latency; automate orchestration with event-driven triggers and state-aware runs; design for resilience with idempotent operations, retries, backfills, and dead-letter queues; maintain lineage and a semantic layer for consistent metrics and discoverability; and continuously monitor performance, cost, and freshness to guide optimization. --- --- title: "How dbt improves your Tableau analytics workflows" description: "Tableau helps you visualize data. dbt ensures it’s clean, consistent, and reliable before it hits your dashboards." url: "https://www.getdbt.com/blog/tableau-dbt" date: "2025-03-20" authors: ["Daniel Poppy"] categories: ["Learn"] --- # How dbt improves your Tableau analytics workflows Pairing Tableau with dbt brings clarity to your analytics workflow. While Tableau helps you visualize data, dbt ensures that data is clean, consistent, and trustworthy before it ever reaches a dashboard. ## Understanding the gap between data and visualization [Tableau](https://www.tableau.com/) is great at visualization and exploration — but it wasn’t built for data transformation. That’s where [dbt](https://www.getdbt.com/product/dbt) comes in. Most analytics teams struggle with inconsistent metrics, unclear business logic, and data quality issues. Tableau handles the presentation layer well, but upstream complexity often lives in disconnected SQL scripts or spreadsheet logic. dbt solves this by transforming raw data into trusted, documented, and tested datasets—before it ever reaches your dashboards. By pairing Tableau with dbt, you create a modern analytics stack with clear ownership: - **dbt** transforms and tests data in the warehouse - **Tableau** visualizes it for decision-makers This separation of concerns reduces duplication, improves trust, and helps everyone move faster. dbt applies proven software engineering practices to data work—like version control, testing, and modular development—so your Tableau dashboards are powered by clean, reliable data. ## The value of dbt for Tableau users ### Consistent metrics across every dashboard Without a centralized transformation layer, teams often calculate key metrics in different ways. For example, a retail company might have marketing define _Customer Acquisition Cost_ one way, while finance uses another formula entirely. This creates confusion and erodes trust. With the [dbt Semantic Layer](https://docs.getdbt.com/best-practices/how-we-build-our-metrics/semantic-layer-1-intro), you define metrics once—in code—and use them consistently across Tableau dashboards. That means everyone, from marketing to finance, works from the same definition. This alignment is critical for decision-making. Executives can compare reports with confidence, knowing the numbers mean the same thing. And when business logic changes, you update it once in dbt—no need to track down every dashboard and update calculations manually. ### Trusted dashboards start with tested data A Tableau dashboard is only as reliable as the data behind it. Without upstream testing, even polished visualizations can hide critical errors. For example, a healthcare organization tracking patient readmission rates must trust their data before making care decisions. dbt helps prevent bad data from reaching your dashboards by introducing automated [data testing](https://docs.getdbt.com/docs/build/tests) into the transformation process. You can validate: - Null values in critical columns - Data ranges and thresholds - Relationships between tables - Custom business rules unique to your organization When a test fails, dbt alerts your team before incorrect data drives decisions, saving time and preventing downstream issues. This regular validation builds confidence. Clinicians, analysts, and business users can focus on insights—not second-guessing the numbers. For the healthcare organization, that means data you can act on, not just visualize. ### Documentation and lineage Understanding where data comes from and how it’s transformed is critical for trust, compliance, and collaboration. A financial services company under regulatory pressure must be able to explain how key metrics are calculated. Without centralized documentation, this often turns into a time-consuming and error-prone exercise. dbt automatically generates [documentation that includes data lineage](https://www.getdbt.com/product/dbt-catalog), table and column descriptions, SQL transformation logic, and test coverage. This gives every stakeholder—from [analysts](https://www.getdbt.com/product/analyst) to auditors—a clear view of how data flows through the system. For the financial services team, this means they can provide regulators with complete, up-to-date lineage showing how metrics are derived from source to Tableau. It also shortens onboarding time for new team members and gives business users the transparency they need to trust the data. ### Version control and collaboration Business logic evolves over time, and tracking these changes is essential. An e-commerce company changing how they calculate "Active User" in their dashboards might face confusion six months later when comparing year-over-year metrics. dbt projects [integrate with Git](https://docs.getdbt.com/docs/cloud/git/git-configuration-in-dbt-cloud), enabling version control for all transformation logic. This integration supports collaborative workflows through pull requests for reviewing changes, branches for developing new features, and history tracking for auditing purposes. When someone proposes a change to an important calculation, team members can review it, test it, and document it—all before it impacts production dashboards. The e-commerce company could use dbt's version control to document when the "Active User" definition changed, making it easy to explain year-over-year differences in their dashboards. This historical record prevents the confusion that often arises when definitions change without documentation. ## How dbt and Tableau work together The partnership between [dbt Labs and Tableau](https://docs.getdbt.com/docs/cloud-integrations/semantic-layer/tableau) has created powerful integrations that connect transformation and visualization more seamlessly. The dbt Semantic Layer Connector lets Tableau users access metrics defined in dbt, ensuring consistent logic across all dashboards. A retail [analyst](https://www.getdbt.com/product/analyst) can connect Tableau directly to the dbt Semantic Layer to pull metrics like “Customer Lifetime Value” without rebuilding them in Tableau. This improves efficiency and ensures consistency across teams. Data Health Tiles surface freshness and quality details within Tableau dashboards. Viewers can confirm when data was last updated, whether tests passed, and click into dbt for more context. A sales manager checking pipeline metrics, for example, can see that the data is fresh and reliable—without leaving Tableau. The integration is rapidly evolving. Soon, teams will be able to export dbt models directly to Tableau, enrich Tableau Catalog with dbt lineage, integrate with Tableau Pulse, and publish metrics across tools more easily. These features will continue to close the gap between trusted data and business insights. ## When to implement dbt with Tableau Not every Tableau deployment needs dbt from day one, but certain signs make it clear when it’s time to level up your stack. If multiple teams are building dashboards with inconsistent metrics, if data quality issues keep surfacing, or if SQL transformations are getting hard to manage, dbt can help. It’s also worth considering when [governance requirements](https://www.getdbt.com/product/governance) increase or your analytics team is growing fast. The ideal time to implement dbt often comes when your data maturity hits an inflection point. Early on, a few dashboards with simple logic may not need a dedicated transformation layer. But as your reporting grows more complex and collaboration expands, centralized transformation becomes essential. If your [analysts](https://www.getdbt.com/product/analyst) are spending more time prepping data than analyzing it—or if no one’s sure which version of a metric is correct—those are strong signals. At that point, Tableau alone may not be enough. Adding dbt gives you the structure and control needed to scale trust and clarity as your organization grows. ## Getting started with dbt and Tableau To [enhance your Tableau deployment with dbt](https://docs.getdbt.com/docs/cloud-integrations/semantic-layer/tableau), start by setting up dbt on the Team or Enterprise tier. Then configure the dbt Semantic Layer in your environment. Install the JDBC driver, download the Semantic Layer connector, and add it to your Tableau environment to establish the connection. Recent improvements have made the integration process more seamless. While there’s some initial setup, the long-term benefits—like consistency, governance, and time savings—are worth the investment. Many teams begin with a targeted pilot project to address a specific pain point, then expand their dbt usage as they see results. Implementing dbt isn’t just about adopting a new tool—it’s about introducing a more collaborative and trustworthy approach to data. Make room for team training and process updates, especially as you roll out version control, testing, and documentation. The most successful implementations involve coordination between data engineers, analysts, and business users. ## Conclusion Tableau is a powerful tool for visualization and business intelligence. But as your data needs grow, visualization alone isn’t enough. dbt brings the structure, testing, documentation, and governance that Tableau doesn’t provide—making your analytics more reliable, consistent, and scalable. When used together, dbt and Tableau offer a complete analytics workflow. dbt transforms and tests your data before it reaches dashboards, while Tableau lets teams explore and share insights with confidence. The result: consistent metric definitions, higher data quality, clearer documentation, and easier collaboration across teams. As the integration between dbt and Tableau continues to evolve, organizations will benefit from even tighter connections between transformation and visualization layers. By adopting both tools, you’re setting your data team up for success—with trusted data, consistent metrics, and a workflow that scales as you grow. ## dbt and Tableau FAQs **Can Tableau connect to dbt?** Yes. Tableau connects to dbt through the dbt Semantic Layer Connector, allowing users to access metrics defined in dbt directly from Tableau. This ensures consistent metric definitions across dashboards—no need to recreate business logic in each workbook. Data Health Tiles can also be embedded in Tableau to show when data was last refreshed and whether quality checks passed, helping build trust in dashboard accuracy. **Can you do ETL in Tableau?** Not really. While Tableau Prep offers basic data preparation, Tableau wasn’t built for complex data transformations. It’s best used for visualization and exploration. For robust, repeatable transformations—like building metrics, cleaning data, and applying business rules—you’ll want a tool like dbt that’s purpose-built for the transformation layer. **Can Tableau be self-taught?** Yes. Tableau is beginner-friendly and widely used by self-taught analysts. Its drag-and-drop interface makes it easy to get started without writing code. There’s a large library of free resources, including official tutorials, online courses, and community forums. Many people learn Tableau by building dashboards with public datasets using Tableau Public. **Is Tableau easier than SQL?** They serve different purposes. Tableau is easier for non-technical users to pick up because it’s visual and doesn’t require code. SQL is more powerful for querying and transforming data, but it requires learning syntax and logic. In practice, many advanced Tableau users benefit from knowing SQL—it unlocks more control over how data is prepared and used in visualizations. --- --- title: "ETL vs data integration: Understanding the differences" description: "How do ETL and data integration work together? Understand the benefits of and differences between each." url: "https://www.getdbt.com/blog/etl-vs-data-integration" date: "2025-03-19" authors: ["Daniel Poppy"] categories: ["Learn"] --- # ETL vs data integration: Understanding the differences Data is often referred to as the oil of the 21st century. It helps organizations understand the business trajectory and gain actionable insights. However, before it can be used, data must be thoroughly collected and processed using techniques like ETL (Extract, Transform, Load) and data integration. Data integration and ETL are techniques used to streamline data workflows within an organization. They help unify data, convert it to a usable form by applying necessary transformations and integrations, and improve its accessibility across the organization. Both techniques are related to one another but have subtle differences. In this article, we’ll explore data integration and ETL in detail. We’ll break down what data integration and ETL really mean, and why data movement, transformation, and unification at scale are crucial for modern businesses. ## What is data integration? Most modern organizations use multiple platforms and services for various purposes, such as marketing, advertising, email, and customer interactions. While each of these services helps improve business efficiency and growth, this architecture means that all generated data is stored in a scattered and decentralized manner. This scattered storage bottlenecks data-related tasks and workflows such as analytics, data science, and reporting. The data integration process aims to collect data from these touchpoints and store it in a different format for improved accessibility. It involves building robust pipelines that collect and combine data and is usually a part of operations like ETL and ELT. It involves the following key steps: - **Source identification: **Identifying and listing all available data sources. These can be relational database systems (RDBMSs), CSV files, or cloud storage for unstructured files. - **Data standardization:** Applying necessary transformations to standardize the data format. - **Pipeline formation: **Designing and optimizing pipeline architecture to collect and integrate the data in a single location. Data integration activity provides a holistic view of the entire organization's operations. It removes [data silos](https://www.techtarget.com/searchdatamanagement/definition/data-silo) by allowing users to access data from any domain or workflow and improves process efficiency and time-to-market. Moreover, a data integration solution design can take multiple approaches. For example, the design may involve a complete **_migration_**, during which the entire data is physically moved to a new centralized location. Or, it may include **_virtualization, _**which involves accessing data from its original location using API endpoints. Virtualization allows users to access data from a single interface without any data movement. ### Benefits of data integration Having a holistic view of your entire data estate has several benefits, such as: - **Breaking down silos:** Integrations improves data accessibility across the entire organization. It provides users with a single interface to access data from multiple sources, eliminating unnecessary inter-team dependencies. - **Improved efficiency for data-related tasks:** Teams can build dedicated pipelines with specific transformations to receive data in a set format. This saves time during data analysis for data science and analytics projects. Moreover, since teams can access any data they want, they are no longer delayed by dependencies on other departments, allowing for faster experimentation and insights. - **Streamlined reporting and analytics:** Integrated data ensures consistency across reports and dashboards. Teams no longer have to manually compile information from various sources, which reduces the risk of errors and improves the reliability of business intelligence outputs. - **Enables data-driven innovation: **With unified access to diverse datasets, organizations can identify patterns, generate new ideas, and deploy AI and machine learning (ML) models more effectively. This fosters innovation and supports smarter decision-making across the business. ## What is ETL? [ETL](https://www.getdbt.com/blog/etl-pipeline-best-practices) stands for **_Extract, Transform, Load_**, and is a standard data migration process used**_ _**across the industry. It consists of three main steps: - **Extract: **The extraction step involves collecting data from one or multiple sources. These sources can include relational databases hosted on different platforms, CSV or text files stored in various cloud storage, or JSON responses from APIs. - **Transform: **Once the data sources are connected, various transformations are implemented as an intermediate step. These transformations help standardize the schema, clean the data (removing duplicates, handling null values, etc.), and create new information or views through aggregation operations. - **Load:** Once the transformations are complete, the data is loaded to a centralized location, such as a [data warehouse](https://www.ibm.com/think/topics/data-warehouse), data lake, or lakehouse. Depending on the access settings, the clean and transformed data is accessible to all teams from a single endpoint. ### ELT: The modern ETL While ETL has been a popular data integration technique for some time, ELT has gained popularity as an alternative in recent years. It involves the same steps as ETL, but shifts the order by applying data loading before transformation, hence the name “**_Extract, Load, and Transform_**.” As data needs grow, [ELT has been recognized as the superior alternative](https://www.getdbt.com/blog/etl-vs-elt), offering various benefits over its counterpart, including: - **Better flexibility:** Since the destination location contains raw data, data analysts and scientists can transform it according to their needs. The data can be modified iteratively for each use case, and no special changes are required to the ELT pipeline. This flexibility allows for quick adaptation to changing business requirements, ensuring teams always have access to the most up-to-date data. It also reduces the risk that the original data will be lost in transformation - a common and frustrating problem with the ETL process. - **Faster data loading:** By offloading complex transformations to the data warehouse, the data loading pipeline becomes more efficient, allowing teams to access the latest data more quickly. - **Increased process efficiency:** Leveraging the power of modern data warehouses for transformations reduces the processing burden on upstream systems. This streamlines the overall data workflow, minimizes resource contention, and enables engineering teams to focus on higher-value tasks instead of maintaining complex ETL logic. ELT is quickly becoming the industry standard for data integration. Although this article focuses more on ETL, everything we have discussed applies equally to both procedures. ### ETL real-world use case To make this more concrete, let’s look at the data integration process using a real-world use case. Take an e-commerce store. Such stores often consist of multiple modules, including a customer chatbot, an inventory management system, a social media profile, and an email marketing system. Each module is often implemented using different platforms and generates data stored across various locations. These diverse data sources hold valuable insights for business reporting and data science initiatives. However, accessing and consolidating data across platforms can be challenging and requires close coordination between teams. For instance, if the data science team needs customer reviews from social media, they must request access and formatting support from the social media team. This leads to delays and inefficiencies. An ETL pipeline streamlines this by collecting and storing all relevant data in a centralized location. It connects to each source system, extracts the data, applies necessary transformations, and loads it into a unified platform. This not only ensures consistency and accessibility but also empowers teams to make data-driven decisions faster without constantly relying on cross-team coordination. ## ETL and data integration: head-to-head Let’s do a one-on-one comparison between ETL and data integration. - **Data integration:** - A specific data transformation technique. - Focuses on data unification and availability, making data usable for a specific use case. - Depending on the business requirement, the integration process can be real-time, a batch job, or a combination of both. - **ETL:** - ETL is a superset of data integration that covers the entire process of transforming data for a specified use case. - Gathers data from one or more sources, transforming it according to business logic, and delivering it to an end system. - It covers three main steps: Data extraction, transformation (which includes data integration), and loading. - Mostly a batch process. ETL jobs are often scheduled outside work hours, so the integration completes before the start of the business day, ensuring fresh data is available for reporting and analysis. ## Breaking silos - Unlock your data-driven potential Data can be a game changer for businesses, but it requires complex processing. Managing transformations effectively is crucial to unlocking data’s true potential. That’s where dbt comes in—a modern data transformation tool that helps bring structure, consistency, and control to your pipelines. If you’re embarking on a journey to streamline your data operations, [dbt](https://www.getdbt.com/product/dbt) is the [control plane](https://www.getdbt.com/blog/data-control-plane-why) you need to unify data transformation efforts across your enterprise. dbt enables you to modernize your data pipelines with features like [version control](https://docs.getdbt.com/docs/collaborate/git/version-control-basics), modular SQL, testing frameworks, and automated data deployment pipelines that [bring a DevOps-like rigor to analytics workflows.](https://www.getdbt.com/resources/the-analytics-development-lifecycle) With dbt, transformations become reproducible and trustworthy, allowing you to shift from reactive data cleaning to proactive data modeling. It supports scalable, governed development by making it easy to track changes, implement peer reviews, and ensure data quality at every step. Beyond just transformations, dbt assists with data integration by allowing teams to write reusable code blocks for common operations, such as standardizing data formats across sources. This makes your pipelines cleaner, more consistent, and easier to maintain over time. Whether you're an engineer, analyst, or business user, dbt empowers you to own your part of the data journey with a unified toolkit. Ready to see the difference? [Book a demo](https://www.getdbt.com/contact) today. --- --- title: "Go fast without sacrificing quality: Everything we announced at dbt Developer Day" description: "Get to know the latest features powering best-in-class developer experience for practitioners building in dbt" url: "https://www.getdbt.com/blog/dbt-developer-day-2025" date: "2025-03-19" authors: ["Alexis Jones"] categories: ["Product"] --- # Go fast without sacrificing quality: Everything we announced at dbt Developer Day dbt Developer Day is in the books! It was a jam-packed virtual event where we showcased new features that help deliver a best-in-class developer experience for the (now, forever, and always) heros of dbt: data practitioners. From the new dbt engine, powered by integrating SDF’s SQL comprehension technology into dbt, to the new dbt-native VS Code extension, the GA of dbt Copilot, and some exciting updates to dbt Core…the features we highlighted all enable dbt developers to embrace speed in their workflows. With the guardrails to ensure that they aren’t compromising data quality in the process. **** Get up-to-speed on everything we announced below, or you can catch the replay—and admire the A+ costuming—[here](https://www.getdbt.com/resources/webinars/dbt-developer-day). ## Good DevEx = better business outcomes While the surface area of dbt has expanded over recent years to include more stakeholders in the data development workflow, [consistent with our mission](https://www.getdbt.com/about-us), our #1 priority has always been and will always be to maintain dbt as the best place for analytics engineers and data practitioners to standardize data workflows. A big part of that means delivering a best in class developer experience so you can empower the teams you support—whether data analysts, business users, or execs—to win with data. And spoiler alert: it turns out that what’s good for the developer is good for the organization. Better DevEx has been proven [time](https://www.atlassian.com/software/compass/resources/state-of-developer-2024) and [time again](https://getdx.com/research/the-one-number-you-need-to-increase-roi-per-engineer/) to equate to better business outcomes. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/8c871aeb2f661bdcfbb46cde3743d2e60d6bee27-2290x1248.png) In our view, data developers need tooling that empowers them to balance these two seemingly paradoxical objectives: to expedite data workflows, while also ensuring data quality. The features we announced at dbt Developer Day strike that balance: - **New dbt engine & VS Code Extension:** Users will enjoy dramatic performance and productivity improvements with this lightning fast new engine that gives them faster feedback, supports rapid iteration, and makes it turnkey to validate analytics code as its being written. - **dbt Copilot:** Users can reduce cognitive load by leaning on AI to automate mundane tasks like creating code, tests, and documentation. Equally important, this embedded AI can harness your data’s full context to automate the creation of refined, well-governed data products. - **dbt Core 1.10:** The next minor version release for dbt Core includes features that promote faster and safer development to help users stay in their flow, including sample mode and stricter validation. ## Get to know the new dbt engine and VS Code extension An entirely new engine will soon power dbt—one that will make development orders of magnitude faster and considerably more cost-effective. And because we know many of you prefer to develop in VS Code, we’re also launching an official dbt VS Code extension to bring these improvements to your local development experience. ![Screenshot of dbt-native VS Code Extension with IntelliSense capabilities](https://cdn.sanity.io/images/wl0ndo6t/main/fe203a665a4a8ab452a6400fa996b071b10736e0-2048x1383.png) ### Faster, smarter, more efficient: Meet the new dbt engine Earlier this year, we [acquired SDF Labs](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs), a next-gen data transformation layer and SQL compiler, and have been hard at work integrating SDF's technology into dbt. The result is a next-generation dbt engine that will fundamentally transform dbt as we know it, now in private beta. 🏁 The implications to developer experience are really exciting: - Lightning-fast parse times. Our demo included a 10,000 model project that parsed in less than a second 🤯. - Rich IntelliSense to autocomplete SQL keywords, as well as suggest available dbt model or column names as you type - Auto-refactor references to model, column names across your project as soon as you make a naming change - Instantly click through to go to another model, ref, or macro from inside a model - Hover over to see available columns and column types within a schema - Preview the data that will be outputted by a given CTE from within a dbt model ✅ The new engine also automatically helps organizations optimize data warehouse costs: - Detect and flag parsing errors (e.g. a missing comma) without hitting the warehouse - Detect and flag compilation errors (e.g. a function doesn't accept the provided parameters or data types), without hitting the warehouse - Soon, dbt will also be able to even better help avoid unnecessary model runs, including in CI, with state-aware and column-aware orchestration And the best part? All of these improvements will happen automatically when you run dbt on the new engine. ### Bringing the new dbt engine to VS Code Of course, this new engine will soon power development in the dbt Cloud IDE. We continue to be as committed as ever to ensuring the dbt Cloud IDE is a robust, accessible, and constantly improving dbt development experience. But we know that many of you do your best work in the familiar local confines of VS Code, and that’s why we’re introducing the official dbt VS Code Extension, now in private beta. This extension is built from the ground up by dbt Labs, for dbt developers—and it’s the best way to take advantage of the new engine while developing locally. It will come with the speed and cost efficiency unlocks made possible by the new dbt engine built right into it, in addition to all the other capabilities you would expect from a first class dbt development experience, such as the ability to explore lineage, preview data, and more. ### What’s next? Both the **new dbt engine** and the **VS Code extension** are currently in **private beta** for select dbt Cloud customers. We'll continue to select more participants for the beta over the coming weeks and months and we’ll be in touch if you’re an eligible candidate. In the meantime, dbt Cloud customers and dbt open source users can express interest in the beta program [here](https://docs.google.com/forms/d/1JElfCGT_fU1HlI-XUPs5SV9_nwNALaxTf0T7a8-RSJA/edit). We're moving quickly towards broader availability, which is planned for our upcoming [dbt Launch Showcase](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase) in late May. ## Disrupting data engineering (again) with AI: dbt Copilot is GA We also announced the general availability of [dbt Copilot](https://www.getdbt.com/product/dbt-copilot), a native AI-assistant that brings context-aware AI to your data workflows so you can deliver higher-quality data faster. In the AI era, speed and quality are essential to staying ahead. dbt Copilot harnesses rich data context—capturing relationships, metadata, lineage, and more—paired with powerful LLMs to automate routine tasks and consistently enforce key ADLC best practices across your dbt projects. With this GA release, dbt Cloud Enterprise customers can now auto-generate documentation, data tests, semantic models, metric definitions, and inline SQL directly within the dbt Cloud IDE with a simple click of a button and natural language prompts. dbt Copilot also supports [Bring Your Own Key (BYOK) for OpenAI or Azure OpenAI](https://docs.getdbt.com/docs/cloud/enable-dbt-copilot#bringing-your-own-openai-api-key-byok), and includes a built-in style guide to ensure consistency. ![gif of dbt copilot generating automated documentation](https://cdn.sanity.io/images/wl0ndo6t/main/a888c629af75c3f011d7b94ace329a8da1e9275c-2143x1293.gif) Early beta users have shared positive feedback, noting improved documentation coverage, faster formatting, and improved query optimization, saving hours on manual work. > “**dbt Copilot has completely changed how we approach documentation and query optimization.** Instead of spending hours manually updating models, I can use natural language to generate tests, infer metadata, and enrich our data models with valuable context. The more metadata we add, the better our entire team benefits, from analysts to executives.” > - **Cody Mclean, Sr. Data Engineer, Hard Rock Digital** [Watch video](https://youtu.be/jbmitdc540M) Read the full GA announcement [here](https://www.getdbt.com/blog/dbt-copilot-is-ga). If you’re a dbt Cloud Enterprise customer, [check out the docs](https://docs.getdbt.com/docs/cloud/dbt-copilot) to learn how to get started using dbt Copilot today. ## Taking dbt Core 1.10 out for a test drive The latest version of dbt Core (1.10) is now in beta. It introduces a couple of powerful new capabilities that will make development not just faster, but also safer. ### 🏎️ Faster: Sample mode You can now opt to use dbt in [sample mode](https://docs.getdbt.com/docs/build/sample-flag) to build just a subset of your data during development or CI, rather than building your entire dataset(s). This will allow you to validate outputs white iterating much more rapidly, and reduce warehouse spend as a result. Sample mode will be particularly helpful if you're dealing with large time-based datasets. Today, dbt Core 1.10 supports time-based sampling for references to any models or sources with `event_time` configured. You can use the `--sample` flag with the `dbt run` or `dbt build` commands to specify a trailing time window (such as the last "3 days"). You can also specify a particular historical time range (such as 2025-01-27 to 2025-01-30). [More in the docs](https://docs.getdbt.com/docs/build/sample-flag). Additionally, you can set a **default sample window** at the environment level, so you don’t have to manually pass the `--sample` flag every run. ![Sample mode in dbt 1.10](https://cdn.sanity.io/images/wl0ndo6t/main/9e2597abe998543525be2a66190f640232f2ddd7-2000x976.png) Thank you to all of the community members who participated in our [Github discussion](https://github.com/dbt-labs/dbt-core/discussions/11200) and Zoom feedback session to help shape this feature! ### 🛡️ Safer: Stricter validation dbt Core 1.10 will also introduce **stricter validation** for project inputs, preventing common configuration errors and ensuring better reliability. In prior versions, dbt allowed any configuration input, including typos or misconfigurations. This could lead to silent failures when models were run. For instance, mistakenly entering `desciption` (which is misspelled) instead of `description` would previously go unnoticed. But no more! Soon, dbt will emit a warning for mistakes such as this one. If you'd like to use a custom configuration, you can still do so, but it will need to be nested within the `meta` config. Sample mode is currently available in the dbt Core 1.10 beta, as well as for dbt Cloud customers on the Latest [release track](https://docs.getdbt.com/docs/dbt-versions/cloud-release-tracks). Stricter validation is on its way soon; keep an eye out for an upcoming discussion in Github! ## But wait, there’s more! A big part of our commitment to data developers is to meet them where they are; across a myriad of data platforms. A few other milestones to share in this regard: **BigQuery** - We continue to invest to make our [BigQuery connector](https://docs.getdbt.com/guides/bigquery?step=1) better than ever. We will soon extend support to Python Models powered by BigQuery DataFrames for advanced analytics and machine learning on BigQuery. - We also now support Workload Identity Federation so users can avoid having to use service keys to authenticate to BigQuery. - Stay tuned for more announcements on our partnership with Google and BigQuery at Google Next, happening April 7-11 in Las Vegas. [Join us!](https://www.getdbt.com/events/summit/googlecloudnext2025) **Teradata** - We introduced our [Teradata adapter](https://docs.getdbt.com/docs/core/connect-data-platform/teradata-setup) in late 2024. After great customer feedback and a rigorous beta program, we’re thrilled to announce that this adapter is now generally available. ## Save the date for May 28 We’re excited about how these innovations will uplevel the developer experience in dbt, helping data teams ship data products faster, while always keeping data quality in check. And we’re just getting started. Be sure to register for our next virtual launch event—[the annual dbt Launch Showcase](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase) happening on May 28—where we’ll introduce even more features designed to empower our users to scale analytics. See you there! --- --- title: "A new era of data engineering: dbt Copilot is GA" description: "Today marks a pivotal shift in data engineering—AI-powered, automated workflows built on deep metadata context are one click away" url: "https://www.getdbt.com/blog/dbt-copilot-is-ga" date: "2025-03-19" authors: ["Chakshu Mehta", "Tom Grabowski"] categories: ["Product"] --- # A new era of data engineering: dbt Copilot is GA Generative AI is redefining how engineers work. In fact, Gartner® predicts that, "by 2028, 90% of enterprise software engineers will use AI code assistants, up from less than 14% in early 2024."¹ Today, dbt is stepping boldly into that future with [**dbt Copilot**](https://www.getdbt.com/product/dbt-copilot), an AI-powered data assistant **now generally available in dbt Cloud**. This is a defining moment for AI in data engineering. One that will not only change how data teams work but redefine the work itself. [Watch video](https://youtu.be/jbmitdc540M) ## AI demands high quality inputs At its core, data engineering has focused on cleaning, structuring, and optimizing raw data to deliver reliable business outcomes. Since the shift to the cloud, dbt has long streamlined these processes, and now with GenAI, we’re entering [the “AI era” for data engineering teams](https://roundup.getdbt.com/p/how-ai-will-disrupt-data-engineering). In this new era, these once routine data engineering tasks are now automated, freeing teams to focus on strategic innovation and building the advanced AI systems of tomorrow. As the pressure for faster AI innovation intensifies, organizations must integrate GenAI into their workflows to stay ahead. Supporting this shift, "Gartner estimates that AI software spend will grow to $297 billion by 2027–representing an annual growth rate of 19%."² The problem is generic LLMs focus solely on code, missing the deep context required to build truly robust data pipelines. For LLM outputs to be useful, it has to be accurate—and the stakes for data quality have never been higher. With dbt Copilot, you’re not just injecting GenAI into your analytics workflows; you’re leveraging the full context of your data—its relationships, metadata, and lineage—to automate routine tasks and consistently uphold key ADLC best practices like documentation, data testing, semantic modeling, and SQL formatting. The result is a refined, governed dataset that serves as a solid foundation for building high-quality analytics and advanced AI systems. dbt is the only solution uniquely positioned to supply your LLM with this level of deep metadata context. For example, if your data warehouse uses a column named `customer_id`, a generic LLM might generate code referencing just `id`. dbt Copilot avoids these mistakes by tailoring outputs to your actual schema. > “dbt Copilot is the future of data engineering. If you aim for truly high-quality, well-governed data—or are building AI systems that depend on trustworthy inputs—you need an AI solution that operates at the critical juncture where data is refined and contextualized. With dbt Copilot, you're laying the foundation for next-generation AI and driving true data-powered success.” > — **Mark Porter, CTO at dbt Labs** ## Embed context-aware AI into every workflow With dbt Copilot now integrated into your analytics workflows, we're modernizing data engineering and empowering teams to deliver high-quality analytics and AI innovations faster than ever before. Your day-to-day tasks as a data engineer can be executed more efficiently. It’s like having a dedicated data intern that standardizes legacy documentation, improves query optimization, checks for SQL syntax errors, enhances metadata compliance, and speeds up migrations, all within dbt Cloud. Our beta users are already cutting hours of tedious work, standardizing processes faster, and unlocking greater value from every data asset. As Cody McLean, Sr. Data Engineer at Hard Rock Digital put it, > "dbt Copilot has completely changed how we approach documentation and query optimization. Instead of spending hours manually updating models, I can use natural language to generate tests, infer metadata, and enrich our data models with valuable context. The more metadata we add, the better our entire team benefits, from analysts to executives.” Today with dbt Copilot, data teams can: [auto-generate documentation](https://docs.getdbt.com/docs/cloud/use-dbt-copilot#generate-resources), [data tests](https://docs.getdbt.com/docs/cloud/use-dbt-copilot#generate-resources), [semantic models](https://docs.getdbt.com/docs/cloud/use-dbt-copilot#generate-resources), [metric definitions](https://docs.getdbt.com/docs/cloud/use-dbt-copilot#generate-resources), and [inline SQL](https://docs.getdbt.com/docs/cloud/use-dbt-copilot#generate-and-edit-code) using natural language prompts, all within the dbt Cloud IDE. By making standardization and YAML development effortless, dbt Copilot ensures your models are secure, rigorously tested, and built to the highest standards—empowering teams to work faster and focus on what truly matters. And now with support for **Open AI** [Bring Your Own Key (BYOK) service](https://docs.getdbt.com/docs/cloud/enable-dbt-copilot), [Azure OpenAI service](https://docs.getdbt.com/docs/cloud/account-integrations?ai-integration=azure#ai-integrations), and a custom style guide (in beta), dbt Copilot is even more flexible, secure, and enterprise-ready. [Watch video](https://youtu.be/Vg7Nu6SKcXE?si=zn0Sn3Nx3cjQv0Is) ## How our customers are using dbt Copilot today Since the beta launch at Coalesce 2024, hundreds of customers have experienced firsthand how dbt Copilot transforms their data workflows. Users are streamlining outdated documentation, enriching metadata, and tapping into dozens of use cases that improve data quality and team efficiency. Here are a few common use cases: ### Standardize legacy documentation and enhance metadata enrichment Writing model descriptions and column definitions can take hours, especially for legacy data where documentation may be missing or outdated. With a click of a button, dbt Copilot contextually generates YAML-based documentation leveraging SQL logic, past queries, and metadata, so teams can instantly improve clarity and maintainability. dbt Copilot helps decipher cryptic column names, making it easier to work with legacy data and maintain consistency across models. ![Copilot gif](https://cdn.sanity.io/images/wl0ndo6t/main/a888c629af75c3f011d7b94ace329a8da1e9275c-2143x1293.gif) ### Accelerate data testing Rather than manually crafting each test case, dbt Copilot uses the context of your dbt models—understanding dependencies, transformations, and schema relationships—to suggest context-aware validation tests. With the click of a button, it adds the corresponding test code directly to your project—ready to run during your builds. This not only speeds up the process but also helps teams detect schema drift by flagging unexpected changes in column structures before they impact production. ### Improve formatting and query optimization Messy SQL slows everyone down. With natural language prompts, dbt Copilot generates SQL inline and then leverages a built-in style guide to ensure it’s formatted correctly—eliminating inconsistent casing, indentation, and redundant syntax. This saves time in code reviews and improves maintainability. dbt Copilot can also refactor legacy queries and suggest cleaner, more efficient SQL, making it easier to optimize older code. ![Semantic Models Best Practices](https://cdn.sanity.io/images/wl0ndo6t/main/551fbc5c2c88cd19630939ba4152211f99673057-1136x714.png) ### Automate semantic layer and metric definitions Defining business metrics across teams is often inconsistent. Based on existing data models, dbt Copilot recommends useful key metrics aligned with business objectives to be used for your semantic models ensuring analysts and stakeholders are aligned on consistent definitions. > "My CFO doesn’t want to see a dashboard—she wants direct answers, like our ARR for the quarter. With dbt Copilot, she can instantly query well-defined metrics using conversational language so she can get the answers she needs instantly and my teams can get out of the business of building dashboards. The future of analytics is AI-powered insights and governed metrics via Semantic Layer, not dashboards. As these tools evolve, organizations that embrace them will lead the way." > — **Josh Carlson, Director of Analytics at Code42** ## What’s new in dbt Copilot With the general availability of dbt Copilot, we’re introducing new features designed to give users greater control while making collaboration even easier: ### OpenAI Bring Your Own Key (BYOK) and Azure OpenAI service By default, dbt Copilot comes with an integrated OpenAI service, making it incredibly easy to get started. For enterprise users seeking additional control, you can opt to [bring your own OpenAI key (BYOK)](https://docs.getdbt.com/docs/cloud/enable-dbt-copilot#bringing-your-own-openai-api-key-byok). Additionally, dbt Copilot now integrates with the [Azure OpenAI Service](https://docs.getdbt.com/docs/cloud/account-integrations?ai-integration=azure#ai-integrations), allowing organizations in the Azure cloud to seamlessly incorporate AI into their workflows. All of these options are built to enterprise-grade standards, ensuring a simple experience while ensuring reliable performance. ### Custom style guide for standardized SQL formatting (in beta) SQL consistency is critical for maintainability, and dbt Copilot now includes **a custom style guide** to enforce best practices across dbt models. Teams can configure style preferences, reducing manual code reviews and ensuring consistency across projects. ## The future of AI in dbt Integrating AI into data engineering isn’t optional anymore—it’s essential. With dbt Copilot, you’re not just adapting to change; you’re setting a new standard for data management. The general availability of dbt Copilot is just the beginning. As AI continues to evolve, dbt Copilot will become more intuitive, proactive, and aligned with how modern data teams operate across the entire analytics development lifecycle. dbt Copilot is available to dbt Cloud Enterprise customers, please contact your sales representative to request access. [Get started with dbt Copilot today](https://docs.getdbt.com/docs/cloud/enable-dbt-copilot) and experience the future of AI-assisted analytics engineering. References: 1. Gartner, [Magic Quadrant for AI Code Assistants](https://www.gartner.com/document-reader/document/5682355?ref=solrAll&refval=456688018&toggle=1&viewType=Full), Arun Batchu, Philip Walsh, Matt Brasier, and Haritha Kjandabattu, August 19, 2024. (Accessible to Gartner subscribers only) 2. Gartner, [Forecast Analysis: AI Software Market](https://www.gartner.com/document-reader/document/5314863) by Vertical Industry, 2023-2027, Anna Griffen and Ina Agamirzian, March 27, 2024. (Accessible to Gartner subscribers only) GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved. --- --- title: "How AI will disrupt data engineering as we know it" description: "It will be hard to compare data engineering in 2024 and data engineering in 2028 and say “those are the same things.”" url: "https://www.getdbt.com/blog/how-ai-will-disrupt-data-engineering" date: "2025-03-16" authors: ["Tristan Handy"] categories: ["Insights"] --- # How AI will disrupt data engineering as we know it _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/how-ai-will-disrupt-data-engineering). _ [Last time I wrote](https://roundup.getdbt.com/p/a-year-of-innovation-in-ai-part-1) I dove into a bunch of AI advancements that have happened over the past year. Reasoning models, chain of thought, inference-time compute, etc. And there’s more to explore there and I need to return to that series. But for this week’s issue, I want to pause on that and talk about AI from a different perspective. I want to think, as rationally as we can about an uncertain future, about how the job of the data engineer will change over the coming 1-2-3 years as a result of AI. I am quite confident these changes will be massive. I think the word _disrupt_ is not at all hyperbole—**I think it will be hard to compare data engineering in 2024 and data engineering in 2028 and say “those are the same things.”** It just turns out that many of the tasks that data engineers do every day are tasks that AI can provide tremendous leverage in. I don’t know what the % efficiency metric will be—20%? 50%? 80%?—but I think it’s totally possible that it’s on the higher end of that range. I think that will be both good for data engineers and good for the companies they work for. Data engineers will have more work to do than ever ([Jevons paradox](https://en.wikipedia.org/wiki/Jevons_paradox) at work), but it will be more strategic, add more value to companies, and will likely see them get raises. Companies will get the higher-functioning, higher-ROI, more accessible data systems that have always seemed out of reach. In this post I want to look at the specific tasks that data engineers spend their time on, and look at how addressable-or-not these tasks are with AI. Let’s dive in. ## The tasks of a data engineer AI doesn’t replace jobs, it automates tasks. So let’s look at the tasks that someone leveled a Senior Data Engineer most commonly spends time on today. Of course, it should go without saying, YMMV: there is no single canonical job description for a data engineer. But I think we can still get close enough to reason about. **What does a Senior Data Engineer spend their time on?** - Create technical artifacts - Landing new data. Building and maintaining automated data ingestion pipelines. - Transforming raw data into bronze, then silver, then gold layers. Includes authoring brand new pipelines as well as refactoring existing pipelines to handle new business requirements. - Defining metrics on top of transformed data. - Writing tests and documentation. - Monitoring costs of data infrastructure and refactor code to optimize performance characteristics. - Reviewing pull requests from peers. - Monitoring production jobs and declare incidents related to either pipeline failures or observability / quality issues. Investigate and resolve those incidents. - Liaise with stakeholders and peers - Answer questions about currently-available data assets like “which data set should I use?” and “can I trust this?” - Collaboratively design changes to existing data assets to accommodate new requirements. Conversations like “what are the edge cases I need to know about when calculating cost of goods sold?”. - Stakeholder enablement and education. - Designing the overall architecture of the DAG, including modularization, team boundaries and ownership, modeling best practices, etc. I’m sure you could find some other things to put on these lists, but I feel like they’re pretty representative. Feel free to tell me what I’m forgetting. ## The role of frameworks and tooling in an AI-centric world Many of the above tasks are already doable with AI. And I want to talk more about that. But before I get there, it’s important to talk about frameworks, and how important frameworks are to an AI-centric world. Claude 3.7 will write you almost any kind of code you could want. You can absolutely build a pipeline from the ground up, building ingestion, transformation, testing, etc. in Lisp. In Assembly. In the style of Guido van Rossum. Whatever. You could even imagine a world in which you had 1,000 distinct pipelines and every one was written in a different language or framework or set of conventions. All reading from and writing to a shared corpus of tabular data. But: just because it is now conceivable to create such a codebase, _is it a good idea?_ The answer is: no. **Obviously not.** Just as a team of humans would have an impossible task of maintaining such a Frankenstein, the heterogeneity would make it intractable for LLMs as well. This _intuition pump_ is helpful to get us to an important conclusion: AI will be more effective as an accelerant when: - a code base is fewer lines of code (less room for error) - a code base is more consistent rather than less consistent: in languages, in coding conventions, in design - a code base uses consistent CI/CD and other developer tooling - a code base uses consistent and well-documented logging / observability - a code base uses well-documented best practices also employed by a large community of users. In general: code bases that are more concise, more homogeneous, and use standard tools that are well-documented in the model training data (i.e. the public internet) will be more comprehensible by AI systems. One of the best ways to make all of these things true at the same time is to use frameworks and open standards. Claude 3.7 knows how to build reliably Airbyte ingestion pipelines because the framework is well documented and there are a lot of examples published. It’s also fantastic at writing dbt code for the same reasons. If you’re able to give it an environment where it can test its own code and validate downstream models as a part of its CoT—code quality goes up even further. Standardized frameworks also emit well-understood error messages, which pushes code quality up further. In short: good frameworks, tooling, and standards are _just as important_ for AI as they are for humans. And the wonderful thing about AI is: it is infinitely adaptable to whatever frameworks, tooling and standards tooling you want to use. No learning curves. Finally the promise of a consistent code base. ## How many of these tasks are already doable? Got it, frameworks are powerful in an AI world. Now let’s look at the individual tasks that data engineers spend time on and try to figure out how tractable they are. In answering this question I am _not_ going to assume massive improvements in model capability. Even with modest improvements I believe all of this will become true. What is fundamentally needed is productization of currently-available models directed at the specific needs of data engineers, not the invention of new frontier tech. ### Creating technical artifacts - **Ingestion pipelines** With nothing but Cursor you can already [vibe code](https://en.wikipedia.org/wiki/Vibe_coding) your way to a working ingestion pipeline from basically any data source with a publicly-available API. You can already add pagination and solve edge cases and inject instrumentation. It’s unclear, though, if this is actually what is needed. I still fundamentally don’t think most data movement code should be written and maintained within the walls of an individual company—AI or no, I still want to hire a vendor or support a community project. Data engineers shouldn’t be spending a lot of time on this problem today and likely shouldn’t be in the future either. When a custom build is required, AI can already do it well; try it yourself in Cursor today. - **Authoring new data transformation assets** If you’re using dbt, data transformation is very soon to become _heavily_ AI-enabled. Whether you’re building models, writing documentation and tests, or defining metrics, this is coming to you _very soon_. We demoed some of these capabilities at Coalesce and will have more to share on Wednesday at our [dbt Developer Day](https://www.getdbt.com/resources/webinars/dbt-developer-day). While we are certainly still in the early stages of where we ultimately want to get to, dbt Copilot is already _very_ good at all of these authoring tasks and there is a very clear path to getting even better. Nick Shrock, in one of his best posts ever, called dbt and tools like it [medium-code frameworks](https://dagster.io/blog/the-rise-of-medium-code). It turns out that medium-code frameworks are extremely well-suited for AI. Having personally used dbt Copilot, I anticipate the time required to author new transformation code for data engineers will drop very significantly. - **Multi-file refactoring** One thing that Cursor now does super-well is stage multi-file edits as a result of a single prompt. You could imagine a similar prompt in dbt: “refactor code in these two parts of the DAG to minimize duplication; combine models where appropriate.” Or: “A new field was added in this data source. Please pull that field all the way through to the DAG into [X] final model.” These types of refactoring tasks are low-creativity but highly time-intensive. Implementing them is product work, not research. The opportunity to get a handle on tech debt with tooling like this makes me giddy. - **Automated incident resolution** Imagine providing the entire log output of a `dbt run` and the associated project code into a context window and getting back a diagnosis and proposed resolution. While we haven’t productized this experience yet, it’s not hard to experiment with this yourself hackathon-style. Imagine a world in which, following a pipeline failure, a full PR was queued up and run through CI, with a full report waiting for you and just ready to hit the merge button. We should anticipate this type of experience for data engineers in the not-too-distant future. How much time are you currently spending on break/fix? Slash it significantly. I’m going to pause there because I’m at risk of boring you. Suffice it to say that I truly believe that a) much data engineering work has already been framework-ized, and b) AI will now make creation of, iteration on, and maintenance of these technical artifacts _far more efficient._ And for the aspects of data engineering that are not yet framework-ized (dbt or otherwise), there will be tremendous gravity towards pulling them into a framework because of the leverage that these types of high-quality AI experiences will provide. ### Liaising with stakeholders and peers There are countless people throughout the business who use data as a core part of their jobs, and data engineers are _constantly_ fielding questions from them. I won’t re-list them all here, but if you’re a data engineer you know the drill. Forever, the hope of “self-service” has been the hope that these data users would not need to lean on data engineers in this way—these interactions inject friction and slowness that neither side wants. This fully actualized self-service has never actually materialized, and the status quo has been frustratingly persistent. But I’m optimistic that we have more of a path today than ever. The easiest thing to do for any technology vendor at the very onset of the AI era was to take all of the domain-specific context that you had and surface it to users in a chat interface. And we did the same thing. It was (and is) quite good—it does a great job of allowing users to ask business questions and answering them with semantic-layer-governed responses. The problem with this approach is that users don’t actually want to interact with dozens of chat interfaces. They don’t want to remember to go to a given tool to get one type of answer and another tool for another type of answer. There will not be 30 chat experiences all with different context. There will be one…or maybe just a few. But likely a single dominant one. This is how [aggregators](https://stratechery.com/aggregation-theory/) work. You likely don’t use a bunch of different search engines—you probably just use one, and it is probably Google. This is how chat will go as well. The problem is, Google could scrape the web and respond to all queries based on that knowledge. But ChatGPT cannot know all of the information you want to ask it questions about (at least, yet). That lack of business context is the problem. That’s where a _context protocol_ comes in. A context protocol—a somewhat new topic in the public AI conversation—is a standardized way for services to provide additional context to models via an open protocol. The most promising one today is called [MCP](https://modelcontextprotocol.io/introduction), but whether or not MCP wins, the awareness/excitement/support for this idea has developed a ton of momentum and I am fairly convicted that _something like this_ will become real and widely-supported. There will be a large number of context providers (every source of valuable enterprise context) and a large number of context consumers (different products with AI capabilities). There is no way to create point-to-point integrations to facilitate this. A protocol will be needed if we are going to see the right type of advancements, and I think it will happen. Imagine that your license to ChatGPT enterprise or Claude Desktop or whatever _already came with_ a connection to all of the metadata about every piece of structured data you had access to. What was there, how trustworthy it was, how suitable it was for the analysis you were describing, etc. I think that, very quickly, you would find yourself asking questions of your friendly AI rather than shoulder-tapping your colleague in data engineering. That’s not to say that the existing relationship would _go away_, but I do think that this would represent a true reset of the working relationship between data engineers and downstream business stakeholders—one that both sides would benefit from. ## Where does that leave us? Over the past two years, critical innovations have been made in foundational AI technology. Chain of thought, reasoning models, inference-time compute, agentic workflows. These are the ingredients needed to build the AI-enabled data engineering future. But they are now here. And open frameworks—from dbt to Spark to Airbyte to others—have become widely deployed. This makes it possible to create great framework-specific AI tooling, both for the commercial stewards of those frameworks (including us), but also by any other vendor. The commercial incentive to innovate here is high, and there could not be more attention on delivering these types of benefits within companies of all sizes. This is going to happen, and data engineering as a profession is never going to be the same. So what? Time to get a new job? Data engineers are obsolete? Hardly. Data engineers, one of the hottest jobs of the last decade, will stay hot. But practitioners will be pushed in one of three directions: towards the business domain, towards automation, or towards the underlying data platform. - **Data platform engineers** will become ever-more-important. They don’t spend their time building pipelines, but rather on the infrastructure that pipelines are built on. They are responsible for performance, quality, governance, uptime. - **Automation engineers** will sit side-by-side with data teams and take the insights coming out of data and build business automations around it. As a data leader recently told me: “I’m no longer in the business of insights. I’m in the business of creating action.” - **Data engineers** that are primarily obsessed with business outcomes will have ample opportunity to act as enablement and support for the insight-generation process, from owning and supporting datasets to liaising with stakeholders. The value to the business won’t change, but the way the job is done will. You’ll hear a lot more from us [on Wednesday](https://www.getdbt.com/resources/webinars/dbt-developer-day) about how we’re making this future a reality for dbt users. I’m excited to disrupt the decade-long status quo and build something better. --- --- title: "Ensuring EU DORA compliance for financial services with dbt Cloud" description: "dbt Cloud helps financial institutions improve operational resilience through data lineage, testing, and collaboration." url: "https://www.getdbt.com/blog/ensuring-eu-dora-compliance-for-financial-services-with-dbt-cloud" date: "2025-03-14" authors: ["Natalie Gladman"] categories: ["Insights"] --- # Ensuring EU DORA compliance for financial services with dbt Cloud The European Digital Operational Resilience Act (DORA) is here, setting new standards for operational resilience in the financial sector. For data teams, this means your data platform must be equipped to meet stringent requirements for operational continuity, resiliency, and data integrity in line with DORA's stringent ICT risk management requirements. ## What is the Digital Operational Resilience Act (DORA)? The Digital Operational Resilience Act (DORA) (EU Regulation 2022/2554) came into effect in January 2025. It establishes a unified regulatory framework for ICT risk management in the financial sector. Its goal is to enhance digital resilience by ensuring that financial institutions can withstand, respond to, and recover from cyber threats and IT disruptions. The regulation applies to financial entities including banks, building societies, investment firms, insurers, payment institutions, crypto-asset service providers, and ICT third-party service providers. DORA introduces mandatory ICT risk management, requiring financial entities to implement robust security measures, conduct resilience testing, and report major cyber incidents. It also tightens oversight on critical ICT third-party service providers, such as cloud computing services, by subjecting them to direct EU supervision to mitigate systemic risks. Additionally, financial institutions are encouraged to share threat intelligence to strengthen the sector’s overall defense against cyber threats. For financial institutions, DORA means greater regulatory scrutiny, increased compliance requirements, and higher costs related to cybersecurity investments. Organizations will need to reassess their ICT risk management frameworks, strengthen contracts with ICT third-party service providers, and improve internal monitoring and reporting capabilities. While the regulation raises compliance burdens, it also standardizes cybersecurity practices across the EU, reducing fragmentation and improving financial stability in an increasingly digital economy. **So how can dbt Cloud help with DORA?** ## Transparent data lineage ### Simplifying audits and compliance DORA demands a clear understanding of your data's journey. dbt Cloud's built-in table and column level lineage capabilities provide a comprehensive view of how data is transformed and used. This visibility simplifies audits and demonstrates compliance with data quality standards. You can proactively identify and address potential risks and ensure the integrity and reliability of your financial reporting. ## Automated Testing ### Ensuring data integrity and business continuity Robust testing is crucial for DORA compliance. dbt Cloud's testing framework allows you to define and automate tests for your data models, ensure data quality, and prevent errors. This proactive approach minimizes downtime, reduces the risk of regulatory penalties, and builds confidence in your data. ## Enhanced Collaboration ### Strengthening data governance Effective collaboration is essential for maintaining data governance under DORA. dbt Cloud provides a centralized platform for data teams to collaborate, share knowledge, and manage data projects. With features like version control and access control, dbt Cloud fosters accountability and transparency to simplify compliance with DORA's requirements. ## dbt Labs is your DORA-ready partner dbt Labs is committed to helping financial institutions achieve DORA compliance. Our services and contractual terms are aligned to applicable DORA regulations, as detailed in our DORA Disclosure documentation (available upon request for Enterprise plan customers and prospects). The Disclosure outlines the specific Articles of the DORA regulations that are applicable to ICT third-party service providers, and further clarifies which additional elements apply to Critical or Important functions, including key clauses related to audits, termination, transition services, and cooperation with authorities. ### Ready for DORA? Visit our [**security page**](https://www.getdbt.com/security) and [**talk to a dbt expert now**](https://www.getdbt.com/contact).​ --- --- title: "What are the steps involved in the data transformation process?" description: "How raw data becomes analytics‑ready through the crucial stages of transformation: discovery, cleansing, mapping and storage." url: "https://www.getdbt.com/blog/data-transformation-process-steps" date: "2025-03-13" authors: ["Joey Gault"] categories: ["Pulse"] --- # What are the steps involved in the data transformation process? [Data transformation](https://www.getdbt.com/blog/data-transformation) is the layer where raw inputs become something useful. It’s where naming conventions get standardized, data types are aligned, and business logic gets applied in a way that scales. Whether you’re producing metrics, feeding a dashboard, or preparing inputs for a model, transformation sits at the heart of modern analytics workflows. In [ELT](https://www.getdbt.com/blog/extract-load-transform) architectures, transformation happens after data lands in the warehouse. That shift has unlocked new patterns for scalability and collaboration—but it’s also introduced complexity. As more teams work with the same raw data, transformation frameworks need to be modular, testable, and transparent by design. This article outlines a four-step framework for managing the transformation process. From profiling raw inputs to delivering structured outputs, each step plays a role in building trusted, analysis-ready data. ## The four-step data transformation framework The data transformation process follows a structured approach that converts one materialized data asset—such as a table or view—into another purpose-built for analytics through a series of SQL or Python commands. This systematic methodology ensures that raw data becomes analysis-ready through four distinct but interconnected stages. ### Discovery and profiling The transformation journey begins with comprehensive data discovery and profiling, where teams assess the structure, quality, and characteristics of their source data. This initial assessment proves essential for identifying anomalies, inconsistencies, and potential issues that require attention during subsequent transformation steps. During this phase, data engineers catalog the organization's entire data estate, examining existing data structures and identifying key attributes. This process involves understanding data types, field relationships, data volumes, and quality metrics. Teams also assess the reliability of data sources, frequency of updates, and any existing data governance policies that might impact transformation approaches. Equally important is understanding the requirements of end users who will ultimately consume the transformed data. Data engineers should interview different stakeholders within the organization to develop a clear picture of their specific analytical needs and determine how to align available data assets with business objectives. This dual assessment of both data characteristics and user requirements forms the foundation for effective transformation design. ### Data cleansing The cleansing stage focuses on correcting inaccuracies, filling in missing values, and removing duplicates to ensure data reliability and accuracy. This step addresses the fundamental data quality issues that can undermine analytical outcomes if left unresolved. Data cleansing encompasses several critical activities. Teams identify and correct formatting inconsistencies, such as standardizing date formats or ensuring consistent capitalization across text fields. Missing value handling requires strategic decisions about whether to fill gaps with calculated values, remove incomplete records, or flag missing data for special treatment in downstream analyses. Duplicate detection and removal present particular challenges in large datasets where subtle variations in records might mask true duplicates. Advanced cleansing processes also address outlier detection, where statistical methods help identify data points that fall outside expected ranges and require investigation or correction. The cleansing process must balance thoroughness with performance considerations. While comprehensive data cleaning improves analytical accuracy, overly aggressive cleansing rules might inadvertently remove valid data points or introduce bias into datasets. Establishing clear data quality standards and validation rules helps ensure consistent cleansing approaches across different transformation projects. ### Data mapping and structure The mapping phase involves aligning data structures according to the needs of target systems through a process known as data mapping and structuring. During this stage, data types may be converted, fields reorganized, and specific business rules applied to ensure compatibility with downstream analytical requirements. Data mapping requires careful consideration of how source data elements correspond to target schema requirements. This process often involves complex transformations where multiple source fields combine to create single target attributes, or where single source elements split into multiple target fields. Business logic implementation during mapping ensures that transformed data reflects organizational definitions and calculation methods. Schema evolution presents ongoing challenges during the mapping phase. As source systems change their data structures or new data sources join the transformation pipeline, mapping logic must adapt to accommodate these changes without disrupting existing analytical processes. Version control and documentation become critical for managing these evolving mapping requirements. The actual transformation execution applies the defined mapping rules, converting data into the desired format while preserving data integrity and maintaining audit trails. Modern transformation tools enable complex mapping logic through declarative approaches that make transformation rules more maintainable and easier to understand. ### Storage and loading The final step involves loading transformed data into centralized data stores, such as data warehouses, where it becomes available for analysis and reporting. This stage requires careful consideration of storage optimization, access patterns, and performance requirements. Storage decisions impact both cost and performance characteristics of analytical systems. Teams must choose appropriate data formats, compression strategies, and partitioning schemes that align with expected query patterns. The choice between different materialization strategies—such as tables, views, or incremental models—depends on factors including data freshness requirements, query performance needs, and computational cost considerations. Loading processes must account for data freshness requirements and system availability constraints. Batch loading approaches offer efficiency advantages for large data volumes but may introduce latency in data availability. Streaming approaches provide near real-time data access but require more complex infrastructure and monitoring capabilities. [Data lineage](https://www.getdbt.com/blog/what-is-data-lineage) tracking becomes crucial during the loading phase, as teams need visibility into how transformed data relates to original sources. This lineage information supports impact analysis when source systems change and helps with troubleshooting data quality issues that emerge in analytical outputs. ## Integration with modern data architectures The transformation process operates within broader data architecture patterns that have evolved significantly with cloud computing capabilities. The shift from traditional [ETL (Extract, Transform, Load)](https://www.getdbt.com/blog/extract-transform-load) approaches to modern [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) patterns has fundamentally changed how organizations approach transformation workflows. In ELT architectures, transformation occurs after data loading, leveraging the scalability and flexibility of cloud infrastructure. This approach enables multiple teams to transform the same raw data according to their specific analytical needs while maintaining a single source of truth for source data. However, this flexibility introduces complexity around consistency management and pipeline coordination. Modern transformation platforms like [dbt](https://www.getdbt.com/product/what-is-dbt) address these architectural challenges by providing structured frameworks that embed engineering best practices into transformation workflows. Rather than requiring teams to build custom infrastructure for version control, testing, and deployment, these platforms provide opinionated approaches that guide teams toward scalable, maintainable transformation architectures. The integration of transformation processes with broader data platform capabilities enables sophisticated governance and monitoring approaches. Automated testing frameworks validate transformation logic and data quality, while comprehensive documentation generation provides visibility into transformation processes and data lineage relationships. ## Operational considerations and best practices Successful transformation implementations require attention to operational aspects that ensure reliability and scalability as data volumes and complexity grow. [Monitoring and alerting](https://docs.getdbt.com/docs/deploy/monitor-jobs) capabilities provide visibility into transformation performance and data quality metrics, enabling proactive identification and resolution of issues. [Version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) represents a foundational requirement for mature transformation workflows. Teams need to track changes to transformation logic, making it possible to identify root causes of data quality issues and roll back problematic deployments when necessary. Effective version control strategies must account for the unique characteristics of data pipelines and the dependencies between different transformation components. [Automated testing](https://docs.getdbt.com/docs/build/data-tests) strategies become increasingly important as transformation systems scale. Unlike traditional software applications, data transformations operate on datasets that change over time, making it challenging to establish stable test conditions. Effective testing approaches must validate both transformation logic and data quality, requiring sophisticated strategies for test data management and assertion design. [Documentation](https://docs.getdbt.com/docs/build/documentation) and collaboration tools support team coordination as transformation systems grow in complexity. Transformation logic that seems clear to its original author may prove incomprehensible to other team members months later. Automated documentation generation helps maintain visibility into transformation processes, but human-authored explanations of business logic and design decisions remain essential for effective knowledge transfer. ## Addressing scale and complexity changes As organizations scale their transformation operations, they encounter challenges around consistency, performance, and maintainability that require systematic approaches to resolve. Maintaining consistency across multiple datasets and transformation processes becomes particularly challenging as different teams develop their own approaches to common transformation tasks. Standardizing core business metrics represents one of the most critical challenges in scaled transformation environments. When different teams generate conflicting reports about the same business metrics, decision-makers lose confidence in data-driven insights. Successful organizations address this challenge by treating metric definitions as critical business assets that require version control, clear ownership, and approval processes for changes. Performance optimization becomes increasingly complex as transformation workloads grow. Teams must balance computational efficiency with development productivity, often requiring sophisticated approaches to incremental processing, dependency management, and resource allocation. Modern transformation platforms provide capabilities for automatic optimization and intelligent workload management, but realizing these benefits requires careful attention to transformation design patterns. The evolution of transformation requirements over time necessitates flexible architectures that can adapt to changing business needs without requiring complete system redesigns. Modular transformation approaches enable teams to modify specific components without impacting entire pipelines, while comprehensive testing frameworks provide confidence that changes don't introduce unintended consequences. ## Future-proofing transformation processes The data transformation landscape continues evolving rapidly, with new technologies and approaches emerging regularly. Organizations must balance the benefits of adopting new capabilities with the stability requirements of production analytical systems. Successful transformation strategies maintain flexibility while establishing solid foundations that can accommodate future changes. Modern transformation platforms continue abstracting away technical complexity while providing powerful capabilities for managing transformation logic at scale. As these platforms mature, they enable organizations to focus more on deriving business value from their data and less on managing the technical infrastructure required to make that data usable. The integration of artificial intelligence and machine learning capabilities into transformation workflows represents an emerging trend that promises to automate many routine transformation tasks. However, these capabilities require careful implementation to ensure that automated transformations maintain the quality and reliability standards that analytical systems demand. Data engineering leaders must consider how their transformation strategies align with broader organizational objectives around data democratization, self-service analytics, and operational efficiency. The most successful transformation implementations combine technical excellence with organizational change management, ensuring that improved data capabilities translate into better business outcomes. The systematic approach to data transformation—encompassing discovery, cleansing, mapping, and storage—provides the foundation for reliable analytical systems. However, success ultimately depends on implementing these steps within broader frameworks that address scalability, governance, and operational requirements. As data volumes continue growing and analytical requirements become more sophisticated, organizations that invest in robust transformation processes will be better positioned to derive competitive advantages from their data assets. ## Data transformation steps FAQs **What are the four stages of the data transformation process and what occurs in each stage?** The four stages are: Discovery and Profiling, where teams assess source data structure, quality, and characteristics while understanding end-user requirements; Data Cleansing, which focuses on correcting inaccuracies, filling missing values, and removing duplicates; Data Mapping and Structuring, where data structures are aligned with target system needs through field reorganization and business rule application; and Storage and Loading, which involves loading transformed data into centralized stores like data warehouses with consideration for optimization and performance requirements. **How does the transformation step differ between ETL and ELT workflows?** In traditional ETL (Extract, Transform, Load) approaches, transformation occurs before loading data into the target system. In modern ELT (Extract, Load, Transform) patterns, transformation happens after data loading, leveraging cloud infrastructure scalability and flexibility. ELT enables multiple teams to transform the same raw data according to their specific analytical needs while maintaining a single source of truth, though this introduces complexity around consistency management and pipeline coordination. **How does the data transformation process improve data quality and analytical outcomes?** The transformation process improves data quality through systematic cleansing that corrects formatting inconsistencies, handles missing values strategically, and removes duplicates. It enhances analytical outcomes by ensuring data reliability and accuracy, standardizing formats across datasets, and applying business logic that reflects organizational definitions. The structured approach prevents fundamental data quality issues from undermining analytical results while maintaining audit trails and data lineage for troubleshooting and impact analysis. --- --- title: "Agentic AI: What's needed from your data?" description: "To act smart, AI needs clean, structured data. Here’s how dbt helps prepare your data for agentic AI." url: "https://www.getdbt.com/blog/agentic-ai-data-requirements" date: "2025-03-04" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Agentic AI: What's needed from your data? According to the [IBM Institute for Business Value (IBV)](https://newsroom.ibm.com/2025-06-10-IBM-Study-Businesses-View-AI-Agents-as-Essential,-Not-Just-Experimental), 70% of executives say agentic AI is critical to their future, and 61% of CEOs report they’re actively deploying or scaling AI agents in their organizations. Agentic AI is no longer an option. It’s a strategic imperative for future-proofing your business. In this article, we’ll break down: - What agentic AI actually is (and isn’t) - Where it’s gaining traction today - What data conditions must be true for it to work - How dbt helps teams build the trusted, structured foundation agentic systems require ## Key features of agentic AI Agentic AI refers to systems that can [act independently](https://hbr.org/2024/12/what-is-agentic-ai-and-how-will-it-change-work) to achieve goals with minimal human input. These agents, often built on large language models (LLMs), [simulate human-like decision-making](https://www.ibm.com/think/topics/agentic-ai), enabling them to reason, plan, and execute tasks in real time. Unlike traditional AI models, agentic AI exhibits goal-driven behavior, autonomy, and adaptability. “Agentic” implies not just intelligence, but _agency _— the ability to assess, decide, and act based on evolving inputs, often without direct instruction. Some key features of agentic AI include: - Autonomy - Reasoning/decision making - Autonomous and continuous learning - Multi-agent collaboration ### Autonomy Autonomy sets agentic AI apart from previous iterations of AI applications. While GenAI-powered chatbots can generate responses and even make recommendations (or reservations), they don’t include enough data or checks and balances to trust them to make important decisions in lieu of humans. In contrast, AI agents go beyond conversation. They autonomously execute multi-step [workflows](https://orq.ai/blog/ai-agentic-workflows), call APIs, access external systems and data, and dynamically adapt to evolving inputs without human intervention. Think less chatbot, more digital teammate. ### Reasoning/decision making AI agents also reason autonomously, without human intervention. They don’t just execute tasks, they [make decisions](https://www.mckinsey.com/capabilities/operations/our-insights/when-can-ai-make-good-decisions-the-rise-of-ai-corporate-citizens), learn from outcomes, reason across time horizons, and can collaborate with other AI agents. Behind the agent is an LLM acting as a reasoning engine, orchestrating tasks, generating solutions, and coordinating models for specialized functions like content creation, recommendations, and visual processing. Techniques like [retrieval-augmented generation](https://www.nvidia.com/en-us/glossary/retrieval-augmented-generation/) (RAG) boost performance by injecting up-to-date knowledge and domain-specific data into the reasoning loop, increasing accuracy and reducing hallucination risk. ### Autonomous and continuous learning Agentic AI doesn’t just follow instructions — it observes, adapts, and improves based on real-world interactions. While traditional and GenAI systems improve through incremental learning — where the model is updated with new data to make better predictions and improve performance — only agentic AI agents engage in [lifelong learning](https://www.xoriant.com/thought-leadership/article/agentic-ai-and-continuous-learning-creating-ever-evolving-systems). For example: An agentic customer support assistant could analyze customer interactions and begin adjusting product recommendations automatically based on patterns it identifies — all without manual updates from your team. ### Multi-agent collaboration In complex environments, agentic systems scale through [multi-agent collaboration](https://www.ibm.com/think/topics/multiagent-system). Each agent specializes in a specific function — say, financial forecasting or marketing ops — and communicates with others through an orchestration layer to tackle broader goals. This modular approach supports scalability and fault tolerance: agents can be added or swapped out without rearchitecting the entire system. ## The state of agentic AI Agentic AI isn’t just theoretical — it’s here, scaling fast. Valued at $5 billion today, the [agentic AI market](https://www.wisdomtree.com/investments/blog/2025/04/21/agentic-ai-the-new-frontier-of-intelligence-that-acts) is projected to hit $50 billion by 2030. Enterprises are already embedding these systems into core business functions to automate decisions, streamline operations, and unlock new efficiencies. According to Microsoft’s [Work Trend Index](https://www.microsoft.com/en-us/worklab/work-trend-index/2025-the-year-the-frontier-firm-is-born), executives cite customer service, marketing, and product development as top areas for near-term AI investment — with HR, finance, and sales close behind. And the ecosystem is maturing. Several major platforms now offer the tools to build and deploy AI agents that reason, act, and collaborate across systems. ### Current applications for agentic AI Here are just a few of the areas where companies are already using agentic AI: - **Customer Service**: Resolving tickets, updating accounts, and escalating issues autonomously (e.g., [Klarna’s ](https://openai.com/index/klarna/)AI agent handles two-thirds of support chats). - **Sales**: Prospecting, personalizing outreach, qualifying leads, and scheduling meetings. - **Marketing**: Orchestrating campaigns, optimizing send times, and generating content. - **HR**: Automating onboarding, answering policy questions, and managing internal mobility. - **Finance**: Auditing expenses, detecting fraud, and generating forecasts. ### Key agentic AI platforms Agentic AI platforms are rapidly evolving, giving businesses the tools to build autonomous agents that reason, act, and collaborate across workflows. Here’s how leading players are enabling this shift: - [**Salesforce Agentforce**](https://www.salesforce.com/agentforce/) embeds customizable, decision-making agents directly into Salesforce workflows. These agents can reason, take multi-step actions, and integrate with tools like Slack and Data Cloud. - [**Microsoft Azure AI Foundry**](https://ai.azure.com/) supports multi-agent orchestration using modular tools like Azure Functions and the Model Router. Paired with [Copilot Studio](https://www.microsoft.com/en-us/microsoft-copilot/agents), businesses can build and deploy agents for [multi-agent collaboration](https://www.microsoft.com/en-us/microsoft-copilot/blog/copilot-studio/whats-new-in-copilot-studio-may-2025/) across Microsoft 365 apps using natural language. - [**OpenAI GPTs**](https://openai.com/index/new-tools-for-building-agents/) allow companies to create tailored agents that reason, act, and integrate with APIs—all without writing code. The Agents SDK and Responses API enable developers to build multi-agent workflows with memory, tool use, and real-time decision-making. - [**IBM watsonx Orchestrate**](https://www.ibm.com/new/announcements/productivity-revolution-with-ai-agents-that-work-across-stack) powers agentic systems using Granite models optimized for enterprise use. IBM’s platform supports prebuilt and custom agents that automate tasks across HR, finance, and IT, with orchestration across Salesforce, Microsoft, and other enterprise tools. But here’s the catch: No matter how advanced your platform is, agents are only as smart as the data they use. To reason effectively, adapt to real-world inputs, and deliver value, agentic AI systems need access to clean, contextual, and trustworthy data. That’s where your enterprise data becomes the catalyst for truly intelligent behavior. ## What does agentic AI need from your data? Low-code and no-code platforms make it easier than ever to build and deploy agentic AI. But here’s the real constraint: agents are only as good as the data they have access to. Before an AI agent can reason, act, or collaborate, it needs something to reason with. And that something isn’t just “data” — it’s governed, structured, contextualized, and accessible data. Without this foundation, even the most advanced AI agents will produce inconsistent or inaccurate results. So what, exactly, does agentic AI need from your data? We believe the five foundational requirements are: - Strong data governance - Structured, transformed data - Semantic clarity and consistency - Governed, accessible interfaces - Feedback loops and observability ### Strong data governance Bad data leads to bad decisions—only now, they happen faster. Agents rely on governed, high-integrity data to reason accurately and comply with internal policies and external regulations. Governance ensures the right people (or agents) have the right access to the right data, and that access is monitored, auditable, and secure. It’s not just about compliance — it’s about ensuring trust in every autonomous decision. ### Structured, transformed data Raw data ≠ ready data. Before agents can reason over data, it must be modeled, tested, and contextualized. That’s where [dbt](https://www.getdbt.com/product/dbt) shines. dbt transforms messy source data into structured, analytics-ready models with [built-in testing](https://docs.getdbt.com/blog/test-smarter-not-harder#why-testing), [lineage](https://docs.getdbt.com/docs/explore/column-level-lineage), and [semantic meaning](https://docs.getdbt.com/docs/use-dbt-semantic-layer/sl-faqs) —exactly what AI agents need to work reliably. ### Semantic clarity and consistency AI agents can’t make smart decisions if key metrics mean different things across your org. Terms like “revenue,” “churn,” or “active user” must be precisely defined — and consistently applied. A [shared semantic layer](https://www.getdbt.com/product/semantic-layer), whether provided through a data governance tool, data platform, or like the one dbt provides, ensures agents reason over consistent definitions, avoiding metric drift across departments. ### Governed, accessible AI interfaces Agents access data via APIs, semantic layers, or metadata endpoints — but access should never mean exposure. Platforms like [Snowflake](https://snowflake.com/), [Databricks](https://databricks.com/), and [Salesforce](https://salesforce.com/) offer native access controls. With dbt’s new [Model Context Protocol (MCP) Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server), AI agents can securely retrieve models, lineage, and semantic context from your dbt project — enabling fine-grained, policy-aware access to production-grade data. ### Validation and observability Autonomous agents improve through continuous interaction — but without monitoring and feedback, they can easily go off the rails. Observability ensures that you know what the agent did, why it did it, and what happened next. Techniques like [testing](https://www.getdbt.com/product/test-and-observe), logging, alerting, and human-in-the-loop review are essential for [catching hallucinations](https://www2.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2025/autonomous-generative-ai-agents-still-under-development.html), avoiding bias drift, and managing error propagation in multi-agent systems. Feedback loops don’t just protect you — they make your agents smarter over time. ## How dbt can help accelerate your agentic AI Agentic AI systems rely on structured, trustworthy, and context-rich data to reason and act effectively. dbt plays a critical role in ensuring your data is [AI ready](https://www.getdbt.com/product/ai). dbt’s powerful capabilities can help you accelerate time-to-value of your agentic AI initiatives further in all five of the above areas. ### Supporting strong data governance dbt encodes [governance](https://www.getdbt.com/product/governance) directly into the transformation layer. By combining modular SQL modeling with version control, testing, and documentation, dbt projects become transparent, auditable systems of record. [dbt data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) provides a holistic view of how data moves through an organization by mapping dependencies between different assets, such as models, sources, and tests. dbt also integrates with platforms like [Atlan](https://atlan.com/dbt/), [Alation](https://www.alation.com/blog/alation-dbt-integration/), and [Secoda ](https://www.secoda.co/glossary/data-lineage-for-dbt)to surface lineage, ownership, and policies. ### Creating structured, transformed data dbt [transforms](https://www.getdbt.com/product/what-is-dbt) messy inputs into clean, analytics-ready models—with built-in documentation, testing, and lineage—giving agents the structure they need to reason clearly and reliably. dbt offers cloud-native scalability with cloud services like [Snowflake](https://www.getdbt.com/blog/data-pipelines-snowflake-dbt), [BigQuery](https://www.getdbt.com/data-platforms/bigquery), and [Databricks ](https://www.getdbt.com/data-platforms/databricks)to enable dynamic scaling, ensuring AI models receive optimized, high-performance data. ### Building semantic clarity and consistency The dbt [Semantic Layer](https://www.getdbt.com/product/semantic-layer) defines shared business concepts like “active customer” or “net revenue” in a governed, reusable format. This ensures agents across teams apply the same logic, avoiding metric drift and conflicting outputs. ### Creating governed interfaces for multi-agent orchestration With the advent of multi-agent agentic collaboration, the need for well-governed interfaces between your data and AI systems has become even more critical. The dbt [MCP Server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) enables this by exposing dbt models, lineage, and semantic context via the open [Model Context Protocol (MCP)](https://modelcontextprotocol.io/introduction). With MCP Server, instead of relying on fragile integrations, agents can query dbt directly in a standardized, governed format. Paired with orchestration platforms like [Kestra](https://kestra.io/docs/use-cases/data-pipelines), agents can trigger dbt runs alongside model training, validation, and reporting tasks to enable fully autonomous, governed workflows. ### Enabling validation and observability While AI agents may be autonomous, they still need supervision. dbt’s native testing capabilities [evaluate the validity of AI responses](https://docs.getdbt.com/blog/ai-eval-in-dbt) before agents act on them. dbt’s built-in tests and logging capabilities catch anomalies, validate assumptions, and support human-in-the-loop review—creating a feedback loop that helps agents learn and improve safely over time. dbt also supports structured evaluation workflows when paired with platforms like [Snowflake Cortex AI](https://docs.getdbt.com/blog/ai-eval-in-dbt), enabling teams to compare AI responses against ground truth and trigger alerts when accuracy drops below defined thresholds. ## Conclusion Agentic AI is moving fast — and so is the demand for high-quality, trusted data to power it. While the long-term potential of AI agents is [still unfolding](https://www.ibm.com/think/insights/ai-agents-2025-expectations-vs-reality), businesses are already seeing real results. dbt has long been the industry standard for building reliable, well-governed datasets. With its support for transformation, validation, lineage, and AI-ready interfaces, dbt is a natural fit for companies looking to scale AI adoption with confidence. Ready to future-proof your AI strategy? [Book a demo](https://www.getdbt.com/demo) to see how dbt can help. --- --- title: "Understanding data transformation frameworks" description: "Learn more about what a data transformation framework is, what it includes, and how to select one that works across your business." url: "https://www.getdbt.com/blog/data-transformation-frameworks" date: "2025-02-17" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Understanding data transformation frameworks It takes work and skill to turn raw data into business insights. Data transformation is an indispensable part of this process that data engineers use to clean, enrich, integrate, and store data from multiple sources in a more accessible format. Doing data transformation well, however, requires more than just writing a bit of code. Low-quality data transformation code results in low-quality data—which, by some estimates, [costs organizations almost $13 million every year](https://www.gartner.com/smarterwithgartner/how-to-improve-your-data-quality). This is why any company doing data transformation at scale needs a data transformation framework. A framework provides tools that bring consistency to the data transformation process, enabling data stakeholders to create, track, verify, deploy, and monitor code across the analytics lifecycle. We’ll take a deeper look at what a data transformation framework is, why you need one, and how to build one with minimal upfront investment. ## What is a data transformation framework? A data transformation framework is a system of processes and tools that enables managing data transformation pipelines across your organization so you can ship new or revised analytics code with both high quality and high velocity. From a pure technical standpoint, teams can create data transformations in a number of languages and run them in a number of ways: - One group of data engineers may write scripts in Python and run them as cron jobs on a virtual server somewhere - Another might write SQL-based transformations, bundle them into a Docker container, and run them on their preferred cloud provider’s container orchestration platform - Yet another might build this code out as stored procedures in their data warehouse The problem with this scattershot approach is that it doesn't provide any centralized visibility or management of a critical corporate asset. ‌Over time, this results in: - **Lack of traceability** of analytics code changes, as some developers may not be using source code control to manage changes - **Lack of code reuse**, which forces multiple teams to solve the same problem multiple times in multiple different ways - **Lack of consistency **as teams take inconsistent approaches to transforming tables and fields, resulting in discrepancies and errors that take time to track down and fix - **Lack of quality** as some engineers run analytics code changes in production with little—and in some cases, no—testing Companies don't end up in this spot for no reason. They get here because they have to manage data over a heterogeneous collection of data sources spread across multiple groups and divisions around the globe. A data transformation framework solves for this by providing a centralized, common, and vendor-agnostic approach to writing, testing, deploying, and operationalizing data analytics code. It works across all of the major data sources and destinations used in today's modern data stacks to provide a consistent approach to creating data pipelines no matter where data lives. ## Components of a data transformation framework There are two fundamental components of a data transformation framework - one focused on methodology and the other on tools: - Data lifecycle management process - Data control plane Let's look at each one of these in detail. ### Data lifecycle management process Many analytics code changes are still performed haphazardly and driven by engineering resources. This results in variable quality of data transformation code across the organization. It also risks data transformation work becoming disconnected from business goals. In addition to a common toolset, companies need a common process. Mature analytics data lifecycle involves all data stakeholders from the get-go and focuses on shipping small, high-quality releases. At dbt, we refer to this process as the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). The ADLC is a vendor-agnostic approach to managing analytics code changes in which all data stakeholders—engineers, analysts, and business decision-makers—work together to ensure the success of each release. The ADLC defines the following stages and processes for creating high-quality data transformations: [**Plan**](https://www.getdbt.com/blog/adlc-plan). Gets all stakeholders on the same page, validates the business case, anticipates downstream impacts, sets out security guidelines, and defines success for the project. [**Develop**](https://www.getdbt.com/blog/adlc-develop). ‌Represents all analytics changes as source code governed by version control, invests in code quality (e.g., by creating [DRY analytics code](https://www.getdbt.com/blog/guide-to-dry) as well as reusable code modules), and institutes a review process to ensure high-quality releases. [**Test**](https://www.getdbt.com/blog/adlc-test). ‌Creates unit and data tests to verify data transformation logic at multiple points during the development and release process. [**Deploy**](https://www.getdbt.com/blog/adlc-deploy). ‌Uses automated processes to test and verify all changes in pre-production environments before making them live for customers. [**Operate and observe**](https://www.getdbt.com/blog/adlc-operate-observe). Test changes continuously in production; the system tolerates and recovers from failure, and engineers can use metrics to continuously measure performance and data quality. [**Discover and analyze**](https://www.getdbt.com/blog/adlc-discover-analyze). Enables analysts and business users to discover the data created by data transformations in a self-service manner, verify its data lineage, and provide feedback. **** ### Data control plane While the ADLC provides a process, it doesn't provide a toolset for implementing it. Processes such as code review, testing, and deployment of changes are easier if you have tools that can work uniformly across all the data stores in your enterprise. Without this, your teams might take varying approaches to implementing data transformation logic that lock you into a number of vendor-specific solutions and make practices such as code reuse more difficult. The next piece of the puzzle in your data transformation framework is a [**data control plane**](https://www.getdbt.com/blog/data-control-plane-introduction). The data control plane is an abstraction layer that sits across your data stack and provides unified capabilities for data transformation, orchestration, observability, and more. ‌It also centralizes metadata across your business, giving you a bird's eye view of everything happening across your data estate. A data control plane provides three major benefits: **Avoid vendor lock-in**. Provides a method for modeling and transforming data whether it’s in [Snowflake](https://www.snowflake.com/), [Databricks](https://www.databricks.com/), [Amazon Redshift](https://aws.amazon.com/redshift/), [Fivetran](https://www.fivetran.com/), [Airbyte](https://airbyte.com/), or elsewhere. A common data modeling platform means you can more easily re-share code and data modeling expertise across use cases. **Make analytics a team sport, with guardrails**. ‌Without a common data transformation framework, you can easily end up with complex data pipelines that only data engineers can run and modify. By utilizing a common modeling solution that leverages transformation languages, such as SQL and Python, you can open data transformation work to a larger number of stakeholders, including analysts and even technically savvy decision-makers. Testing, code reviews, and other quality processes provide a set of guardrails to verify code from contributors before any changes go to production. **Promotes data quality and trust**. A good data control plane can also generate [documentation](https://docs.getdbt.com/docs/build/documentation) as well as [data lineage](https://www.getdbt.com/blog/guide-to-data-lineage) for all data models. This means business stakeholders can discover and verify the origin, quality, and purpose of data in a self-service manner. This results in stakeholders making greater use of data, resulting in better and more timely decision-making. **** ## How a data transformation framework works in practice A good data transformation framework provides a data control plane for all of the data transformation logic and metrics in your business that supports and reinforces the various practices codified by the ADLC. [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) is one such data control plane that’s built from the ground up to manage trusted data at scale across your enterprise. Let’s look quickly at how dbt Cloud implements the various aspects of a data transformation framework [with a quick walkthrough](https://docs.getdbt.com/guides/snowflake?step=1). This explanation assumes you’re using Snowflake—but you can also use a number of [other popular data sources and destinations](https://docs.getdbt.com/docs/get-started-dbt). After connecting to a data source, [such as Snowflake](https://docs.getdbt.com/guides/snowflake?step=4), a data engineer can create a [dbt project](https://docs.getdbt.com/guides/snowflake?step=5#initialize-your-dbt-project-and-start-developing) and [associated Git repository](https://docs.getdbt.com/docs/collaborate/git-version-control). This places all of their code under source control, which enables traceability of changes and will also power their analytics code deployment processes. Anyone who wants to contribute to the project can work on different Git branches, keeping their changes isolated from production code (and other developers) until they’re good enough to ship. Next, the engineer defines dbt models, which define data transformation logic using either SQL or Python code. For example, the following model combines data from the customers, order, and customer_orders tables to create a unified customer order data set with valuable customer metadata attached. This is where they would implement their data transformation logic, including any cleansing, aggregation, and other common data transformations they need to perform. ```sql with customers as ( select id as customer_id, first_name, last_name from raw.jaffle_shop.customers ), orders as ( select id as order_id, user_id as customer_id, order_date, status from raw.jaffle_shop.orders ), customer_orders as ( select customer_id, min(order_date) as first_order_date, max(order_date) as most_recent_order_date, count(order_id) as number_of_orders from orders group by 1 ), final as ( select customers.customer_id, customers.first_name, customers.last_name, customer_orders.first_order_date, customer_orders.most_recent_order_date, coalesce(customer_orders.number_of_orders, 0) as number_of_orders from customers left join customer_orders using (customer_id) ) select * from final ``` Using dbt, data developers can create complex models, even importing other projects and building models on top of other models. They can also create [tests](https://docs.getdbt.com/guides/snowflake?step=12) and [documentation](https://docs.getdbt.com/guides/snowflake?step=13) alongside these models. When the engineer’s ready to commit their changes, [they file a pull request (PR)](https://www.getdbt.com/blog/how-to-review-an-analytics-pull-request) to ask for a code review. Your team can set up dbt Cloud so that this request also runs and verifies all tests associated with the engineer’s changes, providing another level of quality control. Using dbt Cloud, your team can also set up [Continuous Integration (CI) jobs](https://docs.getdbt.com/docs/deploy/continuous-integration) to further test changes before promoting them to production. Once a PR request is approved, it can trigger a CI job that tests the changes against mock data in a non-production environment. This provides added confidence that the engineer’s changes will work as expected in prod. Once the changes are live, analytics and decision-makers can discover the data via [dbt Explorer](https://www.getdbt.com/product/dbt-explorer), using the associated documentation and data lineage to guide their usage. They can then provide feedback on the finished product to drive the next iteration of the ADLC—e.g., by identifying further opportunities for data aggregation or enrichment. ## Implement your end-to-end data transformation framework today A data transformation framework consists of both processes and tools. There’s little you can do to speed up the process component; it takes hard work and time to change old habits. However, using a platform like dbt Cloud, you can implement a data control plane in far less time than it would take you to build your own from scratch. With dbt Cloud as your data control plane, your data teams have a standardized and cost-efficient way to build, test, deploy, and discover analytics code while data consumers have purpose-built interfaces and integrations to self-serve data that is governed and actionable. To see how it can work at your company, [ask us for a demo today](https://www.getdbt.com/contact). --- --- title: "How to build and manage data SLAs for reliable analytics" description: "Define, monitor, and improve data SLAs to build trust, reduce risk, and support business-critical analytics at scale." url: "https://www.getdbt.com/blog/data-slas-best-practices" date: "2025-02-17" authors: ["Joey Gault"] categories: ["Pulse"] --- # How to build and manage data SLAs for reliable analytics A data [Service Level Agreement (SLA)](https://www.atlassian.com/itsm/service-request-management/slas) is a formal commitment between data teams and their stakeholders that defines the expected level of service for data products. These agreements specify measurable targets for data quality, availability, and performance, along with the consequences when those targets aren't met. SLAs serve as contracts that establish accountability and set clear expectations for both data producers and consumers. [Service Level Objectives (SLOs)](https://www.atlassian.com/incident-management/kpis/sla-vs-slo-vs-sli), on the other hand, are the specific, measurable targets that support SLAs. While an SLA might commit to providing "reliable daily sales reporting," the supporting SLOs would define exactly what "reliable" means: perhaps 99.5% uptime, data refreshed by 9 AM daily, and accuracy within 0.1% of source systems. SLOs provide the concrete metrics that teams can monitor and optimize against. The distinction between these concepts is crucial for implementation. SLAs represent business commitments and often include consequences for non-compliance, while SLOs are the technical targets that enable those commitments. Together, they create a framework that translates business requirements into operational metrics that data teams can actively manage. ## The business case for data SLAs At the core of any data-driven organization lies trust. Stakeholders must have confidence that when they need data, it will be available and accurate. Without this trust, organizations inevitably fall back on gut-based decision making, undermining investments in data infrastructure and analytics capabilities. Data SLAs formalize this trust relationship by making reliability commitments explicit and measurable. The business impact of unreliable data extends far beyond frustrated analysts. When executives can't trust their dashboards during critical business reviews, when marketing campaigns launch with outdated customer segments, or when financial reporting is delayed due to data quality issues, the costs compound quickly. [Gartner estimates that poor data quality costs organizations an average of $12 million annually](https://www.gartner.com/en/data-analytics/topics/data-quality), highlighting the financial imperative for better data reliability practices. Data SLAs also enable more sophisticated data governance and resource allocation decisions. When data products have clearly defined reliability requirements, teams can make informed trade-offs between speed, cost, and quality. A daily executive dashboard might warrant a 99.9% availability SLA with four-hour recovery targets, while a monthly compliance report might accept lower availability in exchange for higher accuracy standards. Furthermore, SLAs provide a framework for scaling data operations. As organizations grow and data complexity increases, informal reliability expectations become insufficient. SLAs create the structure needed to maintain service quality while expanding data capabilities across the enterprise. ## Key components of data SLAs Effective data SLAs encompass several critical dimensions that collectively define data service quality. Availability represents the most fundamental component: the percentage of time that data products are accessible and functional. This includes both the underlying data infrastructure and the specific datasets or reports that stakeholders depend on. Freshness, also known as timeliness, defines how current the data must be for different use cases. A real-time fraud detection system might require data latency measured in seconds, while monthly financial reports might accept data that's several days old. The acceptable Service Level Agreement for data freshness varies dramatically based on the business context and decision-making requirements. Accuracy encompasses how closely data reflects reality and maintains consistency across different systems and time periods. This dimension often proves the most challenging to measure and maintain, as it requires understanding both the source data quality and the transformations applied throughout the data pipeline. Accuracy SLAs might specify acceptable error rates, data validation requirements, or consistency checks between related datasets. Completeness ensures that all required data elements are present and properly populated. This includes both record-level completeness (ensuring individual records have all necessary fields) and dataset-level completeness (ensuring all expected records are present). Completeness SLAs become particularly important when dealing with data from multiple sources or when supporting regulatory reporting requirements. Recovery time objectives define how quickly data services must be restored following an outage or quality incident. These objectives should align with business impact assessments: critical operational dashboards might require recovery within one hour, while analytical datasets used for strategic planning might accept longer recovery windows. **** ## Implementing SLOs with modern data tools Modern data transformation tools like [dbt](https://www.getdbt.com/product/what-is-dbt) provide excellent foundations for implementing and monitoring data SLOs. [dbt's testing framework](https://docs.getdbt.com/docs/build/data-tests) allows teams to codify data quality expectations as automated tests that run with each data refresh. These tests can validate everything from basic data integrity (ensuring key fields aren't null) to complex business logic (verifying that revenue calculations match expected patterns). The key to successful SLO implementation lies in making these quality checks an integral part of the data transformation process rather than an afterthought. When data quality tests are embedded directly in [dbt models](https://docs.getdbt.com/docs/build/models), they become part of the standard development workflow. Teams can establish quality gates that prevent poor-quality data from propagating to downstream systems, maintaining SLO compliance proactively rather than reactively. Freshness monitoring represents another area where modern tools excel. [dbt's freshness reporting capabilities](https://docs.getdbt.com/reference/resource-properties/freshness) allow teams to track when source data was last updated and alert when data falls outside acceptable staleness windows. This functionality enables teams to establish and monitor freshness SLOs without building custom monitoring infrastructure. When source data goes stale, [dbt State](https://www.getdbt.com/product/dbt-state) can prevent downstream tables from being rebuilt unnecessarily, preventing wasted compute. Conversely, a fast-changing source table can be throttled with the [`lag_tolerance`](https://docs.getdbt.com/reference/resource-configs/lag-tolerance) configuration, meaning time spent defining SLOs can immediately flow into cost savings by reducing the execution rate of lower-priority tables. For more complex SLO requirements, teams can leverage dbt's integration capabilities with specialized data observability platforms. These tools can consume metadata from dbt transformations to provide comprehensive monitoring across the entire data pipeline, from source systems through final data products. This integration approach allows teams to maintain their existing dbt-based workflows while adding enterprise-grade monitoring capabilities. ## Establishing realistic targets One of the most common pitfalls in data SLA implementation is setting unrealistic targets that set teams up for failure. The goal should be to establish achievable standards that drive continuous improvement rather than perfection. Most successful implementations start with baseline measurements to understand current performance before setting improvement targets. The concept of error budgets, borrowed from site reliability engineering practices, provides a useful framework for thinking about data SLA targets. Rather than aiming for 100% perfection, teams can establish acceptable error rates that balance reliability with operational complexity. A 99.5% availability target, for example, allows for roughly 3.6 hours of downtime per month, providing buffer for planned maintenance and unexpected issues. Different data products warrant different SLA targets based on their business criticality and usage patterns. Executive dashboards used for daily operational decisions require higher availability and freshness standards than analytical datasets used for quarterly strategic planning. Teams should work closely with stakeholders to understand the true business requirements rather than applying uniform standards across all data products. It's also important to consider the interdependencies in data pipelines when setting SLA targets. Downstream data products can only be as reliable as their upstream dependencies, so SLA targets should account for the cumulative impact of multiple processing stages. This systems thinking approach helps teams set realistic expectations and identify the most impactful areas for reliability improvements. ## Monitoring and alerting strategies Effective SLA monitoring requires a multi-layered approach that combines automated detection with human oversight. The goal is to identify and resolve issues before they impact stakeholders while avoiding alert fatigue that can desensitize teams to real problems. This balance requires thoughtful design of monitoring thresholds and escalation procedures. Real-time monitoring should focus on the most critical SLA violations that require immediate attention. These might include complete data pipeline failures, significant data quality degradations, or freshness violations for time-sensitive reports. Teams should establish clear escalation procedures that ensure the right people are notified based on the severity and business impact of different types of incidents. Trend monitoring provides equally important insights for proactive SLA management. Gradual degradations in data quality or increasing pipeline execution times often signal underlying issues that can be addressed before they cause SLA violations. Regular SLA performance reviews help teams identify patterns and invest in preventive measures rather than constantly fighting fires. Documentation plays a crucial role in effective monitoring strategies. Teams should maintain clear runbooks that describe how to respond to different types of SLA violations, including diagnostic steps, escalation procedures, and recovery processes. This documentation ensures consistent incident response regardless of who is on call and helps new team members quickly become effective in maintaining data reliability. ## Building organizational buy-in Successfully implementing data SLAs requires more than technical implementation: it demands organizational change management and stakeholder alignment. Data teams must work closely with business stakeholders to establish SLA targets that reflect actual business needs rather than arbitrary technical standards. This collaborative approach ensures that SLA investments focus on the areas of greatest business impact. Communication strategies should emphasize the business value of data reliability rather than technical metrics. Instead of reporting on pipeline success rates, teams might communicate about decision-making confidence, report availability during critical business periods, or the reduction in time spent investigating data discrepancies. This business-focused communication helps stakeholders understand the value of SLA investments. Training and education initiatives help stakeholders understand their role in maintaining data quality. When business users understand how their data entry practices, system usage patterns, and reporting requirements impact data reliability, they become partners in maintaining SLA compliance rather than passive consumers of data services. Regular SLA reviews provide opportunities to refine targets based on changing business needs and operational learnings. These reviews should include both quantitative performance assessments and qualitative feedback from stakeholders about whether current SLA targets adequately support their decision-making needs. ## Measuring success and continuous improvement The ultimate measure of data SLA success isn't perfect compliance with technical metrics: it's improved business outcomes and stakeholder confidence in data-driven decision making. Teams should track both quantitative SLA performance and qualitative indicators of data trust and usage across the organization. Quantitative metrics might include SLA compliance rates, mean time to recovery from data incidents, and the frequency of data quality issues. However, these technical metrics should be complemented by business impact measurements such as increased self-service analytics adoption, reduced time spent on data validation, and improved confidence in data-driven decisions. Continuous improvement processes should focus on addressing the root causes of SLA violations rather than just fixing symptoms. When teams consistently miss freshness targets, the solution might involve upstream system optimizations, pipeline architecture changes, or revised business processes rather than simply adjusting the SLA targets. The most successful data SLA implementations evolve into comprehensive data reliability practices that extend beyond formal agreements. Teams develop cultures of reliability consciousness where data quality considerations are embedded in every design decision and operational process. This cultural transformation represents the true value of data SLA initiatives: creating organizations where reliable data becomes a sustainable competitive advantage rather than a constant struggle. As data continues to grow in importance for business operations, the discipline of data reliability engineering will only become more critical. Organizations that invest in formal SLA practices today position themselves to scale data operations effectively while maintaining the trust and confidence that data-driven decision making requires. ## Data SLA / SLO FAQs **What is a data SLA (Service Level Agreement)?** A data Service Level Agreement (SLA) is a formal commitment between data teams and their stakeholders that defines the expected level of service for data products. These agreements specify measurable targets for data quality, availability, and performance, along with the consequences when those targets aren't met. SLAs serve as contracts that establish accountability and set clear expectations for both data producers and consumers. **Why are data SLAs important?** Data SLAs are crucial for building trust in data-driven organizations. Without this trust, organizations inevitably fall back on gut-based decision making, undermining investments in data infrastructure and analytics capabilities. The business impact of unreliable data extends far beyond frustrated analysts: when executives can't trust their dashboards during critical business reviews or when financial reporting is delayed due to data quality issues, the costs compound quickly. Gartner estimates that poor data quality costs organizations an average of $12 million annually. **What are some best practices for drafting data pipeline SLAs?** Start with baseline measurements to understand current performance before setting improvement targets, rather than aiming for unrealistic perfection. Use the concept of error budgets to establish acceptable error rates that balance reliability with operational complexity. Different data products should have different SLA targets based on their business criticality: executive dashboards require higher availability and freshness standards than analytical datasets used for quarterly planning. Consider interdependencies in data pipelines when setting targets, since downstream products can only be as reliable as their upstream dependencies. --- --- title: "Bring reliable data to every spreadsheet with the dbt Semantic Layer for Excel and Google Sheets" description: "Learn how to keep your spreadsheets in sync, no matter the sprawl." url: "https://www.getdbt.com/blog/dbt-semantic-layer-excel-google-sheets" date: "2025-02-11" authors: ["Chakshu Mehta"] categories: ["Product"] --- # Bring reliable data to every spreadsheet with the dbt Semantic Layer for Excel and Google Sheets Love them or hate them, spreadsheets aren’t going anywhere. Tools like Google Sheets and Excel remain essential for decision-making across organizations, whether it’s finance forecasting revenue or marketing analyzing campaign performance. But, as your datasets grow, so do the challenges of ensuring reliable, consistent data across systems. Multiple versions of spreadsheets, inconsistent calculations, and manual imports create confusion, inefficiencies, and ultimately, poor decisions. For data analysts and engineers, the struggle is just as real—especially when it comes to complex calculations. Preparing data for spreadsheets often means querying subsets from cloud data warehouses, cleaning dimensions, and manually calculating metrics, especially for complex analyses. Business users need scalable, governed, and easily accessible data to make informed decisions. And data teams need to deliver it efficiently without creating silos, adding pipeline latency, or risking burnout. ## The pain of disconnected workflows in spreadsheets Take an analyst preparing a quarterly sales report. They pull raw data from the warehouse, clean it, aggregate sales by region, and calculate key metrics like average deal size or revenue per salesperson. While the report may look great initially, any new updates to the source data—like adding late-logged sales—render the workbook outdated. Now, the analyst has to re-extract data, redo the cleaning and aggregations, and reapply the calculations, every single time. This repetitive process can take hours and introduces delays, errors, and inconsistencies. For example, if the analyst refines how they calculate customer churn or segment product categories, those transformations must be manually reapplied every time new data arrives. What started as innovative analysis now feels like a repetitive chore. Without a centralized system, their methodology is difficult to track, and any undocumented steps become even harder to reproduce. These disconnected workflows not only waste time but also make it difficult to maintain accuracy or scale as data needs grow. Teams end up working with outdated extracts and repetitive and siloed processes, leading to bottlenecks and frustration. ## A smarter way to work with spreadsheets With the right approach, spreadsheets can remain a powerful and dependable tool, free from the chaos of inconsistent and outdated data. Instead of relying on static tables and manual processes**,** the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) simplifies how organizations access and work with trusted data by serving as a intermediary layer between your data warehouse and spreadsheets, ensuring users are accessing fresh, governed, and pre-defined metrics to maintain consistency and accuracy. ![dbt Semantic Layer](https://cdn.sanity.io/images/wl0ndo6t/main/4b7fb9995b43c05d573731ddb0c66208aaed051c-1296x1024.webp) With built-in integrations for [Google Sheets](https://docs.getdbt.com/docs/cloud-integrations/semantic-layer/gsheets) and [Excel](https://docs.getdbt.com/docs/cloud-integrations/semantic-layer/excel), business users can pull fresh, trustworthy data directly from cloud platforms like [Snowflake](https://docs.getdbt.com/guides/snowflake?step=1), [BigQuery](https://docs.getdbt.com/guides/bigquery?step=1), and [Redshift](https://docs.getdbt.com/guides/redshift?step=1). This eliminates the need for manual imports, rigid OLAP cubes, and complex Multidimensional Expressions (MDX) queries (the language Excel uses to query cubes). Traditionally these methods require extensive infrastructure, pre-aggregations, and frequent refreshes, leading to stale data and high maintenance costs. Instead, the dbt Semantic Layer bypasses MDX entirely, and dynamically translates spreadsheet queries into SQL under the hood, so teams can query even their most complex calculations in real-time without creating extra static tables or writing custom queries. This means no more hardwired definitions or misaligned metrics—just reliable, up-to-date data that teams can trust. Whether calculating revenue growth, analyzing customer churn, or building pivot tables, dbt simplifies workflows and ensures consistency across every tool. ## How the dbt Semantic Layer works for Excel and Google Sheets ### 1. Consistent, always up-to-date metrics dbt Cloud’s transformation capabilities, combined with the [dbt Semantic Layer](https://docs.getdbt.com/guides/sl-snowflake-qs?step=4), provide a unified and trusted foundation for organizational data. dbt Cloud allows teams to transform raw data from cloud data platforms into clean, version-controlled, analytics-ready datasets, while the dbt Semantic Layer centralizes and standardizes _business metrics and definitions_ on top of these datasets, making them accessible across tools like Excel and Google Sheets. Metrics such as “year-over-year revenue growth” or “net retention rate” are defined in the dbt Semantic layer on top of dbt models and accessed across tools. This creates a single source of truth allowing teams to work with accurate and consistent data, whether they’re using Google Sheets, Excel, or any other connected application. And because spreadsheet queries run directly against the warehouse, reports automatically update with the latest data—no more exports or duplicate work. ### 2. Seamless integrations with Excel and Google Sheets dbt connects directly to your cloud data warehouse, eliminating the need for OLAP cubes, MDX queries, or manual imports. With an intuitive built-in query builder, users can easily select pre-defined metrics, apply filters, and retrieve the exact the data they need in seconds. This ensures that pivot tables in Excel or custom formulas in Google Sheets function seamlessly with accurate, up-to-date, governed data. By delivering clean, consistent datasets directly into these tools, dbt empowers teams—technical and non-technical alike—to confidently perform analyses without relying on manual and siloed updates to business logic. ![SI Google sheets](https://cdn.sanity.io/images/wl0ndo6t/main/a4b212d6c3b260e5d881edcc16b0df12edf04741-3800x1662.png) ### 3. Fast insights for complex data needs For teams handling large datasets or complex financial calculations, the dbt Semantic Layer enables high-performance multidimensional queries directly on live warehouse data. Business users can build pivot tables, slice and dice across dimensions, and collaborate with consistent metrics—all without relying on pre-aggregated cubes or static tables. Built-in query optimization features in the dbt Semantic Layer enhance performance with features like: - [Common query caching](https://docs.getdbt.com/docs/use-dbt-semantic-layer/sl-cache), which reduces latency by storing frequently accessed results. - [Pushdown](https://docs.getdbt.com/docs/dbt-cloud-apis/sl-jdbc#query-with-where-filters) computations that offload complex calculations to the warehouse, leveraging its processing power, while filtering - And query pruning which ensures only necessary data is queried, improving speed and efficiency. This approach ensures faster, more accurate insights and reduces the overhead of managing and refreshing static data extracts. Teams can focus on analysis rather than infrastructure, driving better outcomes. ### 4. Governed data that aligns teams With centralized [governance](https://docs.getdbt.com/guides/sl-partner-integration-guide#governance-and-traceability), dbt ensures all teams work with consistent, accurate data, reducing the chaos of spreadsheet sprawl. By providing a single source of truth for business logic, the dbt Semantic Layer eliminates manual processes and repetitive maintenance, allowing allowing anyone in the business, from finance users to analysts, to trust that their reports are aligned across the organization. This governance empowers teams to self-serve their data needs with ease, enabling faster insights and freeing analysts to focus on strategic tasks rather than firefighting data inconsistencies. With dbt, teams can operate with greater confidence and reliability. ## Make quality data accessible to everyone The dbt Semantic Layer redefines what’s possible with spreadsheets. With built-in [integrations](https://docs.getdbt.com/docs/cloud-integrations/avail-sl-integrations), automatic updates, and centralized metric governance, teams are empowered to move faster, collaborate more effectively, make better decisions, and can finally trust their spreadsheets. This is the new standard for trusted, scalable spreadsheet workflows. Book a [demo](https://www.getdbt.com/contact) today! --- --- title: "dbt Labs Surges Past $100 Million in Annual Recurring Revenue, Driven by Significant Adoption from Fortune 500 Companies" description: "More than 5,000 organizations rely on dbt Cloud to power enterprise data practices" url: "https://www.getdbt.com/blog/dbt-labs-100m-arr-milestone" date: "2025-02-05" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Surges Past $100 Million in Annual Recurring Revenue, Driven by Significant Adoption from Fortune 500 Companies **PHILADELPHIA, Feb. 5, 2025** -- [dbt Labs](https://www.getdbt.com/), the pioneer in analytics engineering, announced today that it has surpassed $100 million in annual recurring revenue (ARR), scaling from $2M ARR to this milestone in only four years. Simultaneously, the company has exceeded the 5,000 customer mark, with 85% year-over-year growth in adoption among Fortune 500 companies. This momentum is fueled by dbt’s critical role as the data control plane for enterprise data teams around the world, who rely on the technology to transform data into reliable, actionable business insights. “When my cofounders and I created dbt eight years ago, we were addressing our own frustrations while working with data. In the process, we pioneered the modern practice of analytics engineering, the fusion of software engineering best practices with traditional analytics," said Tristan Handy, founder and CEO at dbt Labs. “It was what data teams needed, and it led us to both this milestone and the path forward for our data control plane. I’m grateful to the dbt Community, our customers, employees and partners for contributing to our progress, supporting our vision, and continuing to rely on dbt.” dbt Labs’ vision for One dbt – a commitment to a single unified dbt experience, regardless of the persona, data platform, or cloud that an organization uses – demonstrates its resolve to deliver the experiences and capabilities users need to build, manage, govern, and maximize the impact of their crucial analytics workflows. This has generated strong demand for [dbt Cloud](https://www.getdbt.com/product/dbt-cloud), the data control plane that supports users across every stage of the [Analytics Development Lifecycle](https://c212.net/c/link/?t=0&l=en&o=4272622-1&h=545613714&u=https%3A%2F%2Fwww.getdbt.com%2Fresources%2Fguides%2Fthe-analytics-development-lifecycle&a=%C2%A0Analytics+Development+Lifecycle). Through dbt Labs’ scalable technology, organizations are able to overcome major hurdles related to data quality, data velocity, and data governance, allowing them to move faster while having far more trust and confidence in their data. The company’s [recent acquisition of SDF Labs](https://www.getdbt.com/blog/dbt-labs-announces-sdf-labs-acquisition) also brings robust SQL comprehension into dbt, supercharging developer productivity, heightening data quality, and optimizing platform costs. These developments are critical as enterprises continue to launch AI projects, building demand for better data quality and carefully governed inputs. In addition to its $100 million ARR milestone, dbt Labs also recently surpassed 5,000 dbt Cloud customers (30% year-over-year growth). The expanding list of organizations that rely on dbt include [Condé Nast](https://www.getdbt.com/case-studies/conde-nast), [Hubspot](https://www.getdbt.com/case-studies/hubspot), [Nasdaq](https://www.getdbt.com/case-studies/nasdaq), and [Siemens](https://www.getdbt.com/case-studies/siemens). dbt Labs’ existing customer base is also investing heavily in the dbt Cloud platform, with 90% year-over-year growth among customers at or above the $100,000 ARR level. The company’s recent innovations, many announced during its [annual Coalesce event](https://coalesce.getdbt.com/), have paved the way for this traction. Among these: [dbt Copilot](https://www.getdbt.com/blog/introducing-dbt-copilot), a new [visual editing experience](https://www.getdbt.com/product/develop) and [cross-platform dbt Mesh](https://www.getdbt.com/blog/introducing-cross-platform-dbt-mesh). “With dbt, we’ve empowered teams by providing the tools and techniques needed to take ownership of their data, which has led to faster workflows, improved data quality, and higher engagement across the board,” said Brian Lloyd-Newberry, AVP Architecture, [Cox Automotive](https://www.coxautoinc.com/) Data Portfolio. “This shift has driven significant successes, enabling our teams to deliver greater value with confidence and precision.” Sunny Pachunuri, Head of Data Engineering and Platform at [Endpoint](https://www.endpoint.com/), added: “dbt Cloud is the data control plane at the center of my organization’s data strategy. Since adopting dbt, we have increased our productivity by 75% and reduced costs by almost 80%, and are far more agile,” he said. “It has completely changed the way we create data products, ship them in the organization, and build trust with stakeholders. We are invested in growing our data strategy in step with dbt’s innovation roadmap.” dbt Labs also continues to cultivate relationships with strategic industry partners including AWS, Databricks, Google Cloud, Microsoft, Snowflake, and Salesforce. In fact, the company recently announced [a new collaboration](https://www.getdbt.com/blog/dbt-labs-and-salesforce-announce-strategic-partnership) to tightly integrate Salesforce Data Cloud AI automation and analytics solutions with dbt Cloud, enabling customers to maximize the collective potential of dbt, Salesforce and Tableau platforms. For these organizations and the hundreds of others in the dbt Labs partner ecosystem, close collaboration is critical to delivering the best experience to customers. “At Slalom, we’re passionate about empowering organizations to unlock the full potential of their data, and dbt Labs has been an invaluable partner in that journey,” said Ryan Gifford, Managing Director of Data+AI at [Slalom](https://slalom.com). “Their innovation in analytics engineering has transformed how data teams operate, delivering the scalability and governance that modern enterprises need. We’re eager to continue working alongside dbt Labs to shape the future of our joint customers and partners, helping them achieve meaningful outcomes with their data.” dbt Labs has also continued to expand its corporate footprint, with customers located in 43 countries (27% year-over-year growth) and new offices in Austin and Dublin. In addition, the company is concentrating expansion efforts in APAC. In-market employees in Japan are now joining dbt Labs teams in Australia and New Zealand to support the growing demand from the region’s customers and partners. The industry is celebrating dbt Labs’ trajectory and impact, with the company recently being named to the [2024 Deloitte Technology Fast 500](https://www2.deloitte.com/us/en/pages/technology-media-and-telecommunications/articles/fast500-winners.html) list (based on revenue growth rate) for the third consecutive year. In addition, it has been named finalist in the inaugural [2025 Tech Innovation CUBEd Awards](https://www.thecube.net/awards), to be announced in February 2025. “dbt Labs accelerates the time to value for data practitioners worldwide," said Matt Miller, venture partner at Sequoia Capital and dbt Labs Board of Directors member. “dbt simplifies the process of transforming and understanding complex sets of data. The product works across all the cloud data warehouses as a simplifying layer and brain of sorts. dbt is fueling new business analytics, AI models and the AI apps changing the world. Its future couldn't be brighter.” For more information on dbt Labs, visit: [https://www.getdbt.com/](https://www.getdbt.com/) . **About dbt Labs** Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 50,000 teams using dbt every week. To learn more about dbt Labs, visit [getdbt.com](https://www.getdbt.com/) and follow us on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/). --- --- title: "Building reliable data products with dbt Cloud and SYNQ" description: "Instabee migrated critical models into dbt and reduced processing time from 8 hours to 1 hour for key models." url: "https://www.getdbt.com/blog/building-reliable-data-products-with-dbt-cloud-and-synq" date: "2025-02-04" authors: ["Hrishi Kulkarni"] categories: ["Partnerships"] --- # Building reliable data products with dbt Cloud and SYNQ ## **Data at Instabee** Instabee simplifies the way consumers and merchants ship and receive parcels. With a presence in six countries, and serving thousands of online merchants, including ASOS, Zalando, Inditex, and H&M, Instabee is on track to become the leading European e-commerce platform, reaching more than 45 million consumers across Europe. Instabee is taking a technology-first approach to online shipping and near real-time data is used for key processes and decisions – from operational data such as understanding lead time, volume, and forecasting to financial data on what is revenue-generating versus not. ### Key challenges - **Missing or inaccurate data** in dashboards used by terminal managers frequently impacted key dashboards and operational KPIs, putting merchant retention at risk - Instabee had no **central documentation** or visibility into data assets, code, and how they connect, which created a lack of trust ### Key wins - By combining dbt tests with SYNQ anomaly monitors, Instabee **significantly reduced detection time** and now resolves most operational issues within 5 minutes - Established clear data **product ownership and a live health overview**. Teams across business functions, like finance and operations, are directly notified of relevant data issues - Migrating critical models into dbt **reduced processing time** from 8 hours to 1 hour for key models, enhancing efficiency and enabling faster insights for the business > "If we lose too many merchants due to poor data quality, we lose business—and then our jobs. It’s that simple." – Paul Flynn, former Head of Data ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/6f6c13da8642d98408a4940d1356aaef1c53af5a-1600x807.jpg) ## Adopting dbt Cloud and SYNQ to build reliable operational data products After the Instabox and Budbee merger, Instabee wanted a best-in-class, modern data platform and adopted Snowflake as the data warehouse, Fivetran for ingesting data, Tableau for BI, dbt Cloud for data modeling and documentation, and SYNQ for data observability. > "We have data people working across the spectrum – from analysts doing ad-hoc and strategic analysis, analytics engineers building data models, data scientists running ML models and forecasting algorithms, and tech teams relying on data. Our platform needed to be able to support all of these stakeholders." – Josefin Gruvander, Analytics Engineer ## **Building reliable, well-architected, and documented data models in dbt** The team knew right away that they wanted dbt Cloud as part of the stack. Today, all central reporting and data modeling rely on dbt Cloud. Dozens of jobs run everything from the daily production run to forecasting models. Instabee has taken a deliberate approach to data modeling, creating consistent layers in the data architecture to make sure sources, staging models, and data marts are distinguishable. “This makes it faster to add new models and easier to reason about potential root causes and see how different tables connect in the lineage,” Josefin says. dbt Cloud is the control plane to develop, deploy, and document data models. dbt Explorer is the source of truth for understanding and exploring data models. “We are rigorous in documenting our data models and fields in dbt metadata and use dbt Explorer across the company to expose documentation. This helps everyone relying on data go to one place to find the information they need,” Josefin says. As a result, Instabee has significantly improved ‌data trust and transparency for everyone outside the data team. The development workflow has sped up drastically since adopting dbt Cloud. The team relies heavily on features such as version control to review new changes, SQLFluff for linting for faster development, CI/CD to make sure changes are tested before deployment, and dbt Explorer’s built-in lineage to understand dependencies. "We recently moved a 'black box' critical finance model into dbt. Since then, everyone can see all dependencies in the lineage and we’ve reduced the time it takes to build the model from 8 hours to 1 hour helping us save money and reduce the time to insight," Josefin says. ## **Using data products to manage business-critical data** “Data products have become the lens through which we evaluate and reason about our most important business processes in data. We group data products into areas such as BI and Finance and can instantly see if there are any errors on or upstream of data products,” Josefin says. "The Data Product overview in SYNQ is the first page we open each morning to check if all the nightly runs have run successfully or if there are any errors across dbt and SYNQ anomaly monitors impacting our key data products," Josefin says. > “Each data product has a priority ranging from P1 to P3 based on its importance and we use that to decide how urgently we treat issues. With this at hand, the data team brings transparency to the business – from discovering and understanding data assets in dbt Explorer to seeing a live overview of the health of data products in SYNQ.” - Josefin Gruvander, Analytics Engineer Data products are closely tied to ownership. The analytics engineering team is notified of issues on core models in the data warehouse. But ownership isn’t limited to just the data team. The finance team is notified in the #finance-data-quality-monitor Slack channel if there are issues with finance data products. They’ve also extended it to operational use cases where key models rely on manual input data from spreadsheets for fuel data – input issues on these spreadsheets trigger not_null or unique dbt test errors and are routed directly to the operations manager responsible for the spreadsheets. ## **The data product reliability workflow with dbt Cloud and SYNQ** Early on, the data team at Instabee knew they wanted data tooling on par with what they had in engineering – especially when detecting and resolving issues fast and learning from incidents. **Reduced time to detection:**_ _“We take testing seriously and learned that the best way to catch both ‘known unknowns’ and ‘unknown unknowns’ was to combine dbt tests with automated SYNQ anomaly monitors. Today, we run more than 1,000 dbt tests each day and combine that with 600 SYNQ anomaly monitors running on key sources and tables every 30 minutes. Combined, these help us be the first to know about issues in most cases,” Josefin says. **Reduced time to resolution:**_ _The restructured dbt architecture makes it easier to trace back issues to source systems and model our data. The combination of dbt Cloud and SYNQ has significantly sped up debugging workflows. “Especially SYNQ’s column-code lineage and the ability to select multiple columns gives us a really good idea about how everything is connected,” Josefin says. > "We are now able to solve most issues within 5 minutes of learning about them—something that could have taken us hours in the past. This is a significant factor in us retaining key merchants." – Paul Flynn, former Head of Data **Learning from incidents:**_ _The team relies on SYNQ’s incident management functionality as a knowledge base to follow up and track issues across their data stack. “Some incidents tend to reoccur so it shortens our time to resolve and ability to mitigate issues when we have a log of previous issues and incidents,” Josefin says. **Incidents managed in SYNQ span both dbt test and model errors, and SYNQ anomaly monitors.** "We typically declare a handful of incidents in SYNQ each month. Over time, this has become our knowledge base with valuable information on how we tackled issues the last time they occurred," Josefin says. > “Since adopting dbt Cloud and SYNQ, we rarely have data issues impact of operational KPIs, our team has been freed up from the majority of firefighting and we have created the transparency we needed – from documentation to cross-team ownership.” - Josefin Gruvander, Analytics Engineer Want to learn more about data at Instabee? [**Watch their talk from the 2024 Data Innovation Summit: Creating a culture of shared ownership & 5-minute issue resolution times**](https://www.youtube.com/watch?v=JN4dMu61K20) --- --- title: "One dbt: Data collaboration built on trust with dbt Explorer" description: "How to use dbt Explorer to bridge across different stakeholder’s needs, building trust and encouraging collaboration." url: "https://www.getdbt.com/blog/dbt-explorer-collaboration-trust" date: "2025-02-03" authors: ["Alexis Jones", "Roxi Pourzand"] categories: ["Product"] --- # One dbt: Data collaboration built on trust with dbt Explorer At dbt, we believe a large part of trust is aligning teams that are invested in data. This can be challenging, as different data stakeholders approach data from different perspectives. Consider a developer, Dante. He's a data producer. He really wants to know how he can better understand, troubleshoot, and improve the quality of his data pipelines so that he can effectively serve his stakeholders. On the other side, we have Kathy. She's a consumer of data. She's really wondering how she can gather context, build trust in the data, and reuse it so she can make decisions quickly. Even if you’re using a platform like dbt to build, transform, and harden data pipelines and code, that doesn’t mean that your data is always easy to maintain or draw insights from. Used as part of One dbt, [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) bridges this gap by supporting data discovery, transparency, and collaboration - not just for data engineers but for _all_ data stakeholders. We’ll discuss how dbt Explorer fits into a modern analytics workflow and the key features that enable teams to collaborate on data, regardless of their needs or perspectives. ## dbt as the standard Before we get into the details of dbt Explorer or even dbt Cloud, it's useful to talk about the future that we envision for dbt as a standard. Last year at Coalesce, we shared the vision for what it means for dbt to be the standard. By “standard,” we mean the best and chosen framework for data transformation. Standardizing means that different teams aren't solving the same problems in slightly different ways. We can actually share solutions, align faster, and collaborate more seamlessly because we all speak the same language, the language of dbt. And that's regardless of your cloud provider data platform, what team you sit on, or whether your team uses dbt Cloud or dbt Core. We call this vision One dbt. With One dbt, for example: - A product team in Seattle can deploy and run dbt on [AWS](https://aws.amazon.com/) - Meanwhile, their colleagues in engineering over in Madrid run dbt on [Azure](https://azure.microsoft.com/en-us/) - One dbt operating in the cloud environment or environments that make the most sense for the business - A data science team running [Databricks](https://www.databricks.com/) can directly reference a dbt project that the finance team manages over in [Snowflake](https://snowflake.com) - A central data team can build data models in a CLI using DBT Core, and those data models can then be investigated and built upon by a downstream marketing ops team that uses dbt Explorer and the new visual editing experience in dbt Cloud It’s all just One dbt - one central platform disseminating knowledge throughout the business. ### The Analytics Development Lifecycle (ADLC) Providing better collaboration across data stakeholders first requires making sure that everyone is working together as part of the same team. That’s why dbt has heavily promoted what we call the [**Analytics Development Lifecycle (ADLC)**](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). ![ADLC](https://cdn.sanity.io/images/wl0ndo6t/main/948eb5cda47eacf1bf78a2268c0666be43706ce5-4581x2126.png) The ADLC is our recommended and standardized approach for a mature analytics process. It’s a **vendor-agnostic framework** designed to help organizations of any size mature their analytics workflows. The ADLC encourages collaboration among various stakeholders. It’s designed to help data producers, data consumers, and - ultimately - the business ship and use trusted data products at speed and scale. The ADLC has eight distinct phases: - [Plan](https://www.getdbt.com/blog/adlc-plan) - [Develop](https://www.getdbt.com/blog/adlc-develop) - [Test](https://www.getdbt.com/blog/adlc-test) - [Deploy](https://www.getdbt.com/blog/adlc-deploy) - [Operate and Observe](https://www.getdbt.com/blog/adlc-operate-observe) - [Discover and Analyze](https://www.getdbt.com/blog/adlc-discover-analyze) The ADLC borrows heavily from the Software Development Lifecycle (SDLC), which became popular in the early 2000s to help cross-functional teams work better together with more agility, velocity and ultimately impact to the business. The SDLC was very successful at breaking down the barriers the industry had built between software engineers who were building applications and the IT professionals who maintain the systems that those applications ran on. The new approach gave both roles a standardized, repeatable framework by which to work better together. It’s high time that analytics professionals had a similar revolution - one that accelerates and hardens data workflows for the data engineers, analysts, and data stakeholders who turn business requirements into data-driven reality. ### The data control plane However, the ADLC isn’t enough. Powering it requires a single, uniform way to access your data. The modern data stack looks like an eye chart, with a myriad of solutions - data orchestration, observability, data catalogs, semantic stores, etc. - springing up over the pats decade. While all this has been great progress for the industry and our maturity, all of these add-ons ultimately create [data silos](https://www.getdbt.com/blog/how-dbt-can-help-solve-4-common-data-engineering-pain-points). Centralizing these metadata silos will make or break your analytics workflow. We believe that the solution that accomplishes this is a data control plane. A [**data control plane**](https://www.getdbt.com/blog/data-control-plane-introduction) sits across your data stack, unifying capabilities for orchestration, observability, and more. dbt’s data control plane centralizes this metadata across the business, giving you signals on what's happening in your data estate - all supercharged with AI. The data control plane helps you understand: - Is your data fresh? - Is your data platform cost-optimized? - Is everyone running from a common understanding of how business metrics are defined? While the ADLC is vendor-agnostic, the data control plane is a vendor-backed technology solution built to embrace various phases of the ADLC. A good data control plane implementation should have three defining characteristics: 1. It should be **flexible and cross platform** so that it can power distributed teams, help organizations avoid vendor lock-in, and manage data platform costs 2. It should make data **streamlined, accessible, and governed** to more types of users - not just your data engineering team 3. It must **produce trustworthy outputs** - i.e., stakeholders need to understand where data comes from, how to improve it, how to troubleshoot it, and trust that the data that they're receiving is fresh and error-free. ## How dbt Explorer bridges the gap To trust something, you have to understand it. Consider, for example, cooking. You trust a recipe because you understand the ingredients, the cooking process, and the expected outcome. You might not know every chemical reaction, but you know the basics of heat, timing, and seasoning well enough to trust what comes out of it. (This analogy only works for us for cooking. We still haven’t figured out baking.) Data is the same way. For an organization to foster a culture of data collaboration, all the people expected to collaborate around the data - both the producers and consumers - need to understand where the data comes from, how it's used, and what its current quality is. Teams also need a standardized way to troubleshoot problems and fine-tune the overall workflow. This is where [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) comes in. **dbt Explorer is the catalog for your dbt ecosystem**. It can help not only your producers of data, but also your consumers of data. Its goal is to help all data stakeholders discover existing assets, view lineage, be able to troubleshoot and optimize your pipelines, and build knowledge and context about the data estate. ### dbt Explorer collaboration features dbt Explorer has many features that help foster this part of the process foundationally. Let’s dig into a few of the key ones in detail. #### Resource pages [Resource pages](https://docs.getdbt.com/docs/collaborate/explore-projects) are a foundational part of Explorer because every asset has a page. This can be an important starting point when you're doing discovery - whether it's a dbt model, a source, or a metric. ![Resource pages](https://cdn.sanity.io/images/wl0ndo6t/main/4789a3b07570bf0fbef93a4a8e3844b45c31fa68-2048x1466.png) Here you can understand the contextual information about an asset, like what it means. Through its description, you can validate its health. It shows overall health scores (which you can see at the top of the page here in the image) as well as specifics on quality and test results (which you can see in the middle). Most importantly, these resource pages help you view the [data lineage](https://www.getdbt.com/blog/guide-to-data-lineage) around the asset so you can see where the data comes from and in general how it flows and its dependencies. #### Auto-exposures One of our newer features, called [auto-exposures](https://docs.getdbt.com/docs/collaborate/auto-exposures), integrates natively with your BI tools to add dashboards that are actually built off of your dbt models. These are generated automatically and added to your data lineage. ![Auto exposures](https://cdn.sanity.io/images/wl0ndo6t/main/112cb86597681fcf929580b3891923d1e4768f3f-2048x1150.png) This is an incredibly valuable feature because you can see in what downstream assets the models you have are consumed in. This helps you improve assets that are high visibility. After all, you don't want to break your CFO's dashboard - you want to make it better! You can also identify what’s not consumed in any way downstream and either remove or find a different use for them. #### Model query history Continuing our theme of understanding, [model query history](https://docs.getdbt.com/docs/collaborate/model-query-history) is another important feature that helps with trust because it tells users how frequently models are being consumed. This is visible, not only in the Performance section of our interface, but also through Lineage Lenses, where a heat map allows you to zoom in and out on the hotspots. ![Model consumption (performance)](https://cdn.sanity.io/images/wl0ndo6t/main/45bdd65cfc861d98ce059a7d3789dc6ccd48a57a-1874x1010.png) ![Query count lineage lens](https://cdn.sanity.io/images/wl0ndo6t/main/814d188f24e64d7d720911e619dea5067037910f-2048x1153.png) Similarly to auto-exposures, this can help you understand what are your most popular assets, so you can continue to maintain them and make sure the quality is good. Also like auto-exposures, it can help you prune unused models and save costs. Model query history is currently supported for Bigquery and Snowflake, and we're working on Redshift and Databricks next. #### Data health tiles Finally, to tie this story on understanding together, we have [data health tiles](https://docs.getdbt.com/docs/collaborate/data-health-signals) and in-app trust signals. Data health tiles are tied to your exposures, since that’s the unit of how things are being consumed in dashboards and reports for your downstream assets. But these are incredibly useful because they're just iframes you can embed directly into the dashboard where they're being consumed. ![Data health tiles](https://cdn.sanity.io/images/wl0ndo6t/main/32a3fff7520a709e729689bce31bf13754d6b354-1492x644.png) Data health tiles meet users where they are, signaling to them whether the asset is healthy. In the application, we also have high level trust signals based on a number of factors, such as test status, freshness, and usability. In other words, whether you're in application or outside, you'll be able to understand the relative health of the data you're looking at. ## See dbt Explorer in action A picture’s worth a thousand words. A video may be worth even more. To see more of dbt Explorer in action, [view the full webinar](https://www.getdbt.com/resources/webinars/one-dbt-data-collaboration-built-on-trust-with-dbt-explorer) showing how dbt Explorer works as part of One dbt to drive a mature analytics workflow and enable data discovery, transparency, and collaboration. --- --- title: "The power of a plan: How logical plans will impact modern data workflows" description: "Investigating the benefits that come from using a compiler to power your development workflows." url: "https://www.getdbt.com/blog/how-logical-plan-impact-modern-data-workflows" date: "2025-02-02" authors: ["Grace Goheen", "Elias DeFaria"] categories: ["Insights"] --- # The power of a plan: How logical plans will impact modern data workflows _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/the-power-of-a-plan-how-logical-plans). _ If you work with SQL, you are used to working with a compiler – you just might not know it yet. You've probably seen a compiler error message from your warehouse like this: ```sql $ dbt run --select my_model [...] 22:29:56 Database Error in model my_model (models/marts/my_model.sql) 001044 (42P13): SQL compilation error: error line 4 at position 11 Invalid argument types for function 'DATE_ADDDAYSTOTIMESTAMP': (TIMESTAMP_LTZ(9), NUMBER(1,0)) ``` This error is because the database's SQL compiler received invalid input. Compilers are programs that translate high-level language into a logical plan (read more in [our post about the key technologies behind SQL Comprehension](https://docs.getdbt.com/blog/the-levels-of-sql-comprehension)). Historically, this capability has been constrained to the data warehouse. But the error above wasn’t just in a warehouse! It was surfaced in dbt - a data transformation tool. This presents an interesting dichotomy: - Your **data warehouse** produces logical plans internally for one query at a time - Your **transformation tool** knows the dependency graph of every query at once (your DAG!) Imagine if you could bring the compiler into your development workflow. You'd know the implications of your changes as you worked, instead of when you hit run. We believe combining the powers of SQL Comprehension and DAG Comprehension will make data practitioners more effective. It will [move us up the stack](https://www.getdbt.com/about-us/values#we-believe-in-moving-up-the-stack). To do this right, you need the right tech - an accurate, multi-dialect, performant SQL compiler. Today we’ll break down three major benefits of a development workflow powered by a compiler. 1. **Validate**: write provably correct SQL queries, saving time and money by catching errors in SQL _before_ running them against a cloud warehouse 2. **Analyze**: with access to comprehensive, type-aware metadata you can understand how the data flows from ingestion to consumption, powering experiences like precise column-level lineage, description propagation, and more 3. **Optimize**: improve performance on both the query and engine level ## Validate: Will my SQL _really_ work?[​](http://localhost:3000/blog/the-power-of-a-logical-plan#validate-will-my-sql-really-work) To validate your SQL code is to guarantee it will run and process data successfully (as described in [this post on the levels of SQL comprehension](https://docs.getdbt.com/blog/the-levels-of-sql-comprehension) that unlock different types of validation). The compiler can validate everything except the data itself. This means by the time the compiler has produced a logical plan, it has by definition also validated the SQL. But most compilers validate one query at a time, in isolation. **The opportunity in a developer experience powered by a compiler is to validate all queries that make up a data pipeline.** What does this mean for you? Let’s start with a time saver - precise error messaging _before_ executing your query in a cloud warehouse. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/746f6b0ce3a15182b476033b6b8616af678a4ea3-1456x401.jpg) A compiler understands all mechanisms that can change the data type of a column, including functions like UDFs, operators, and implicit conversions. In the query above, the compiler looks and says, “hey nobody told me what to do if I see a `round()` function with a boolean passed in, I’m out of here!” and reports an error message. This is a simple example - but most of the time you're not referencing a boolean directly, you're referencing a column. If you can analyze all queries at once and understand their dependencies, you can determine when modifying a column in one model will break something downstream. This is called _impact analysis._ Let’s say we have two models: ```sql -- model_1.sql select some_numeric_data as my_column from {{ source('my_source', 'my_table') }} ``` ```sql -- model_2.sql select round(my_column) as rounded from {{ ref('model_1') }} ``` The SQL in `model_2` is valid - it takes `my_column` which contains numeric data and rounds it. But what if you tried to change the datatype in `model_1` to a boolean? Maybe it's a 1/0 indicator and you want it to be converted to true/false instead. ```sql -- model_1.sql - select some_numeric_data as my_column + select cast(some_numeric_data as bool) as my_column from {{ source('my_source', 'my_table') }} -- model_2.sql select round(my_column) as rounded from {{ ref('model_1') }} ``` The updated SQL in `model_1` is still valid. `model_2`, however, is now broken — something you’d only discover when you run the DAG. By having full awareness of both your DAG and each query's logical plan, you can understand the impact of your change during development and prevent errors in all downstream dependencies. Once we fix our DAG to make sure the SQL is valid, what’s next? Time to analyze it. ## Analyze: Generating precise column-level lineage and beyond[​](http://localhost:3000/blog/the-power-of-a-logical-plan#analyze-generating-precise-column-level-lineage-and-beyond) Now that our compiler has validated the SQL and produced a logical plan for each query in our DAG, an analyzer can step in to extract valuable metadata. During analysis, user-defined metadata (descriptions, tests, classifications) is combined with the Logical Plan to power data catalogs, automated data governance, smart caching, and more. ### Precise Column-Level Lineage[​](http://localhost:3000/blog/the-power-of-a-logical-plan#precise-column-level-lineage) By coupling SQL Comprehension and DAG comprehension with one another you could unlock: - a full and accurate understanding of the column dependencies in your data pipeline, informing debugging workflows and better development - even _slimmer_ CI builds by only building downstream models if they depend on the exact columns that were changed - downstream impact analysis by knowing if changing a column in one model will break a downstream dashboard What enables all of those experiences? _Precise_ column-level lineage. Column level lineage traces the changes to your columns throughout your DAG to give practitioners a deeper view into how their data is being transformed. In the example below, we can see all of the column selections and modifications that lead to the `last_activity_at` column in `fct_users`. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/711c57a38d35faa9c03d6cdbafaccae9f594f7e4-1456x930.jpg) Historically the state of the art for column level lineage has been derived from parsing (level 1 SQL comprehension, what we have in dbt Explorer today). Generating this reliably has been slow and required a lot of additional work on the backend. But remember that the logical plan already has information about every column and every type in your DAG, and can be generated in a snap. This makes it an even better input for CLL: **using the logical plan results in completely accurate lineage of not just every column but even the properties of every struct, variant, or object inside a column.** ### Beyond CLL: Information Flow Theory and Metadata propagation[​](http://localhost:3000/blog/the-power-of-a-logical-plan#beyond-cll-information-flow-theory-and-metadata-propagation) By leveraging principles from [information flow theory](https://en.wikipedia.org/wiki/Information_flow_(information_theory)), the compiler can track not just the physical or logical operations on columns, but also how data “flows” from one stage to another and how different transformations might affect its classification. For instance, if a particular column is marked as containing sensitive information at its origin, the compiler can propagate that classification across subsequent transformations in the DAG. This gives organizations a systematic way to maintain compliance requirements, enforce security policies, and deeply understand data provenance—making it far easier to see exactly which downstream datasets or reports are derived from a given source column and ensuring that sensitive information is handled appropriately across the entire data pipeline. When we combine CLL with the metadata associated with a given column (name, description, whether or not it’s tested, etc.), we can propagate this metadata all the way from ingestion to consumption. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f8e6fd8bb32ee6dd6413c170fc8b286e512f1c08-1456x814.jpg) Propagation lets you put your metadata to use: - **Naming propagation:** faster development - rename a column in one place and have it updated in every SQL query that references it - **Test propagation:** save compute by removing duplicative tests - if you’ve already validated `column_a` is unique, don’t re-check for uniqueness unless your SQL query changes that column - **Description propagation:** DRY-er code and more complete documentation - define a column description once and it flows everywhere that column is referenced ## Optimize: Right query, right place[​](http://localhost:3000/blog/the-power-of-a-logical-plan#optimize-right-query-right-place) Logical plan analysis opens up several doors to optimize our workflow. Looking at this through the lens of a compiler, which has been disaggregated from the compute engine itself, we can explore two distinct approaches to query optimization. The first is to optimize the query for the engine of choice. The second is to optimize _the choice of the engine_. ### Query optimization[​](http://localhost:3000/blog/the-power-of-a-logical-plan#query-optimization) All modern widely-adopted OLAP query engines come with impressive query optimizers. However, the expressiveness of SQL can be a burden here, as code written with the best of intentions can accidentally undercut these optimizations. Instead of relying on organizational knowledge to avoid these pitfalls, compilers can automatically warn and even modify queries to ensure they produce optimal queries. Let's look at an example in action - Query Pruning. Query pruning allows you to make sure that your warehouse only scans the subset of data referred to in your query - ie if you're looking at just one day of data in a table it doesn't need to scan the whole table. Through static and dynamic analysis of the query, your data platform can scan drastically smaller subsets of the data (i.e. micropartitions) while still computing the same result. Type conversions, however, can sometimes prevent your data platform from pruning your query. For example, take the following query: ```sql select * from {{ ref('my_orders') }} where to_varchar(order_date, 'yyyy-mm-dd') = '1992-01-31' ``` At first glance, this makes sense. Your date on the right hand side is rendered as a [ISO-8601 string](https://xkcd.com/1179/), so you want to make sure the `order_date` values are formatted the same way. Unfortunately, this cast to `varchar` will be a major performance hit. The direction of the cast in this query is backwards - by casting the entire column’s values to `varchar`, Snowflake can’t use its date-specific pruning capabilities. It’d be better to cast the `'1992-01-31'` string to a date; and, in fact, left to its own devices, Snowflake would implicitly do just that, meaning the following query is significantly more performant: ```sql select * from {{ ref('my_orders') }} where order_date = '1992-01-31' ``` That's the type of query optimization you'd want your compiler to flag - helping you get out of the way and let the data platform optimizer do its thing. Let’s look at another example with CTEs. [Thanks to our friends at SELECT](https://select.dev/posts/snowflake-ctes), we know that referencing a CTE more than once prevents another type of query pruning: column pruning. Consider this query, where we `select *` from a model (an “import CTE”). Using it in two different CTEs will take longer than necessary to run on large datasets: ```sql with american_sales as ( select * from {{ ref('sales') }} where region = 'AMERICA' ), -- Calculate the total sales for all products total_sales_america as ( select sum(amount) as total_sales from american_sales ), -- Calculate the total sales for each product product_sales_america as ( select product_id, sum(amount) as product_total_sales from american_sales group by product_id ) ... ``` As you can see, `total_sales_america` only needs the `amount` column, and `product_sales_america` needs both `amount` and `product_id`. If we only had the first CTE (`total_sales_america`), Snowflake would recognize that the `*` in the import CTE was overkill and only scan the `amount` column from the `sales` model. It would _prune_ the rest of the columns from its table scan (that's how column pruning gets its name!). But in this case, Snowflake will scan _all_ columns – even though `product_sales_america`'s column requirements are a superset of `total_sales_america`'s. Heeding the compiler’s warning, we can improve the query’s performance by specifying only the columns we need: ```sql with american_sales as ( select product_id, amount from {{ ref('sales') }} where region = 'AMER' ), -- Calculate the total sales for all products total_sales_america as ( select sum(amount) as total_sales from american_sales ), -- Calculate the total sales for each product product_sales_america as ( select product_id, sum(amount) as product_total_sales from american_sales group by product_id ) ... ``` Static analysis of the logical plan opens up a plethora of opportunities to optimize a single query, but like we mentioned before, databases are already pretty good at this. Where this gets even better is when you can understand all queries at once, even if they’re written for different engines. ### Engine optimization[​](http://localhost:3000/blog/the-power-of-a-logical-plan#engine-optimization) Five years ago it was rare to hear a data team operating on multiple distinct data platforms. Today, not so much. This is in large part fueled by the adoption of open table formats like Iceberg, which unify the data storage format across warehouses and unlock an abundance of opportunities to build [cross-platform data workflows](https://www.getdbt.com/blog/introducing-cross-platform-dbt-mesh). But today, these workflows are missing a crucial component: a multi-dialect compiler. This compiler allows data teams to fine-tune their DAG execution. They can place their queries on a sliding scale between cost and performance based on their priorities and use case. This can even be done on a single engine, dynamically selecting the optimal warehouse size for a query based on its data load history and plan complexity. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/03e90f744c15b9fa4b7d8d94921db30da7db7e08-1456x811.jpg) The average dbt user runs all their development queries, unit tests, and CI checks against the same massively scalable engine they use to run their DAGs in production. However, most of these are querying relatively small datasets, especially when using the upcoming [sample mode](https://github.com/dbt-labs/dbt-core/discussions/11200) in dbt Core. Using a nimble single-node engine like Apache DataFusion for these workflows would make them faster and more cost-efficient, since on small datasets with low complexity, these engines often perform better and can remove network overhead. Imagine introducing a new model and testing it on a speedy single-node engine - how would you guarantee it behaves the same when run in production? You could manually inspect the output of potentially thousands of rows, or you could **guarantee identical behavior by having an intelligent compiler produce the same logical plan for the query** that would be created for the production engine. Without guaranteed conformance, you’d have to _literally run all your queries on all engines_ to be certain you were getting the same data - not a very likely workflow, so instead you’re probably accepting poor data quality as part of your process. Multi-dialect validation is a requirement of an engine that will reliably power cross-platform data workflows. ## Conclusion[​](http://localhost:3000/blog/the-power-of-a-logical-plan#conclusion) A system that can analyze both the DAG of transformations and the logical plan of each transformation is more powerful than systems that can only understand one or the other. We are very excited about what compiler-driven experiences are going to power across the ecosystem. --- --- title: "Building the next-gen dbt engine: How SDF levels up data tooling" description: "So dbt Labs acquired SDF Labs. What does that mean, exactly?" url: "https://www.getdbt.com/blog/building-the-next-gen-dbt-engine" date: "2025-01-30" authors: ["Jason Ganz", "Jeremy Cohen"] categories: ["Product"] --- # Building the next-gen dbt engine: How SDF levels up data tooling Two weeks ago we shared the news that [dbt Labs has acquired SDF Labs](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs). We heard the same reaction from many of our customers, partners, and community members: 1. This is very exciting! 2. What does it mean, exactly? So today, we wanted to take some more time to explain. ## What will this do? With the integration of SDF under the hood, dbt will be both much **faster** and **significantly more cost-efficient** — while unlocking new metadata use-cases like true column-level lineage. As a standalone tool, SDF made modern data development easier than ever, thanks to some seismic technical innovations. Over the past week, we began peeling back the layers on those innovations for you in a series of posts, sharing what makes them possible, and what they can unlock in your data workflows. In the first post, we covered [SQL comprehension](https://docs.getdbt.com/blog/the-levels-of-sql-comprehension) — why it matters, and how it can be done at _three distinct_ levels of precision. ![The 3 Levels of SQL comprehension](https://cdn.sanity.io/images/wl0ndo6t/main/8b8009e8e2479a36acfb20ed9cceb5c3af792c91-3288x1580.png) Then we followed up by diving into [why achieving the highest level of SQL comprehension is a devilishly complex technical challenge](https://docs.getdbt.com/blog/sql-comprehension-technologies), which requires several distinct layers of underlying technology. In this weekend’s Roundup [we unpacked the power of a Compiler](https://roundup.getdbt.com/p/the-power-of-a-plan-how-logical-plans) — how the logical plan enables it, and what it might do for data transformation workflows. A lot of technical explanations! But now you have the context to know why this matters. ## dbt should know about SQL Thanks to the strength of the dbt Community, over the past 9 years, dbt has achieved widespread adoption and defined the modern data development experience. And as dbt has scaled to be the standard worldwide, we’ve had the opportunity to learn — collectively and in public — about a lot of things that have worked and are worth keeping, and some things that are ready for an update. We also get the privilege of rethinking foundational choices from 2016 with the benefit of 2025 technology. We’ve been specifically following the topic and tooling around SQL Comprehension for some time now. This is a problem we’ve been _excited_ to solve so that we could elevate the dbt developer experience further. But we didn’t just want to solve it. We wanted to do it right - we should have immediate wins for dbt users and real technical depth. And we needed to know that we were introducing a durable solution that is going to stand the test of time. ## A new engine for dbt Last year, we met the team at SDF, and we knew they had something special. Their approach to SQL comprehension operates at _all three_ of the levels of comprehension we shared in the first post above: 1. SDF is a parser, with syntactic support for several major dialects (and more on the way) 2. It’s a compiler, capable of precise validation and calculation of column-level lineage that’s fast at scale 3. And it’s an executor, leveraging Apache DataFusion’s to enable local development workflows SDF was built by a hyper-talented team of people with world-class **expertise in the technologies required to enable SQL comprehension at every level. They know how to build an engine of the technical depth that befits being the industry standard. And it’s _these same people_ who are building a new engine for dbt. Importantly, this will all be done under the same code authoring layer that we’ve all spent the last decade building together — the one learned by practitioners the world over and adopted by tens of thousands of companies: the One dbt standard that enables [collaboration across dbt Cloud and dbt Core](https://www.getdbt.com/blog/uniting-core-and-cloud-with-one-dbt). Right now the dbt Labs team is heads down on this integration to make it the best it can be. We’re motivated to start sharing progress soon - [tune into our Spring Launch event on March 19th](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase) for more info. We’re excited about just how much this new engine will unlock when it’s ready: many things we’ve all dreamed of having in dbt, soon to be within our grasp. Hold on folks - this is going to be a fun one. --- --- title: "January dbt Community update" description: "Catch up on January's dbt Community highlights: AMAs, Meetups, Slack discussions, and more. Join the conversation today." url: "https://www.getdbt.com/blog/january-2025-dbt-community-update" date: "2025-01-28" authors: ["Kathryn Chubb"] categories: ["Community"] --- # January dbt Community update Welcome to the January edition of the dbt Community Update, your monthly roundup of all things happening in the [dbt Community](https://www.getdbt.com/community). Highlights included an insightful AMA with Paige Berry and Lauren Benezra, six in-person [dbt Meetups](https://docs.getdbt.com/community/spotlight), and a ton of great discussions across Slack. Are you ready for the recap? Let’s get started. ## Community Slack AMA Each month, we host a live Ask Me Anything (AMA) event in the [#dbt-community-ama](https://getdbt.slack.com/archives/C025ZN1L679/p1716308109072499) channel on Slack, bringing together analytics engineers, data enthusiasts, and industry experts for open and honest conversations. This month, we featured Paige Berry, Lead Data Analyst at dbt Labs, and Lauren Benezra, Lead Analytics Engineer at dbt Labs. They shared their journeys, experiences, and advice on navigating the world of data. Here are some of the key takeaways: 1. **Diversity in data careers: **Both speakers debunked the myth that a specialized degree is required to break into data. Drawing from her own journey, Lauren highlighted how her background in applied math took her from biotech to building impactful data systems. Paige emphasized leveraging individual passions, encouraging aspiring data professionals to view their unique experiences as an advantage. Their advice? Start where you are, and let curiosity guide you. 2. **The critical need for backups: **Sharing lessons from past data challenges, the duo underscored the importance of disaster management. Paige illustrated this with a candid anecdote: “I once deleted ZIP codes from a database—something you only let happen once before you prioritize backups.” 3. **Teamwork drives data excellence:** Effective collaboration was a recurring theme. Lauren shared her role in empowering analysts by creating clean, reusable data models, while Paige underscored how this partnership accelerates insights: “Clear roles and strong teamwork enable us to transform complex SQL into actionable stories.” 4. **The value of continuous learning: **Both speakers championed mastering foundational tools like SQL and Python, but they didn’t stop there. They encouraged practicing with real-world datasets and refining the skill of storytelling with data to stand out in the field. 5. **Facing impostor syndrome head-on:** Acknowledging the common struggle, Paige and Lauren shared their strategies for building confidence. From documenting achievements to mentoring others, they offered practical ways to overcome the doubt that can creep into even the most seasoned professionals. ### Catch the full conversation If you missed the live AMA, don’t worry—you can catch the full recording here. [Watch video](https://youtu.be/4Zb5d91ibow?si=gCrYfFQEWS-3hv_U) ### Get ready for another AMA in February Our next AMA is just around the corner. [Register now](https://www.getdbt.com/resources/webinars/community-ama) to join the conversation live, and don’t forget to join the [#dbt-community-ama](https://getdbt.slack.com/archives/C025ZN1L679/p1716308109072499) Slack channel to stay in the loop. ## January dbt Meetups In January we had six dbt Meetups: Paris, Austin, Raleigh, Tokyo, Florianópolis, and Taipei. The Paris dbt Meetup on January 22nd was a great success, filled with energy and excitement from over 60 attendees who braved the winter cold, rain, and wind of that day to attend. dbt Labs’ Jeremy Cohen spoke about how the dbt Labs product team leverages dbt to analyze feature adoption and identify key challenges faced by its users. He also shared more about dbt Labs’ [acquisition of SDF Labs](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs) and led a Q&A session. ![Paris meetup 1](https://cdn.sanity.io/images/wl0ndo6t/main/1385032fbc6b09cd7f7b0965b7f59f2f5dbcfa83-4032x3024.jpg) ![Paris meetup 2](https://cdn.sanity.io/images/wl0ndo6t/main/974a75bf3b8ae3a8176dab3b6650513ca8431ddd-4032x3024.jpg) These events continue to be a cornerstone of the dbt Community, bringing members together to share knowledge, network, and collaborate. Stay tuned for even more Meetup opportunities in February by checking out our [Meetup page](https://www.meetup.com/pro/dbt/). ## What’s something you wish you’d known when starting out? The Community team recently launched a new weekly program called Local Data Exchange, where they post a question prompt in all 70+ #local channels. This month, our Community Engagement Manager, Bolaji, asked “What’s something you wish you’d known when starting out?”. Here are some of our favorite responses. ![Local data exchange 1](https://cdn.sanity.io/images/wl0ndo6t/main/d0353e3a2da11418841cdf72b9782b78d8bf4e7d-960x540.png) ![Local data exchange 2](https://cdn.sanity.io/images/wl0ndo6t/main/e14705c889562d1db56173eaa27e0402e825f784-960x540.png) ![Local data exchange 3](https://cdn.sanity.io/images/wl0ndo6t/main/b75c7f133dc34f9779c915a5916abf2c7fdb2e20-960x540.png) ![Local data exchange 4](https://cdn.sanity.io/images/wl0ndo6t/main/2ef1cef31be73cb47e2b051e91449ad70778e905-960x540.png) ## Slack channel spotlight: #memes-and-off-topic-chatter If you haven’t joined the dbt Community Slack yet, you’re missing out. It’s where analytics engineers, data enthusiasts, and practitioners at all experience levels connect in real time. With channels ranging from beginner-friendly discussions to deep dives into niche technical topics, there’s a space for everyone. Sometimes we all need a break from debugging SQL or perfecting our dbt models. Enter [#memes-and-off-topic-chatter](https://getdbt.slack.com/archives/C0VLNUUTZ), the channel that reminds us that data professionals have a great sense of humor too. Whether you’re looking for a quick laugh, a clever take on the latest data trend, or just some light-hearted banter to break up your day, this channel delivers. It’s also a great way to discover that analytics engineers share more than just a love for clean data—we share a love for good laughs too. ![Panda meme](https://cdn.sanity.io/images/wl0ndo6t/main/2669a32b8c4c0f99498b7d823ec0e37154d7ff5f-642x570.png) ![Interesting dilemma meme](https://cdn.sanity.io/images/wl0ndo6t/main/02411dd50c38b7fc1f1ae56085ca288f10802ae4-479x400.png) ![Daily active users meme](https://cdn.sanity.io/images/wl0ndo6t/main/7a8e8ced39290f1b44175a51d7e13f6d18088f14-613x450.png) So, what are you waiting for? Join the [dbt Community Slack](https://www.getdbt.com/community/join-the-community), hop into #memes-and-off-topic-chatter, and start sharing your favorites. Who knows? Your meme might just be the one that goes viral (at least in our Slack). ## Community announcements We’ll wrap up this month's update with some of the exciting announcements that are regularly posted in our [#announcements](https://getdbt.slack.com/archives/C0VLZM3U2/p1715777392876319) channel on Slack. ### Upcoming events - In case you missed the news, we recently acquired SDF Labs. Tune into the [Accelerating dbt with SDF webinar](https://www.getdbt.com/resources/webinars/accelerating-dbt-with-sdf) to learn more about the acquisition - [Register for our next Community AMA](https://www.getdbt.com/resources/webinars/community-ama) in February - Ongoing [Cloud Demo with Experts](https://www.getdbt.com/resources/dbt-cloud-demos-with-experts/) in North America, EMEA, and APAC-friendly times - Add [dbt Events](https://www.addevent.com/calendar/Tb314369) to your calendar ### Upcoming dbt Meetups We’ve got a busy month coming up with five [in-person dbt Meetups](https://www.meetup.com/pro/dbt) scheduled. If you’re looking for opportunities to learn with fellow members of the dbt Community, and have fun while doing so, join us at one of the sessions listed below: - 🇪🇸 Barcelona | Thursday, February 13th, organized by [Spaulding Ridge](https://www.linkedin.com/company/spaulding-ridge-llc/) - 🇹🇼 Taipei | Wednesday, February 19th, organized by community members [Karen Hsieh](https://www.linkedin.com/in/karenhsieh/), [Laurence Chen](https://www.linkedin.com/in/humorless/), [Allen Wang](https://www.linkedin.com/in/allenwangs/), and [LI KUAN LIAO](https://www.linkedin.com/in/li-kuan-liao-03a619117/) - 🇧🇪 Belgium | Thursday, February 20th, organized by community members [Sam Debruyn](https://www.linkedin.com/in/samueldebruyn/) and [Lise Kerckhove](https://www.linkedin.com/in/lise-kerckhove-data-analytics/) - 🇺🇸 Chicago | Thursday, February 20th, organized by [Analytics8 | Data & Analytics Consultancy](https://www.linkedin.com/company/analytics8/) - 🇯🇵 Tokyo | Friday, February 21st, organized by community member [Shinya Takimoto](https://www.linkedin.com/in/shinya-takimoto-2793483a/) There are so many exciting things going on in the dbt Community, and we can’t wait to see you all there. If you haven’t yet, [join the community](https://www.getdbt.com/community) today. --- --- title: "Building a data team from the beginning" description: "Daniel Avancini discusses how fast-growing Indicium went from Brazilian beach town to global data consultancy." url: "https://www.getdbt.com/blog/building-a-data-team-from-the-beginning" date: "2025-01-26" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Building a data team from the beginning _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/building-a-data-team-from-the-beginning). _ Daniel Avancini is the chief data officer and co-founder of [Indicium](https://www.indicium.tech/)—a fast-growing data consultancy started in Brazil. There are a lot of data consultancies around the world, and a lot of them do great work. What has been so fascinating about Indicium’s journey is their HR model. Rather than primarily hiring experienced professionals, they decided to go hard on training. They built a talent pipeline with courses and an internal onboarding process that takes new employees from zero to 60 over a few months. The result has been phenomenal and Indicium delivers great client outcomes, but most importantly, they're building skills for hundreds of brand new data professionals. Data is a hard field to break into because fundamentally you can't do the real thing unless you have access to data. So any company investing in building scalable hiring and training processes for analytical talent is one to be excited about. **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways from this episode ### Can you give a little bit of an introduction to you and to Indicium? Yeah, sure. So I'm the co-founder and CDO of Indicium. We're a data consultancy. Now we're based in New York, but we also have a presence in Latin America and Brazil where we started. We mostly focus on the data stack and new data stack tools. We've been helping business companies use modern data platforms and move to new data stack tools, including dbt, for about seven years. We are a young company, but not that young in the modern data stack world. ### Tell me about your journey starting in Latin America to expanding internationally. We started in a small city in Brazil called Florianopolis. It's like a tech center, like San Francisco. There are many new companies there, but it's not a business data consulting space. We really started from the beginning; we started with smaller, mid-sized regional companies, really trying to find something that made sense. So we pivoted a lot in the beginning on how we could deliver value. ### You skipped the step “we wanted to start a company.” What was the original idea? I was working for a startup in agricultural hardware machinery. Nothing related to data or services in general. My cofounder was managing a surfboard manufacturing factory. ### That is wild. I love that. You can come to data from so many different backgrounds, including surfboard manufacturing. It's a beach town, so there's a lot of surfing there. We realized at the beginning that there were a lot of technology platforms for marketing analytics, data intelligence, SaaS tools. But when we talked to anyone that was making decisions, no one was really using that data. Our first insight was there’s a need for someone in this market to bridge the gap and bring all this really great data to companies in a more organized way and in a value-driven way. At first that was our goal. We were not focused on building a data platform consultancy. But as we grew, we found out that it's harder than we thought. We needed to do a lot of foundational work, especially on smaller companies. All they had were Excel spreadsheets. Databases, SQL databases, and a lot of Excel spreadsheets, and a lot of the complex analysis we wanted to do, was just not ready. We helped them build platforms and foundations for these companies. ### And it took off like a rocket ship? What's a sense of your scale that you want to share? Yeah, we're at around almost 400 employees right now. So we're pretty big for this market. ### What has allowed you to become successful at the scale that you have been? I think what really helped us scale is that since the beginning, we have really focused on building our own teams and our own capabilities to scale. Even before we started using dbt or any of the modern data stack, we already thought, because we are in a smaller market, we couldn't compete for data engineering talent. So at this age in Brazil, probably the same in the US, it was a very competitive market for data engineering in general. There was a data engineering talent pool in Brazil, but it was expensive. But there were a lot of training programs on the internet. There are a lot of data camps, Udemy, Coursera. There's so much good stuff there. But maybe there's a lack of curation, right? People want to work in this area, but how do they start? What do they have to do? So we really focus on building that talent pool, our own talent pool right from the beginning. We were at the university just bringing good people, good talents for engineering, from economics, from business. “Hey guys, do wanna work with data?” Look at this program; it's free. Just go there and train. And then we would hire maybe one or two of the best ones. We would bring like 10 or 15 people. They would come to the program, we would hire the two or three best ones. Maybe four years ago we were starting to grow faster and we needed more people. We needed a more stable source of talent. First, we built our own analytics engineering course. We found dbt and realized this is the way we grow because we don't need to hire experienced data engineers. We can hire experienced marketing analysts and train them. ### It's such a consulting hack, right? I've been excited to hear your story because I think it is so parallel to our own story. We were doing a similar thing in that we were hiring people with no data experience. We can grow much faster because we can hire analysts in general. We can train anyone. It was so hard to train Airflow and Spark at that time. But if we use dbt, we can just teach these guys how to work with data analytics. And so I built the Analytics Engineering Formation, our first course. And what we did in this course wasn't only dbt. We trained about dimensional modeling, ETL, a lot of foundational analytics work that we weren't seeing when we're trying to hire people. But everyone at that time wanted to be a data scientist. But that's not the work. For every data science, you're going to find 30 data analytics engineers because there’s so much more work with analytics. We've trained more than a thousand people with this course in the past four or five years. And a lot of those people, a lot of these talents we hired, so they would do the course and then we're like, yeah, we have an open position. Do you want to work for us? And so we started really hiring from this course for the analytics engineering profession. And it really worked. And we still use the same course today for our own team, Everyone has to do the course so they understand what we do. There's a practical exam, so they need to build their own data warehouse with dbt by themselves. ### My best guess is that there's probably a million or so humans in the world that have used dbt pretty regularly. In the grand scheme of things, that’s not a big number when you're a giant consulting organization and you have a huge hiring pipeline. Building a practice that puts dbt at the center of it can work really well, but you have to really build the business model around it. But I would argue dbt is not the only one. If you think about data science, about data engineering, all these other data professions, it's really hard because there's no undergrad. People don't graduate on airflow engineering. Everything they use at work, they learn after they start working. ### Yeah. And so the point is you have to build a talent pipeline that teaches people how to do the stuff as opposed to expecting it to already exist. ### One of the things that people don't fully understand, unless they've been through this journey, is that it is an unbelievable level of investment to do what you've done. Consulting businesses don't generally raise a ton of venture money. It's a real strategic investment, but two, it's a real risk. ### If for some reason this doesn't work, that's a giant problem for Indicium, I would imagine. And that means that if you're going to make this type of investment, you have to feel like you have control over the technology that you're choosing. I imagine that it would be very hard for you to make this type of investment in something that was not not open source. Is that a true statement? Probably yes. Especially because a lot of these tools, I can't really pay for the tool when I'm educating and when I'm teaching. Maybe if I have a partnership, but yes, for a lot of the work we needed to use some kind of open-source tool for this work. ### That makes total sense. I didn't even think about it from a seat's perspective. Let's say that you were gonna use Amplitude or something like that. You would have to figure out how to get whatever, 100 people per semester access to Amplitude and that requires partnership. And also we had to build our own courses because if I needed to use market courses like Udemy, I would have to pay for all these courses for all of these students and then it becomes too expensive. So we had to invest a lot of our time just building our own training programs and our own training. ### What other tools did you incorporate into the standard training? So what we did after a few years is instead of just training dbt and analytics engineering, we created another program we call the Lighthouse program. When we open positions for analytics engineering, we get all kinds of people just because they are engineers. Then we're like, what kind of engineering? “I'm a chemical engineer.” Okay, but do you know anything about data? “No, but I'm an engineer.” We had so much work on teaching because it's such a new market. People don't know what the work is. A lot of the undergrads. They still don't understand what an analytics engineer does. So the idea of the program was be a lighthouse. Like, I'm going to show you the best career for you. After a person joins the program, we're going to tell you, you're going to be a data engineer because of their competencies. ### It's like the sorting hat in Harry Potter. And is that about skill sets or personality or interests or what? Yeah, I really looked into skill sets, personality, and we did some personality traits tests. ### Okay, so tell me what's the personality of a data scientist versus an analytics engineer? Okay, that's a good one. So what I did on this, I look into being very innovative, like looking to innovation, new things. You want to build new things. I want to build new things all the time, but I also want to build reliable things. The new things side is for data scientists, like experimenting, experimenting, building new things. On the other side of the spectrum, data engineers. So I usually put the analytics engineers kind of in the middle. Like I want to build stuff, but I also want to have reliable pipelines. And I want to build things that are closer to business. And I want to understand the value of what I'm building. Do you prefer to bring to make something new, but unstable or do you prefer to have something that works every time? Just that question would filter the personalities for these professionals really well. ### I really identify with that so much. I'd be curious to hear where you fall in this spectrum, but I am a deeply impatient person and so I can't stay on one thing too long. I love making pipelines and getting them to a certain point. But then I'm like, okay, let me try something else where I'm learning about the business. Having this like bi-modality, I think is what keeps me forever engaged in this work. Yeah, you should probably ask Matheus, my co-founder, because he always says the same thing, but I'm really closer to the data scientist when I build new things. My background is in economics and statistics and data science. Our CTO is an engineer. He's a data engineer. He's angry if something doesn't work. ### Okay, so you developed a course. You developed an ability to funnel people into what? Data scientist, data engineer, analytics engineer? Now we also have data analysts, so more in the BI, it's like an analytics engineer, with a deeper BI knowledge. Analytics engineering, data science, AI engineering, data engineering, and we are also adding a data consultant career. We have all these tracks. And this program is now a six-month program, and we are paying for them to study. So that's also very also risky for us. ### I'm glad you said it's risky. One of the things that I think you don't recognize until you run a consulting business is that it is terrifying to face attrition. Attrition is the thing that kills your business. People will quit, life. Things happen. This is the world. ### But when somebody quits, it's not only revenue walking out the door, but it is also your investment in them as a human walking out the door. One reason why I think it is very rare for companies to invest in people is that they are going through this J curve where they are not making money at first. Then you're slowly working your way out of it. ### I don't know if you've done the math, but I think you could figure out how much you'll pay back for the training you've put in. For some of these programs, it's not only training. One of the reasons we built the program is that people need to work on something. They need practice. We put them in internal projects. We have phases. So the first phase is the foundational phase, just learning. We teach them databases, APIs, cloud computing. These things are important for anyone who works with data now, but not important for someone who just graduated from college or is changing careers. Then we have theory, or the data journey, that's where they train on dbt and data engineering techniques. And then we gradually put them into projects, into real work. They can be a copy of someone, they can work on their own projects, and sometimes they can actually work on real projects. We can even charge them for some clients. Most times we are able to pay for the program with the work they do inside the program. So we have a break even before they graduate from the Lighthouse. ### I really feel like so much innovation is business model innovation. To me, what you're describing is a new business model that lets you invest in more people who know how to do great data work. This is very cool. Yeah, that's what we thought. You can always invest in technology. But in our work in consultancy, it's really humans, right? How can you get the better humans and how can you get them to stay at your company? How can you keep them, right? We've been very intentional on how to create the social structures that you would find in working in an office. We have a very low attrition rate. If you compare to the market, we are like six, seven, eight times lower than a competitor in our attrition rate. And that's compared to Brazil, which has a lower attrition rate than the U.S. If you compare to the US, it's like 20 times lower than the U.S. --- --- title: "dbt Labs on dbt: How our training team uses dbt Cloud" description: "See how our training team uses dbt Cloud to connect training data to business impact and prove the ROI of customer education." url: "https://www.getdbt.com/blog/how-our-training-team-uses-dbt-cloud" date: "2025-01-24" authors: ["Kyle Coapman", "Damaris Lasa"] categories: ["Learn"] --- # dbt Labs on dbt: How our training team uses dbt Cloud For those of us on customer education teams, we know how to make an impact on customers, but how do we tell that story to the greater business that isn’t directly working with customers or processing the feedback? How do we go beyond NPS and completion rates to demonstrate value to the business and real ROI? We’ll talk through how we’ve tackled this at dbt Labs to communicate the impact of our on-demand learning. ## Measuring the impact of on-demand learning As a training team, we launched our first [on-demand learning experience ](http://learn.getdbt.com)through an LMS in October 2020. In our world of customer education, there are some essential tools to launching this _one to many offering _for educational content: - An **LMS** is for building educational content, enrolling learners, and monitoring course progress. - **Credentialing software** issues credentials for course completion and passing certifications. - **Survey software** collects learner responses throughout the course and at the end of the course. As we saw signups and course completions take off, we started to monitor our work across the following metrics: - **Quality** - NPS on courses and training sessions - Completion rate of on-demand courses - **Reach** - Enrolled — started or attended at least one learning experience - Trained — received the dbt fundamentals badge - Certified — passed the developer exam - **Revenue** - How is training impacting our GTM efforts? - What direct revenue are we generating from the site? (Note: N/A as all content is currently free) These metrics allow us to know where we can improve and communicate the ROI of our training programming. ## Connecting learning data to business outcomes As a team that doesn’t directly generate revenue, we’ve often faced the challenge of finding effective ways to demonstrate the return on investment of our programming. From talking to other customer education leaders, this is a common challenge in our space. As outlined above, a customer education stack has multiple tools with various data sets. In isolation, this data doesn’t tell a compelling story as it is typically used to drive improvements on the content itself, _not the larger business impact._ The impact above that is hardest to measure is **how is on-demand training impacting our GTM efforts?** The impact comes when we can connect the learner experience to business outcomes. - How does course completion for customers correlate with churn? - How do we use learner enrollment data to drive pipeline for marketing? - How does learner engagement impact win rates for our sales team? One method for doing this is the all too common approach of exporting to CSV, creating multiple tabs in an Excel workbook, and joining it all together with `vlookups` (or index + match if you prefer that like me) and aggregations. Every time we run our report, we need to get the latest version of the data, bring it into the Excel workbook, and run the same analysis. If we want to compare changes over time, we are then likely comparing multiple versions of the same notebook. This is a ton of complexity as we change versions and iterate on our reporting. ## Bringing it all together with dbt Cloud Luckily, we work at dbt Labs with customers, so we know what the dream state can look like. Here is what we can build with dbt Cloud: - Model our data to be in the shape that we want (not stuck in in-app reporting and CSVs) - Enrich existing business data sets with learner data (I see you Salesforce 👀) - Automate the refreshing of our reporting on a cadence of our choice - Robustly manage the changes to our logic over time with version control To get there, we need some world-class, modern data practices. Let’s talk through that journey together. Let’s take one of our questions above > How does learner engagement impact win rates for our sales team? After doing some digging with the data team and using dbt Explorer, I’ve discovered the data sets that I need 1. Salesforce data - all the information about our sales cycle, specifically Land opportunities where are talking to a new potential customer. 2. LMS data - all the data about learner engagement, specifically sign-ups and completions Next, let’s go on a quick analytics development cycle. ## Mise en Place: Prepping ingredients for data excellence Like a professional chef, we need to get all our ingredients in front of us and prepared. For us, that means getting all of our application data into a data platform like Snowflake or Databricks. Luckily, our data team has already loaded (and modeled) our Salesforce data. We used Fivetran to grab the data from our LMS and load it directly into our Snowfake instance. This is similar to exporting a CSV and uploading it to a database, however, Fivetran helps us run this on a schedule without the manual effort and clicks. Now we have everything we need to cook up some insights. ## How do you _wish_ your data looked: slicing and serving When we look at _any raw_ data, we notice that it’s actually pretty messy and not very ready for business analysis. There are all sorts of cryptic IDs and fields that were likely designed to make the application work well but is less optimized for analytics. This is where **staging** your data comes in — we call this “cleaning the data to make it look like how we wished it looked”. This is like the food prep we do to chop our onions and dice our carrots before we really get cooking. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/6baf43eed264d73f540b824a1affcac4215fb6b4-1560x606.png) Mapping original columns to match our style guide, making them more descriptive, and creating a new column that we wish we had. Once I’ve staged the data, then we can start to create the business concepts that I want. For learner data, I’ve created two models: - `dim_all_learners` - a comprehensive model where each row is a learner with information like `is_trained` and `courses_complete` - `fct_training_enrollments` - a single table that shows one row for each unique learner enrolling in each unique course with information like `is_complete` and `percent_complete` By publishing these two models, we’ve already created a single place for our team to go for learner information. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/71d55f72b78f6da46bbf49d0df76468d0658fcf4-2566x366.png) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/13f3f1f9014f2a7ccee90cf54e488d0a654ede89-1600x216.png) However, we still haven’t quite told the story of business impact yet. ## Enrich your business data Our data team has created a model called `fct_opportunities` that lists every single sales opportunity in the history of the business. Coming back to the question: > How does learner engagement impact win rates for our sales team? We need to show _learner engagement_ mapped against _win rates_. Let’s take the following approach: - How does learner enrollments in dbt Learn before close correlate with win rates? - How does learner completion of dbt Fundamentals before close correlate with win rates? We need to map these learner stats against each of the opportunities. Here is how we did this - The `fct_opportunities` model already has `email_domain` extracted as a column - We add an `email domain` column to the `dim_all_learners` model along with `enrolled_date` and `trained_date` - Using `email domain` (think vlookup) we can then roll up the number of enrolled learners and completed learners before close date. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/167977c0df4b46615b894e87d1981f76f76c8baa-1718x220.png) Now these stats live on the `fct_opportunities` model in our data platform. Keep reading for what we discovered. ## Automate refreshing of reports The spreadsheet workflow I outlined above is extremely cumbersome. The beauty of SQL in dbt is that we can write all this logic once and refresh all the models on cadence that we choose. In dbt, we do the following: - Push changes to our production environment with a review from our data team. By moving our code to production, the rest of the org knows this data can be trusted - Extract / load data on a cadence using Fivetran and other loading tools (no more CSV exports!) - Transform data on a cadence to refresh the models like `dim_all_learner` and `fct_opportunities` (no more checking and re-writing formulas!) This means we can show up to work on Monday morning with a coffee and our reporting is up to date. ## Manage changes like a boss We might update our logic as our business and learning program evolves over time. If we update our logic in spreadsheets, it might look like the following: - training_win_rate_analysis_jan_2024_v1.xlsx - training_win_rate_analysis_apr_2024_v2.xlsx - training_win_rate_analysis_jun_2024_v3.xlsx Oof! That’s really hard to navigate across files and compare cells to see how logic has changed over time and wrangle / review the changes in a robust way. dbt Cloud leverages git for version control and here is the impact: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a3d48b5b5171709499b4e97268722c4191b8a3a3-1962x1004.png) 1. We can compare changes in logic between any two points in time. Think of this like a highly scalable version of version history in Google Sheets. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/6967356708c77f6d0b8c3bafcd3971b450ecf08a-2526x2166.png) 1. We can ensure the right people review the code before changes get pushed to production. _Just send it_ doesn’t have to be the way we work with the right governance in place. Git can be intimidating at first, but with a little practice, you can unlock this super power. ## Turning insights into action Through this data workflow, we were able to show a correlation such that if just one or more learners simply enroll in [dbt Fundamentals](http://learn.getdbt.com/), our win rate increases by 4% on average. Similarly, if we can get two or more learners to complete the fundamentals course, our win rate jumps by an additional 1% on average. There are undoubtedly other factors at play here, but being able to point to this correlation speaks to the ROI of our on-demand learning efforts. Why does that matter? Three things: - Prospects can learn dbt and see the value of dbt Cloud before buying - Sales can leverage our one to many content to win more deals with less time on calls - Training gets more enrollments and completion to drive our KPIs It’s a win-win-win for the everyone. **Data speaks volumes—make sure yours is telling the right story. **Unlock the full impact of your customer education efforts with dbt Cloud. Automate reporting, connect learning data to business outcomes, and prove ROI with confidence. [Start learning today](http://learn.getdbt.com/). --- --- title: "Data integration software: How to choose" description: "Learn what to look for in modern data integration tools—from automation and transformations to security and scale." url: "https://www.getdbt.com/blog/data-integration-software" date: "2025-01-23" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data integration software: How to choose With data pouring in from dozens of systems, modern businesses need a streamlined way to bring it all together. That’s the role of data integration software. It connects your sources, centralizes data in warehouses or lakes, and helps create a single, trusted view of the business. As data volumes grow, managing integration pipelines manually becomes a liability. Manual workflows introduce delays, increase errors, and limit scalability. Data integration tools solve this by automating workflows, catching issues early, and helping systems scale efficiently. But with so many platforms on the market, how do you choose the right one? In this guide, we’ll break down what to look for in a modern data integration tool—and how to evaluate your options based on your team, tech stack, and data goals. ## What is data integration software? Data integration software helps unify data from multiple systems into a single, consistent view. These tools power [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) or [ETL (Extract, Transform, Load)](https://www.getdbt.com/blog/etl-vs-elt) workflows, moving data from sources like databases, APIs, and SaaS tools into centralized storage platforms like data warehouses or lakes. Most integration platforms come with built-in connectors for cloud services, RDBMS, file systems, and more — making it easier to automate complex data workflows. They also offer transformation capabilities for cleaning, standardizing, and formatting data along the way. Many tools now offer low-code or no-code interfaces, enabling data engineers and less technical users alike to design and deploy pipelines quickly — without needing to write extensive code. ## Selecting the right data integration software With dozens of options on the market, choosing the right data integration software can feel overwhelming. These key considerations will help you find a solution that meets your organization’s needs today and scales with you tomorrow. ### Understand your business requirements Before evaluating features, clarify your organization’s specific needs. Where is your data located? What formats are involved? Do you need to centralize in a data warehouse or [data lake](https://www.databricks.com/discover/data-lakes)? Understanding your architecture and use cases will help you identify the must-have capabilities. Also consider your future roadmap. If you expect rapid growth or data volume increases, prioritize tools with strong scalability. ### Prioritize ease of use Deploying new software can be challenging for developers, as each comes with a learning curve. Easy-to-use, intuitive software means developers will take less time getting used to it. Some of these quality of life features include: - Support for familiar languages (like SQL or Python) - Low-code/ no-code capabilities - Clear, up-to-date documentation User-friendly tools reduce ramp-up time, prevent errors, and accelerate deployment. ### Review transformation features A major benefit of data integration software is built-in transformation logic. These pre-built modules support common tasks like deduplication, outlier handling, and null detection — saving time and reducing complexity. Make sure the platform’s transformation library covers your specific needs. Gaps in functionality may require custom development, increasing time to value. ### Look for automation capabilities Automation is critical to scaling your data workflows. The right platform should support: - Automated data ingestion - Scheduled transformations - Event-based pipeline triggering This ensures your data stays fresh, accurate, and consistently delivered. ### Assess scalability As your data grows, so must your pipelines. Evaluate whether the tool can scale compute and memory resources based on workload. [Scalability](https://www.getdbt.com/blog/common-challenges-to-scale-data-operations) ensures performance remains consistent even during heavy data loads or increased user activity. ### Ensure data security Data integration must be secure by design. Choose software that includes: - End-to-end encryption - [Role-based access control (RBAC)](https://www.ibm.com/think/topics/rbac) - Logging and audit trails These features help protect sensitive information and support compliance with industry regulations. ## Benefits of using data integration software While it’s possible to build integration pipelines manually, modern data integration tools offer major advantages in speed, scale, and quality. Here’s how dedicated software can help your team work smarter and faster. ### Save time and reduce complexity Data integration platforms accelerate development by combining built-in connectors, drag-and-drop interfaces, and workflow automation. - Prebuilt connectors make it easy to integrate with popular data warehouses like [Snowflake](https://snowflake.com/), [Databricks](https://databricks.com/), and [BigQuery](https://cloud.google.com/bigquery). - Built-in scheduling and automation eliminate repetitive tasks, freeing teams to focus on analysis and innovation. - Intuitive interfaces reduce development time — even for complex pipelines or less-experienced developers. ### Improved data quality By standardizing collection and transformation workflows, data integration tools reduce manual error and enforce consistency across datasets. - Real-time validation and anomaly detection help teams identify issues early. - Automated formatting ensures uniformity across sources, improving downstream analytics and trust in data. ### Support scalability Data integration software is built to grow with your business. - Dynamic scaling allocates resources based on workload, ensuring consistent performance during peak usage. - Support for diverse data sources and destinations lets teams expand without overhauling their pipelines. ### Enhance collaboration Centralized, well-governed data unlocks better teamwork. By breaking down data silos, it reduces dependencies and enables seamless collaboration between teams. - A unified integration layer ensures that all teams work from the same reliable data. - Shared access reduces silos, speeds up handoffs, and helps business users make decisions faster. ## Get the most from your data with dbt A good data integration tool breaks down silos and brings all your data into one place, so everyone across the organization can access it easily. It provides a single data control plane for your data that enables collaboration and speeds up analytics, reporting, and AI workloads. dbt is a [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) that eliminates data silos by supporting data integration, orchestration, observability, cataloging, semantics, and more. dbt enables: - A **flexible, cross-platform** approach to data integration and transformation - A **collaborative** platform that standardizes data integration tooling and makes data development more accessible, streamlined, and governed to more types of users - Creating **trustworthy outputs** that are tested, documented, and discoverable by all data stakeholders Using dbt, data developers can publish high-quality outputs that bring together data from multiple sources, combining them into a consistent, compliant, and governed dataset that fulfills a specific business purpose. Data stakeholders can then easily [find these datasets](https://www.getdbt.com/product/dbt-catalog), learn how to use them, and see at a glance where the data came from. This breaks down barriers to data access, accelerating the delivery of new data products. Try it yourself - [sign up for a free dbt account today](https://www.getdbt.com/signup) to see how dbt simplifies data integration. --- --- title: "Why your AI will fail without a semantic layer" description: "Is your data AI-ready? Learn why AI systems fail without constraints and how a semantic layer can help." url: "https://www.getdbt.com/blog/why-your-ai-will-fail-without-a-semantic-layer" date: "2025-01-23" authors: ["Drew Banin", "Tom Grabowski"] categories: ["Insights"] --- # Why your AI will fail without a semantic layer ## Why AI systems fail without constraints Imagine a retail company using an AI system powered by an LLM to evaluate product sales performance. You ask the system to determine _“What was the total adjusted revenue for Product X in 2023?”_ to inform strategic decisions. Without proper safeguards, the AI system might query raw data from various tables across a database and attempt to generate a result that is likely inaccurate. Why? Because _“adjusted revenue”_ wasn’t clearly defined in your system. Does it account for discounts, returns, or currency fluctuations? If the underlying data contains inconsistencies—like mismatched definitions of "revenue" or incomplete product details—the system might confidently generate recommendations based on flawed calculations. This could lead to misguided decisions, like overinvesting in a poorly performing product or underestimating demand for a successful one, costing the company time and money. Now, consider a slightly different culprit. Instead of a vague question, the problem lies with an undefined metric in the data itself, like a column labeled `"RevAdj_2023"`. Without clear metadata or context, the AI system cannot reliably interpret or use it. These ambiguities force the system to rely on incorrect assumptions, further compounding the risk of errors and unreliable outputs. This isn’t just a minor inconvenience. It can result in poor decisions, wasted resources, and a loss of trust in AI systems. We know that businesses are eager to integrate LLMs into their operations, whether for chatbots, smarter analytics, or creating entirely new innovations. But to realize the potential of these systems, they must be implemented with proper safeguards. ## How a semantic layer solves this problem That’s why AI systems need a **[semantic layer](https://www.getdbt.com/product/semantic-layer).** It’s a centralized framework that defines key metrics and business logic, embeds metadata, and provides business logic and context for the data your AI system (or any other downstream system) queries. ![Semantic layer graphic](https://cdn.sanity.io/images/wl0ndo6t/main/4b7fb9995b43c05d573731ddb0c66208aaed051c-1296x1024.webp) The semantic layer enforces guardrails, ensuring the AI system queries only approved, governed, and contextualized metrics. It maintains consistency and ensures metrics and business logic are applied accurately across teams and systems. In the earlier example, if the system receives a request like _“What was the total adjusted revenue for Product X in 2023?”_ the semantic layer flags the query as invalid because no such metric exists. Instead, the semantic layer enables the system to steer the user toward valid, pre-defined metrics, for example: - **Revenue adjusted for discounts and returns for ProductX (2023)** - **Adjusted revenue growth for Product X (2023 vs 2022)** - **Revenue of active accounts for Product X (2023)** This ensures the AI system provides accurate insights while prompting clarifications like: _“There isn’t a metric for ‘2023 total adjusted revenue for Product X.’ Did you mean total revenue adjusted for discounts and returns for Product X in 2023? Or something else?”_ Similarly, as in the earlier example with `"RevAdj_2023"` the semantic layer provides important context by embedding metadata and clear definitions into the data pipeline. `"RevAdj_2023"` could include metadata and clear documentation that explains exactly what revenue adjustment covers—like the metric name, a description, calculation logic, and usage guidelines. This ensures there’s no ambiguity. Without these constraints, even the most advanced AI systems can deliver inaccurate outputs. But with a semantic layer in place, you’re creating a single source of truth for your data—centrally managed and accessible to both business users and AI systems. This becomes the foundation for driving successful and trustworthy AI initiatives. ## Why a semantic layer is key to AI success To succeed, your AI systems require these critical essentials: ### 1. Consistency for reliable insights Without a semantic layer, your data sources might have conflicting definitions or inconsistent calculations. For example, one team might define "revenue" as gross sales, while another subtracts discounts and returns. A semantic layer aligns metrics and business logic to a single, consistent definition, ensuring AI systems always query trustworthy data. ### 2. Governance to protect sensitive data and ensure consistency Effective governance is about more than just securing sensitive data—it’s about ensuring consistent, accurate, and trustworthy insights across the organization. A semantic layer enforces governance by: - Restricting access to sensitive metrics, ensuring the right people access the right data. - Tracking changes to metrics and logic with a clear audit trail. - Preventing unauthorized access to sensitive data. For example, the semantic layer can prevent an HR team from accessing finance metrics or stop a customer-facing AI agent from exposing sensitive client data. Without these controls, your AI systems risk producing outputs that are not only inaccurate but also legally or ethically non-compliant. Governance also ensures consistency when business logic or metric definitions change. Imagine the executive team decides to update the definition of "total adjusted revenue" to include discounts. Without a semantic layer, this definition change would have to be accounted for within any possible downstream system (a BI tool, an LLM, etc) that might query that metric, adding tedious overhead and unnecessary risk; it's just a matter of time before teams who query that metric get conflicting answers, resulting in confusion and compromised trust. A semantic layer gives organizations a place to define metrics once—centrally—and ensures updates are automatically applied across _any and all_ connected systems, so AI interfaces, LLMs, and users always work with the latest, consistent, approved definitions. ### 3. Context for smarter decision-making AI systems need metadata and logic to understand relationships between tables, interpret the meaning behind columns, and apply correct business logic. For example, does "customer churn" refer to a canceled subscription, or does it mean inactivity over a certain period? Without clear definitions, AI systems can make incorrect assumptions, leading to flawed outputs. A semantic layer solves this by embedding these essential elements directly into the data pipeline, helping uncover the “why” behind the data by: - **Defining relationships between data elements:** A semantic layer links tables (e.g., connecting "Customer ID" in "Customers" to "Transactions") so the AI system understands relationships like how purchases relate to customers or revenue to products. By defining these relationships explicitly, you can feel confident that joins between tables will _always_ be performed correctly. - **Embedding metadata for clarity:** Metadata defines what each field means, how it’s calculated, and how it should be used. For example **Metric Name:** customer_churn_rate **Description:** "The percentage of customers who canceled their subscription within the last 30 days." **Calculation Logic:** count(churned_customers)/count(total_customers) **Usage Guidelines:** "Only use for customers with active subscriptions in the last quarter." - **Standardizing business logic:** It embeds rules and calculations like "'revenue'= price - discounts - returns" to prevent mismatched definitions across teams. By embedding relationships, metadata, and logic, the semantic layer gives AI systems the context needed to deliver accurate insights. This makes it possible to handle complex queries like, _"What’s the average revenue per customer who purchased a specific product last month?"_ ### 4. Speed and scalability for faster adoption When data queries are slow and complicated, users abandon AI systems and turn to data teams for quicker answers, stalling AI adoption. The semantic layer changes this with [**smart caching**](https://docs.getdbt.com/docs/use-dbt-semantic-layer/sl-cache#:~:text=The%20dbt%20Semantic%20Layer%20allows,platform%27s%20built%2Din%20caching%20layer) and a centralized **metric store**. Instead of scanning raw tables for every query, AI systems pull precomputed, validated metrics, delivering faster results. It also streamlines scaling by letting teams reuse standardized, governed metrics across projects. This removes the need to rebuild logic for every new initiative, speeding up AI adoption while ensuring accuracy and reliability. **** ## Build AI the right way: Start with a semantic layer Even the most advanced AI systems will fail without the right foundation. A semantic layer ensures your data is consistent, governed, and full of the context AI systems need to deliver meaningful results. Before kicking off your next AI project, ask yourself: **Is my data ready for AI?** If not, it might be time to prioritize a semantic layer—because the success of your AI strategy depends on it. The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) translates dbt models into well-defined business metrics that build the foundation for clean, reliable, and AI-ready data. It’s designed to seamlessly integrate with your dbt workflow, ensuring your data is not only accurate and governed but also aligned with your business goals. Ready to build AI the right way? [Schedule a demo](https://www.getdbt.com/contact) of the dbt Semantic Layer today. --- --- title: "What's new in dbt Cloud - January 2025" description: "The latest new features and integrations landing in dbt Cloud." url: "https://www.getdbt.com/blog/whats-new-in-dbt-cloud-january-2025" date: "2025-01-22" authors: ["Alexis Jones", "Sara Gawlinski"] categories: ["Product"] --- # What's new in dbt Cloud - January 2025 Hello and happy new year! We’re only a few weeks in, and already we can tell 2025 is gonna be a good one. Many exciting new features to share, and the big news of our acquisition of SDF Labs! Looking forward to keeping the good times rolling with our upcoming dbt Cloud Launch Showcase event—[register here](https://www.getdbt.com/resources/webinars/2025-dbt-cloud-launch-showcase)—that’s coming to an internet browser near you later this spring. For now, keep reading to learn what’s new in dbt Cloud (and you can catch up on our last "What’s New" post from November [here](https://www.getdbt.com/blog/whats-new-in-dbt-cloud-november-2024)). 👇 ## A new standard for SQL comprehension with SDF In case you missed it, we [recently announced](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs) that we have acquired SDF Labs! This is a monumental step forward not just for dbt, but for the analytics industry. SDF will bring SQL comprehension to the dbt engine to supercharge the developer experience: 100x faster performance, the ability to validate code in dev to boost data quality and optimize compute costs, and rich metadata to improve lineage and enable nuanced governance use cases. Definitely be sure to [sign up for our upcoming webinar](https://www.getdbt.com/resources/webinars/accelerating-dbt-with-sdf) hosted by dbt Labs CEO Tristan Handy and SDF’s CEO to hear their perspective on the bright future ahead for dbt and our users. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/7fd27823713b544422d057b3ca0e8856192eaa39-3600x1890.png) ## dbt Copilot 🔑 **Azure OpenAI bring your own key (BYOK):** Enterprises can now use their Azure-hosted OpenAI API keys with dbt Copilot, providing full control over LLM usage, costs, and compliance. This allows data teams to leverage dbt Copilot AI capabilities like automated documentation and SQL generation securely within their Azure environment, combining the productivity of AI with enterprise-grade security and governance—perfect for organizations with strict data requirements. [Read the docs](https://docs.getdbt.com/docs/cloud/account-integrations?ai-integration=azure#ai-integrations) to learn more, and reach out to your sales rep if you’re interested in getting involved in the ongoing dbt Copilot beta. ## dbt Semantic Layer 📊 **Sigma integration:** Business users can now query metrics defined in the dbt Semantic Layer directly within Sigma's analytics platform (in Preview), ensuring consistent metric definitions across the organization. [Check out Sigma’s docs](https://help.sigmacomputing.com/docs/configure-a-dbt-semantic-layer-integration) for details on how to configure the integration, or watch the demo video below. [Watch video](https://www.loom.com/share/10d5b69e42ed4580bc6f048792849fd3?sid=39620572-1dd4-40cd-ad75-54644b88a78f) 📝 **Alias argument in the dbt Semantic Layer query API:** The dbt Semantic Layer query API now includes an `alias` argument, allowing you to rename metric column names directly in your query results. This feature simplifies downstream integration by aligning column names with your team’s naming conventions or tool-specific requirements, improving consistency and clarity in reporting workflows. To see how to use the `alias` argument in your queries, check out the [JDBC API docs](https://docs.getdbt.com/docs/dbt-cloud-apis/sl-jdbc#query-metric-alias) and the [GraphQL docs](https://docs.getdbt.com/docs/dbt-cloud-apis/sl-graphql#query-metric-alias). 👥 **New grouping options via the query_with_all_group_bys API endpoint:** The new `query_with_all_group_bys` endpoint simplifies analysis by returning all valid dimensions for a given set of metrics. This helps data teams quickly identify grouping options for downstream tools to streamline data exploration. Check out examples of how to use this endpoint in our [documentation](https://docs.getdbt.com/docs/dbt-cloud-apis/sl-jdbc#query-by-all-dimensions). 📖 **MetricFlow now compiles SQL using CTEs:** SQL generated by MetricFlow now uses Common Table Expressions (CTEs), improving performance for multi-metric queries and making the SQL easier to read. You can use the `compile` flag on a query request to inspect the updated SQL structure. Read more in our [documentation](https://docs.getdbt.com/docs/build/metricflow-commands#additional-query-examples). ## Deploy 🪄 **Customize compare runs:** Users can now change the default `state:modified` into whatever they want as they compare changes with Advanced CI. This enables you to tailor how comparisons are run, with the ability to exclude certain models or tags or run further downstream each time. Read more about [dbt compare custom commands](https://docs.getdbt.com/docs/deploy/job-commands#compare-changes-custom-commands). ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/bf05c8cd5a2f7008643d7aa7a423b7515b499e84-2000x437.png) 🖇️ **Group by PR:** The CI interface has been updated to include a new ‘Group by PR’ tab that shows each PR with its name and status, ensuring that users merge the correct changes into production. 🧺 **SQL linting in CI jobs:** You can now enable SQL linting in your CI jobs, using [SQLFluff](https://sqlfluff.com/), to automatically lint all of the modified SQL files in your project as a run step before your CI job builds. This feature is now GA for dbt Cloud Team or Enterprise accounts who are running on [release tracks](https://docs.getdbt.com/docs/dbt-versions/cloud-release-tracks). Refer to [the documentation](https://docs.getdbt.com/docs/deploy/continuous-integration#sql-linting) for more information. 📸 **Snapshots improvements:** A series of snapshots improvements just landed for customers running on release tracks: - The `dbt_valid_to_current` [config](https://docs.getdbt.com/reference/resource-configs/dbt_valid_to_current) lets you set a custom indicator for the value of `dbt_valid_to` in current snapshot records (like a future date). This makes it easier to assign a custom date, work in a join, or perform range-based filtering that requires an end date. - The `hard_deletes` [config](https://docs.getdbt.com/reference/resource-configs/hard-deletes) and `new_record` [method](https://docs.getdbt.com/reference/resource-configs/hard-deletes) let you track deleted records by inserting a new row for the deleted state with a `dbt_is_deleted` flag. 🗓️ **Orchestration for Tableau auto-exposures:** Last year, [we launched auto-exposures for Tableau](https://www.getdbt.com/blog/coalesce-2024-product-announcements), giving users the ability to automatically populate and visualize downstream Tableau exposures directly in dbt Explorer. Now, you can take it one step further by configuring dbt to automatically _orchestrate_ your pipeline and refresh those same Tableau dashboards when new data is available upstream, simply as part of running dbt. Reach out to your account team today to join the private beta. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/c41ef96e0994bcca9599700e73d05b500890fd7a-2752x1538.png) ## Observe 🤳 **Model notifications:** Now GA, model owners can receive email alerts on job status (success, warnings, fail) in real-time—as the job is running—so they’re always first to know of any issues and can keep pipelines humming. [Read the docs](https://docs.getdbt.com/docs/deploy/model-notifications) to learn more. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/9250c7cb686038df0bc36c935a8b6a6b80cd793e-1793x330.png) ## Platform ☁️ **Support for Azure deployments:** Now in GA for both America and European regions, dbt Cloud multi-tenant can be deployed natively on Microsoft Azure and [procured directly in the Azure Marketplace](https://azuremarketplace.microsoft.com/en-us/marketplace/apps/fishtownanalyticsinc1621444423835.dbt_cloud?tab=Overview)! This is in addition to the existing support for [AWS deployments](https://docs.getdbt.com/docs/cloud/about-cloud/tenancy), bringing the same powerful dbt Cloud experience to even more data teams—regardless of your choice of cloud. Available for dbt Cloud Enterprise customers, [read the docs](https://docs.getdbt.com/docs/cloud/about-cloud/tenancy#available-features) to learn more. 📀 **Release tracks:** Say goodbye to manual dbt version upgrades for good. With release tracks, you have an automated and flexible way to manage your dbt version upgrades across your dbt Cloud environments. With three flexible options, you can ensure that you’re always enjoying the latest (or very recent) capabilities without any added maintenance overhead. [Read the blog post](https://www.getdbt.com/blog/introducing-release-tracks-for-dbt-version-upgrades) to learn more. 🌑 **Dark mode:** You asked, we listened! Dark mode is now available in Preview for Developer accounts (with access for Team and Enterprise plans rolling out over the coming weeks). Just jump to your profile settings and toggle the theme to “Dark” or you can configure the setting directly in the IDE. [Read the docs](https://docs.getdbt.com/docs/cloud/about-cloud/dark-mode) to learn more. 🏠 **New account homepage:** The new account homepage gives you quick insights on data delivery, project performance, and more. You can easily “favorite” your projects to navigate to them with ease, and get a birds’-eye view of data quality by sorting your projects by test or documentation coverage. Check it out and customize your account homepage today. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0669d802a48cf8a4aa46911b915a4a09722df989-2630x1504.png) ## Onto the next one Hitting “post” on these "What’s New" blogs is so rewarding because we absolutely love to see your responses and feedback to new features. Keep them coming! And be sure to [sign up for next week’s webinar](https://www.getdbt.com/resources/webinars/accelerating-dbt-with-sdf) where we’ll dive deep into all things SDF + dbt and answer any questions you may have about the upcoming integration. See you there! --- --- title: "Data pipeline automation: Why it’s important" description: "See how modern teams automate pipelines with dbt to deliver accurate, reliable data, without sacrificing speed or control." url: "https://www.getdbt.com/blog/data-pipeline-automation" date: "2025-01-17" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data pipeline automation: Why it’s important Getting data into production isn’t just about moving it from point A to point B. Before stakeholders can trust and act on it, data needs to be integrated, transformed, and tested — not just for [freshness, but for accuracy, consistency, and reliability](https://www.getdbt.com/blog/data-quality-metrics). Otherwise, outdated or broken data can lead to poor decisions, broken dashboards, and downstream cleanup. If you’ve worked in software engineering, this might sound familiar. [Software deployment pipelines](https://www.pagerduty.com/resources/continuous-integration-delivery/learn/what-is-a-deployment-pipeline/) automate the process of verifying, promoting, and monitoring any change to production systems. These pipelines improve quality and shorten time to release by reducing human error and building quality checks into every point of the deployment process. Data pipeline automation brings the same discipline to [analytics engineering](https://www.getdbt.com/resources/the-analytics-development-lifecycle). Instead of running SQL scripts manually or relying on brittle workflows, automated pipelines transform raw inputs into trustworthy outputs every time new data lands. In this article, we’ll break down what data pipeline automation is, why it matters, and how modern data teams can build scalable workflows that deliver trusted, production-grade analytics — no matter where your data lives. ## What is data pipeline automation? Most business decisions can’t be made directly on raw data. Before it’s useful, data needs to be extracted, cleaned, and transformed into a format that’s fast to query, consistent to interpret, and trustworthy enough to use. Historically, this was done manually. Analysts or data engineers would run SQL scripts by hand, often from their local machine. Unsurprisingly, this introduced a long list of problems: - **Lack of reuse: **transformation code lived in notebooks or local scripts, hidden from other teams. No one could improve it, reuse it, or even find it. - **Inconsistent results: **new data only showed up when someone remembered to run the pipeline. There was no guarantee that reports were up-to-date or aligned across teams. - **No quality controls: **changes weren’t reviewed, tested, or versioned. Errors went unnoticed until they showed up in a dashboard — or worse, in a boardroom. - **Zero governance: **there was no audit trail. Data might be pulled into spreadsheets, emailed as CSVs, or manipulated outside approved systems, creating security and compliance risks. Data pipeline automation solves these challenges by orchestrating the flow of data from source systems to destination systems. Pipelines are written in a shared language like SQL or Python, [version-controlled in Git](https://docs.getdbt.com/docs/deploy/deploy-environments#git-workflow), and executed on a schedule or in response to events. Teams often use [orchestration tools](https://www.getdbt.com/product/integrations#orchestration) (like Airflow, Prefect, or Dagster) to manage dependencies, track runs, and enforce testing before changes reach production. Automating data pipelines solves several of the shortcomings of manual data transformation: - **Repeatability**. The data transformation process runs with almost no human intervention, reducing potential errors. When problems are fixed, they’re incorporated into the pipeline’s code so they don’t recur in future runs. - **Discoverability**. Code is run on a centralized server accessible by the data engineering team, and is checked into Git so others can find and understand it. - **Reliability**. Team members can see when a pipeline was last run, so they know that the data they’re seeing in their dashboards is the latest reality. If an issue with the pipeline occurs, the pipeline can issue an alert so that the data engineering team can rapidly assess and fix the issue. ## Components of data pipeline automation The steps in a data pipeline automation workflow differ slightly depending on whether you’re using [ETL or ELT](https://www.getdbt.com/blog/etl-vs-elt) - Extract, Transform, and Load vs. Extract, Load, and Transform. Modern data pipelines typically follow an ELT pattern: Extract, Load, and then Transform your data inside a centralized platform (like a cloud warehouse or lakehouse). This approach is more flexible than traditional ETL because raw data remains accessible and can be transformed multiple times for different use cases. Here’s how an ELT pipeline maps to the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) — and how automation keeps the process scalable and trustworthy. #### Orchestration An orchestration tool runs your transformation code (like SQL, Python, or dbt models) [on a schedule](https://docs.getdbt.com/docs/deploy/job-scheduler) or in response to an event. This code can take multiple forms, such as an [AWS Lambda](https://aws.amazon.com/lambda/) function, a [Docker container](https://www.docker.com/resources/what-container/), or a [dbt model](https://docs.getdbt.com/docs/build/models). The orchestration platform will track a data pipeline’s runs and report on each runs’ pass/fail state, providing debugging logs in the case the pipeline encounters an error. #### [Extraction](https://www.geeksforgeeks.org/what-is-data-extraction/) The data is exported from one or more source systems. The method used to extract will differ depending on the source system and may be extracted via direct query (e.g., a periodic SQL query), API calls, a scheduled export to an object storage system like [Amazon S3](https://aws.amazon.com/s3/), a CSV or JSON export from a partner, web scraping, etc. You might filter, parse, or reformat this data before loading, depending on your toolchain or business logic. #### Loading Raw data is then ingested into your warehouse or lakehouse, typically unchanged. This allows for flexible, modular transformation downstream. #### Transformation Now comes the dbt magic. Transformation turns messy raw data into clean, analytics-ready models. At dbt Labs, we recommend a three-layer modeling approach: - **Staging**: Standardize raw inputs. Rename columns, clean formats, and align data types. - **Intermediate**: Aggregate or refine the data into business logic — e.g., calculating revenue per order or user behavior by session. - **Mart**: Define reusable, domain-specific outputs like `customer_lifetime_value` or `daily_active_users` that feed dashboards and ML models. **** #### Testing Before shipping updates, run [automated tests](https://docs.getdbt.com/docs/building-a-dbt-project/tests) to validate that your models behave as expected. dbt makes it easy to add these checks directly in your [DAG](https://www.getdbt.com/blog/dag-use-cases-and-best-practices). #### Review With your code version-controlled (e.g., in Git), you open a pull request for peer review. Another analytics engineer reviews your changes, ensuring they follow naming conventions, logic standards, and won’t break anything downstream. This second set of eyes is the key manual touchpoint in the system. By instituting a review gate, teams ensure that another expert has reviewed a set of changes for quality and conformance to best practices. #### Deployment A software deployment pipeline uses a multi-stage Continuous Integration (CI) process that tests code changes in pre-production environments before promoting them to the live environment. [A data pipeline deployment does the same](https://docs.getdbt.com/docs/deploy/deployments), running data tests in [isolated environments](https://docs.getdbt.com/docs/deploy/deploy-environments). This is typically where you test for edge cases, known issues detected in past releases, etc. If a failure occurs, the pipeline orchestrator issues an alert for engineers to inspect, debug, and fix the issue. If the tests pass, they’re promoted to production and made available to stakeholders. The automatic execution of tests is one of the primary benefits of data pipeline automation. It ensures that a mandated level of testing occurs with every code change. This greatly reduces the number of errors shipped to production. #### Operate and observe The orchestration platform, besides running data pipeline code, will also implement observability tools, such as logging, [data health](https://docs.getdbt.com/docs/explore/data-tile), and data quality metrics. #### Discover and analyze Once the pipeline finishes, models are available in your warehouse and registered with a centralized data catalog. Analysts and business stakeholders can find this data, experiment with it, and use it to create reports and drive business decisions. ## The need for consistency in data pipeline automation Modern data teams often work in technical silos. One team builds stored procedures in SQL. Another scripts transformations in Python and deploys with Docker. Elsewhere, someone still relies on a Cron job duct-taped together years ago. This fragmented approach might work at a small scale. But as data volumes grow and teams multiply, inconsistency becomes a liability. It also complicates data governance, as you have no way to obtain a 360-degree view of all of your data. Inconsistent pipelines erode trust and make it harder to scale. That’s why more organizations are adopting [dbt](https://www.getdbt.com/product/dbt) to centralize transformation workflows. With modular SQL models, built-in testing, and native CI/CD, dbt gives every team a shared framework to build production-grade data pipelines — faster, safer, and with more confidence. ## Building and optimizing data pipelines with a single data control plane The solution to this scattershot approach to data pipeline automation is a [**data control plane**](https://www.getdbt.com/blog/data-control-plane-introduction), an abstraction layer that sits across your data stack and provides unified capabilities for orchestration, observability, cataloging, semantics, and more. A data control plane spans your entire stack. It connects platforms, pipelines, and people while centralizing key capabilities: - **Orchestration**: Trigger and monitor pipeline runs across tools. - **Observability**: Track freshness, performance, and lineage in real time. - **Governance**: Enforce standards for testing, access, and quality. - **Discovery + documentation**: Help teams understand, trust, and reuse data. With a control plane, you can manage data pipelines from one interface—no matter how many tools, teams, or platforms are involved. **dbt is the control plane for data collaboration at scale. ** It gives you a unified framework for transforming, testing, deploying, and documenting your data. Whether you’re operating across clouds or scaling across teams, the dbt platform helps you ship trusted data—faster, safer, and with less overhead. **Want to see how dbt powers production-grade data pipelines? **[Book a demo](https://www.getdbt.com/contact) to explore how the dbt platform enables reliable, scalable automation across your data stack. --- --- title: "DataOps vs data engineering: Which one do you need?" description: "Confused about DataOps vs data engineering? Learn what they mean, how they differ, and when to use each in your data workflows." url: "https://www.getdbt.com/blog/dataops-vs-data-engineering" date: "2025-01-14" authors: ["Daniel Poppy"] categories: ["Learn"] --- # DataOps vs data engineering: Which one do you need? Data is at the heart of how modern organizations operate. It contains crucial insights that allow businesses to understand and improve their operations. But data alone isn’t enough — it needs to be collected, transformed, and delivered in a way that’s reliable and scalable. That’s where data engineering and DataOps come in. While these terms are often used interchangeably, they serve different purposes. Data engineering focuses on building the infrastructure to move and transform data. DataOps takes it further by applying DevOps-style practices to streamline, automate, and manage those workflows. In this article, we’ll explore how DataOps and data engineering compare, where they overlap, and how teams can apply principles from each to build better data systems. ## What is data engineering? [Data engineering](https://www.getdbt.com/blog/what-is-data-engineering) is the foundation of modern data infrastructure. It focuses on building and maintaining the systems that collect, store, and move data across an organization. From defining architecture to managing ingestion pipelines, data engineers ensure teams have access to reliable, scalable data for analysis and decision-making. Traditionally, data engineers focused heavily on building integration pipelines — connecting sources, writing ETL jobs, and centralizing data. But as the data landscape has evolved, so has the role. Today, data engineers focus more on designing scalable data architecture, enabling self-serve analytics, and supporting data reliability across the business. One major shift: The rise of [analytics engineering](https://www.getdbt.com/blog/what-is-analytics-engineering). Modern teams now split responsibilities across specialized roles. Data engineers typically manage ingestion and infrastructure, while analytics engineers transform data into business-ready assets using tools like dbt. This aligns with the ELT model — extract and load raw data into a centralized platform, then transform it using modular, version-controlled code. ### Key components of data engineering The scope of data engineering spans several core responsibilities: - **Data architecture: **Designing the high-level architecture for how data moves and is stored—whether in a data lake, warehouse, or hybrid setup. This includes choosing tools and creating abstractions that support self-service across teams. - **Data discovery and exploration: **Identifying and understanding source systems and formats. Engineers work with structured and unstructured data across relational databases, NoSQL systems, APIs, and files like JSON, CSV, and YAML. - **Pipeline creation: **Building scalable pipelines that extract and load data into a central repository. Engineers write code and manage orchestration to ensure data flows reliably and efficiently. - **Business logic and transformation: **Writing and maintaining logic to clean, standardize, and structure data—often in collaboration with analytics engineers. This logic may be implemented in SQL, Python, or both, depending on the pipeline and business needs. ## What is DataOps? DataOps is an emerging discipline that applies [DevOps](https://aws.amazon.com/devops/what-is-devops/) and [agile principles](https://agilealliance.org/agile101/) to the world of data engineering. While traditional data engineering focuses on building and managing pipelines, DataOps brings automation, collaboration, and continuous improvement to every step of the data lifecycle. Think of it as DevOps for data: It standardizes workflows, introduces CI/CD, and emphasizes testing, observability, and fast iteration. The result is faster, more reliable, and more scalable delivery of data across the organization. Where data engineering builds the infrastructure, DataOps improves how that infrastructure is managed and deployed. It encourages shared ownership across teams—connecting data engineers, analysts, and business stakeholders in a collaborative, agile loop. The goal? To establish a mature [**Analytics Development Lifecycle (ADLC)**](https://www.getdbt.com/resources/the-analytics-development-lifecycle) where data products are versioned, tested, deployed, and monitored just like software. ### Key components of DataOps DataOps adds several critical capabilities on top of traditional data engineering: - **Workflow automation and orchestration: **Automated scheduling and orchestration ensure pipelines run consistently and on time. This minimizes manual errors and keeps data flowing reliably across teams and systems. - **CI/CD for data pipelines: **DataOps applies [Continuous Integration and Continuous Deployment (CI/CD)](https://docs.getdbt.com/docs/deploy/continuous-integration) to data workflows. With version control, automated testing, and staged deployment, teams can ship changes faster and with less risk. - **Monitoring and data observability: **DataOps emphasizes the health and performance of pipelines — not just their output. With data quality checks, lineage tracking, and proactive monitoring, teams can quickly detect and fix issues before they impact downstream users. - **Agile collaboration: **DataOps encourages agile, cross-functional collaboration between data engineers, analytics engineers, and business users. Shared tooling and iterative development help teams respond faster to changing requirements. ## When should you use data engineering vs. DataOps? **Data engineering** is foundational. It’s the discipline responsible for designing your data architecture, building pipelines, and enabling the movement of data across your organization. Whether you’re a startup or an enterprise, you need strong data engineering to collect, store, and prepare your data for use. **DataOps**, by contrast, is focused on operationalizing those engineering efforts. It introduces automation, CI/CD, observability, and agile collaboration to improve the speed, quality, and scalability of data delivery. It doesn’t replace data engineering — it enhances it. ### When is each one relevant? - **Data engineering is essential. **Every organization that works with data needs engineering. It’s especially critical for defining your data stack, establishing architecture, and enabling initial analytics. - **DataOps is increasingly necessary. ** For early-stage teams, DataOps may feel like a "nice to have". But as your data volumes grow and complexity increases, DataOps becomes a must-have. It reduces manual work, increases deployment confidence, and ensures your pipelines can scale with your business. Together, DataOps and data engineering form a complementary practice: - **Data engineering** lays the groundwork by building pipelines and infrastructure. - **DataOps** ensures those systems run smoothly, scalably, and continuously — with minimal human intervention. The takeaway: start with strong data engineering. But as your needs evolve, layering in DataOps principles will be key to maintaining speed and trust in your data operations. ## Bridging data engineering and DataOps with dbt Data engineering and DataOps are both essential to building a modern, scalable data practice. Data engineering moves raw data into centralized storage like a warehouse or lakehouse, enabling complex transformations that turn it into trusted, usable information. DataOps builds on this foundation with automation, testing, and collaboration — ensuring pipelines are scalable, agile, and production-ready. **dbt bridges the two.** It empowers teams to build, test, and document reliable data pipelines using proven software engineering practices: - **Accelerated development** with [dbt Fusion](https://www.getdbt.com/product/fusion), which validates SQL models before they run, and [dbt Canvas](https://www.getdbt.com/product/canvas), a visual builder for analytics workflows - **High-quality deployments** powered by [testing](https://docs.getdbt.com/docs/build/data-tests), version control, and [data orchestration](https://docs.getdbt.com/docs/deploy/continuous-integration) - **End-to-end governance** with auto-generated docs and visual lineage - **Data discovery and exploration** via [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) - **AI-assisted development** through [dbt Copilot](https://www.getdbt.com/product/ai), which helps users generate SQL, documentation, and models using natural language With dbt, data engineers can apply DataOps best practices from day one — enabling faster delivery, fewer errors, and more confident decision-making. **Ready to operationalize your data workflows? [Try dbt today](https://www.getdbt.com/signup).** --- --- title: "dbt Labs acquires SDF Labs to accelerate the dbt developer experience" description: "New tech will significantly improve performance, enhance developer ergonomics, and bring rich new metadata to dbt." url: "https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs" date: "2025-01-14" authors: ["Tristan Handy"] categories: ["Company"] --- # dbt Labs acquires SDF Labs to accelerate the dbt developer experience I am not generally an excitable person. I do not dance. I try to avoid hyperbole. And yet. And yet! It is very, very hard for me to avoid literally jumping up and down as I share this news with you. The TL;DR: today, I have the pleasure of announcing that **dbt Labs has acquired [SDF Labs](https://www.sdf.com/)**. The two teams are already working side-by-side to bring SDF’s SQL comprehension technology into the hands of dbt users everywhere. SDF will be a massive upgrade to the very heart of the dbt user experience moving forward. It will enable faster dbt project compilation (~2 orders of magnitude), amazing developer experience (think: type-ahead in your IDE of choice), the highest-fidelity lineage on the market, and much more. Let me take a sec to share the story of how we got here. Because I think it’s an interesting one. **** ## A standardized way to author SQL data pipelines From the very beginning, we wanted writing dbt pipelines to feel as simple as writing SQL. Just write a select statement, in your dialect of choice, and dbt would take care of all of the fiddly bits. This desire arose directly from our focus on empowering more humans—not just highly technical data engineers—to author production-grade data pipelines. So we started with models that were just SQL. Then it was obvious that we needed some dynamic ability in our pipelines. Declare variables, create dynamic relationships between nodes, etc. So we added Jinja. Initially for just a few functions: `ref()` being the first and most important, but later `config()` , `var()`, and others. These functions were a gateway drug, though, and we quickly became convinced of how important custom macros would be. Users could define their own functions, and as long as they output syntactically valid SQL, they could do whatever they wanted! And thus, dbt utils and the entire [dbt package hub](https://hub.getdbt.com/) was born. Community-wide code reusability for the win! Through the ensuing years, dbt became an ever-more-complete SQL authoring framework. One of the biggest steps we ever made was to push all materializations into _user-space_. This allowed advanced users to take complete control of the SQL authored by dbt. **From that point forwards, if a data transformation was expressible within a given data platform, dbt users could author it.** Throughout this entire journey, however, _dbt could not actually understand the SQL its users were writing!_ This remains true today. dbt is an incredibly powerful framework for users to author data pipelines, but it fundamentally treats the SQL that you author as text and leaves the evaluation of that text to the database. This gives the user ultimate control—the framework doesn’t get in the way—but it also asks the user to do more work. For a long time, this felt entirely normal and rational. We had, for a long time, thought of dbt as “Ruby on Rails for SQL”. ## Borrowing from software engineering (again) For those of you who didn’t do web development in the mid-2000’s, Ruby on Rails (RoR) is a web development framework that made it incredibly straightforward to develop web applications. It combined the power of two (!) declarative languages (HTML and CSS) with a templating system (.erb, or ‘embedded ruby’) to enable developers to do things on the web that they could never do with HTML alone. RoR did not, itself, understand the HTML its developers authored; that task was left up to the browser. The analogy is pretty direct, and it worked for a long time. Everything you can do with dbt today—the entire control plane, from transformation to testing to catalog to orchestration—is powered by this paradigm. dbt helps you author data pipelines using SQL; Ruby on Rails helps you author web applications using HTML and CSS. But RoR is no longer the dominant framework for building web applications today. Over the last decade, React has taken that mantle. And React _does_ understand the HTML/CSS/Javascript its users write—and it does that in a very clever way. Rather than writing to a given browser target, React developers code against an intermediate abstraction and the framework takes on the responsibility of ensuring browser-level compatibility. Gone are the days where web developers have to struggle with browser compatibility. The level up in capability for web applications from the days pre-React and post-React is pretty dramatic. (Think of the web interfaces you used in 2014 and compare them with what you’re using today!) These types of epochal changes in developer tooling can make _massive_ differences in the products that can be built on top. ## Enter: SDF We’ve known for years that this type of step function change was coming for the world of SQL. And different teams have taken some decent swings at it along the way. Unfortunately, every proposed solution had tradeoffs that we considered unacceptable. Limiting the dynamism of Jinja. Unpleasant syntax. Forcing developers to use some intermediate proprietary language. Etc. In our opinion, the right solution just hadn’t emerged yet. That is…until a little company called SDF Labs [came out of stealth in the summer of 2024](https://blog.sdf.com/p/announcing-sdf-general-availability). SDF’s story is fascinating. Founded by a father/son duo (Lukas and Wolfram Schulte, CEO and CTO respectively), and with a core team of database researchers from Microsoft Research, Meta, and others, they are among the most qualified humans on the planet to think about the problem of highly reliable SQL comprehension at scale. Wolfram, in fact, was hired at Meta to build the system that tracked PII throughout all data pipelines at the company across over a million tables. Talk about being battle-tested—I can’t imagine there are many (any?) more demanding use cases for this technology anywhere. A few years later, the two decided to take the lessons learned from this work and build a company around it. They recruited a team of the best talent in the world and got to work building in stealth for two years, emerging in June of 2024 with a fully-functioning and dbt-integrated product already in production use by customers. Fatefully, Benn Stancil introduced me to Lukas on the day before their public launch. It was clear to me after a single conversation that this was the future of dbt. ## How does SDF work? So how does SDF actually…work? SDF is a high performance toolchain for SQL development packaged into one CLI; a multi-dialect SQL compiler, type system, transformation framework, linter, and language server. It is written in Rust, highly parallelized, and designed for scale. The toolchain is powered by a state-of-the-art development in SQL understanding. SDF represents each SQL dialect (Snowflake, Redshift, BigQuery, etc.) as a complete [ANTLR grammar](https://www.antlr.org/index.html) with definitions for all datatypes, coercion rules, functions, scoping intricacies and more. Unlike dbt historically (which has treated SQL as strings), SDF sees objects and types and syntax and semantics. In the same way that virtual machines (VMs) emulate physical hardware, SDF emulates the SQL compilers native to the data platforms you use. ![SDF architecture diagram](https://cdn.sanity.io/images/wl0ndo6t/main/36741d46d1935129825ae060989f2877ca37e62d-2018x1184.png) The result is magical: at every point in time _the entirety of the data warehouse is fully defined and statically analyzed as code_. A complete understanding of SQL allows the SDF engine to faithfully emulate cloud data warehouses in their behavior and provide that feedback _before execution_ and catch breaking changes as part of development rather than after deployment. Best of all, integration is easy. SDF has adopted dbt’s syntax, configuration, libraries, and Jinja natively, as part of the SDF runtime. As a result, for most dbt projects _there will be no code changes required to take full advantage of SDF’s capabilities!_ ## The power of a new paradigm Ok cool, all of this has been interesting. But I'm sure you're wondering...as a user, _what does it actually get me?_ The answer is: quite a lot. Let’s start with the first, most basic benefit. SDF parses and compiles dbt projects really, _really_ fast. Because it’s built in Rust, it simply runs faster than Python. As a result, SDF compiles the same dbt project multiple orders of magnitude faster than dbt Core. If you’re working in a large dbt project, this will _meaningfully_ impact your productivity. Next is developer experience. There are many things that will eventually go into this bucket, but here are two great examples. First, SDF’s ability to understand SQL means that it can power IntelliSense in your IDE of choice. With every keystroke, SDF understands what you are typing and can automatically suggest what comes next, including suggesting table and column names. Second, because SDF understands your SQL, it can detect errors without connecting to the remote database. Troubleshooting all of a sudden becomes _far faster_, as errors get caught as you are typing, not when you do a `dbt run`. Third is lineage. SDF has both the highest-fidelity and most high-performing SQL parsing on the market. And lineage and metadata is, of course, at the heart of the entire data control plane. Understanding how tables and columns flow throughout your entire data estate is what SDF's technology was originally built to do, and it has been proven out in the most complex data environments on the planet. Finally is local execution. It is common for the workflows of software engineers to run development environments on a local machine, then for higher environments to be cloud-based. The local development environment gives software engineers speed and control that are important in the very tight, iterative development cycle. But that’s not how it’s worked in data in the past. Most modern data platforms cannot be ‘run locally’, but that’s one of the superpowers of building a logical plan from the SQL query: you can take that logical plan and execute it in a local environment. And that’s exactly what SDF does in development, making the developer experience that much more responsive and delightful. The benefits above are only the start; the ability to deeply understand the SQL authored inside of dbt pipelines will fundamentally transform the experience of every dbt user. ## Gimme the details! What do I get? All of the benefits laid out above are being realized in production today by SDF’s existing customers. But it will take some work to get all of these capabilities integrated into dbt, and this won’t happen overnight. Our first goal is to get SDF’s SQL parsing capabilities integrated into dbt. These capabilities will enable meaningful improvements to the dbt developer experience, and we want everyone to have access to them. While SDF won't be included as part of the Apache 2.0 code base, _we plan to make meaningful parts of SDF’s capabilities available to all dbt users_—whether you’re using dbt Core or dbt Cloud. As we work through integration details we’ll share more about how this will work. But **if you use dbt today, you’ll be able to use this new tech**. In a few weeks, we're hosting a webinar between me and SDF's Co-founder and CEO to share more and answer any of your burning questions. [Be sure to register](https://www.getdbt.com/resources/webinars/accelerating-dbt-with-sdf). In the meantime, you can learn more about this acquisition and what it means for the bright road ahead for dbt by [reading our press release](https://www.getdbt.com/blog/dbt-labs-announces-sdf-labs-acquisition), the [follow-up blog post](https://www.getdbt.com/blog/advancing-the-data-control-plane-vision-with-sdf-and-dbt) written by dbt Labs Chief Customer Officer Ryan Segar, and SDF's [acquisition announcement blog post](https://blog.sdf.com/p/dbt-labs-has-acquired-sdf). Beyond that, all I can say for now is that the technology that SDF has built is foundational to the entire data control plane, and you should anticipate seeing it show up in more and more dbt experiences over the coming 12 months. We’ll share more in public as soon as we can. In returning, for a moment, to my statement from the very beginning of this post: I am just so very excited about this development and what it will mean for dbt users everywhere. This is the type of step function change that doesn’t come along in an industry very often, and it is an absolute privilege to be able to share this with the entire dbt Community. --- --- title: "Advancing the vision for the data control plane with SDF and dbt" description: "Improve productivity, reduce costs, and eliminate risks with dbt Cloud." url: "https://www.getdbt.com/blog/advancing-the-data-control-plane-vision-with-sdf-and-dbt" date: "2025-01-14" authors: ["Ryan Segar"] categories: ["Company"] --- # Advancing the vision for the data control plane with SDF and dbt Earlier today, we announced that dbt Labs has acquired SDF Labs. If you missed the announcement, you can learn more about it from Tristan [here](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs). Since launching into GA in mid-2024, [SDF](https://www.sdf.com/) has quickly emerged as a powerful tool in the data analytics space. The technology offers a multi-dialect SQL compiler and software toolset that results in remarkably fast developer feedback. This means that users can write and deploy data models orders of magnitude faster as compared to dbt Core alone. This extremely responsive development experience not only boosts developer productivity and pipeline velocity, but allows teams to “shift left” to catch data quality issues earlier, supports improved governance, and helps organizations optimize data platform costs. What’s more, integrating SDF into dbt brings additional rich metadata into the “consciousness” of dbt. This will give dbt Cloud’s data control plane improved table- and column-level awareness, enabling users to confidently address new advanced use cases, such as for PII and PHI. ![list of 5 primary ways SDF's tech will drive business outcomes for dbt customers](https://cdn.sanity.io/images/wl0ndo6t/main/f566743689ab3f7e4e782b217538b74419f1e96a-2180x500.png) Folding this technology into the dbt engine will give our users unprecedented data velocity and efficiency while organizations can enjoy improved data quality and optimized data costs. We look forward to integrating the SDF team and technology into our company and are excited to work with them in cementing dbt as the industry standard for data transformation. ## Advancing our vision of a data control plane Last year at Coalesce, we [introduced our vision for dbt Cloud as a data control plane](https://www.getdbt.com/blog/dbt-labs-unveils-ai-innovations-and-new-features-to-improve-collaboration-and-multi-platform) that supports users across every stage of the [analytics development lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle)—regardless of their title, technical aptitude, chosen data platform, or where they build and consume data. Central to this vision is empowering all users to collaborate on data. For a data control plane to truly become a central component of any organization’s data stack, it needs to be built in a way that meets all users where they are, regardless of the role they play in the organization. SDF helps us advance this vision in a number of ways. First, when developers are more efficient, it makes the platform that much more sticky. The performance and data quality benefits SDF will bring to the dbt workflow will help developers do more with dbt and cement it as the common framework for how data work gets done. **** Second, enabled by capabilities like the [visual editing experience](https://www.getdbt.com/blog/coalesce-2024-product-announcements) and [dbt Explorer](https://www.getdbt.com/product/dbt-explorer), as more collaborators standardize on dbt, SDF’s tech will supercharge their efficiency. SDF helps everyone in the collaboration workflow “shift left,” making it more efficient for them to come to a common understanding on what's happening and what needs to get done. Less time dissecting why data pipelines don’t look right, more time delivering value to the business. And third, SDF will add a new type of detailed metadata to the dbt consciousness—that is, metadata about the deep semantic understanding of the SQL itself—which enriches data lineage (both table- and column-level) and opens up an array of new complex use cases to be solved with dbt. The ability to traverse the data pipeline with this detail allows organizations to confidently embrace nuanced governance use cases (for example, that require tracing of PII or PHI to ensure compliance), helping business leaders build data trust and reduce business risk. ## Enhancing the business value of dbt While the combination of SDF and dbt brings a lot of productivity benefits to developers, the improvement in pipeline performance and data quality will deliver meaningful downstream impact to the business. Customers will be able to improve data quality while building trust in data through these new capabilities. Data quality is driven by the producer of data assets and trust is the primary interest of those that consume the data. Since good data quality leads to high trust, integrating SDF’s tech into dbt will ensure everyone can make decisions with confidence. Customers will also benefit from how SDF will help them drive cost optimization. SDF and dbt will provide organizations the visibility and tools required to optimize infrastructure, operational, and people costs in data. “Shifting left” to catch data quality issues earlier in the development process—immediately as the code is being written—not only makes developers more productive, but organizations can avoid unnecessary warehouse compute by validating data models without materializing anything in the warehouse. This capability not only improves velocity of trustworthy data products, but does so in a way that optimizes data platform costs. ![image of SDF's data development workflow](https://cdn.sanity.io/images/wl0ndo6t/main/1de04baee4a91eef1e6355a38fb1fefef9b111f0-2180x1242.png) ## The road ahead Over the next several weeks and months our team will be working to bring the SDF functionality into dbt. In the meantime, be sure to [join our upcoming webinar](https://www.getdbt.com/resources/webinars/accelerating-dbt-with-sdf) where Tristan and SDF’s co-founder and CEO will talk more about the benefits of incorporating SDF’s tech into dbt. We look forward to working with you, our customers, as we bring these new capabilities to your workloads. We hope you are as excited about this next chapter as we are, and look forward to working closely with you to help unlock the value and impact that SDF and dbt can have in your organization. --- --- title: "dbt Labs Acquires SDF Labs to Introduce Robust SQL Comprehension into dbt and Supercharge Developer Efficiency" description: "Technology integration will dramatically improve platform performance, developer efficiency and stakeholder trust." url: "https://www.getdbt.com/blog/dbt-labs-announces-sdf-labs-acquisition" date: "2025-01-14" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Acquires SDF Labs to Introduce Robust SQL Comprehension into dbt and Supercharge Developer Efficiency **PHILADELPHIA**, January 14, 2025 – [dbt Labs](http://getdbt.com), the pioneer in analytics engineering, has acquired [SDF Labs](http://sdf.com), the team of former Meta and Microsoft engineering leaders behind SDF, the next-generation data transformation technology. The acquisition will integrate SDF’s powerful multi-dialect, dbt-native SQL comprehension capabilities into dbt, delivering orders of magnitude improvements to dbt performance and enhancing the developer experience with new levels of efficiency, data velocity and data quality. SDF Labs, which emerged from stealth in June 2024, has focused its efforts on building a state-of-the-art toolset and framework to address one of the analytics industry’s biggest complexities – compiling and understanding SQL that users write, regardless of platform. Its technology, built on the Rust programming language and natively integrated into dbt, solves this challenge at lightning speed and at scale. SDF validates the SQL code a user is writing, immediately as it’s being written. This real-time feedback allows developers to embrace modern development accelerants like code completion and content assist as well as pinpoint errors and ensure data quality far earlier in the development process. This expedites data velocity, boosts data quality, and makes organizations much more efficient in their analytics practices. SQL comprehension also adds a new layer of detailed metadata to dbt’s table- and column-level lineage for enhanced data classification, enabling organizations to accomplish nuanced governance use cases. These capabilities will now be part of dbt. “We are acquiring SDF to bring SQL comprehension into dbt and usher in a new era of ‘what’s possible’ for analytics: supercharging developer productivity and heightening data quality, all while optimizing data platform costs,” said Tristan Handy, founder and CEO of dbt Labs. “SDF’s technology will bring a massive upgrade to the heart of dbt and the dbt user experience. This isn’t an incremental improvement to dbt; it’s a step-function change, and I’m very excited to work alongside the team to get this technology into the hands of dbt users everywhere.” dbt Labs has a strong track record of building experiences on top of and around the data workflow to solve problems at enterprise scale (as with [dbt Mesh](https://www.getdbt.com/product/dbt-mesh)), democratize data analytics to more types of collaborators (via capabilities like [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), [dbt Explorer](https://www.getdbt.com/product/dbt-explorer), [dbt Copilot](https://www.getdbt.com/product/dbt-copilot) and the recently announced [visual editing experience](https://www.getdbt.com/blog/dbt-labs-unveils-ai-innovations-and-new-features-to-improve-collaboration-and-multi-platform)), all designed to help organizations mature their analytics practices and embrace the Analytics Development Lifecycle ([ADLC](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle)). For the more than 50,000 teams who use dbt every week, these capabilities are critical to transforming data into reliable business insights. With this acquisition, dbt becomes an even more robust, reliable, and delightful platform to power enterprise data practices. As part of the acquisition, the SDF Labs team will become a part of dbt Labs. SDF Labs CEO Lukas Schulte is energized by the opportunity to introduce these powerful capabilities to even more organizations and members of the dbt Community. “Bringing SDF and dbt together is going to completely transform the dbt user experience with unprecedented levels of speed, accuracy, and velocity,” said Schulte. “The SDF Team and I are so excited to magnify the impact that our technology can have by powering the data control plane that sets the standard for the future of data analytics.” dbt Labs’ key industry partners also are eager to see the impact this new set of capabilities will have on users in the coming year as dbt integrates SDF into the ecosystem. “As demand for data intelligence grows, high-quality, real-time data is a must-have. dbt Labs is a key partner in our data warehousing business, which has grown 150% in the past year,” said Roger Murff, Head of Product Partnerships and Ecosystem at Databricks. “dbt developers can now build faster, with richer metadata, on the Databricks Data Intelligence Platform, and we look forward to seeing what our joint customers build in the new year.” Christian Kleinerman, EVP of Product at Snowflake, added: “dbt empowers enterprise customers to quickly launch on Snowflake, supporting customers’ long-term success as they build governed data and AI products on our platform. With this acquisition, we are thrilled that dbt developers will be able to more rapidly innovate with richer metadata, accelerating their value from Snowflake’s AI Data Cloud.” For dbt Cloud customer Bilt Rewards, this acquisition strengthens the role of dbt in their organization’s data practice and introduces new potential. "dbt is a critical component in our data stack that empowers our teams to streamline how data is managed and distributed across our organization. We’re excited about the addition of the SDF capabilities into dbt and the productivity improvements it will bring to our dbt development,” said James Dorado, VP Data Analytics, Bilt Rewards. “This acquisition positions dbt Labs as a best-in-class data development platform and will help us accelerate our development and ship data products faster.” For more detail on dbt Labs’ acquisition, read Tristan Handy’s blog post at [http://getdbt.com/blog/dbt-labs-acquires-sdf-labs](http://getdbt.com/blog/dbt-labs-acquires-sdf-labs). **About dbt Labs** Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today, there are 50,000 teams using dbt every week. --- --- title: "Data engineering at Snowflake" description: "Rahul Jain gives us a look inside at the data work happening at Snowflake." url: "https://www.getdbt.com/blog/data-engineering-at-snowflake" date: "2025-01-12" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Data engineering at Snowflake _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/data-engineering-at-snowflake-w-rahul). _ Rahul Jain is a data engineering manager for Snowflake's internal data organization. He joins Tristan to discuss the Indian tech scene, Iceberg, streaming, AI, and how Snowflake’s data team does data work. **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways from this episode ### There's this funny thing when you are the user of a software product, but you work for the company that makes that software product, then you have this dual role. You have to be a data engineering manager, but then you also have to explain and advocate for the platform that you're using. ### How do you balance these responsibilities? Are you mostly spending your time delivering data outcomes to the business, or are you mostly spending your time on stages in front of audiences? That's one of the reasons I joined Snowflake. Before joining Snowflake, I was a Snowflake customer. My team was implementing a data platform on Snowflake. It may sound a little cliche, but when I got introduced to the Snowflake platform, it was love at first sight. The ease of use and so many things. At Snowflake, my core responsibility is building data products and data-driven solutions, which help Snowflake’s internal businesses across different verticals. But additionally, one of the roles I play here is talking about use cases in the platform. I work closely with the marketing team, sales engineering team, and sales team. I give many keynotes, breakout sessions at global events and with the developer community I'm close to. ### That's explicitly a part of your role? That's not part of my role. My role is evolving into that. If you say on the books, the title I own, it's not part of my role. But I love doing it. And the leadership here is very, very appreciative. If you are a proud user of something, whether it's a tech product or any day-to-day utility product, then you automatically try to market it. What I do here at Snowflake, those use cases which I build, I just go and talk about it to the world, to the data community. ### Let's do it. Tell me, how does Snowflake do data engineering? First of all, before I jump into it, I just want to mention that I'm here at my personal capacity. This is not sponsored by Snowflake. Since we’re a cloud data platform, we take data very, very seriously. This is my sixth organization where I've been working in the last 14 years. And it is truly a data-driven organization. We practice data. We live and breathe data. Not only the data engineering team, all the functions, be it sales, sales engineering, marketing, finance, workplace, every function tried to have this data-driven mindset. My team is a horizontal team within Snowflake. And my team supports different verticals-GTM, finance, legal, and other verticals. We have a centralized repository of data where all the data which belongs to Snowflake comes to a single-tenant, single platform. And then based on the domain, when I say domain, you can call it the verticals, we cater to them. Most of the time we create data models. My team spends 80% of the time in analytics engineering creating data models, common data models, some aggregations, and then using data quality, observability, and data governance. We then share this with these business units so that they can create their own analytics if they have their own analytics team. Or sometimes we are engaged in enriching their source system data. We reverse it here or push back this golden data to their source systems, let's say Workday, Salesforce, ServiceNow, Jira, these sources. ### Snowflake has been using dbt for a long time. I'd be interested to hear if you feel like your use of dbt is different or novel based on your unique role in the ecosystem. I'd be curious to hear if there's other tooling that you use in your stack that's worth talking about. Yeah, so the stack is very big, but we’ve used dbt since the beginning, especially for data modeling and analytics engineering purposes. We are very, very satisfied users of dbt. And especially my team, they spend almost 70% to 80% of their time writing models in dbt and deploying them. ### What do you think about the talent market for dbt in India? It's still a new-ish product. There's, probably deep benches of talent in India using different ways to get similar jobs done. Do you have a hard time sourcing DBT talent in India or do you think there's a lot of it there? As you said, India is majorly heavily relying on Spark, Informatica and ETL. dbt is picking up especially with niche companies or new-edge tech companies. But I would say I still find some difficulty sourcing talent. ### Do you assume you're going to have to train people when you bring them into your team? I do, but the learning curve is very, very easy because there’s a lot of documentation available. When somebody new joins the team, there is a four-hour dbt workshop. I ask them to go through it on the first day. ### One of the interesting things about India is that often, and this is not true everywhere, people are worried about their budgets. This makes open-source tools like Spark and dbt more popular. But what's interesting is that I think it's fair to say Informatica is pretty freaking expensive. And yet there's a huge base of Informatica expertise in the country. Who is using Informatica, and how do you square these two things? When you talk about India's cost sensitivity, that’s true. But you need to understand the developer community in India, most of the time, is working for global companies head offices in the U.S., Europe, or Australia. India is not the revenue-generating entity. Thus these decisions of whether to have Informatica or dbt or Snowflake. These still are done in the headquarters where the company exists initially. Informatica is expensive, but this is getting paid from the headquarters. ### Changing gears, I think that you are on record as talking about Iceberg in public settings. Iceberg is an open-table format that has kind of taken off in popularity over the past two years. ### What do you think is driving customer interest in Iceberg? One is the interoperability, which results into the no-vendor lock-in mindset. This is a fast-evolving ecosystem, right? If customers want to be agile, then they are looking for some middle ground where they can think about switching the platform or the processing engine they are using currently. Open-table formats like Iceberg give you that kind of flexibility. You can store data in the open-table format and use a processing engine like Snowflake or Databricks to process it. You save money on storage, but it may cost more because you need the knowledge to use it and then keep it up-to-date. ### I think most people agree that larger companies are mostly driving the market. The main benefit they hear is that it's more flexible and can't be stuck with just one platform. ### There's a misconception that if you store data in the Iceberg table format, that you've done it. That's what everyone's talking about. But in fact, storing data in the Iceberg format is only a part of the game. The next phase is like, well, where's your catalog? ### This is where things get complicated. Snowflake made an announcement about Polaris at Snowflake Summit. There's internally managed, like Snowflake's managed catalog, and then there's externally managed catalogs. And I'm just curious if you could help us figure out the differences between these things and the advantages and the limitations. You said it right. Storing data in Iceberg format is one thing, but unless you have a catalog, you will not be able to query the latest data or keep track of the latest and all the atomic properties if you want to leverage it, right? When this Iceberg table format started getting traction, each platform or like Snowflake, Databricks, they all started creating their own catalog. And you need to understand what a catalog is not. A catalog is not storing the actual data. It is just a pointer to the data, which is stored somewhere in the cloud in Iceberg format. A catalog is just keeping the pointer to the latest data or the latest files. You can think of it as metadata. It keeps track of metadata. Now, where do you keep this metadata? One way of doing this is you keep this metadata with Snowflake in the Snowflake Managed Catalog. You need not worry about the UI or the console or how you and your team will view the catalog. ### And if you're using Snowflake Manage Catalog, could you point Athena to Snowflake Managed Catalog also, or is it just to be used by Snowflake? It is just to be used within Snowflake. From the Snowflake processing engine, you can only query the Snowflake managed catalog. That's why Snowflake came up with another concept of open-sourcing the catalog, or the externally managed catalog. It is currently in the incubation stage. It's called Polaris. If you don't want to have your catalog managed by Snowflake, you can manage your own catalog. You can bring that code base and create and manage your own catalog in your own infrastructure using the Polaris capabilities. But in this case, you need to take care of the wrapper, the front-end UI you want to put in front of Polaris. ### I really did appreciate the clarity that came from both Databricks and Snowflake in 2024 standing up on stage and both saying open-table formats are a big deal and we care about Iceberg. I think it's really meaningful that both CEOs got up on stage and said that. I think it's a pretty reliable indication of where the industry is going. ### Do you have any expectations on a performance difference when people use Snowflake native storage versus Snowflake managed Polaris? I think it's very obvious, right? If the data is stored inside Snowflake, the native data storage performance will be faster for obvious reasons. They will always be faster than the externally managed catalog and then stored in the Iceberg format. ### I think of Snowflake's history with AI and LLMs as having two distinct phases. There's the pre-Sridhar phase and the post-Sridhar phase. And then the post-Sridhar phase is more like Cortex. Do you think that's an appropriate way of thinking about this? Yeah, definitely. Sridhar comes with a lot of experience in artificial intelligence, especially in Semantic Search. He's a technologist, by his work in the past with Google, with his own startup. The moment Sridhar joined Snowflake, all of a sudden Cortex came into picture. The core philosophy of Snowflake is simplicity, the way the platform was built. Cortex functions, whether it's machine learning powered functions or LLM functions, these things are so simple to use. And there is so much excitement within Snowflake about these functions. That was the shift which happened post-Sridhar, where everybody is empowered to use these large language models, not directly but in the form of SQL functions. And there is lot of talk about how to expand that and create more. ### Since Cortex, have you seen adoption of AI in the Snowflake platform accelerate? 100%. Sometimes I think we are doing too much inside the company. Everyone, not just the data team, but also the project management and all the other non-technical teams, can write SQL. We still have to figure out internally a lot of use cases which will impact at scale. But still these Cortex functions we are using heavily internally. ### Do you have any dbt pipelines that are just end-to-end Snowflake dynamic tables? I have not personally gone all in on dynamic tables, but I'm curious if you've pushed it hard. We are right now in that phase where we are moving some of the pipelines, which were managed through Airflow Decks to complete dynamic tables using dbt. And I would say we are not completely there with end-to-end pipeline using dbt using dynamic tables inside dbt. But one of the initiatives we are currently doing where we are migrating from those Airflow Decks to dynamic tables and on dbt itself. Those are more from the master data management side. ### It's so interesting, many of the things that we think about in the industry are, there's an underlying consensus that we're fundamentally talking about batch. I didn't say Airflow, but you added it to the conversation. This makes sense because if you put everything in dynamic tables, then suddenly there's no orchestration. ### Have you spent much time thinking about how you provide an observability? What happens if it fails? Yeah, so you're right. That's why I get very practical and I tell my team the same thing. If someone is coming to you because this is new, real-time or streaming, this sounds great, right? But do you really need it? What's the end use of it? Who is consuming that data? Do they really need real time stuff? If there is no business impact, keep things in batch mode because you can observe them very well. And these are more stable. I will be very transparent, we have not built anything concrete to observe the dynamic table, data processing using dynamic tables. ### I don't think you're alone there. We as an industry are still early. ### Let me ask you the question to close out the podcast. What is something that you hope is true of the data industry over the next five years? Data literacy is something which is increasing and especially with the democratization of LLMs. People are taking data seriously. And a lot of tools are evolving very fast. if you can interact with the data using copilots, that puts a lot of focus on data-driven products and the data industry. That's why I'm very, very hopeful for the next five years. --- --- title: "The Analytics Development Lifecycle: Discover and Analyze" description: "You’ve published your data project to production. But can your stakeholders find, use, and refine it?" url: "https://www.getdbt.com/blog/adlc-discover-analyze" date: "2025-01-03" authors: ["Sara Gawlinski"] categories: ["Learn"] --- # The Analytics Development Lifecycle: Discover and Analyze Data isn’t any good if your stakeholders can’t find it. Splunk estimates that as much as [55% of a company’s data might be dark data](https://www.splunk.com/en_us/form/the-state-of-dark-data.html) - i.e., data that lies dormant. [In this series](https://www.getdbt.com/blog/adlc-plan), we've been looking at the different stages of the [Analytics Development Life Cycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle), a framework for a mature analytics workflow modeled after similar processes in the software development world. The last phase, the Discover and Analyze phase, is the phase that delivers business value to data consumers. It's also the phase that we have found to be the one that's often the most immature in many companies today. Without a solid Discover and Analyze phase, the hard work you've put into developing high-quality data products may go to waste. This phase is also critical for providing and processing feedback, which kicks off another run of the ADLC and fosters a culture of continuous improvement in data. We'll look at the key attributes of a successful Discover and Analyze phase, and how to tie back the input from this phase into the start of a new loop of the ADLC. ## The practices in the Discover and Analyze phase The ADLC unites Data development and Operations into a single development cycle that emphasizes shipping small analytics code changes with high quality. It breaks down the artificial barriers that have grown between data product development and data management by creating a single process in which all data stakeholders work together to plan, develop, test, deploy, use, and manage data changes. A successful Discover and Analyze phase enables two key things: - Discovering data sets, dashboards, and metrics; and - Using these data artifacts to answer questions These answered questions can then themselves serve as a basis for another round of the ADLC. This is what makes the ADLC a loop —analysts create insights in this phase that data engineers can then productize and make available to a wider audience. The goal is to encourage experimentation and exploration while also promoting maturity and productization of the overall data system. The key practices in the Discover and Analyze phase include: - Discovering and operating on data - Leaving feedback - Requesting and delegating access - Ensuring data is accurate - Ignoring implementation details Let’s look at each of these in detail. ### Discovering and operating on data Data discovery is challenging within a sprawling enterprise. The data that users need is often split over hundreds of thousands of sources. This can make it hard —or impossible — to locate unless users know where to look. Lack of discoverability is one of the causes of dark data. Dark data costs money —not only from lost business opportunity, but due to the compute, storage, and personnel spend required to transform and maintain it. After publishing a data set, your stakeholders need a way to find it. A common solution is some form of data registry or repository that can catalog data from all sources across the company. Tools such as [dbt Explorer](https://www.getdbt.com/product/dbt-explorer), part of dbt cloud, provide this out of the box. After publishing a data model to production, stakeholders can search for it via a simple unified interface. They can also find any documentation that accompanies the model. ### Leaving feedback As mentioned above, the ADLC is a loop. That means you need mechanisms for decision-makers to provide feedback to data engineers and analysts. These feedback mechanisms can take multiple forms. For example, you may hold regular brown bag sessions or productive briefing meetings to gather feedback one on one. You might also provide internal support forms or access to a ticket issuing system where data stakeholders can log feature requests. A common problem in data analytics in this stage is that most such requests go through a central data engineering team. The team quickly becomes a chokepoint for data issues. Tools for data discovery and documenting data sets can help with this by providing stakeholders with the ability to self-service answers to specific questions. Additionally, the ADLC emphasizes that roles aren't static, but flexible. Using modern data modeling tools that leverage widely understood technology such as SQL, different people may wear the data engineer, analytics engineer, and decision-maker hats at different times. This means that a wider range of people can develop and update analytics code than ever before. When developing feedback mechanisms, keep in mind this variability and create processes that allow all applicable stakeholders to capture, track, and act on feedback. Facets of a successful feedback mechanism include: - Ensuring requests are routed to the data product owner(s) - Setting SLAs around review and action for requests - Tracking SLAs to ensure stakeholder feedback is being considered and incorporated into future releases ### Requesting and delegating access To be fair, one reason companies don’t make data more generally available is that not all data **should** be generally available: - Some internal data may contain Personally Identifiable Information (PII) that should only be viewable by a subset of employees under strict protocols - Other data may contain Intellectual Property (IP) or other sensitive internal information - Many companies will also have to limit exposure of customer data to comply with local data handling laws, such as the [General Data Protection Regulation (GDPR)](https://gdpr-info.eu/) in the European Union Along with making data discoverable, you need a system for requesting and delegating access. By default, users should only see relevant metadata when conducting a data search. You should also establish mechanisms for requesting and either approving or rejecting access to data. Role-Based Access Control (RBAC) supports granting access to data based on a user’s business function. Tools such as dbt Cloud [support granting access to models using RBAC](https://docs.getdbt.com/docs/collaborate/govern/model-access), simplifying permissions management by managing permission sets via business functions instead of on a person-by-person basis. ### Ensuring data is accurate Just because somebody can _find_ data doesn't mean they can _trust_ it. Data consumers need proof that a data set is both timely and accurate. Data engineers and analysts can help provide this reassurance to stakeholders through mechanisms such as defining [data quality metrics](https://www.getdbt.com/blog/data-quality-metrics). Tools such as the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) can publish these statistics as standardized metrics so they're centrally available to everyone. Additionally, tools such as dbt's automatically generated column-level [data lineage](https://docs.getdbt.com/docs/collaborate/column-level-lineage) can show the origins of data in a data set. This gives data stakeholders increased confidence in the data as relevance and provenance. ### Ignoring implementation details Finally, many of the products can be hard to use for data consumers without a deeper knowledge of what the data is, where it lives, and its various idiosyncrasies. A mature analytics process should avoid this by delivering data products that _just work_. An analyst, for example, shouldn't have to understand the structure of a dozen or more distributed tables to grab data for a quarterly sales report. All of the relevant data and data structures should be documented in a data model and easily accessible to anyone with a basic knowledge of SQL. ## How to implement the ADLC Over time, the ADLC becomes a repeatable process that your teams can use to ship well-scoped data changes with high velocity and high quality. Getting there, however, requires more than just having good processes in place. It requires a data platform that simplifies working with data. dbt Cloud is your control plane for data that makes supporting the ADLC easy. Its platform features support data developers and their stakeholders across various stages of the analytics development lifecycle to make data analytics a team sport. It also provides the trust signals and observability features required to ensure all data outputs are accurate, governed, and trustworthy. Want to learn more? [Contact us for a demo ](https://www.getdbt.com/contact)today. [Watch video](https://youtu.be/_t8hGUXo8PA?si=zxFmUZkMC2DAB8z4) [Watch video](https://youtu.be/cXrSsu23wmI?si=MhM-9K7n_TrcfUIB) --- --- title: "The Analytics Development Lifecycle: Operate and Observe" description: "Analytics projects don’t end once you ship a change to production. Here’s how to ensure your code keeps working as expected." url: "https://www.getdbt.com/blog/adlc-operate-observe" date: "2025-01-02" authors: ["Sara Gawlinski"] categories: ["Learn"] --- # The Analytics Development Lifecycle: Operate and Observe The [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) is a methodology you can use to bring higher quality and repeatability to your analytics projects. With the ADLC, you can use [a DataOps approach](https://www.getdbt.com/blog/what-is-dataops) to ship data products to production more frequently, reducing expensive rework and errors. In the previous installments of this series, we focused on the **Data** phase - on how to create, test, and deploy new analytics code changes. In this and the next article, we're moving on to the **Ops** phase —operationalizing your analytics code. Below, we’ll dive into how the Operate and Observe phases ensure your code is available, performs well, and is truly free of errors. ## Practices in the Operate and Observe phase In the past, data engineers would work in isolation, deploying large analytics code changes in a haphazard fashion. The ADLC, patterned off of the [Software Development Lifecycle (SDLC)](https://aws.amazon.com/what-is/sdlc/), uses a DataOps approach to ship smaller changes in an iterative and incremental fashion. Both sides, Data and Ops, are critical to any successful analytics project: - **Data**: Obtain stakeholder agreement, develop maintainable and reusable code, test your work, and deploy it automatically to staging & production - **Ops**: Observe code as it runs in production, alert and respond to errors, and ensure data is both discoverable and well-governed The Operate and Observe phases are critical phases that ensure your code not only works but continues to work as expected and remains performant over time. Best practices in the Operate and Observe phase include: - Provide always-on analytics - Test in production - Catch errors before customers do - Tolerate and recover from failure - Choose your own metrics and measure them religiously - Don’t overshoot Let’s dive into each of these in detail. ### Provide always-on analytics Older data systems frequently required data engineers to take them offline in order to batch import data. At the speed of business and volume of data we deal with today, however, this won't fly. Taking down a report or data-driven application for making critical business decisions or supporting real-time tasks such as [fraud detection](https://www.getdbt.com/blog/data-transformation-examples) could cost your business —and your customers —money. Modern software systems are designed with an always-on architecture. Our data system should be, too. For analytics, this usually means having multiple data sources and performing replication to secondaries as new data trickles into the primary. Any required downtime for an analytics system should be kept to a necessary minimum and outside of core business hours. ### Test in production Earlier in our series, we emphasized [the importance of building unit, data, and integration tests](https://www.getdbt.com/blog/adlc-test) to vet analytics code changes before they made it into customer’s hands. [A multi-environment deployment process](https://www.getdbt.com/blog/adlc-deploy) that verifies changes against test data reduces the risk of shipping data transformation code that results in incorrect data or downtime. However, pre-production environments can never be exact replicas of production. To protect consumer privacy, pre-production data must be either mock data or a subset of real-world data that's been anonymized. Additionally, there are other data workloads running constantly in production that aren't running in pre-production during testing. The bottom line is you shouldn't just test in preproduction —you should be testing in production as well. [Testing in production](https://www.browserstack.com/guide/testing-in-production) is a technique from software that continues running tests on a live environment with real data. By testing within the context of other workloads and incoming, real-world data, you can uncover hidden issues, edge cases, and performance concerns. Most major tech companies use some form of test in production to validate and monitor their solutions post-deployment. Using dbt Cloud, you can easily [set up the tests you already wrote alongside your models](https://docs.getdbt.com/docs/build/data-tests) to run in prod as well. ### Catch errors before customers do Testing in production helps catch errors before they land in your stakeholders’ laps. Whether a missing data field value, an incorrect value, relational cardinality violation, or other data anomaly, catching issues before customers reduces confusion and miscommunication. Testing in production, however, is only one part of catching errors. You can catch more errors, more quickly, by implementing multiple safety measures, including: - Dashboards for observing metrics - Alerts the system can throw on any error it detects - Automatic incident creation - Tools - such as logging and [data lineage](https://www.getdbt.com/blog/guide-to-data-lineage) - for analyzing issues and finding their root cause - A mature process around incident triage and response, including a distributed, around-the-clock staffing model and on-call rotation for support - Staff who have clear ownership of and appropriate training in resolving data issues quickly Tools like dbt Cloud provide some of these functions - such as job notifications, model notifications, and webhooks for communicating failure status - out of the box. This enables adding reliability to your overall data control plane with little additional engineering overhead. ### Tolerate and recover from failure Part of being “always-on” is not going down due to an undetected failure. As your analytics system matures, consider building in additional measures that enable detecting errors, issuing alerts, and restarting once data engineers have deployed a fix. dbt Cloud bakes this in by enabling [job retries from the point of failure](https://docs.getdbt.com/docs/deploy/retry-jobs). This can be orchestrated via the dbt cloud dashboard, command line, or even automated through an API endpoint. This enables building even more sophisticated error recovery mechanisms that attempt retries after automated error resolution. ### Choose your own metrics and measure them religiously There are a number of metrics to measure the general reliability of your data systems, including uptime, availability, latency, and throughput. There is also a wealth of possible metrics for [assessing data quality](https://www.getdbt.com/blog/data-quality-metrics), which you can use to improve reliability over time across your entire data estate. It's important to define the metrics that matter to you and your team as part of the ADLC process. Preferably this is something you've done during [the Plan phase](https://www.getdbt.com/blog/adlc-plan). Definitions should include both how these metrics are calculated plus their acceptable thresholds. It’s also important, not just to monitor them, but to centralize and document their definitions so they’re accessible to all data stakeholders. You can also improve your overall data reliability by integrating with third-party tools that specialize in data quality. [dbt Cloud integrates with multiple third-party platforms](https://www.getdbt.com/product/integrations) that provide advanced data quality monitoring services, such as ML-based anomaly detection, code impact analysis, and root cause analysis. ### Don’t overshoot Software engineers like to warn that we shouldn't let the perfect be the enemy of the good. Once you have general stability in your analytic systems, it can be tempting to eliminate every last issue. This is almost always counterproductive. Adding an additional [nine of reliability](https://www.splunk.com/en_us/blog/learn/five-nines-availability.html) carries exponentially higher costs. This is encapsulated in [the 10x9 Rule](https://blog.alexewerlof.com/p/10x9): For every nine you add, you increase reliability 10x - but at 10x the total cost of your solution. Determine when you’ve achieved “enough” reliability and when you should rely on your systems and processes to resolve previously undetected errors quickly. Aim for progress, not perfection. ## Conclusion Ensuring quality doesn't end once you ship something to production. Constant testing and monitoring are required to verify your changes work as expected in a real-world context. The good news is that this creates a virtuous cycle. As you encounter issues in the real world, you can anticipate and test for them in subsequent releases. This makes your analytics code higher quality and more resilient with every release. In the [final installment of the series](https://www.getdbt.com/blog/adlc-discover-analyze), we'll look at how to make sure stakeholders can find and benefit from new data products, as well as provide feedback for future releases. [Watch video](https://youtu.be/TWUZxyF-u-Q?si=oN7p4NWIPuCAIPx9) --- --- title: "Building an ETL pipeline: Best practices" description: "Turn raw data into reliable insights with proven ETL and ELT best practices—plus how dbt helps teams scale quality and trust." url: "https://www.getdbt.com/blog/etl-pipeline-best-practices" date: "2024-12-31" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Building an ETL pipeline: Best practices Data transformation is the heart of data quality, turning raw data from multiple sources into high-quality information that business decision-makers can mine for insights. At the heart of the data transformation process is the data pipeline. Most data engineers [are managing and maintaining dozens of such pipelines](https://www.montecarlodata.com/blog-data-quality-survey) across dozens or more data sources. Running these with high performance and high quality requires implementing a set of best practices to manage changes, facilitate collaboration, and push changes to production safely. The traditional approach to implementing data pipelines is ETL—Extract, Transform, and Load. We’ll look at both [ETL](https://www.getdbt.com/blog/extract-transform-load) and its modern successor, [ELT](https://www.getdbt.com/blog/extract-load-transform) (Extract, Load, Transform), and look at the key best practices you should implement to ensure high-quality data transformations. ## What is an ETL pipeline? An ETL pipeline is an automated workflow that moves and processes data. It takes raw information from various sources and transforms it into useful business insights. ETL stands for Extract, Transform, and Load. These three steps form the backbone of most data processing workflows. - **Extracting data** means pulling information from source systems. This could include databases, APIs, files, or streaming data sources. - **Transforming data** involves cleaning, validating, and reshaping the raw information. You might filter out errors, combine datasets, or calculate new metrics. - **Loading data** means storing the processed information in a target system. Usually this is a data warehouse or analytics database. [Data transformation is just one part of the broader ETL process.](https://www.getdbt.com/blog/data-transformation-vs-etl) While transformation focuses on changing data structure and content, ETL encompasses the entire workflow. ### Example: converting sensor data into a table Let’s say that your sensors generate timestamps, measurements, and device IDs throughout the day. The extraction step pulls this raw data from sensor systems. The transformation step adds proper units, readable labels, and standardized timestamps. The loading step stores the clean data in your analytics database. [Modern tools like dbt can automate these pipelines completely](https://www.getdbt.com/blog/automating-your-etl). You write transformation logic once, then the tool handles scheduling and execution. This automation ensures consistent data processing without manual intervention. ## ETL vs. ELT: Identifying the “best” workflow for data [ETL and ELT](https://www.getdbt.com/blog/etl-vs-elt) are sometimes used interchangeably these days. However, historically, they take slightly different—but significant—approaches to working with data. Both processes take raw data from one or more sources and apply a series of [data transformations](https://www.getdbt.com/blog/data-transformation) to create a new dataset in a target system. The difference lies in when the transformation occurs: - ETL transforms data, often semi-structured or unstructured, before loading it into the target system - ELT, by contrast, loads the raw data, often from other relational or data warehousing systems, into the target system and then transforms it in-place For the most part, in a cloud environment, you’ll want to use an ELT workflow, which is what most organizations are doing today. There are a few reasons for this: - Scalability - Flexibility - Cost efficiency - Data democratization ### Scalability Compared to ETL, ELT scales better because it takes advantage of the processing power of cloud-native systems like Snowflake and Redshift. Once data is loaded, these systems can transform it efficiently at scale. ELT also generally handles complex datasets better—and faster—than ETL. In ETL, the larger the dataset becomes, the harder it is to manage, as the data must be fully transformed before it can be stored in its final form. ### Flexibility In ETL, the source data is transformed en route, and only the transformed version ends up in the data warehouse. This can lead to a couple of issues: - If the original data isn’t preserved anywhere, it’s lost forever. That can have severe implications for compliance and data governance, especially if errors are discovered in the pipeline. - The new data format might not be the ideal format for all teams. ELT preserves more flexibility and accountability by loading the original data first. This guarantees that you always have a copy of the original. It also enables using the data for different use cases. ### Cost efficiency ETL has historically required specialized hardware and software to implement. By contrast, ELT makes use of the native capabilities of cloud systems, standing up compute resources (like virtual servers) when transformation work is needed and shutting them down immediately afterwards. ### Data democratization In ETL, no one except the data engineering team typically has access to the original data until it gets loaded into the data warehouse in its final form. In ELT, the original data is immediately available in the data warehouse. This enables other teams to discover it and build their own data pipelines immediately, without waiting on the centralized [data engineering](https://www.getdbt.com/blog/what-is-data-engineering) team to act on a support ticket. #### Related reading: - [How to build scalable data pipelines with Snowflake and dbt](https://www.getdbt.com/blog/data-pipelines-snowflake-dbt) - [Data orchestration vs ETL: What's the difference | dbt Labs](https://www.getdbt.com/blog/data-orchestration-vs-etl) ## ETL pipeline use cases ETL pipelines power data-driven decisions across every industry. They transform raw information into actionable business insights at scale. - **Business intelligence and reporting.** ETL pipelines consolidate data from CRM, marketing, and financial systems into unified dashboards. This enables comprehensive performance tracking and regulatory compliance across departments. - **Data migration and warehousing.** Organizations use ETL to transfer legacy systems to modern platforms while transforming data formats during transit. This ensures data integrity and standardization before loading into target systems. - **Batch analytics processing.** ETL pipelines process large datasets during scheduled intervals, transforming raw data into structured formats. This provides reliable insights for strategic business decisions and historical reporting. - **Operational data integration.** Pipelines automatically sync customer service platforms with inventory management and billing systems. This eliminates manual data entry and creates seamless workflows across departments. - **Financial services.** Banks use ETL for fraud detection by analyzing transaction patterns and regulatory reporting requirements. These pipelines transform complex financial data into standardized formats for compliance auditing. - **Healthcare.** Hospitals integrate patient records from multiple systems into unified electronic health records. ETL processes clean and standardize medical data before loading into centralized databases. - **E-commerce.** Retailers synchronize inventory data across online and physical stores through nightly ETL processes. These pipelines transform product information and pricing data into consistent formats across all channels. ## Best practices for ETL and ELT pipelines In most modern data stacks, ELT is the recommended approach—but there are a few cases where ETL still makes sense, including: - You’re dealing with a small dataset that requires complex transformations - You have to integrate with a legacy or third-party system that works better with an ETL approach - Scenarios such as Internet of Things (IoT), where you want to perform transformations to combine data from different formats before storing it, or you need to de-duplicate entries Whether you use ETL or ELT, however, there’s a set of best practices you can follow to improve the quality, reliability, and performance of your data transformations. #### Related reading: - [Analytics engineering: Six best practices for success | dbt Labs](https://www.getdbt.com/blog/analytics-engineering-six-best-practices) - [Data quality best practices: Six essential principles for analytics and AI | dbt Labs](https://www.getdbt.com/blog/data-quality-best-practices) - [Data transformation: Six critical best practices | dbt Labs](https://www.getdbt.com/blog/data-transformation-best-practices) ### Have a workflow process in place Many data pipelines—even today—are driven primarily by engineering and managed on an ad hoc basis. There are few processes in place to establish SLAs, version pipelines, and respond to data pipeline issues. That leads to: - Data that diverges from business requirements - Untested (or undertested) data changes pushed to production - Broken pipelines and outdated reports that lead to impaired trust in data - Slow decision-making velocity We originally created dbt to address what we saw as deficiencies in the data workflow process. We sought to bring some of the best practices from software engineering - version control, deployment pipelines, testing, and documentation—into the data world. Today, we’ve expanded that mission by championing what we call the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). Like its counterpart, the Software Development Lifecycle (SDLC), the ADLC aims to break down the barriers between data team members—data engineers, analysts, and business decision-makers. It ensures each new data pipeline or pipeline improvement aligns with core business objectives and has measurable key performance indicators (KPIs). From a quality standpoint, the ADLC focuses on creating small, well-tested changes that engineers promote to production with high velocity. Once changes are in production, they’re monitored, operationalized for performance, and made available for analysts and business users to discover in a self-service manner. ### Use version control and enable collaboration A major source of errors in ETL and ELT pipelines is the lack of a central repository for a data pipeline’s source code. At a minimum, all pipelines should be tracked and versioned using a [version control system](https://docs.getdbt.com/docs/collaborate/git/version-control-basics). That guarantees that all data changes are represented as code and can be discovered, changed, or even reverted as needed. Besides enabling tracking changes, version control enables collaboration among data engineers, analytics engineers, and other technically savvy data users. It provides a single source of truth that people can discover, modify, and submit for approval. That enables more people across the company to contribute to building and maintaining data pipelines. ### Build a culture of quality around data workflows [According to Monte Carlo](https://www.montecarlodata.com/blog-data-quality-survey), half of all companies surveyed said that 25% of their revenue is impacted by poor data quality. That speaks to a need to build quality controls and measures into every part of the ADLC. Part of building a culture of quality means that engineers should be [creating tests](https://www.getdbt.com/blog/adlc-test) for all new data pipelines and for any changes they implement. Equally important, however, is implementing a version control-based process for enforcing quality checkpoints: - Engineers should work in their own source control branches, separate from the “main” branch containing production code - After testing their changes locally, engineers [create a pull request (PR)](https://docs.getdbt.com/docs/deploy/continuous-integration) to push code into the main branch - The PR triggers a manual review of their code, as well as an automatic run of any tests they’ve written - Once approved, a Continuous Integration/Continuous Deployment (CI/CD) process further tests and deploys their changes automatically to production ### Keep separate environments for each part of the dev process No one should be developing changes against production data. At a minimum, data pipeline development should have three sets of environments: - One or more **dev** environments that engineers can test their changes against locally - A **staging** environment where the CI/CD pipeline runs all tests against before deploying to production - The **production** environment, against which only fully tested and approved changes run #### Related reading: - [Building a robust data pipeline with dbt, Airflow, and Great Expectations](https://www.getdbt.com/blog/building-a-robust-data-pipeline-with-dbt-airflow-and-great-expectations) - [Automating your ETL: A guide to improved efficiency](https://www.getdbt.com/blog/automating-your-etl) ### Develop and maintain a style guide Part of enabling collaboration on data pipelines means creating code that’s readable and understandable by others. A [style guide](https://docs.getdbt.com/best-practices/how-we-style/0-how-we-style-our-dbt-projects) brings clarity and consistency to all data pipeline code by laying out a set of conventions that all engineers should follow. ### Document all workflows Most data pipelines go undocumented, which makes it harder for others to understand: - How to modify them - How to use the resulting data - Where the data was derived from and how it was calculated [Documentation](https://docs.getdbt.com/docs/build/documentation) enables pipeline maintainers to understand how and why certain decisions in the code were made. It also makes it easier for data consumers to have confidence in the resulting datasets, which increases overall data trust. ### Monitor data pipelines in production The [Operate and Observe](https://www.getdbt.com/blog/adlc-operate-observe) phase of the ADLC acknowledges that a data pipeline isn’t “done” just because the code has shipped. Monitoring ensures that all new data pipeline code is running smoothly by testing in production, reporting errors, and tolerating and recovering gracefully from failure. It also ensures code is meeting defined performance metrics and is answering users’ requests in a timely fashion. #### Related reading: - [Building a data quality framework with dbt](https://www.getdbt.com/blog/building-a-data-quality-framework-with-dbt-and-dbt-cloud) - [Data pipeline automation: Why it’s important | dbt Labs](https://www.getdbt.com/blog/data-pipeline-automation) ### Optimize your pipelines based on your platform All of the above advice applies to any approach to crafting ETL and ELT pipelines, regardless of the language or tools you use. Your toolset will usually support additional features that make developing and shipping revisions to data transformations faster. In dbt, for example, you create new data transformations by creating [a new data model](https://docs.getdbt.com/docs/build/models). Within dbt, you can optimize your pipelines for better performance and lower cost in a number of ways: - Using the [Defer to Production](https://www.getdbt.com/blog/optimize-costs-with-dbt-cloud-defer-to-production-feature) feature, engineers only need to rebuild models that they’ve changed locally - Engineers can also use [rerun from point of failure](https://docs.getdbt.com/docs/deploy/retry-jobs) to quickly test a fix to their in-development code - dbt supports different approaches to [materializating your data](https://docs.getdbt.com/docs/build/materializations) to optimize both performance and storage costs ## Enabling collaboration on ETL and ELT pipelines at scale Building out the infrastructure to support these processes for ETL and ELT pipelines from the ground up requires significant engineering investment. This is made even more challenging by the heterogeneous nature of most corporate data environments, which may be using hundreds of different data sources. dbt can serve as your company’s data control plan for managing ETL and ELT pipelines across the enterprise: - It’s natively interoperable across various cloud and data platforms, so you’re never locked-in - Its platform features support data developers and their stakeholders across various stages of the analytics development lifecycle to make data analytics a team sport - It provides the trust signals and observability features required to ensure all data outputs are accurate, governed, and trustworthy. To learn more about managing data pipelines at scale with dbt, [ask us for a demo today](https://www.getdbt.com/contact). ## FAQs about ETL pipelines **What's the difference between ETL and ELT workflows?** ETL (Extract, Transform, Load) transforms data before loading it into a destination system. ELT (Extract, Load, Transform) loads raw data into the warehouse first, then transforms it in place. ELT is better suited to modern cloud data platforms like Snowflake, Redshift, and BigQuery, offering scalability, better performance, and the ability to preserve raw data for compliance and future reuse. **Why is version control important for data pipelines?** Version control brings software engineering rigor to data workflows. It provides a single source of truth, enables collaboration via branching and pull requests, and ensures every change is tracked and auditable. Teams can test updates in isolation and review changes before deployment—improving reliability and reducing risk. **What role does documentation play in ETL/ELT pipelines?** Documentation is crucial for understanding how to modify pipelines, use the resulting data, and trace data lineage. Without proper documentation, it's difficult for team members to understand how and why certain decisions in the code were made, leading to maintenance challenges and reduced confidence in the data. Well-documented pipelines reduce maintenance headaches and ensure teams don’t lose institutional knowledge over time. **What are the benefits of using ELT instead of ETL?** ELT offers several advantages: - **Scalability**: Transformations leverage the power of modern cloud warehouses - **Flexibility**: Raw data is preserved, enabling multiple use cases and auditability - **Cost efficiency**: Compute runs on-demand, only when needed - **Data democratization**: Analysts and other teams can self-serve from raw data without waiting on data engineers **How can teams build a culture of data quality?** Start with automated testing for all new models and pipeline changes. Enforce quality through code reviews, pull requests, and CI/CD workflows. Maintain dev and production environments separately, and ensure all transformations are documented. These practices build confidence in your data and make quality a shared responsibility. **How do I optimize data pipelines for better performance?** Use platform-native features to streamline development and reduce costs. In dbt, for example: - Use **Defer to Production** to only rebuild models you’ve modified - Use **Rerun from failure** to quickly recover from errors - Choose materializations (e.g., tables vs. incremental models) based on usage and freshness needs - Continuously monitor production jobs and tune based on query patterns and bottlenecks **How does dbt support data pipeline best practices?** dbt turns SQL into a version-controlled, testable, and collaborative workflow. It supports best practices like: - **Version control** with Git - **Automated testing** for data quality - **Modular development** and reusable code - **Lineage and documentation** for transparency - **Job orchestration and observability** With dbt, teams manage transformations like software—improving trust, performance, and maintainability. --- --- title: "The Analytics Development Lifecycle: Deploy" description: "How you can borrow techniques from software development to launch analytics code changes safely in production." url: "https://www.getdbt.com/blog/adlc-deploy" date: "2024-12-30" authors: ["Sara Gawlinski"] categories: ["Learn"] --- # The Analytics Development Lifecycle: Deploy For too long, the data world has shipped analytics code changes to production in an ad hoc manner. The result is often data errors, downtime for users, and wasted engineering effort. It’s time we got smart about shipping pipeline changes to production. Using techniques borrowed from software engineering, we can take an approach to deploying changes that reduces human error and builds in numerous quality checks throughout the process. In our latest series of articles, we’ve been covering how you can use the different stages of what we call the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) to ship high-quality analytics code. In this installment, we’ll look at the principles behind an automated approach to deploying data changes from developer’s boxes and into data consumer’s hands. ## The basics of CI/CD If you’re familiar with software engineering best practices, you may be familiar with the basic tenets of Continuous Integration and Continuous Deployment, or CI/CD. For those who aren’t, here’s a quick refresher: **Continuous Integration** builds and tests changes to code as soon as it’s checked into a specific location of a source code [version control system](https://docs.getdbt.com/docs/collaborate/git/version-control-basics). When developers think they’re ready to shop their changes, they merge their changes from the branch of source code they’ve been working on into one that triggers a series of tests and sanity checks against their work. **Continuous Deployment** pushes these changes automatically to production. To do this, the process first pushes them to pre-production environments, where another series of tests are run against testbed data. It then runs any migration procedures required to make the changes live, using telemetry to ensure everything’s operating within expected parameters. CI/CD builds in numerous safeguards to the deployment process, such as: - Submitting all changes to review by another team member before going live - Testing changes in multiple environments before production release - Rolling back a change quickly and automatically if a deployment exhibits problematic behavior - Scoping changes to the smallest possible unit of release to limit the scope of potential errors It turns out that these processes from software engineering map nicely to analytics code systems like dbt. [dbt’s models](https://docs.getdbt.com/docs/build/models) capture all analytics code changes in text as SQL or Python code. This means we can capture all changes in source control, [validate them with testing](https://docs.getdbt.com/docs/build/data-tests), and employ automatic processes to [push vetted changes to production](https://docs.getdbt.com/docs/deploy/continuous-integration). ## Principles of the ADLC Deploy phase The exact steps of your Deploy phase will reflect the specific needs of your company. However, the following are universally attributes of any mature analytics code deployment process: - Managed in multiple environments - Size of change chosen by developers - Triggered by source code merge - Automated deployment - Downtime-free - Automated rollbacks ### Managed in multiple environments Pushing changes directly to production is a surefire way to break existing data workflows and applications. That's why, in the software world, dev teams deploy their changes through multiple environments, such as staging and testing, before pushing them live to end users. We can emulate this best practice in the data world by creating multiple environments - at a minimum, a pre-production staging environment - to test our changes before pushing them live. The test environments contain mock data that emulates our real-world data as closely as possible. Using this approach, data engineers can safely expand the scope and availability of their changes in a safe and controlled manner: - While developing the changes, data engineer can work on their own dev boxes/environments and isolated source code branches, ensuring their changes don't interfere with those that others are making - When ready to test, they can ship their changes to a shared test environment that may also contain other pending changes - Only after changes are thoroughly tested and vetted will engineers then make them available for decision-makers and other data stakeholders Setting up isolated environments with enough realistic mock data to make testing worthwhile takes some upfront investment. However, this approach pays dividends in the long run. Defects are much more expensive to fix in production than they are to fix earlier in the development lifecycle. This means that, the earlier in the ADLC that you catch an error, the cheaper it is to fix. ### Size of change chosen by developers One hallmark of an immature analytics workflow is that teams will typically ship large changes - changes consisting of dozens of tables or updates to existing models. The problem is, the more you change, the more likely you are to inject errors. By scoping their changes to smaller units - such as adding only a few tables or changing a couple of fields - data engineers create change sets that are more tightly scoped and easier to test. Instead of pushing out a huge batch of risky changes simultaneously, they can use the quality gates built into the ADLC - code reviews, automated testing, pre-production deployments, etc. - to vet a smaller set of changes in isolation. ### Triggered by source code merge A key benefit of [source code control](https://docs.getdbt.com/docs/collaborate/git/version-control-basics) is that data engineers can work in their own branch or repo fork, isolating their changes from other changes in progress. If their changes aren't ready to promote your production, they can checkpoint their changes to their own branch, safe in the knowledge that they won't break anything in production or anything else that might be in flight. Once they’re ready to promote their code, they can create a pull request (PR) to merge their branch into the main production branch. After successfully passing code review, the release process commences, triggering the next stage - an automated deployment. ### Automated deployment An automatic deployment immediately vets a set of changes for release to production. This increases the velocity of analytics code deployments, reducing the time and cost involved in deploying. Automation ensures that the same steps are taken to verify and release every deployment. It removes human variability from the equation, reducing human error and improving the repeatability of the release process. Once a PR is approved, changes will run against a pre-production environment, such as staging. If any tests or other integrity checks fail, the change is pushed back to data engineers for resolution. After verifying the change, the release process runs any and all required tooling - such as data migration - to make the change live for all data consumers. Building an automated deployment process for analytics code changes can take time to develop and perfect. Platforms such as dbt Cloud [that support CI/CD processes out of the box](https://docs.getdbt.com/docs/deploy/continuous-deployment) can greatly reduce the overhead of building custom deployment pipelines for your data plane. ### Downtime-free These additional processes may sound to some like unnecessary overhead. However, the goal of this process is to eliminate the far costlier overhead of production downtime. In an immature analytics workflow, teams may frequently push out changes that haven't been properly vetted. This can result in pushing changes with basic errors - such as null values in required fields - that cause reports to break or data pipelines to cease functioning. To be sure, your company wastes time and dollars any time data engineers must scramble to put Humpty Dumpty back together again. You may lose just as much or more time and money, however, through delayed business decision-making - or, even worse, inaccurate business decisions based on bad data. A mature analytics workflow builds repeatability and quality control into the deployment process, incorporating lessons learned from past deployments to detect errors before they impact users. ### Automated rollbacks Pre-production testing is a great way to identify and resolve foreseeable data errors. Production systems are complex, though. While the goal is to eliminate all errors in production, it's impossible to predict every edge case. When this happens, it's critical to have an automated rollback mechanism for changes. Teams should identify a subset of their test bed as [smoke tests](https://www.techtarget.com/searchsoftwarequality/definition/smoke-testing) that get run on production data on a regular basis. If the system detects an error, this should trigger alerts and notifications. Data engineers can then revert their changes in production while they identify the root cause. ## What’s next Pushing your changes to production doesn't mean you're done with the ADLC. In the next installment of our series, we’ll move to the operations phase of the process, which is critical to ensuring your changes run smoothly. [Watch video](https://youtu.be/7Zo9YOsDbNM?si=czACekUaVn7Fdh3E) --- --- title: "The Analytics Development Lifecycle: Test" description: "High-quality data requires a test-driven culture. Here’s how it fits into a mature analytics workflow." url: "https://www.getdbt.com/blog/adlc-test" date: "2024-12-26" authors: ["Sara Gawlinski"] categories: ["Learn"] --- # The Analytics Development Lifecycle: Test Data quality errors are your company’s worst enemy. At best, they undermine people’s trust in the data that drives the business. At worst, they can provide false information that leads to erroneous - and costly - business decisions. The [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) aims to create a mature analytics workflow that produces high-quality, frequently updated data with every iteration. A key part of delivering that quality is not just creating tests, but fostering a test-driven culture as part of the ADLC. We’ll explore how testing fits into the ADLC, the types of tests you should be creating, and how to best manage your testing efforts for maximum positive impact. ## Test in the ADLC The ADLC is a variation of the [Software Development Lifecycle (SDLC)](https://aws.amazon.com/what-is/sdlc/) that focuses on shipping new or revised data products. Like the SDLC, it breaks down artificial barriers between the different personas that deal with data, treating it as a single, unified process. ![ADLC loop](https://cdn.sanity.io/images/wl0ndo6t/main/948eb5cda47eacf1bf78a2268c0666be43706ce5-4581x2126.png) In the ADLC, the different personas that handle data - the engineer, the analyst, and the decision-maker—work together to plan, develop, test, deploy, monitor, and use new data products. The process focuses on creating small, well-defined changes and shipping frequently. We’ve covered how the [Plan](https://www.getdbt.com/blog/adlc-plan) and [Develop](https://www.getdbt.com/blog/adlc-develop) phases of the ADLC work. These phases help ensure quality by ensuring that: - The work done accurately captures business requirements (Plan); and - All data changes are captured in code, and that code is clean, readable, and reusable (Develop) The Test phase creates assets that validate that your assumptions about your data and analytics code are correct before pushing a change to production. By testing your data, you can identify issues early in the development lifecycle, preventing expensive rework and downtime down the road. A good Test phase involves: - Writing tests for every data asset you own - Running tests before they’re merged into production - Continuously testing production data to detect anomalies ## Types of tests in the ADLC Let’s first look at the different types of data tests you’ll want to focus on writing: - Unit tests - Data tests - Integration tests ### Unit tests Unit tests validate small functional portions of your data models and transformations to ensure correctness. They validate your logic on a small set of static inputs before running it on actual data. In data pipelines, this means validating your SQL modeling logic’s correctness. [dbt Cloud supports developing unit tests](https://docs.getdbt.com/docs/build/unit-tests) alongside your [SQL models](https://docs.getdbt.com/docs/build/models) and running them on demand. You don’t need to create a test for every single transformation. However, you should always aim to create unit tests when you have: - SQL with custom logic - Reported defects (to verify the fix and prevent regressions) - Edge cases - High criticality models, such as organization data sets where a defect could have wide-scale negative impact For example, this test in dbt verifies that a routine for verifying email addresses captures known edge cases, such as malformed addresses and invalid domain names: unit_tests: - name: test_is_valid_email_address description: "Check my is_valid_email_address logic captures all known edge cases - emails without ., emails without @, and emails from invalid domains." model: dim_customers given: - input: ref('stg_customers') rows: - {email: cool@example.com, email_top_level_domain: example.com} - {email: cool@unknown.com, email_top_level_domain: unknown.com} - {email: badgmail.com, email_top_level_domain: gmail.com} - {email: missingdot@gmailcom, email_top_level_domain: gmail.com} - input: ref('top_level_email_domains') rows: - {tld: example.com} - {tld: gmail.com} expect: rows: - {email: cool@example.com, is_valid_email_address: true} - {email: cool@unknown.com, is_valid_email_address: false} - {email: badgmail.com, is_valid_email_address: false} - {email: missingdot@gmailcom, is_valid_email_address: false} ### Data tests [Data tests](https://www.getdbt.com/analytics-engineering/transformation/data-testing) validate that data transformations are running correctly against the actual data. They verify that: - The data is current - The model is sound - The transformed data is accurate Data tests usually start by testing basic assumptions about unique and non-null fields (e.g., primary keys), accepted values, and relationships between data. Once you’ve nailed those aspects, you can move on to more proactive tests that focus on verifying freshness and looking for domain-specific problems. For example, if a customer can only have one active subscription to a service, verifying that there aren’t records that violate this constraint. As with unit tests, you can specify these tests using dbt Cloud and run them with the [dbt test command](https://docs.getdbt.com/reference/commands/test). ### Integration tests Whereas unit tests test one small unit of functionality, integration tests operate against the entire application or project. They ensure your solution works end to end and not merely in isolation. In the software world, this might involve calling a REST API and ensuring that the REST API endpoint, associated authentication procedures, underlying data stores, connected APIs, etc. all work. In data, you’ll use it most often to test [packages](https://docs.getdbt.com/blog/unit-testing-dbt-packages), reusable units of analytics code that multiple projects leverage. In dbt, you can keep unit, data, and integration tests separate by placing them in separate subdirectories. That enables running them at different points of the ADLC. ## When to run tests Anyone who’s creating or updating analytics code - i.e., who’s wearing [the engineer hat](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle#stakeholders-of-the-adlc) - is responsible for creating or updating the associated tests. The engineer should make sure to run unit and data tests on their local machine prior to check-in. As discussed in the [Develop phase](https://www.getdbt.com/blog/adlc-develop), engineers should work in their own source control branches. When ready to push to production, they should cut a [Pull Request (PR)](https://www.getdbt.com/blog/how-to-review-an-analytics-pull-request). Another engineer should review their changes - yet another quality control measure - before approving the merge. The PR should also automatically trigger a run of any associated tests for the change against non-production data in an isolated environment. If the tests don’t fail, the PR should prevent the merge to production until the issue’s resolved. [dbt Cloud supports running tests automatically](https://docs.getdbt.com/docs/deploy/continuous-integration) against a staging schema when it detects a PR has been opened or submitted in your Git provider. You can see this run in either the dbt Cloud dashboard or directly on the PR page of your Git provider, along with any errors that resulted. ![Run](https://cdn.sanity.io/images/wl0ndo6t/main/b2f4aa07d28da40af2c2e8b1c1164c3e6a8e85f0-1562x810.png) ## Tips for managing testing Here are a few more tips to get the most out of data testing: **Developing a culture of testing**. It’s easy to throw testing by the wayside because you’re busy and you just wanna get something out the door. As our CEO Tristan Handy has written, “The desire to skip writing good tests and move on to the next task is always present and must be balanced via accountability mechanisms like code reviews, linting, and test coverage metrics.” Get everyone on board with testing as a matter of habit. Set a bar where testing is required for a change and enforce it during PR reviews so that team members hold each other accountable. **Keep the scope of work small**. This is a central tenant of the ADLC that’s critical in testing. The larger a change, the harder it is to verify its functional correctness. Conduct training on [properly scoping PRs](https://www.getdbt.com/blog/how-to-review-an-analytics-pull-request) so that all submitted changes contain enough new logic to be useful - but not so much that you can’t verify its accuracy. **Determine your level of test coverage**. Decide how much of your analytics code should require testing. In the software field, most teams aim for [around 70-80% test coverage](https://learn.microsoft.com/en-us/answers/questions/778016/test-coverage-definition-unit-testing). You may need less depending on the complexity of your code. Once you have a metric for test coverage, monitor it over time to ensure you’re hitting your goal. [The dbt Cloud dashboard Recommendations page](https://docs.getdbt.com/docs/collaborate/project-recommendations) shows you your overall test coverage as a percentage of how many of your models have defined tests. ![Recommendation page](https://cdn.sanity.io/images/wl0ndo6t/main/809c3b9b25db0eb471297fa19555e487b5ac981b-1562x844.png) **Fix or retire “flaky tests.” **A [flaky test](https://www.datadoghq.com/knowledge-center/flaky-tests/) is one that fails intermittently, usually due to some network or environmental condition, or just poorly written logic. Ignoring flaky tests is dangerous because it can foster “alert fatigue,” leading people to tune out and ignore real errors. Either identify the cause of a flaky test and fix it or remove it from your test suite altogether. ## Conclusion The ADLC creates high-quality data sets by making small changes over a series of rapid iterations. Testing verifies quality by making assertions about the state of your data and analytics code. Since its inception, dbt has supported creating a test-driven culture by building support for testing directly into both dbt models and [dbt Cloud](https://www.getdbt.com/product/dbt-cloud). With dbt Cloud as your data control plane, your data teams have a standardized and cost-efficient way to build, test, deploy, and discover analytics code. In our next installment of this series, we’ll look at how you can leverage dbt Cloud to implement a CI/CD-style approach to [deploying analytics code](https://www.getdbt.com/blog/adlc-deploy) safely to production. [Watch video](https://youtu.be/TWUZxyF-u-Q?si=BGvrGwjkz6WdWdQD) --- --- title: "The Analytics Development Lifecycle: Develop" description: "Why the Develop phase of the Analytics Development Lifecycle is so critical to quality releases." url: "https://www.getdbt.com/blog/adlc-develop" date: "2024-12-23" authors: ["Sara Gawlinski"] categories: ["Learn"] --- # The Analytics Development Lifecycle: Develop In the past, data teams and stakeholders didn’t pay much attention to the quality or reusability of analytics code. These days, more teams are realizing how properly written, reviewed, and standardized analytics code contributes to more frequent and higher-quality releases. Patterned after the Software Development Lifecycle (SDLC), the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) provides a model for implementing changes to your data system. In our previous installments in this series, we looked at [how the ADLC impacts planning](https://www.getdbt.com/blog/adlc-plan). In this article, we’ll see how you can change the way you author analytics code to improve reliability and increase data product velocity. ## Stages of the Develop phase The ADLC provides a framework for a mature analytics workflow. It aims to solve issues with low data code velocity, inaccurate results, and impaired trust that continue to plague too many analytics projects. Data stakeholders use the Analytics Development Lifecycle to develop and ship small changes to analytics code in a short timeframe, repeating the full process for every significant change to the system. When done well, the ADLC yields better collaboration, scale, velocity, correctness, and governance. The SDLC is built around a DevOps approach that treats developing and managing software applications as part of the same unified process. Similarly, the ADLC is built around a DataOps model that coordinates the efforts of analytics code developers with the data engineers, [analytics engineers](https://www.getdbt.com/blog/analytics-engineer-vs-data-analyst), and business analysts who manage and use analytics data. ![ADLC loop](https://cdn.sanity.io/images/wl0ndo6t/main/948eb5cda47eacf1bf78a2268c0666be43706ce5-4581x2126.png) The Develop phase is where engineers - dedicated data engineers, analytics engineers, or even business stakeholders with technical chops - turn analytics use cases into deployable data products. It’s a critical phase, as it has an outsized impact on the overall quality of the final solution. An effective Develop phase in the ADLC consists of the following components: - Code first - Adhere to a style guide - Prioritize functionality over performance - Invest in code quality - Use code reviews - Use standards to avoid lock-in Let’s look at each phase in detail. ### Code first In the past, analytics data changes existed in a mix of Excel macros, SQL scripts, programmatic code, stored procedures, and visual tools. Most of these existed only on engineer’s laptops and were run by hand, with engineers performing tweaks and corrections as needed. These processes weren’t repeatable or discoverable. They often didn’t result in high-quality or fast releases, as the knowledge needed to run them resided in someone’s head. In the ADLC, all business logic impacting data should be captured in code. All code should be: - Editable by multiple people with different tools - programmatic text editors, Integrated Developer Environments (IDEs), etc. - Checked into [version control systems](https://docs.getdbt.com/docs/collaborate/git/version-control-basics) to enable collaboration and prevent conflicts - Broken down into composable units for reuse - Deployable via an automated process, a.k.a. [CI/CD](https://www.redhat.com/en/topics/devops/what-is-ci-cd) This code first approach ensures that all analytics code changes are: - **Discoverable**: Others can find and modify the code as needed - **Traceable**: Data stakeholders can see when a change was made and who made it - and revert it if necessary - **Reusable**: Other data stakeholders can find and reuse general-purpose solutions in their projects - **Repeatable**: The same process can be used to develop and ship any analytics code changes - **Tool agnostic**: Analytics code developers can use any development tools that fit their workflow Developing a code first strategy can take time and effort to create, configure, and deploy across an organization. Using a platform like dbt Cloud, which is built from the ground up on a code first philosophy, provides the necessary tooling and infrastructure out-of-the-box, shortening the time required to transition to a mature approach to developing analytics code. ### Adhere to a style guide Putting all code in common version control repositories makes it easier to find and maintain. However, code can be hard to read and maintain if everyone’s using different coding conventions. A style guide provides consistency in code formatting and conventions (names of variables, use of whitespace, etc.) across everyone who touches analytics code. That makes code easier to read - which makes it easier for those who didn’t write it to pick up and maintain. You can also enforce standardization automatically via mechanisms such as linting code - e.g., running [sql-lint](https://github.com/joereynolds/sql-lint) on all SQL code. Tools like the [dbt Cloud IDE](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/develop-in-the-cloud) can assist standardization with features such as syntax highlighting for SQL, code formatting, and linting. ### Prioritize functionality over performance [Premature optimization](https://stackify.com/premature-optimization-evil/), said Sir Tony Hoare, is the root of all programming evil. In the initial stages of coding, engineers should focus on implementing their business use case versus fine-tuning for optimal performance. This doesn’t mean engineers shouldn’t consider performance _at all_. The design of a data solution will have the largest impact on how quickly it runs. It means, instead, **not chasing small efficiencies at the project’s outset**. That time’s better spent on iterating with stakeholders to ensure the solution fits their requirements. In Hoare’s words, “We should forget about small efficiencies, say about 97% of the time.” Encourage engineers to focus on the requirements first. Then, at the tail end of the project, the y can implement the final tweaks needed to get the most out of the system at scale. ### Invest in code quality Quality analytics code requires a process that builds quality into the entire development lifecycle. Part of that is keeping code clean and maintainable. A few key practices here include: - **Write DRY code**. [DRY, or Don’t Repeat Yourself](https://www.getdbt.com/blog/guide-to-dry), is the principle of factoring out common code to reusable modules. This prevents you from duplicating code unnecessarily, which can inject defects. It also enables teams to work more quickly, as they can use tried-and-tested procedures for common operations rather than develop new code from scratch. - **Define common metrics**. In analytics, creating common metrics is an indirect form of reusability where you centrally define critical business metrics, such as revenue, that others can leverage in their own solutions. Tools like the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) simplify defining and deploying centralized metrics. - **Write documentation and in-line comments**. Use a tool like dbt Cloud that supports [writing documentation in code](https://docs.getdbt.com/docs/build/documentation) to document method parameters, field definitions, assumptions, and other important aspects of your analytics implementation. Any authoring platform you adopt for implementing the ADLC should have a mechanism for defining and sharing reusable code. For example, dbt Cloud supports defining [packages](https://docs.getdbt.com/docs/build/packages) - standalone dbt projects with models, macros, and dependencies that other teams can reference from their own dbt projects. ### Perform code reviews Code reviews ensure that every proposed analytics code change is seen by a second set of eyes. Code reviews have been shown to provide multiple benefits: - Reduces defects in shipped code. In one study cited in the classic software engineering book _Code Complete_, introducing code reviews reduced errors in one-line maintenance changes from 55 percent to 2 percent. Others have seen reductions in errors of up to 80 percent. - Provides accountability for enforcing practices such as style guide compliance, testing, documentation, and writing DRY code. - Increases shared team knowledge of data product solutions and their underlying code. Creating code reviews is easy to set up once you have a version control system in place. Engineers create a branch in the source control system in which they make their changes. When they’re getting ready to ship, they create a [pull request](https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/proposing-changes-to-your-work-with-pull-requests/about-pull-requests) to merge their changes with the main branch. Potential reviewers are notified that a change requiring review is pending. [Some additional best practices for code review](https://www.getdbt.com/blog/how-to-review-an-analytics-pull-request) include: - Size your pull requests to represent a single unit of work. - Set up automated testing in your version control system and require that tests pass before a pull request can be approved. - Use a common pull request template across all teams and engineers ([we shared ours here](https://docs.getdbt.com/blog/analytics-pull-request-template)). A code review is a critical quality gate in a CI/CD deployment system. It ensures that a change has been fully vetted and tested before it’s made available to data stakeholders. **** ## Conclusion Creating data that stakeholders can trust requires a process that builds in quality at every stage. By standardizing the way your team develops analytics code, you can ship smaller, higher-quality changes more quickly than you could with a manual, ad hoc analytics process. Good code, however, isn’t the only thing you need. In the next installment of our series, we’ll look at how to use testing in the ADLC to validate quality prior to shipping your changes to stakeholders. [Watch video](https://youtu.be/iw1OW1W_BqQ?si=DK5GiPCSPCvwUlVZ) --- --- title: "The intersection of UI, exploratory data analysis, and SQL" description: "Hamilton Ulmer from MotherDuck discusses the technologies driving data visualization today." url: "https://www.getdbt.com/blog/intersection-ui-exploratory-data-analysis-sql" date: "2024-12-22" authors: ["Daniel Poppy"] categories: ["Insights"] --- # The intersection of UI, exploratory data analysis, and SQL _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/the-intersection-of-ui-exploratory). _ Hamilton Ulmer is working at the intersection of UI, exploratory data analysis, and SQL at MotherDuck, and he's built a long career in EDA. Hamilton and Tristan dive deep into the history of exploratory data analysis. Even if you spend most of your time below the frontend layer of the analytics stack, it’s important to understand trends in both the practice of data visualization and the technologies that underlie that practice. All of it deeply shapes the space that we operate in. **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways from this episode ### If you're like other people who started their data science careers in the 2010s, you probably ended up doing all different parts of the analysis pipeline. Over time, it seems like you and your career have become more focused on the data visualization part. Is that fair to say? When I joined Mozilla, we had what was then considered pretty big data. It was also very complicated, messy stuff that the browser was generating. We were using a lot of that telemetry to basically calculate the numbers for the business as well. Data visualization for me has always been a means to an end, and that means understanding the data that powers the business and the product. And so I think those interests are intrinsically connected. Many data visualization and exploratory data analysis projects focus on the end of the analysis process, the presentation layer. This includes things that you might put in a slide deck. But they're having to work with extremely messy, highly nested data generated by the web browser, which possibly runs in a semi-degraded state. Not all the data that it sends is good data. That first mile is‌ way more interesting to me. And it's maybe the genesis of my interest in data tools in general is that first mile problem, not the last mile problem. This is a place where exploratory data analysis is ‌especially valuable. But I think the way that people think about EDA is more in the middle or toward the end. That first part is‌ really critical. ### Exploratory data analysis is used to answer the question, can I trust this data? What do I need to do to it to get it to a state where I can trust it so I know what I can expect from it. Absolutely. I think this is the fundamental trade-off in a lot of analytics tools. The people that make those tools or the libraries that we use are often made by technical people, oftentimes people with a research background that have to do data cleaning. But the financial value comes from the end-user experience of dashboards. And so there's always been this trade-off‌ with EDA tools between these two things. That’s the case of Polaris, which is a research project out of Stanford in the early 2000s. These researchers wanted to‌ figure out how to make EDA interactive and exploratory. This was during a time when computers were just starting to get better at this kind of work and data was being generated. And those researchers at the Polaris paper were groundbreaking for analytics. Those researchers went on to found Tableau. And the killer use case, wasn't the first mile. It was the last‌, because the economic buyer of the tool cared a lot about understanding what was going on with their business. So a lot of the focus for EDA tools has been on BI rather than the thing that‌‌ really vexes data practitioners—cleaning the data. Everyone likes the joke that 80 percent of the job is cleaning up the data. So if you look at a model of data work, it's largely about trying to correct problems with data collection as early as possible to figure out what you can possibly say about the business down the line. That's‌ a high-value thing, but it's hard to sell that‌ to people. And that's why I think BI tools focus on the presentation layer. ### Can you talk a little bit about how over the last 20 or so years the data visualization industry has evolved? Are we operating at a higher level of abstraction than we used to be? Imagine it's the 1970s and you're a statistician doing research. You find that you can put your tables of data into the computer to combine and show them quickly. Before, you had to do it by hand. If you've ever read anything by Edward Tufte about historical data viz, you could tell someone that you had a pencil and paper and drew, had to figure out where to put the points to show the aggregation, right? Really time-intensive. John Tukey was this really famous statistician. He's‌ the person doing exploratory data analysis. He said something that isn't controversial, which is that you should look at the data before doing a statistical analysis. This wasn't easy to do without computers. It was part of a movement to bring computation to statistics that became‌ the whole point of the field after a certain point. And then in the early 90s, spreadsheet software—Excel—became the most important data tool ever created. They began adding charts to their spreadsheets, and that was really a great early form of data visualization, probably the most popular form, right? You have this other cross-current here where the browser became the medium for interactive data visualization, and not some desktop app that a bunch of people have to write C code. That‌ was probably the biggest expansion of the labor market in data visualization. And that's where you see D3 becoming one of the most important entry points for those people to become essentially front-end engineers. D3’s premise was that building high-level primitives for data visualization wasn't possible without getting a mid-level connection to the browser APIs. D3 is still widely used. It's not used in the same way as it was in the past, but it's still very widely used. I use it every day for all of the data visualization tooling I build, just because it has so many helpful things that I don't want to build myself at this point. So the web became important for the medium of data visualization. And really for analytics tools as well. Most of the BI tools moved to being web-based in some way. And so then you had this other cross-current in the last 20 years, which is the tech boom, internet companies, things like that. You had a huge influx of technical PhDs in the industry. You had all of these people bringing their analysis tools that they used in their research. The scientific Python computing stack was something that you might've‌ toyed with in grad school and used for your research. And now you're bringing it to your job because it's a tool, you know, R is another example of this. This is sort of where I enter the story‌ is as a statistician with R. ### My guess is that probably everybody listening here is familiar with the name DuckDB, but you should probably do a little bit of an overview. DuckDB is an in-process database. The closest analogy would be something like SQLite, which I don't think is a fair comparison because DuckDB does so much more. SQLite is this tiny transactional database that's‌ the most important piece of software ever made. I don't think that's a stretch to say that our lives are powered by thousands of SQLite databases on all of our devices. Our browsers all have individual SQLite databases powering them. It's an incredible thing when you can just have your database as a file somewhere and then whatever process can just query that directly rather than having a database run on its own independent server. And so DuckDB is‌ like that but for analytical queries, not transactional ones, the kinds of queries that your audience is quite familiar with. The project started‌ in the late 2010s. Hannes Muehleisen and Mark Roosevelt, who are two researchers at a research institution in Amsterdam called CWI, which previously was probably best known for being the place where Python was invented. The influence for DuckDB was the workloads that PhD data scientists were bringing to industry. Crunching down CSV files and Parquet files has always been a bit of a challenge. Mark and Hannes realized they could build a database to make data analysis easier. And so that was sort of the genesis of DuckDB. But as they began working on it and applying some of the most cutting-edge ideas in analytical databases to the project, they began to realize it could do more and more stuff. And that's ‌why I joined MotherDuck because I'm part of this movement of people that care about data visualization that have discovered DuckDB, and really can't look back. That divide we were talking about, the front end being all JavaScript and the back end being who knows what. If you can make that back-end DuckDB, and you can do incredible things with it, you can actually determine the future. The queries you need to run on the front end effortlessly update your UIs. And so it reduces the latency of interactions for really complex things. ### We talked about the difference between the BI world and scientific computing world before. One of the interesting differences there is that scientific computing does not typically speak SQL. Is DuckDB capable of doing some of these scientific computing functions or does it not need to? What's changed? In 2015 there were people that would scoff at the idea of writing SQL, maybe they adopted BigQuery and discovered writing SQL actually wasn't a big deal. There was a period, especially in the 2010s where people weren't sure if SQL was going to survive. The environment was different then, but one thing that ended up happening was more stuff moved to SQL rather than less stuff moving to SQL. I was at Mozilla when we bought BigQuery. It was like a breath of fresh air, being able to write SQL, a query that I can analytically verify myself and just have it do the thing was really special. Industry has moved more towards SQL. That said, DataFrames are amaz ergonomically quite amazing, especially our ecosystem with dplyr for data transformation is like a really elegant, nice way to work with data. Databases like DuckDB can do things in dplyr and it will write the DuckDB query for you which is really nice. ### I really desperately love dplyr. It is maybe my favorite of all our packages. I feel like sometimes there's this religious war between SQL people and not SQL people. What's interesting about this too is a lot of these analytical SQL dialects are starting to also address some of the same types of problems. I think I mentioned BigQuery before, like writing BigQuery SQL and working with arrays in BigQuery and things like that became easier to do some of those hard data manipulation things in SQL. And looking at DuckDB, something that Mark and Hannes care a lot about is just the ergonomics of writing SQL. So they've made their own extensions to SQL to make it easier to do the sorts of things that you would see in dplyr. SQL is complicated because it's kind of a ancient programming language that has stood the test of time. It's not Latin. It's something else. The spec for SQL is like thousands of pages and there's not a single database on the planet that actually implements all of the spec. It's one of these bizarre situations we're in where the dialects differ in some critical ways. If you know one dialect of German, you could speak the other one kind of with other people. It's similar with SQL. This lack of control over the language itself, does two things. One, it frustrates everyone because it's much easier to go to R or Python where everything is well defined and the grammar is small. There's not thousands and thousands of keywords to implement in these languages. But also, it's an area of innovation and opportunity for other database engines, and I think the DuckDB creators have seen that. And so things like list comprehensions, which are so useful in Python, you can actually do in DuckDB SQL. Function chaining, you can also do that in DuckDB Python and or DuckDB SQL It's an interesting area of innovation, and I think it also upsets a number of people as well. There are people who think SQL needs to be kind of this boring thing that everyone knows and stop innovating. I'm much more of an experimentalist. Given that there is no actual SQL standard beyond the bare minimum that everyone implements. Why not innovate? ### Can you talk a little bit about how, how is it possible that DuckDB runs locally. Why can't I run Snowflake locally? To understand this, you have to understand a bit of the historical trends in the tech industry around the concept of big data. So 15 years ago, tech companies began generating large amounts of data as they had users use their applications. Facebook's a great example. The amount of data that needed to be processed in order for you to understand it was much larger than what one computer on its own could even do. You could go buy a desktop tower, put it on your desk, get all the data on it and attempt to do something with it, but it was going to be extremely slow, right? The disk space wasn't big enough. You didn't have enough memory to do interesting things. Computers weren't good enough. Computers have recently become good enough, I think. And that's the TLDR. But because they weren't good enough at that period, we began to see great research out of Google around MapReduce and about splitting up computation across a bunch of rented out computers in the cloud, a bunch of cheap, low powered machines, and this fanning out the computation to a bunch of machines that weren't on your computer was the way that you could actually do anything meaningful. And that idea took hold in industry. As more companies became data-driven, this was the way to do it. And this is actually part of the core thesis of MotherDuck. It's 2024 now, and I'm using a modern MacBook Pro. And when I switched from SQLite to DuckDB, querying a really large, 10 gigabyte dataset, it was instantaneous. The first time that I ran a query in DuckDB, I thought, is it broken? I don't think this is possible. MotherDuck’s core thesis is that our computers have actually gotten fast enough to handle those workloads that back in 2011, we could not have done on a desktop tower.And that's really magical. Computers have gotten good enough to do what was previously considered very big data or big data. The definition of big data has shifted over time to be increasingly larger. And so the things that were big data when Snowflake and BigQuery were created might not necessarily be big data today. They might be something that you could process on one really beefy computer that you rent from AWS. And so I think that's why DuckDB is becoming really popular, by the way, is because computers have caught up. ### There is an architecture, the MPP, Massively Parallel Processing architecture, that became initially popular in the 2000s. You could reasonably say that Snowflake and BigQuery are also versions of an MPP architecture, certainly evolved from earlier versions, but there's probably some overhead involved in. But that all of a sudden can be faster and it can also be distributed to your local machine. Am I getting that right? I think that's the most concise story for DuckDB and why it's so successful. The fact that it can run in a process anywhere means it's running everywhere. Companies are adopting DuckDB as point solutions all over the place. And you don't hear about it all the time, but it is absolutely happening. So it may not be the total solution for their data today, but it's oftentimes the best solution for individual pieces. I think I saw somebody joke that DuckDB may single-handedly prevent global warming caused by the JVM. --- --- title: "The Analytics Development Lifecycle: Plan" description: "Here’s how to perform planning as part of the ADLC." url: "https://www.getdbt.com/blog/adlc-plan" date: "2024-12-21" authors: ["Sara Gawlinski"] categories: ["Learn"] --- # The Analytics Development Lifecycle: Plan Are all your data assets versioned, tested, and easy to support and maintain? For most companies that work with data, the answer is “no.” We have a plethora of tools for managing data. What’s needed is an analytics practice—not just a set of tools, but a workflow that enables delivering well-governed data projects that accelerate data delivery, tune data quality, and optimize compute costs. dbt Labs advocates an approach we call the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). In this article, we’ll look at the first stage of the ADLC, Plan, examining what exactly it entails in an analytics workflow and how doing it well helps you deliver better data, faster. ## The key to planning in the ADLC The ADLC works similarly to the [DevOps](https://www.atlassian.com/devops) process in software engineering. It’s a rapid and highly iterative process you repeat for each change you make to your data system. In other words, planning in the ADLC isn’t a “once and done” activity. It’s also not an endless process that mires you in requirements hell, preventing progress. Rather, it’s a succinct phase you use to ensure what you’re delivering conforms to what your data stakeholders need. This prevents expensive rework down the line. The length of the planning phase will be variable, depending on the size of the change. Every planning phase, no matter its length, should deliver these key benefits: - Gets everyone on the same page. In other words, identifying all major stakeholders—engineers, analysts, business users, decision-makers—up front and ensuring everyone’s understanding of the business case, key metrics for the project, etc. are in alignment - Identifies tooling needs early - Estimates resources more accurately - Ensures security by making discussions of data visibility, access rights, compliance, etc. part of the process from the beginning, instead of tacking them on later as an afterthought Done well, ADLC planning ensures a high-quality and secure data product delivered with minimal rework. **** ## Stages of the planning phase There is no one-size-fits-all approach to the Plan phase. Every team needs to implement a version of the process that meets their business needs and meshes with their organizational culture. Similarly, there's no single planning process that works out of the box for everyone. However, the following stages will provide a good jumping-off point for most teams from which to build their own unique process: - Create and validate the business case - Create your implementation plan - Get stakeholder feedback - Create a test plan - Anticipate downstream impacts - Plan for maintenance - Determine access levels - Implement larger changes in small pieces Let’s look at each one of these stages in detail. ### Create and validate the business case In the past, many data changes were driven, not by the business, but by engineering. This means that most new data deliverables were measured mostly in technical terms (e.g, data throughput). It’s hard to excite business stakeholders with technical metrics. This approach also divorces data projects from the company’s larger business objectives. That risks shipping data projects [that no one ever uses](https://www.getdbt.com/blog/analytics-engineering-for-everyone). To avoid this, all data changes should be based around a business case—i.e., what the change does, who it’s for, and the quantifiable benefit you expect to see. They should further tie back to business Key Performance Indicators (KPIs) or Objectives and Key Results (OKRs). For example, instead of focusing on improving database performance, focus on how that helped the team reduce average resolution time by 25%, improving customer satisfaction scores. Not every minor change needs to go through this process. Your team should define a threshold of work above which a change needs a solid business case before proceeding. ### Create your implementation plan Once everyone’s agreed on the business plan and the need for the work, identify how you’ll put your proposal into action. This includes where you obtain your data, the inputs and outputs of the data product, and what code and architectural assets you’ll need to accomplish it. A key part of implementation planning is identifying what you need to build versus what you can reuse. Wherever possible, aim to adhere to the DRY (Don’t Repeat Yourself) principle. Find a way to leverage existing code and also make your work available to others who might need it—e.g., [by using packages to promote reuse](https://docs.getdbt.com/docs/build/packages). ### Get stakeholder feedback Once you have an implementation plan, run it by your stakeholders for final approval. Use whatever format—Slack, email, a recorded meeting, a ticketing system, etc. —to capture approval and address outstanding issues. To prevent the planning phase from drawing out, set deadlines for accepting feedback. Once all feedback is in, make any changes and proceed to another round of sign-offs, if necessary. ### Create a test plan [Testing](https://docs.getdbt.com/docs/build/data-tests) is a critical part of any data change. It ensures you get the outputs you expect from your data transformations relative to the process’s inputs. Good data tests should ensure your data transformation code works for both normally accepted inputs and fails gracefully on edge cases or unexpected inputs (e.g., null values for required fields, malformed text strings). Good testing also includes running tests against pre-production (historical or mock) data so you can validate its functionality before making it live for users. Identify any data sets you need for testing environments as part of your planning process. ### Anticipate downstream impacts A common problem in the data world is breaking changes that impact data consumers you may never even know existed. One day, your data engineering team changes the format of a text field or eliminates a column from a table. The next day, a critical forecasting report on which the sales team depends fails to refresh hours before an all-hands meeting. If you’re making changes to existing models, perform an impact analysis of your changes before you exit the planning phase. This involves using [data lineage](https://www.getdbt.com/blog/guide-to-data-lineage) to discover what other downstream data products, reports, and applications depend on your work. Once you’ve identified your consumers, notify them of the intended change so you can work together on a migration plan. This may involve asking them to update their reports after release or [versioning your data models](https://docs.getdbt.com/docs/collaborate/govern/model-versions) to give consumers time to transition. ### Plan for maintenance In the past, most data transformations involved exporting data to a CSV file and [wrangling it in a spreadsheet](https://www.getdbt.com/blog/data-transformation-tool-choosing). Much of this work was disposable—no one cared if a spreadsheet formula broke a month or two after delivery. A high-quality, reliable data system is different. Any change you make to data is a commitment to future users. Before you exit planning, identify what tests, metrics, and alerts you’ll use to monitor your code’s behavior in production. Identify who will own the data transformation code going forward. If it isn’t you, work with the eventual maintainers to ensure they understand the implementation. ### Determine access levels According to IBM, [the average data breach costs a company USD $4.45M](https://www.ibm.com/reports/data-breach). Such threats [can come from inside](https://www.cisa.gov/topics/physical-security/insider-threat-mitigation/defining-insider-threats) the company just as easily as outside. You need to incorporate security with every release—even if it’s “just” an internal project. Think early about your data and what you need to do to keep it safe. Which groups or individuals need access? Which should be restricted? (E.g., should vendors have access?) Are you handling sensitive information—customer’s Personally Identifiable Information (PII), company secrets—that requires additional scrutiny and governance? After identifying your data’s security needs, decide how you’ll enforce them. Tools you might use here include: - [Using project permissions and role-based access control (RBAC)](https://docs.getdbt.com/docs/cloud/manage-access/enterprise-permissions) to grant or deny access automatically - Establishing a process to administrate access requests for sensitive data - Applying data classifications to your tables and columns so that you can identify and remove customer data as needed per regulations such as [GDPR](https://gdpr-info.eu/) ### Implement larger changes in small pieces Finally, if you find your change becoming too large and unwieldy, consider breaking it up into multiple releases. Don’t try to boil the ocean with large changes. Instead, leverage the iterative nature of the ADLC to break it down into smaller, well-tested components. For example, you may plan a complex change with six different models. Instead of releasing all six simultaneously, develop one or two (along with their tests) and put them through a full develop/test/debug/release cycle. Then, move on to the next model, repeating until you’ve implemented the full business case. Tackling large changes in multiple releases keeps you from getting bogged down in technical issues. It also lets you get faster and more frequent feedback from your stakeholders. ## Conclusion A good workflow is just one part of an analytics practice—you also need powerful tools to enable it. dbt Cloud is your data control plane, delivering multiple tools that make the ADLC planning process easy to implement: - Built-in support for defining data transformation [models](https://docs.getdbt.com/docs/build/models), [tests](https://docs.getdbt.com/docs/build/data-tests), and shared metrics via the [dbt Semantic Layer](https://docs.getdbt.com/docs/build/build-metrics-intro) - [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) for finding and leveraging existing data and data transformation code - [Data lineage](https://www.getdbt.com/blog/guide-to-data-lineage) to see all downstream dependencies, simplifying impact analysis Learn more about how dbt Cloud can transform how you do data—[ask us for a demo today](https://www.getdbt.com/signup). [Watch video](https://youtu.be/_Uu6atDQgaY?si=2nYSck8gTqM9vUVv) --- --- title: "How Moderna uses dbt Mesh to foster collaboration and streamline data engineering" description: "Discover how Moderna uses dbt Mesh to connect data teams, enhance governance, and deliver life-saving medicines with precision." url: "https://www.getdbt.com/blog/moderna-dbt-mesh" date: "2024-12-21" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # How Moderna uses dbt Mesh to foster collaboration and streamline data engineering Delivering life-saving medicines worldwide requires more than scientific precision—it demands sophisticated data infrastructure. For Moderna, this meant finding new ways to connect data across multiple platforms and teams. Their implementation of** **dbt Cloud with** **[dbt Mesh](https://www.getdbt.com/product/dbt-mesh) enabled faster data velocity without data quality compromises. ## Managing data across a global pharmaceutical supply chain ### From pandemic response to data complexity When Moderna became a household name during the pandemic, their data needs exploded. With **over 5,000 employees across 17 countries** and **millions of mRNA vaccine doses** being manufactured and distributed globally, they faced unprecedented scaling challenges. The stakes were uniquely high. Vaccines require precise temperature control and timing, with **costly consequences for any miscalculation**. A single delayed shipment or temperature deviation could impact thousands of doses—and the patients waiting for them. ### Three barriers to efficient data flow Moderna’s data team identified three challenges affecting its data operations: #### Data accessibility and availability Moderna needed to break down data silos and enable teams across manufacturing, supply chain, and distribution to self-serve on accurate, real-time data to make decisions. #### Data governance and compliance As a licensed pharmaceutical company, Moderna needed to ensure that every data point was traceable, secure, and compliant with regulations across multiple jurisdictions. Data transparency, cleansing, and standardization were also essential to produce data everyone could trust. #### Scalable infrastructure Beyond immediate needs, Moderna needed a** **future-proof** **data platform that could scale to support increasingly sophisticated analytics, data science initiatives, and emerging AI applications while still managing costs. ## Tackling data challenges with five data mesh principles Moderna embraced data mesh principles to solve these challenges and implemented them with the help of dbt Cloud. ### Data domain ownership Moving away from a centralized data team, Moderna organized its data teams around business domains. Each domain had their own dbt project, giving them **autonomy** to support their use cases. ### Data as a product Moderna used dbt Mesh to join data from disparate data warehouses and create** purpose-built data sets**. dbt Cloud features like data lineage and out-of-box documentation created a common foundation among teams. ### Self-service data platform With accurate and readily available data, domain experts could leverage their own tools to deliver valuable insights autonomously, enabling **faster development cycles**. ### Federated data governance While teams gained autonomy, **automated checks and guardrails were set in dbt Cloud**—such as enforced metadata standards—to ensure consistent data quality. ### Data and platform discoverability The Mesh architecture made it easy for **teams to discover and use data** across platforms: from data lakes to Redshift. **** ## Validating the new data infrastructure in a supply chain project ### Preventing vaccine shortages and waste A critical project put this new architecture to the test. The data team was tasked with building a visibility dashboard for the supply chain team. Success meant **ensuring optimal vaccine delivery** to pharmacies—preventing both costly overshipping and dangerous shortages. ### Connecting three types of domain data across two different platforms The solution required integrating data from three disparate business domains: 1. Supply chain operations managing inventory and distribution 2. Shipping logistics tracking delivery status and routes 3. Manufacturing data monitoring production schedules and output This data was scattered across their Athena and Redshift environments. ### Streamlined data engineering workstream Without dbt Mesh, Moderna would’ve had to duplicate the data across the two environments. This would have led to redundancy, increased costs, more pipelines, and a loss of data lineage. Instead, Moderna could combine ‌data from different domains and environments into a single dbt project. **This streamlined software engineering, maintained the data lineage, and ensured the project’s timely delivery.** ## Three data learnings from Moderna To those in similar data journeys, Sri Kamireddy, Principal Cloud Architect, Data and Analytics at Moderna, shares three key principles: 1. **A data platform is the foundation of any organization**. Without it, there’s no data success 2. **Strong data governance and security** are necessary to foster efficient collaboration 3. **A scalable infrastructure helps keep costs under control** **** ## Transforming pharmaceutical DataOps Moderna's adoption of dbt Mesh enabled it to unite critical data across platforms and decrease data silos while maintaining pharmaceutical-grade governance. The result was faster development cycles, improved data quality, and reliable vaccine delivery to patients worldwide. Find out how dbt Cloud can accelerate your digital transformation and streamline data operations—[contact us for a demo today.](https://www.getdbt.com/contact) Watch Moderna's session at AWS re:Invent to learn more about how they optimized vaccine delivery with dbt Mesh. [Watch video](https://www.youtube.com/watch?v=Vb6qSMN2AoU) --- --- title: "How to scale analytics at your organization: Insights from data leaders" description: "Explore practical strategies from data leaders for scaling analytics through transparency, aligned teams, and optimized tools." url: "https://www.getdbt.com/blog/how-to-scale-analytics-at-your-organization" date: "2024-12-20" authors: ["Daniel Poppy"] categories: ["Learn"] --- # How to scale analytics at your organization: Insights from data leaders Data is as central as any other pillar of your business. Scaling data analytics to an adaptable, enterprise-wide data architecture demands more than raw technical skill. Organizations need an integrated approach to analytics to manage data complexity at scale. Our guide, [How to do analytics at scale: 10 tips from data leaders](https://www.getdbt.com/resources/data-leaders-analytics-at-scale), is a curated collection of insights from seasoned data leaders, each sharing practical strategies and guiding principles that have proven essential to scaling analytics at their organizations. From fostering transparency with non-technical stakeholders to structuring teams with DevOps principles, these tips reveal how organizations can avoid common pitfalls and instead empower teams across the business to work with data. Here’s a sneak peek. ## Tip one: Start with business impact, and design your team accordingly In the past, data teams were often isolated from the business side, focused solely on collecting data while leaving its interpretation to others. Today, though, **integrating data teams into specific business domains transforms their rol**e: they become close, strategic partners within the business. Embedding data experts in each domain allows them to align directly with the business context, ensuring that data is both relevant and actionable. When data teams are embedded, they know exactly which data to collect, how to transform it effectively, and how best to present insights that directly solve business needs. This proximity to business users unlocks a more substantial, measurable impact on business outcomes, driving data’s value far beyond simple collection. _–Raman Singh, Engineer Manager, Analytics, Symend_ ## Tip two: You’re more versatile than you think When it comes to scaling data organizations, data leaders should know that analytics engineers are so versatile. They have the business context of analysts, but they’ve also picked up these more technical engineering skills. They can be a bridge between data analytics and software engineering, but it goes further than that. As you scale,** lean on your analytics engineers and their skillset to flex into infrastructure problems that might traditionally call for a DevOps engineer**. Analytics engineers have the skills to solve their own infrastructure problems. Embrace moving your data team up the stack. _–Katie Claiborne, Founding Analytics Engineer, Duet_ ## Tip three: Lean on DevOps principles The first thing to stop doing is thinking about technology. It’s not about technology. It’s about people, processes and then technology. The first thing to address is the problem of how disconnected people can start working in a more connected way. We had multiple teams using different technologies, but more importantly, the ways of working were different. We had teams working in waterfall, we had teams working in agile, working in one-week sprints, other teams working in three-week sprints. How do we get these people all on board into the same ways of working? DevOps—or its spinoff DataOps—is a must have nowadays. It’s easy to build new solutions. What’s hard is to maintain and scale those solutions in the long run. If you don’t have a solid DevOps or DataOps process in place, you’re not going to go far. You need to change the culture to embrace DevOps. And then obviously technology comes into the picture. We wanted something that is code based because for us CI/CD was non-negotiable. Why? Because it embeds reliability into the release process and ultimately enables scaling. On top of that, even though our scattered data teams were using different technological stacks, all of them had something in common—everyone knew SQL. But again, it’s not only about technology. It’s really about getting people and process in place and then thinking about the technology. It takes a lot of convincing and many after-hour meetings (especially if your teams are spread across four different time-zones). But eventually you will start to gain momentum and start scaling. Once we had a solid foundation in place, it became easier and easier for us to go to a new country and say, “These are our ways of working and here’s why you should accommodate these in your lifecycle.” There was a clear tangible benefit. Now we have all the workstreams deploying their production- grade pipelines on top of our data platform. The tech leads of each of these teams come together in the same sprint planning sessions, in the very same stand-up sessions, and in the very same sprint retrospective sessions. We guarantee that we are aligned on the roadmap and execution, and that we don’t reinvent the wheel. **Even though we’re a large team made up of many smaller project-specific workstreams, we deliver together as one unified team every two weeks**—not every six months—and this in itself was a big mindset change. _–João Antunes, Lead Engineer, Roche_ [Download the guide](https://www.getdbt.com/resources/data-leaders-analytics-at-scale) for the rest of the tips. As you dive into this guide, consider these elements of a roadmap for the processes, tools, and mindsets that can transform your organization’s approach to analytics at scale. **** --- --- title: "One dbt: Accelerate data work with cross-platform dbt Mesh" description: "How One dbt and dbt Mesh provide a consistent interface for working with data, no matter where in your business it lives." url: "https://www.getdbt.com/blog/one-dbt-cross-platform-data-mesh" date: "2024-12-19" authors: ["Jeff Mills"] categories: ["Product"] --- # One dbt: Accelerate data work with cross-platform dbt Mesh Organizations often find it difficult to get a truly complete view of their data when it’s spread across multiple platforms. That’s the reality for most companies these days, however. Some of your data may be stored in [AWS](https://aws.amazon.com/), while the rest is split between GCP and on-premise servers. Despite this, you still need to provide a unified experience for working with data for all of your data producers and consumers, no matter where your data lives. In this article, we’ll explore how [One dbt](https://www.getdbt.com/blog/coalesce-2024-product-announcements) provides this by bringing data under a single, unified data control plane, enabling you to build powerful, cross-platform data mesh networks. ## Using the Data Control Plane to implement the Analytics Development Lifecycle The [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) is a cyclical process that borrows from the Software Development Lifecycle. The ADLC helps organizations work with their data, operationalize it, and discover new types of data with which to repeat the process. This approach solves questions that your data consumers may have about data origins or process rigor. However, implementing the ADLC is a multidimensional process. And that’s where the data control plane comes into play. The [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction) isn’t a dbt-specific term. It’s comprised of components that dbt believes are critical to implementing the ADLC and creating a unified experience for all data users within an organization: - Orchestration - Observability - Catalog - Semantics - Transformation dbt calls this the Active Metadata Layer. It’s a layer within your data pipeline that lies between your ingestion and your analytics or AI/Business Analytics deployments. We see this layer as critical to data transformation and management. This transformation is basically a declarative statement of how the business thinks about data and how it should be organized to answer the business’s questions. It’s that metadata about your data that becomes valuable. Things like: - When will this pipeline run and how fresh is the data? - Where did the data come from and what is its [lineage](https://www.getdbt.com/blog/guide-to-data-lineage)? - Have the metrics being used been calculated properly and vetted? ## Providing a unified experience for all data product users In the last 15 years, I haven’t seen a single customer in the data and analytics space that has all of their data on a single platform. It’s always in more than one place. That presents some hard challenges. Getting a comprehensive view of your data that includes the interdependencies between your transactional and your analytics data can be a tough nut to crack. _It’s not just the data team itself that works within your data estate._ Your data engineers, analysts, or marketing and revenue operations teams all may have varying levels of technical skills, but are domain literate and need to work with the data, too. Historically, users of dbt come by a number of means. By far, the most common one is dbt Core, which is open source. That doesn’t mean that there aren’t teams who weren’t using dbt Cloud right from the start. Sometimes, those people found dbt Cloud more suitable because they wanted managed deployments, or the teams they work with weren’t as technically advanced, etc. Maybe they wanted lineage, catalogs, or metrics so that they could all work faster and have more accurate results. Whatever the case may be, there are many reasons that dbt users might be using Core, Cloud, or even both. If you’re already using dbt Cloud, dbt Core may not seem appealing because all of the benefits of Core are also included in Cloud. Yet some organizations find themselves in a position where it makes sense to use both. One dbt is meant to address precisely that. ## Demonstrating effectiveness with a hybrid approach Yannick Misteli from Roche was in [exactly that position](https://docs.getdbt.com/blog/dbt-squared). Yannick said that after investing in a data platform, the team at Roche needed to show that it was being used and valuable, so they used dbt Core. They were then able to leverage dbt Cloud afterward to scale that impact across the globe, covering more data products within their organization. This approach allows you to steadily build trust in data and data teams, while still scaling efficiently. Simultaneously, you can leverage that Active Metadata Layer to pass metadata to your AI tooling to put results into an appropriate business context. This also enables you to ship data products faster. Features like the [visual editor](https://docs.getdbt.com/docs/cloud/visual-editor-interface) and [dbt Assist](https://www.getdbt.com/blog/introducing-dbt-assist) in dbt Cloud considerably lower the ramp up time for using dbt, making onboarding more contributors, including less technical users, doable safely and quickly. Lastly, using dbt Cloud allows for reducing duplication and reusing things like jobs to make the most cost-effective use of your compute for data transformations. ## Hybrid dbt Core/Cloud in practice I’ve typically found that central data teams will be using dbt Core. At the same time, domain expertise teams closer to the business side—teams like marketing, sales, and people operations—were taking those data products and using them on other platforms. Think Google Sheets and Excel. Since these teams weren’t using dbt, there was no governed connection for this. After implementing a hybrid deployment, those teams were instead provided tools to use that data and those data products within the envelope of dbt. That way, the data team maintains visibility into what those other teams are doing with the data That makes building a multi-step mesh across data domains possible. It also gives domain teams a wider, world-class set of tools to work with, thanks to dbt Cloud. ## Merging a sandwich shop and circus Take the theoretical example of a merger between the Jaffle Enterprises and the Cirque du Jaffle businesses. Although the example is intended to be humorous, the challenges it presents are often found in real-world organizations. Whether it be a merger or acquisition, combining two data teams, or a platform migration, the environments quickly become complicated. In this example, the Jaffle shop using dbt Cloud had all of their sales data in Redshift to provide data products to their BI applications, notebooks, and machine learning. On the other hand, the circus was using concessions, ticket sales, and circus personnel data via AWS S3 buckets. They utilized dbt Core with Athena and the AWS Glue Data Catalog to provide data products to their BI applications. In other words, these teams operated differently despite both using dbt. ## What does a hybrid architecture actually look like? The primary use case for a hybrid architecture is to have a foundational dbt Core project, with downstream teams using dbt Cloud projects that are specific to their domain. That gives them access to the visual editor, dbt Explorer, and other useful tooling. In this way, you can have your domain teams still utilizing those assets from your dbt Core project, but with the niceties of dbt Cloud. But…what if your dbt Core and dbt Cloud versions don’t match? That problem is why dbt is creating the new [“Compatible” release track](https://docs.getdbt.com/docs/dbt-versions/cloud-release-tracks) beginning this month. Each monthly “Compatible” release will match the open-source version of dbt Core and adapters at release time. For Enterprise organizations that are a little slower moving and want some extra assurance, the “Extended” release delays this by a month. dbt enables this approach with dbt Mesh. [dbt Mesh](https://docs.getdbt.com/best-practices/how-we-mesh/mesh-1-intro) is a pattern for collaboration across multiple data projects aligned to business domains. Rather than having one huge [dbt project](https://docs.getdbt.com/docs/build/projects), you can have a collection of smaller, domain-oriented projects. Each of those projects can build on each other with guardrails in place via software engineering-like interfaces. Those interfaces are defined using [contracts](https://docs.getdbt.com/docs/collaborate/govern/model-contracts), [versioning](https://docs.getdbt.com/docs/collaborate/govern/model-versions), and [access controls](https://docs.getdbt.com/docs/collaborate/govern/model-access). This means domain teams can maintain control of their data pipelines and reference other team’s projects with confidence that nothing will break. Using dbt Mesh, you can integrate dbt Core and dbt Cloud projects as follows: - First, prepare your core project for access through dbt Mesh - Second, mirror each dbt Core “producer” project into dbt Cloud - Lastly, create and connect your downstream projects to your dbt Core project using dbt Mesh In the first step, you leverage dbt Mesh to configure your public models to serve as interfaces for your downstream projects. After that, mirroring each project in dbt Cloud enables you to connect those to the dbt Core project in dbt Mesh, ensuring that changes in Core are inherited in dbt Cloud as part of your Mesh Architecture. That’s pretty much it! ## Iceberg and dbt Mesh for full multi-platform support If dbt doesn’t support a platform well, then users of the platform can’t adopt dbt. It’s in everyone’s best interest to continue adding support for more platforms and empower users. It’s for precisely that reason that dbt Labs moved to support Synapse, Fabric, and Teradata. The new adapter for [Amazon Athena](https://aws.amazon.com/athena/), too, gained support for this purpose: customers like [Moderna](https://www.youtube.com/watch?v=Vb6qSMN2AoU) had their entire analytics stack built on AWS, and Athena was a key component of that. Alongside Athena, Moderna also used Redshift and Iceberg format tables in S3. [Apache Iceberg](https://iceberg.apache.org/) is an open table format standard for storing data and accessing metadata—this means that it’s agnostic to data platforms and compute engines, so you have more flexibility in how and where you access it. Given these advantages, Iceberg is a hot topic in the data community. That’s why [dbt now supports it](https://docs.getdbt.com/blog/icebeg-is-an-implementation-detail). All it takes is a single line of code to materialize your [dbt models](https://docs.getdbt.com/docs/build/models) in Iceberg format. So how does Iceberg fit in with dbt Mesh? dbt customers with a need to integrate the data estates of multiple companies often found that, to share data between their dbt projects, they needed to consolidate to one platform or replicate that data. This is costly and time-intensive. By supporting Iceberg in dbt Mesh, you can now have cross-platform references using the same data in a common format without the need for application or movement. Simply changing the table type in your SQL configuration and changing the access configuration on the model is enough to make this data usable within queries in your other project. ## Conclusion As organizations grow and change, their data landscapes become large and diverse. Capabilities like Iceberg and cross-platform data mesh can help you scale the impact of your data operations. Support for Iceberg is just getting started, too, with more integrations for things like catalogs on the horizon. Watch our on-demand webinar [One dbt: Accelerate data work with hybrid deployments and cross-platform dbt Mesh](https://www.getdbt.com/resources/webinars/one-dbt-accelerate-data-work-with-cross-platform-dbt-mesh) to learn more about scaling your data operations in a cross-platform way. --- --- title: "Introducing release tracks for dbt version upgrades" description: "Upgrade to release tracks for more flexibility in how you get automatic dbt updates." url: "https://www.getdbt.com/blog/introducing-release-tracks-for-dbt-version-upgrades" date: "2024-12-17" authors: ["Jeremy Cohen"] categories: ["Product"] --- # Introducing release tracks for dbt version upgrades Today, we’re announcing release tracks as the new-and-improved way to manage dbt version upgrades across your dbt Cloud environments. Now, when you configure your dbt Cloud environments to one of the three release tracks, your dbt versions will be upgraded automatically—ensuring that you’re always enjoying the latest (or very recent) capabilities without any added maintenance overhead. The three release tracks are: - Latest (GA, available to all accounts): Always run the latest and greatest version of dbt, updated daily (this was formerly known as “versionless” dbt). - Compatible (in Preview, available to Team and Enterprise accounts): Version of dbt that includes all changes from final **dbt Core OSS** releases, updated once per month. - Extended (in Preview, available to Enterprise accounts): The previous “Compatible” version (one month delay), updated once per month. dbt Cloud customers can configure different release tracks across different environments to support a range of patterns for development, testing, and deployment. Release tracks mark the next evolution of our platform, helping our customers simplify maintenance, enjoy ready-access to our latest capabilities, and more seamlessly support hybrid dbt Core and dbt Cloud architectures. ## How it started: Delivering latest and greatest dbt with “versionless” So, how did we get here? As an open-core company, historically we had different development cadences for “core” dbt functionality and our dbt Cloud platform. Unfortunately, this meant that: - dbt Cloud customers would wait for core dbt functionality to land in final releases of dbt Core - Then, they’d try upgrading some of their environments, from v1.X to v1.Y - And [hope it worked](https://docs.getdbt.com/blog/upgrade-dbt-without-fear) When we introduced an exciting new capability into the dbt framework—like model contracts in v1.5, or “retry from point of failure” in v1.6—we had to tell our dbt Cloud customers that they needed to upgrade their environments before they could access those features. This wasn’t a great user experience, and it was difficult to scale across many teams. It blocked valuable new capabilities from reaching the people who would most benefit from them, for months or years. And it just isn’t the way you expect SaaS to work. [This is the problem we set out to solve a year ago.](https://github.com/dbt-labs/dbt-core/blob/main/docs/roadmap/2023-11-dbt-tng.md#the-next-six-months-stability--unit-testing) Observing that the plurality of dbt projects were still running versions from 6-12 months ago, Grace and I wrote: "There is a lot of value locked up in the features we’ve already released in 2023, and we want to lower the barrier for tens of thousands of existing projects who are still on older versions." Enter: [automatic dbt upgrades](https://www.getdbt.com/blog/dbt-cloud-launch-showcase-2024), most recently known as “versionless” dbt, and generally available since May 2024. When enabled, this configuration _automatically_ upgrades dbt environments to the latest and greatest version of dbt, including all new functionality from dbt-core and adapters. **** **Today, 85% of dbt Cloud customers are running “versionless” dbt. **They get access to the latest features, fixes, and performance enhancements—automatically, without the manual overhead of coordinating upgrades across all their users and teams—and always before those features are available in dbt Core. Plus, they get access to [new](https://www.getdbt.com/blog/announcing-advanced-ci) and [exciting](https://www.getdbt.com/blog/introducing-cross-platform-dbt-mesh) capabilities that are exclusively available in dbt Cloud. To enable this, we made strong commitments to our customers about compatibility, and put in the work to see those commitments through: [decoupling our adapter interface, introducing behavior-change flags, and making our robust testing and release pipelines more robust](https://docs.getdbt.com/blog/latest-dbt-stability). Thanks to that work, the number of functional regressions in core dbt functionality has decreased by 60+% year over year. This is vital work for defending dbt’s position as a standard for the industry, and its place at the heart of mission-critical workflows for some of the largest data organizations in the world. The strong organic adoption of “versionless” dbt validated our work on stability and ease of upgrades. We’ve consistently heard that customers value the ability to stay up to date with feature releases, bugfixes, and security updates. Still, some of those customers asked for a bit more control over when and how to receive those ongoing updates—trading off some of the latest-and-greatest functionality, in favor of predictability and Core/Cloud compatibility. ## How it’s going: New naming and more flexible options Based on feedback from our customers, from today forward, “versionless” is being crowned with a new name (or really an old one… third time’s the charm): “Latest.” This change is in response to customer confusion, and also better reflects how this configuration option really works. (“Versionless,” like “serverless,” is a misleading term; there’s always a version/server, so the real question is, do you need to manage it yourself?) And, we’re broadening beyond “Latest” and including two new configuration options—“Compatible” and “Extended”—as part of a coherent set of “[release tracks](https://docs.getdbt.com/docs/dbt-versions/cloud-release-tracks).” Compatible is the version of dbt that includes all changes from final dbt Core releases, updated once per month. Extended is the previous cycle’s Compatible version (on a one-month delay), updated once per month. We’re broadening to release tracks because there are certain scenarios in which our customers don’t always want the hot-off-the-press version of dbt; there are real use cases for deploying slightly older versions: - With Compatible, you can support a hybrid pattern (combining dbt Core + dbt Cloud) for developing and deploying in the same project. This use case requires a consistent set of functionality in dev and deployment environments, regardless if the team is on Core or Cloud. This is a step towards [our vision for “One dbt”](https://www.getdbt.com/blog/coalesce-2024-product-announcements), as a turnkey way to ensure interoperability for hybrid deployments. - With Extended, you can run production environments on versions that have been in the wild for more than one month. - By using both in combination, you can try newer versions in a lower (dev or test) environment (on Compatible) before they land in your production environment (on Extended). Below is a table of generalized customer architecture recommendations, by environment, based on your priorities: ![Table of recommended architectures based on various customer priorities](https://cdn.sanity.io/images/wl0ndo6t/main/ecc2d9568e8f9590dc55d1f71799a022514c94c3-3226x1338.png) It’s important to note that choosing Compatible or Extended means that you’ll get a slightly older version of dbt—but you’ll never be grossly out-of-date with your version upgrades. Your friends on Latest will have access to some good stuff sooner—and you can too, in a separate environment or [user-level override](https://docs.getdbt.com/docs/dbt-versions/upgrade-dbt-version-in-cloud#override-dbt-version) to try out a beta feature and give us feedback. It’s also important to note that it will no longer be possible to pick a years-old dbt Core version and get stuck on it, left behind while missing out on all the features and fixes of the subsequent years. _This is a good thing._ ## Understanding the Compatible and Extended release tracks By aligning to the functionality available in the most recent open source releases of dbt Core and adapters, the Compatible release track meets the needs of customers with hybrid Core/Cloud deployments, or customers who (for whatever reason) prefer less-frequent release cadences. The Extended release track—which will always include _the exact same versions_ as Compatible, just one month later—will meet the needs of enterprise customers who want the ability to test changes in development or staging environments, before those changes land in production environments. These release track options make it possible for customers to get automated upgrades on a less-frequent cadence—after that version has been live in dbt Core for a period of time, and later than “Latest.” ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/79a084d3954fe2abb23efc48b39d3cfaa8807772-1256x704.png) To offer one example: Microbatch incremental models are an exciting new feature [just released in dbt Core v1.9 last week](https://www.getdbt.com/blog/dbt-core-v1-9-is-ga). _This feature has already been available on the Latest release track for months_. The same goes for improvements to snapshots, state:modified, foreign-key constraints in model contracts, and more. Release tracks enable you to opt into trying new functionality on your terms—and to participate in [discussions](https://github.com/dbt-labs/dbt-core/discussions/10672) and feedback sessions that help shape the product and framework—with fewer barriers than before. Now that microbatch models have become available in a final release of dbt Core (v1.9.0), they have also become available in the subsequent release of the Compatible track. Both Compatible and Extended are now available in **Preview** for eligible customers. For the very first release (December 2024), they will be identical, and they both include all functionality from dbt Core v1.9—meaning that the functionality they contain is very close to what’s in “Latest.” Over the coming months, they will diverge: - **Latest** will include the very latest changes as we make, test, and release them—just the same as how “versionless” dbt has worked all year long. The “Latest” release track will include early access (always opt-in) to all new framework features _before_ they land in dbt Core v1.10 prereleases or final releases. - **Compatible** will have its next update in mid-January, including fixes in dbt-core v1.9.X patches and any adapter updates. It will not include new dbt Core _features_ until the v1.10 release ~6 months from now. - **Extended** will have its next update in mid-February, unless there's a critical security issue we need to hotfix. ## Looking ahead The evolution of our platform to release tracks puts us on a path to be able to regularly and continuously deliver new framework capabilities that dbt Cloud customers can use immediately. And to do this in a way that our customers never even have to think about—it just works. We’re going to keep investing in the reliability of our framework and platform, across all release tracks—proactively detecting functional and performance regressions before they roll out to customers, and [ensuring ongoing compatibility with more of the popular open source dbt packages](https://github.com/dbt-labs/dbt-core/discussions/10668). If you’re not yet on release tracks, or are interested in using Compatible or Extended, below are important milestones: - Latest release track is currently GA to all dbt Cloud customers. Compatible (Team & Enterprise) and Extended (Enterprise) release tracks are currently in Preview. We will also be introducing self-managed rollbacks for Business Critical customers in February 2025. - All release tracks are expected to be GA in March 2025. - Once all release tracks are GA, dbt Core v1.7 in dbt Cloud (which had its End Of Life date extended) is no longer going to receive any updates. At that point, we will be encouraging all customers to move to release tracks. Finally, I’d like to extend my personal gratitude to the dbt Cloud customers and dbt Community members who have taken the time to speak with us about dbt upgrades over the past year. This is a big, nuanced topic—and it is thanks to your good-faith engagement and thoughtful feedback that we have been able to get this far. I’m excited about the balance we’re striking with release tracks— providing our customers with more control and flexibility, while ensuring they stay close to the latest and greatest that dbt has to offer. --- --- title: "A decade of data evolution and 2025 predictions" description: "A fascinating 2024 is a precursor to a big 2025." url: "https://www.getdbt.com/blog/2025-data-predictions" date: "2024-12-16" authors: ["Tristan Handy"] categories: ["Insights"] --- # A decade of data evolution and 2025 predictions _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/reflections-and-predictions). _ I’ve been writing this newsletter since September of 2015. This will be the 10th year I’ve had the opportunity to reflect on a year gone by and make predictions about the year ahead. ## Data science and the rise of Python In the early years (**2015-2017**), the data ecosystem was dominated by data science. Data viz, developments in the Python and R ecosystems, posts demoing statistical techniques wrapped up in open source packages, strategies for winning Kaggle competitions, and a community dominated by highly technical people (even if mostly they were just building ETL pipelines in notebooks 😆). ## The fall of Hadoop and the return of SQL From **2018-2019**, attention was more focused on the fall of Hadoop, the return of SQL, the advent of analytics engineering and the dbt community, and the rise of data ops. Many of the developments in the space were led by data and data infra teams inside of large digital natives, most especially Airbnb, Uber, and Netflix. Each of those three companies contributed very meaningfully in both open source code and best practices to the data ecosystem that was to be built next. ## The modern data stack boom From **2020-2022**, attention was focused around the modern data stack and the rapid growth of cloud data platforms (including Snowflake’s IPO, which gave us public data about the size of the financial opportunity). Conversation in the community was about new companies started, new fundraising events, and who was going to win what categories. Categories that had once been sleepy became the subject of much attention, and other categories were created from nowhere. Data team sizes grew quickly as companies were flush with cash from the COVID and ZIRP boom, and they spent a ton of time updating best practices and incorporating new tooling into their stacks. Data architecture slides went from having 5-8 logos to 30+ logos. ## 2023: A year of reckoning In **2023**, everything changed, very quickly. Inflation drove rates up. Very quickly, all parts of the economy become concerned with the chances of a recession. While the recession never materialized, an immediate pull-back in investment forced a reckoning across the software space. Cloud earnings growth dropped significantly, and downstream of that, almost all software companies started missing quarters. Cue layoffs, often impacting data teams, from across big tech, growth companies, and the enterprise. As a result, attention overnight shifted from improving best practices and platforms to delivering near-term business value. At the same time, ChatGPT was launched in late 2022 and all the sudden it was no longer clear to either software buyers or software investors what categories of software would be helped and which would be hurt by the coming AI wave. So everyone stayed on the sidelines during 2023. The “big 5” data platforms (SNOW, DBRX, Azure, AWS, GCP) continued to chug forwards, supported by an underlying exponential: the S-curve of the move from on-prem to cloud for enterprise data. The larger climate certainly impacted the speed of this move, but this megatrend was resilient to the underlying macro. ## The big shifts of 2024 But in **2024**, things changed again. Here are what I consider to be the biggest and most salient things to happen in the data industry over the past year. 1. **Macro stabilized.** The predicted recession didn’t happen. Layoffs dissipated, and companies began thinking more strategically and long-term about data (among other things). It was not a return to 2021, but it was a return to stability. Enterprise CIOs and CDOs were again thinking about how to drive their organizations into the future. Venture dollars were still more anemic, so data companies targeting early adopters and SMBs struggled, but if you targeted the enterprise there was renewed customer demand and a path to growth. 2. **Iceberg won.** It’s hard for me to tell whether this development was more driven by customers’ desire to avoid lock-in, or more driven by Ali Ghodsi’s maniacal focus on driving the Lakehouse vision, but this was the year that open table formats broke out. The coming out party was in the two back-to-back weeks of Snowflake and Databricks summits when both CEOs made strong public commitments to Iceberg. Immediately, the topic of Iceberg and open table formats became salient to CDOs—I have never seen such a seemingly-esoteric topic go from 0 to 60 in executive interest so fast. Over the next six months, the hyperscalers followed suit, releasing features that made it meaningfully easier to work with open table formats. If Iceberg felt like it was leading in June, by December it is clearly the winner. 3. **AI shifts from a headwind to a (modest) tailwind.** In 2023, AI was a headwind to data. In 2024, that changed. It became clear that AI and unstructured data didn’t somehow replace the need for structured data and the associated data pipelines—rather, AI became another downstream use case for existing data technology. EL companies like Fivetran and Airbyte reported an acceleration of growth as a part of customers’ AI initiatives. Data platforms grew their native AI capabilities to bring AI directly into the hands of current data practitioners (think: Snowflake Cortex). Many data products shipped experiences bringing AI into the workflow of data practitioners (think: dbt Copilot). At this point it is clear that a) AI will only make data more critical, b) data practitioners’ work will both change and be accelerated, but not _disrupted_ by AI (at least…not in the foreseeable future). 4. **Consolidation is happening.** M&A ramped up in the space this year, and my indicators are that this has even accelerated from H1 to H2. Data companies that raised in ‘20-’22 are running low on cash and many do not have the needed traction to raise another round. I have personally gotten half a dozen inbounds from companies looking for M&A outcomes just over the past month or two, and I get to see even more of this through my angel investing. It is happening. But consolidation is not just about M&A. Many players in the space are beginning to expand into each others’ lanes organically as well. Data clouds building native EL. Observability, lineage, and catalog all smashing into one another. It feels like plate tectonics: slow, but inevitable. And we are headed towards Pangea. The question is: which companies have earned the right to become true platforms? How many will there be? And will they all just be roughly carbon copies of one another or will there be meaningfully different visions on display? 5. **The big players got religion on semantics.** SNOW, DBRX, Tableau, and more all introduced semantic layers of varying levels of maturity. If there are three primary use cases for semantic layers (internal analytics, embedded analytics, and AI), these initiatives were either mostly or completely focused AI. This was an important step towards increasing industry awareness / adoption of a long-term important technology. ## 2025 and beyond: Predictions for what’s next Here’s how I believe that all of this translates into 2025: 1. **Open table formats get implemented, fast.** Few companies use Iceberg in prod today. We have pretty good data on this; widespread adoption is taking some time. But between dbt Labs, EL vendors like Fivetran, the data clouds, and the [hyperscalers](https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tables-buckets.html) all building features to make implementation easier, we’re going to start to see this line go up quickly. 2. **The rise of utility compute.** Fivetran recently released the ability for customers to write Iceberg tables without needing to pay for any underlying compute (on platforms like Snowflake and Redshift this had previously incurred a non-trivial cost). This quote from the post is fascinating: “When you build a specialized engine for a specific workload, you can be more efficient, because you can rely on special characteristics of your workload. For example, when Fivetran built our data lake writer service, we were able to make it so efficient that we can simply absorb the ingest cost as part of our existing pricing model. Ingest is free for Fivetran data lake users.” This will happen more often with the move towards Iceberg: purpose-built engines will run very specific workloads in a highly-optimized way. I am calling this “utility compute.” It is not a substitute for the engines that allow for processing chunky production workloads, but it is an ability for vendors to create very significant optimizations specific to their own very particular workloads. I would expect to see more products do what Fivetran did over the coming year. 3. **A diversification of compute environments, but a unified layer on top.** If open table formats encourage customer choice and enable multiple purpose-built compute engines to proliferate, there still needs to be a single pane of glass into the entire data estate. Previously, it was common for CDOs to say something like “We’re a [Snowflake/Databricks/etc.] shop.” That is much more rare today: already, over half of all enterprises use multiple data platforms. So: where does that single pane of glass shift to? Is it the metadata catalog (i.e. Unity)? That doesn’t feel right to me, although from what I can tell that’s Databricks’ bet. Is it the user-facing catalog (i.e. Alation/Colibra/Atlan)? That doesn’t feel right to me either. Watch this space; I believe this is the biggest real estate being fought over in 2025. 4. **Acceleration of consolidation, the rise of end-to-end platforms, and the focus on end-to-end user workflows.** Every data platform is going more end-to-end. Azure has Fabric, which is an integrated bundle. Databricks is going end-to-end through both acquisitions and organically built products. GCP has cared about this for years; it is what motivated the Looker acquisition. AWS and SNOW are moving in this direction as well, which is the _most_ interesting indicator for me, because both companies have always been notoriously against doing this in the past. AWS has its approach of selling individual puzzle-pieces and allowing developers to fit them together, and Snowflake has been (to their credit!) very ecosystem focused, letting partners solve all of the adjacent problems that weren’t fundamentally about delivering a compute offering. So: everyone is going wide, going integrated. My read: this is a recognition that the axis of competition is moving from ‘owning the workload’—making it really hard for customers to move workloads from one platform to another—to ‘owning the user’. If open table formats are giving customers more choice as to where those workloads run, it is important for vendors to re-establish strategic power in other ways. So: go wide, own the user experience for the end-to-end data workflow. I don’t know that this industry shift is exclusively positive or negative for practitioners; I do consider it basically inevitable. I don’t know about you, but I’m honestly feeling optimistic about the future going into 2025. The thing I care most about in data is the ability to make consistent progress as an ecosystem, and not be stuck in a world of cyclicality, rebuilding the same set of technologies over and over again like we have for 30+ years. The best ways to make that happen are open source and open standards, replacing cyclicality with consensus and steady forward progress. This is how software engineering has made progress for decades. The fact that we have coalesced around a standard way of storing data, and a catalog to manage transactional consistency on top of that, is just incredibly good news for our collective future. Everything else I wrote above is downstream of that. I hope you’re ending your year this year with some optimism as well. See you in 2025 :) --- --- title: "Learning dbt as an analyst" description: "Discover what two data analysts wish they'd known when starting with dbt. Learn how to simplify workflows and build trust in data." url: "https://www.getdbt.com/blog/learning-dbt-as-an-analyst" date: "2024-12-13" authors: ["Rachael Gilbert", "Chris Fiore"] categories: ["Insights"] --- # Learning dbt as an analyst Hello, we’re Chris and Rachael, two data analysts at dbt 👋. Today, we're reflecting on the things we wish we had known when we first learned dbt. Unsurprisingly, we think dbt is pretty great. But it wasn’t always smooth sailing. To be frank, when we first encountered dbt, we each felt confused and uncertain. We knew SQL and were semi-comfortable with data modeling concepts and git, but dbt’s technical terminology induced some serious impostor syndrome. We didn’t identify as engineers, or even pretengineers. YAML? Jinja? Unit tests? We had gotten used to transforming our data via simple internal tooling and cron jobs. We didn’t quite understand what all the dbt hype was about. Despite our initial hesitation, we ultimately each had to take the plunge. For Rachael, the tipping point was moving to a new startup with “the modern data stack”. For Chris, his team had reached the limits of their legacy infrastructure and realized they needed a more robust solution. In learning dbt, it took time to connect the dots, to grok project structure, and to wade through the breadth of resources and find the ones we needed. But once everything clicked, it was transformative. Things like environment separation went from intimidating concepts to common sense, and our workflows improved dramatically in a way we hadn’t anticipated. We wished we had done this _years_ earlier. This post is all the specifics on what we wished had clicked earlier for us and made this journey smoother. If you’re an analyst who doesn’t quite “get” dbt yet, maybe this is for you. We’ll first share what makes dbt valuable to us as analysts, then we'll dig into how some of the main concepts all fit together—no engineering background required. ## Why dbt helps us as analysts When learning dbt, we found the biggest challenge to be in revising _how_ we thought about data transformation. Instead of viewing queries and analyses as one-off scripts, dbt pushed us to learn best practices around scalable architecture and a robust data ecosystem. This paradigm shift transformed how we worked as data analysts. It brought structure, reliability, and automation to our transformations, unlocking value in four key ways: ### 1. Less manual work, more impactful analysis _**Before dbt**:_ We were running queries manually or at set times, exporting CSVs, and passing cleaned data to stakeholders. We spent so much time just **wrangling** data, taking bandwidth away from deeper analysis. _**With dbt**:_ Our transformations are automated, version-controlled, and repeatable. If we need to adjust a calculation, fix a data issue, or update a data table, we update the relevant model in our repository and let dbt handle everything else downstream. No more manual pipeline dependency management required. **_Example:_** We build a “customer lifetime value” model in dbt. When marketing wants an updated report, all we do is run dbt—it pulls in the latest data, recalculates everything, and updates the results in the report. ### 2. Transparent business logic (that we can explain) **_Before dbt:_** We often lost business logic to messy SQL scripts, spreadsheets, dashboards, or spaghetti notebooks. If someone questioned a number, tracking down its source felt like detective work and became a huge time suck. **_With dbt:_** Every data transformation is documented, version-controlled, and transparent. The logic behind metrics lives in one place—our dbt project. This makes explaining “how the sausage gets made” straightforward and accessible to anyone with access to our repo or dbt. **_Example:_** When a stakeholder asks, “How did you calculate monthly active users?”, we now point them directly to a tested and documented dbt model in our repository instead of scrambling through five different queries or a stale notebook. ### 3. Collaboration without chaos **_Before dbt:_** SQL scripts often lived across personal folders, BI tools, and Slack, making collaboration difficult and error-prone. **_With dbt:_** Our whole team works in the same version controlled repository. Every change is reviewed through pull requests, and we can easily roll back changes when needed. No more wondering who changed what—or where the most recently agreed upon business logic lives. **_Example:_** When we work on the same report, we can build models in parallel, knowing dbt will merge them smoothly once approved. ### 4. Greater accuracy and reliability **_Before dbt:_** We validated results manually, using spot-checking or a gut feeling to ensure our data looked right. Sometimes we didn’t realize when things were broken until our stakeholders flagged it in a report. Human error is hard to avoid. **_With dbt:_** CI functionality lets us understand what we are changing, so we can fix bugs before they go live. Automated tests as we deploy ensure that our data meets specific criteria—whether it’s checking row counts, unique values, or even more complex custom logic. We know immediately if and where something breaks. **_Example:_** We add a test ensuring that every transaction has a valid customer ID. If the test fails, dbt tells us exactly where the issue is—before it makes its way into a report. Now our team can proactively roll out a fix without manually inspecting our data pipeline piece by piece. ### Four main reasons we’re now so big on dbt - **Automation**- We’ve eliminated time spent on repetitive tasks, manual queries, and data testing. - **Transparency**- Everyone can see and explain how every metric is calculated. - **Collaboration-** No more duplicative queries, lost scripts, or reading .sql files in Slack. - **Correctness-** We know our data is more accurate with built-in testing and checks. ## How we do our work in dbt Hopefully by now, we’ve painted a picture of why we've found dbt so helpful in our analyst work. dbt has a lot of functionality, and we’re certainly not going to cover it all today. But we want to chat about the three main [ADLC stages](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) we execute on most as analysts in dbt and the relevant terminology, to help build a better mental model of the product (caveat: we use dbt Cloud). ### 1. Develop All the folks using dbt Core plus many in dbt Cloud are used to a command line (CLI) experience when writing code. That's one of the main interfaces for developing (or building tables/views i.e. models) in dbt. However, it’s not the only one. While Chris tends to invoke dbt commands in his local CLI with VS Code, Rachael prefers to run commands and build models in the dbt Cloud IDE. She finds it to be an easier interface coming from other GUIs like RStudio. dbt is also launching a new no-code interface (Visual Editor) coming next year for other types of users who are less comfortable with SQL. ### 2. Deploy Once we develop and test a model, we deploy it to our production environment. With the dbt Cloud orchestration functionality (or another product if you’re using dbt Core), we can also use a job to keep it up to date. There are many ways to run jobs. Sometimes it can be as simple as setting them to run at a certain time, but sometimes we want them only to run after other jobs complete, or run when we merge a change. ### 3. Explore Data architecture gets complex with time. It’s hard for us to remember every model or column name, the code that builds them, or what they represent. Thus, we're very often poking around dbt Cloud Explorer to view model and column-level lineage for upstream/downstream dependencies, reference documentation, or peek at data health. ### 4. And more (…but not for today) We haven’t even scratched the surface of how we do testing, how we scale our architecture across projects with Mesh, how we dynamically codify metric breakdowns with the Semantic Layer, and more. We acknowledge there's much more analyst-on-dbt ground to cover in future posts. ## What’s next We’ve covered a lot of content in this short post. We've recapped why we find dbt valuable, explained how we actually use it in our main day-to-day, and hinted at more to come. But ~~Rome~~ dbt models weren’t built in a day. From here, if you want to learn more, we recommend the [dbt Fundamentals](https://learn.getdbt.com/courses/dbt-fundamentals) course. Or if you’re more of a learn-by-doing type of person, [dbt Cloud has a free plan](https://www.getdbt.com/pricing) where you can jump in and start trying things out. We’d love for you to catch us in the [community slack](https://www.getdbt.com/community/join-the-community) to say what else you’d like to see as well. --- --- title: "dbt Core v1.9 is GA" description: "Upgrade to take advantage of microbatch incremental strategy, improvements to snapshots, and much more." url: "https://www.getdbt.com/blog/dbt-core-v1-9-is-ga" date: "2024-12-10" authors: ["Grace Goheen"] categories: ["Product"] --- # dbt Core v1.9 is GA Today, we're excited to announce that [dbt Core `v1.9` ](https://github.com/dbt-labs/dbt-core/releases/tag/v1.9.0)(named after Dr. Susan La Flesche Picotte) is GA. Since we started this journey in 2016, dbt has quickly become the standard in data transformation, with over 50,000 teams worldwide relying on dbt to build, test, and document their data products. dbt Core `v1.9` includes improvements large and small designed to help data teams work more like software engineers—that is, in a way that's modular, scalable, repeatable, and governed. I'll dive into more details in this post, but the highlights of `v1.9` include: The big hits: - 🤏 New **microbatch** incremental strategy to optimize your largest datasets—transform your event data in discrete periods with their own SQL queries, rather than all at once - 📸 New **configurations and spec for snapshots** to make them easier to configure, run, and customize The smaller stuff: - 👯‍♀️ **Improvements to state:modified** behaviors to help reduce the risk of false positives - 📄 **Document your data tests** by adding `description`s - and more! And adapter-specific features: - 🧊 **Standardizing** **support for Iceberg**, with standard configs on more adapters to materialize dbt models in Iceberg table format (read more [here](https://docs.getdbt.com/blog/icebeg-is-an-implementation-detail)) Check out our [upgrade guide](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.9) for more information and read on for a deeper dive into the community efforts and conversations that helped bring these features to life. ## Marriage: It’s what brings us together If you attended my [2024 Coalesce talk](https://www.youtube.com/watch?v=DC9sbZBYzpI), we’re married now ;) To symbolize our continued commitment to open source and discuss the ways our relationship has grown over the years, I hosted a very special ceremony in which I renewed my vows with the dbt Community. And, it’s Vegas, so of course Elvis was there to officiate. ![Grace and Elvis at Coalesce](https://cdn.sanity.io/images/wl0ndo6t/main/5402cc6e89a7b0b94bb2533a44ae970ab2fa783a-4000x2667.jpg) dbt Core was created _eight_ years ago—a tool and framework to standardize the way we all do data transformation. Because our co-founders decided to make dbt Core open source—following the pattern set by the best CLI tools and programming languages—the dbt community was able to grow organically and collaboratively. dbt enables data practitioners to work like software engineers—to automatically handle dependencies, test and document our data models, and version control our code. But the real power of dbt is providing a **standard framework** for doing this work. dbt gives us a common language that empowers us all to share our solutions and build together. dbt isn't just a tool, it’s a community. We are a group of people who have chosen to believe in a viewpoint—and then watched as that viewpoint transformed data work across the industry. Like any relationship that’s lasted this long, we’ve grown and changed together. dbt, the company, has grown from an analytics consulting practice with an open source tool, to [an open core business with a real path to long-term sustainability](https://www.getdbt.com/blog/next-layer-of-the-modern-data-stack) (one that can fund the ongoing development of that open source tool). dbt, the product, has gotten a lot more [mature](https://www.getdbt.com/product/dbt-cloud) and [stable](https://www.getdbt.com/blog/seamless-scalability-effortless-upgrades-the-enhanced-dbt-cloud-platform). dbt, the community, has expanded to [over 100,000+ members](https://www.getdbt.com/community), all across the world. Eight years and 50,000 weekly active dbt projects ‌later, **dbt Core is the industry standard for data transformation; and we're committed to living up to the trust organizations worldwide have shown in us**. What does that mean? It means: - dbt Core will remain licensed under Apache 2.0. - the dbt framework will continue to be shaped by a collaborative effort between you (the community) and us (the maintainers) - when we add something new to the standard, we are committing to the long term. We must be intentional about _how_ and _when_ we do it. You can be confident that, once added, it's there to stay. ## Extensibility is what powers the community So: We take our responsibility of owning the standard very seriously, and we aim to be really intentional when that standard is updated or expanded. One of the best things about dbt is that it is flexible and extensible. If you have a problem to solve, you can use the tools dbt gives you—yes, including jinja—to implement a solution. Have you ever written a custom materialization, overridden one of the built-in core macros, nested some DML in a post-hook, to solve a niche problem for your data team? I certainly have. The ability to do this type of “customization” is _by design_. A mantra for our team, borrowed from [the programming language Perl](https://en.wikipedia.org/wiki/Perl#Philosophy), is “make the easy things easy, and the hard things possible”. We _want_ you to find creative solutions, unblock yourself, iterate rapidly, and share your code snippets to unblock others. We don’t want to be a bottleneck for you getting your day-to-day problems solved. Providing an out-of-the-box solution for the entire gamut of problems you might encounter when doing data work wouldn’t be feasible or recommended. dbt the framework should be **opinionated** and have **clear scope**. _And_ the **extensibility** of dbt can and should be leaned on to solve a whole host of niche problems. We also recognize that extensions have their limitations, and sometimes, the _right_ strategy is to go from making something "possible" via custom extensions... to making it "easy" with an out-of-the-box solution built into the dbt Core standard. So how do we know when it’s time to make that transition? We rely on a few signals: - When we see a ton of people upvoting an issue in our GitHub repos - When we see multiple open source dbt packages trying to close the gap - When we see people leaving dbt’s framework to solve this problem, rather than hacking within it - When we see technical advancements in the ecosystem, a hard thing becoming easier These things tell us a capability is ready to graduate to become an **official part of the dbt standard**. A lot of the features available to you **now** in dbt Core `v1.9` are things that community members have been discussing and experimenting with for years. Thank you for your efforts, for your thoughtful conversations, for helping us shape the future of dbt. Let’s get into it. ## What’s new in dbt Core `v1.9` `v1.9` delivers two new major features to the dbt framework, as well as a long list of smaller improvements. For each of the “big ones,” let’s take a look at not just what the feature is, but how community efforts shaped how the feature was prioritized and built. ### Microbatch incremental models ![Microbatch incremental models (a brief history)](https://cdn.sanity.io/images/wl0ndo6t/main/97fca5b2ce5026a1212924568fee94464be5247b-1920x1080.png) The `incremental` model materialization was first introduced in **2016** in `v0.4.0`. It is a [foundational](https://docs.getdbt.com/docs/build/incremental-strategy) part of how we think about optimizing your datasets that are too large to be dropped and recreated from scratch every time you do a `dbt run`. The `incremental` materialization is an advanced strategy best used on models that are large enough that you don’t want to do a full refresh on every run. Instead of reprocessing an entire dataset every time a model is run, incremental models process a smaller number of rows with new data, and then append, update, or replace those rows in the existing table. If anything goes wrong or your schema changes, you can run in “full-refresh” mode, by running the same simple query that rebuilds the whole table from scratch. In **2018**, an experimental materialization called `insert_by_period` was created as a _further_ performance optimization for datasets that are _so _massive that **rebuilding the whole table in a single query is just not possible without running into warehouse timeout issues**. Instead, `insert_by_period` processed event data in discrete periods _with their own SQL queries_, rather than all at once. This is why extensibility is so powerful; we can experiment, we can let a solution bake, and we can then ask the question, “Should this become a _real part_ of dbt?” ([as Joel did in **2021**](https://github.com/dbt-labs/dbt-core/issues/4174)). While the original experimental materialization was only supported in Redshift, it was expanded this year to work for other data platforms, including [BigQuery](https://github.com/dbt-labs/dbt-labs-experimental-features/pull/42) and [Databricks](https://github.com/dbt-labs/dbt-labs-experimental-features/pull/44). This experiment allowed us to iterate on batched-based processing and, in the meantime, unblock dbt users who needed this kind of approach. However, it lacked the official seal of support from dbt project maintainers. And limitations existed—it required copy/pasting macros into your own project, batches had to be run in serial, and there was no individual logging for each batch. It was time for this feature to be built directly into dbt Core. We are excited to announce that starting in dbt Core `v1.9`, you can use the brand-new `microbatch` `incremental_strategy` to break up your massive datasets into smaller, bite-sized batches that dbt can process individually, making your workflows faster, more reliable, and easier to manage. Simply set your `incremental_strategy` as `microbatch` and write your SQL for a single “batch” of data. **dbt will then evaluate which batches need to be loaded, break them up into a SQL query per batch, and load each one independently.** Batches can be run concurrently, and batch size can be `hour`, `day`, `month`, or `year`. Check out [our docs](https://docs.getdbt.com/docs/build/incremental-microbatch) for more information. ### Improvements to snapshots ![Snapshots (a brief history)](https://cdn.sanity.io/images/wl0ndo6t/main/429ef75a75f5c3c660520f9c9332ddbbf18cf352-1920x1080.png) Snapshots were first released as `dbt archive` in `v0.5.1` back in **2016**, just two days shy of dbt's 6-month anniversary. Snapshots provide a solution for capturing changes to your data as type-2 Slowly Changing Dimensions, so you can “look back in time” at previous states of your tables to incorporate historical data into your analyses and measure trends over time. While they were originally declared within the `dbt_project.yml`, in 2019 they moved to a dedicated snapshots folder and were defined within a special “jinja block”. (Oh `{% snapshot my_snapshot %}`!) We built snapshots a long time ago, and the issues that the community has opened over the past ~7 years signaled to us that the original snapshots hadn’t really kept up. Part of our responsibility of maintaining the standard includes admitting that we don’t always get it right. In early 2023, our DX advocate for dbt Core (Doug Beatty aka “[Timestamp Doug](https://github.com/dbeatty10/Phippy-Goes-Fact-Finding)” for those in the know) opened the discussion “[Problems & (Potential) Solutions for Snapshots](https://github.com/dbt-labs/dbt-core/discussions/7018)” pulling from a long list of issues and comments describing the ways that our current implementation of snapshots just wasn’t cutting it. And so, I’m happy to announce that dbt Core `v1.9` incorporates a number of improvements that make snapshots **easier to configure, run, and customize, including:** - New simplified snapshot specification: snapshots can now be configured in a YAML file, which provides a cleaner and more consistent set up - New `snapshot_meta_column_names` and `dbt_valid_to_current` configs: allow you to customize the meta fields that dbt automatically adds to snapshots - `target_schema` is now optional for snapshots: when omitted, snapshots will use the schema defined for the current environment, meaning you can maintain environment-aware snapshots - New `hard_deletes` config: get more control on how to handle deleted rows from the source - And more Check out [our docs](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.9#snapshots-improvements) for more information. ### The smaller stuff It’s not _just_ about the big stuff. The little things matter too, and we tackled a lot of them in `v1.9`. To name a few: - set your foreign key constraints using `ref` - document your data tests by setting a `description` - less false positives for `--select state:modified` Check out [our docs](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.9#quick-hits) for more information. ## I'm just a ~~girl~~ PM, standing in front a ~~boy~~ community, asking ~~him~~ them to love her <3 On a personal note, I feel so lucky to be building this product with all of you. I get to log on to my computer every day and interact with thousands of “internet friends” who want to help make dbt the best it can be. This community is so special, and it’s truly an honor to be a part of it. To everyone who participated in what we built this year, and all of you who are helping shape what we build next, thank you. Speaking of… I want to hear from you! What do you struggle most with when using dbt today? What custom solution have you had to implement multiple times that you wish dbt offered out-of-the-box? Big or small, I want to know what’s on your wishlist of dbt features. [There are so many ways to participate in this community](https://docs.getdbt.com/community/resources/oss-expectations): - Upvote and comment on GitHub issues, or start a GitHub discussion or discourse post when something’s not-so-clear-cut - Join us on Zoom for feedback sessions to help us design features that _feel_ like dbt - When you find a way to solve that unique problem, share it—in a blog post, at a dbt meetup, or by talking at Coalesce - If you really want to get into the weeds, contribute code back to one of our open source repos, for one of our issues tagged `help_wanted` and `good_first_issue`, and our engineering team will work with you to get it over the finish line You don’t have to write code to contribute to the dbt open source community. Sharing your different approaches and what you’ve learned in practice is how we all move up the stack. I said my vows to you on the Coalesce stage, and I’ll say one again here: We vow that dbt is not dbt without you, the community. Whether you use dbt Core or dbt Cloud or some mixture of the two… your thoughts and opinions matter to us. **Because this _is_ a relationship, and we want to build dbt with you for a long time to come.** --- --- title: "Making data movement as reliable as electricity" description: "Iceberg, unstructured data, and the data infrastructure needed for AI, with Fivetran's cofounder and COO Taylor Brown." url: "https://www.getdbt.com/blog/making-data-movement-as-reliable-as-electricity" date: "2024-12-08" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Making data movement as reliable as electricity _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/making-data-movement-as-reliable). _ Fivetran recently passed $300 million ARR and has over 7,000 customers globally. Taylor Brown, the cofounder and COO of Fivetran, joins the show to talk about Fivetran’s moat, the impact of AI on the data ingestion space, and open table formats and catalogs. **Listen and subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways from this episode ### The Fivetran mission is to make data movement as reliable as electricity. Is that right? Taylor Brown: The thinking behind that mission statement is that when we think about Thomas Edison and what he did, he spent all this energy bringing electricity into the house to power light bulbs. And then what happened after they had electricity in the house was an explosion of additional innovations—hairdryers and washing machines and all the electronics. One of the biggest challenges to innovation is just access to data. And BI, as we thought about it, was really the light bulb of the modern data stack. I think especially with AI, the innovation set to come is still to be defined. We're at that light bulb moment still. ### You've got to have the use case that drives the original infrastructure, but then who the hell knows what the infrastructure is going to be used for next. Exactly, exactly. In the last 10 years, it's been a lot of light-bulb kind of BI stuff. And the last two years have been this new, fun, more exciting innovation around AI, which makes my life more fun. And, you know, I think what we're doing is more interesting. ### You guys are big time. Where's the business? You guys have hit some milestones recently. We recently passed 300 million in ARR. We have over 7,000 customers now globally. And we're growing at a great clip right now. The last two years were not maybe the best years for Fivetran. It was just a challenging time in the market. And I think there were a lot of folks who pulled back on any sort of innovation. We've seen a resurgence of growth for ourselves this year, which has been great. A lot of folks are really starting to feel more confident in the market, which ultimately ends up in more spend on innovation, which means we all see more investment in data. ### One of the interesting things about this space that you're in is that everybody thinks that they can build data pipelines. And in an environment where somebody up the org chart is looking to save money, there's probably somebody lower down in the org chart that says, “Screw it. I'll roll that myself.” ### Is that a conversation that you've had over the years? It's a conversation we've had and a conversation we continue to have. Especially when you think about the modern data stack where you do more ELT to extract the data and load it with a small amount of transformations directly into your cloud data warehouse. I think a lot of folks that are at senior levels at organizations say, “Hey, you're not even doing the hard part. You're not doing the transformation part. Why would I ever use a tool for that?” When you get into the details, there's a lot of complexity, as you pointed out earlier, to moving data effectively, doing it accurately, doing it at scale, making sure that you don't miss any data, doing it incrementally instead of batch. There's all this complexity that we put into making it so that we can have this highly reliable replication and copy of your data within the warehouse. That's a challenge that we have to face with buyers on a constant basis to help them understand why this is cheaper, better, faster than having their own team build it, where the quality can be all over the place. A lot of times engineers don't really want to do this. They see this as a shitty job for them to do. ### Moving data from point A to point B is ’t how you build a career. It's a demotion, really. We got this these pipelines and need you to go do this instead of doing this really important mission-critical stuff over here. ### Do you have types of metrics that you show? We have a lot of metrics that we show. We have uptime metrics. We're working on a bunch of latency metrics right now. I’d say for your average company, we can probably do it better and faster than you can. ### But sometimes that one data engineer doesn't want to hear that. One hundred percent. There's two aspects of this. For the really big companies like Facebook, data really is their business. Building the infrastructure around it is their business as well. If you're outside that, in the build-versus-buy scenario, you're going to end up with something that's better, faster, more reliable, because you have this crowdsource effect. We have 7,000 plus customers who are using the same infrastructure. We're able to really battle test it over a large number of customers and an even larger number of connectors. And we're going to catch all those edge cases, right? The flip side of that is that you have some engineers who think they can still do it better or faster, or just want to have control over the overall pipeline. There's going to always be preference in any stack. And so we certainly see that, but we try to point to more of the objective outcomes that folks see when they set up Fivetran versus building it themselves. What ends up winning ultimately is that customers just try out Fivetran and see how great it is compared to building it. ### There are two different motions that dbt is brought into an organization. ### One of them is that somebody in the central IT org gets religion and then they push it out to the business units. ### And the other way is that the central IT org is on some different version of the world and they don't get religion and yet one of the business units does. Then they adopt dbt and like a shadow IT are constantly trying to convince the central team to support this. ### Do you see this? We definitely see that same pattern. Most of these large organizations have a central data warehousing approach. They try to have a centralized approach towards data integration. When we see that, we typically have to go through central IT. There are times where you have a central IT team that has built something, but then as you said, you have a separate team—often the marketing team—with their own warehouse doing their own thing because they are moving so quickly. We get success there and we move our way into helping the central IT teams. Every company has a different pattern for adoption, and we try to fit into as many different ways as we can. But since we touch a lot of the critical infrastructure for them, it's much harder to do shadow IT for that. There's just so much oversight on making sure that data is secure, protected, following governance, all that kind of stuff. ### You have 500 connectors now. Do you make money from connectors number 21 through 500 or is it important to say that you have a thousand connectors in a year? We do make money on the last 200 connectors we've added. Now the amount of money we make per connection is certainly lower. A lot of the systems of record for these older companies are on-premises systems. For them, a lot of the core information that they need to put into their cloud data warehouse comes from those particular systems. More and more of these companies are starting to adopt additional cloud systems around that. Maybe they'll add Workday or they have Salesforce. They'll add Jira and they start to add some other cloud systems and those also have value to them. They may not be quite as valuable as their bedrock of data that they have in these older systems. I think the newer companies don't have that on-premises system of record problem, because everything they do is in some sort of cloud system. Fivetran, for example, almost everything is in a cloud system that we decided to buy and that we run everything on top of. For us, and for a lot of our cloud-native customers, it is spread across multiple different sources. We need to have every one of those different connections. And so that's where valuable for a customer to have a single platform that they're getting all of their data. ### Even if they theoretically could buy three different tools and combine the list of connectors together, folks really want to buy one data ingestion tool, right? If you think about it, having three different tools, having three different support systems, having three different account managers, you need to train the team three different times on each of those things. And so if you can pick a single tool and be a standard across the organization, it just makes it a whole lot easier. We have a new SDK coming out that’s in private preview right now for building custom connectors. Because one of the challenges we've faced is even though we have 600 connectors and we're building 100 or 200 connectors a year, there's just endlessly more. There's 6,000-plus SAAS connections. There's something close to 30,000 APIs available and different business-to-business applications. We're never going to get to 30,000, but if our customers have a platform that they can build on top of that has 70% of the replication built in and the core functionality is there, you know, that's where I think we start to really help our customers use us as the single platform for all their data movement. ### A lot of smart people over the years have said ingestion is a commodity. But you folks are empirically proving that there is a really good, defensible business to be built in this category. How do you think about the Fivetran moat? It's a question we've talked about for years, and certainly a lot of our early investor conversations ask this same question. Like, is this defensible? Right? Why doesn't one of the hyperscalers build all of the same stuff? Our hypothesis was, which I think has turned out to be true, is that while yes, building these connections in theory is easy, in practice, there are 10,000 edge cases that happen and you only really get to a hardened state over many years and multiple different customers using the same code base where you run into all these same edge cases or different edge cases. And so the defensibility is really time and bug fixes over a long period of time against the same code base. The type of problems that we run into versus the type of problems that say an Amazon might be focused on is that an Amazon engineer is focused on the kingdom that they're building within, which is their own kingdom. Ours is completely focused on everything outside of our kingdom. We don't control anything. We're just dealing with APIs and databases and all the other things that we don't own. And so it's a very different problem, and it takes a very different skillset and a very different group of engineers. And that's what we've optimized heavily on over the last 12 years. It's a combination of who's on the team and then also just a never-ending bug fix. When you set up a Fivetran connector, it's been hardened by thousands of customers and it's going to work. ### George wrote a blog post called, “[How Do People Use Snowflake in Redshift?](https://www.fivetran.com/blog/how-do-people-use-snowflake-and-redshift)” ### It posited things like maybe we don't need to use massively parallel processing engines (MPP) for everything. And maybe vendors will supply their own compute for the workloads that they're responsible for. Have you guys gotten any blowback from this? So far, we haven’t gotten that much blowback from it. George is obviously extremely bright and he has a insatiable appetite for reading and thinking about technology. I think the combination of both of those ends up leading to him being quite a visionary thinker. He thinks a lot about the data space, a lot about all the way down to the database level. There are a lot of people who are like, “Hey, we need to use a cloud data warehouse because we have so much data.” But when you look at the actual data and the amount of data on average being run in Redshift, for example, it's not that big. Our laptops have improved significantly over the last five years, 10 years. They can probably run a lot of this compute at the same or faster speed without spending any cost, right? And so these are controversial observations because they go the opposite direction of what we've been saying for a long time. Ultimately, we care about what is right for customers and where the industry is going. And if something's right for our customers, even if we don't want it to happen, it's going to happen. And so it's better to just face reality and figure out how to live within this new world. Data lakes are a great example of what he was talking about in that blog post. He said you can use your own laptop instead of using a cloud-based warehouse. ### I wanted to use the blog post to talk about Fivetran Data Lake Service. Can you tell listeners what that is and how is it’s different from the way that Fivetran worked in the past? When we first started, it was integration with Redshift. We’d just take your Salesforce data and put it into Redshift in a very automated way. This includes the first sync of data, creating the tables, putting it all into your warehouse, and updating all that data. We effectively own that first layer of data within say Redshift. And then Snowflake came along. The big innovation there was the separation of compute and storage, where you have this now elastic ability to grow both compute and storage within the cloud. And that was really the advent of the modern cloud data warehouse. I think that’s 100 times better than the previous version, which is the on-premises data warehouses. Many customers want to be able to use their own S3. They don't want to have to take all the data in S3 and put it into Snowflake's S3, then create on top of it. A lot of customers and people have been thinking about this for a fair amount of time. There was a previous version of just loading it into S3, which I’d call a data lake version one. This version was more like a data swamp where you just put a lot of data in and then you spend all your time trying to understand what data is in there and changing it to make it logical. The next version of this came through open file formats like Iceberg and Delta. These formats take the organizational style that you get in a data warehouse and use it in a data lake. And you have DDL statements and updates, inserts; it's organized in a logical way. So you get the best of both worlds having this organized data warehouse within your data lake. And then you can put different query engines on top of that. And there are a few things that had to happen for this evolution to happen. There were a lot of large customers who had a ton of data within data lakes who wanted to access this within downstream warehouses but didn’t want to move the data. They were already paying for the storage here once. They don't want to pay for the storage again. That customer-first approach really pushed data warehouses to now start to support this concept. And I think that also drove the innovation from the open-source Iceberg community to then build these capabilities and for folks to start to adopt them. All of these things have come together in the last year. Now customers can load data directly into Iceberg in S3 and Fivetran Data Lake Service effectively does that for them. So instead of loading into Redshift or Snowflake or Databricks directly, we can load to a customer's Iceberg instance. ### And this all relies on an open catalog, right? Are you folks using a particular catalog to support this? Yes, that’s a big part of it. Once the data is in the warehouse, then the question is, well, how do you query within Databricks, Starburst, Athena or Redshift. You need to understand the actual metadata there. And so that's where these open-source catalogs have come out. Polaris is one of them. We’ll also support Unity from Databricks. I believe this will become the sort of postmodern data stack or the modern datalake stack or something that everyone moves to over the next few years. but I think there's still a lot to sort of be figured out around how to make this more of a turnkey type offering for the ecosystem. ### Is it your experience that there are more data leaders who are Iceberg and Delta curious than those who are using it in production today? Yeah, the early adopters and the ones who drove the initial innovation is the phase that we're at right now. We are seeing a fair amount of folks using this service within Fivetran, but it's not everyone yet. I think part of that is because many people are not ready to use new technology right away. They will wait a while and then use it. In the conversations I've had this year with data leaders, they're all thinking about it. One reason is that they want to be able to use data for many different things after they move it to a certain place. There are also some costs to this. It might be cheaper for them to just load it into their own S3 bucket using cheap compute, rather than loading it directly into a warehouse. ### I really agree with what you're saying on the turnkey part of this. If you are a data engineer and you try to roll out your own Iceberg support today, it’s really non-trivial. ### We were able to ship some dbt functionality at Coalesce 2024 that you just flip a flag and all of sudden your model outputs to Iceberg. I think that’s the type of stuff that's gonna have to happen across the ecosystem to make this widely adopted, which I'm very excited about. Yeah, totally. It’s very hard to roll it yourself. I mean, it's very hard on the ingestion side and then it's very hard on the bronze, silver, gold side. There's still a lot of pieces. I think what we've done helps the first part of it. What you've done helps the second part of it. There's still more around the catalogs and all of that. I think it’ll come together and it’ll be exciting, but it's still somewhat early days. ### Snowflake popularized the notion of separation of storage and compute. I think about this as the separation of compute and compute. Just like multi-modality was never really a thing. You had to pick an engine and go all in on it because otherwise you were moving data around all over the place. And that's just not the case anymore. Yeah, totally. In one sense, it's interesting because you'd say, well, this is probably worse for warehouses like Snowflake because they're getting less lock-in, right? At the same time, I think it's better in a way because customers don't necessarily want to. ### Yeah, make that case to me. I can't see it. The customer wants to have all the data within their own data lake. It’ll force people, companies like Snowflake to innovate a lot and continue to drive that customer value in the things that customers really care about. From what I can tell, Snowflake is doing all the right things, focusing a lot on the AI layer on Cortex and building out the key functionality that customers want. If they do this right, they'll get more jobs over time. Many customers already have their own data lake strategy and asked for Snowflake to help them query a lot of data. And so you forego this old world lock-in for a new world to compete on the things that customers really care about. And that's what makes a business much more lasting. ### Let’s pivot to AI. If Fivetran is now landing data in a data lake, do you have any visibility into what people do with it? Are you able to observe folks using this in AI workloads? Only through talking with them. Using the Edison analogy, we don't know what they're plugging into their outlets. We just know they're using energy; they're using the data that we're moving through it. The AI industry for B2B is still pretty early in a way. Early on, we were building an internal chat bot. Let's make it super easy. Let's pull the data from all the different sources that we have, like Slack, our internal Wiki, our Docs, and our email and a bunch of other places. And let's just pull those in together and then make those available. We started talking to a couple of different vendors and the vendors asked us to send all our data in a CSV. And we're like, “What do mean send us your data in a CSV?” We were just so surprised that it sounded very similar early days with BI. We've found that many people thought AI was its own industry and the infrastructure for it was its own industry. And the way we think about it more is that your BI stack and your overall data platform is the foundation that you build your AI on top of. Now we have a lot of customers who have been successful in building out various AI platforms on top of the data that Fivetran delivers. And that is where I think things really start to get interesting. That’s when companies really think about them as a singular platform, like what we did for our internal chatbot. I think where a lot of people are sort of going sideways is they're not just like thinking about reliable access to their own data. The difference between what companies can do within OpenAI and what they can build with their own data is that their own data gives them a competitive advantage. That's the thing that only they can access. A lot of folks are ’t thinking about it at that level yet. ### We have not yet unlocked enough downstream use cases to make the infrastructure that both of us are powering have the level of attention on it that it needs to get to the 11 nines of reliability that S3 promises. ### One of the things that's exciting to me about AI is that it is going to drive a lot more attention onto the quality of the infrastructure that Fivetran and dbt are providing. I completely agree. Again, I think we’re still in early days where folks are still tinkering with it. Folks are investing a lot but haven't had real gains from it. And I think once it starts to get more traction over the next year or so with the actual applications that companies are building on top of data with AI, that's when the pressure starts to build around the infrastructure underneath it. And that's where it really starts to harden. I'm just not sure we're there yet. And I think that's where we are seeing folks who are building on top of Fivetran infrastructure being successful with this. I can't talk a whole lot about it, but OpenAI is building on top of Fivetran. That's a pretty good AI use case. And now there's a lot of other companies as well. It comes back to, as you said, it has to be reliable, it has to work, it has to scale out. ### We’ll often get asked about unstructured data when we're in conversations with folks on the topic of dbt and AI. My answer is generally no. People are ’t transforming data from customer call WAV files or reading PDFs. Are you playing in the unstructured data world? We're starting to. We just recently added support for PDFs. A lot of folks had a massive SharePoint with tons of emails. And that's the first step into it. AI allows you to make unstructured data more structured. You can take all of this data from your email, for example, that's quite valuable to you, and apply some of the same concepts we've done in BI successfully now. When you apply the right embedding and model on top of this unstructured data, then you can do a whole lot more with it. We're transcribing a lot of our sales calls into text to see what we can learn. And then those are fed into our internal chat bot that then helps us train and helps our internal team ask questions of like, “How does ” ### I’d bucket the things that we've talked about so far as Fivetran for AI, but there's this whole other bucket of AI for Fivetran. How’s Fivetran’s product going to change as a result of AI? Yeah, so it's funny because when AI really started to take off, we sat down with our CTO Meel Velliste, who's very smart, PhD in machine learning. And we said, “If AI is going to put us out of business, let's be the first to do it.” We built an AI app that we can point to APIs. It will read the documentation and make a full application or a full connector for us. And then we have a human who goes and looks at it, mostly an analyst versus an engineer who then reviews and tweaks t. And then that's how we're building so many of these long-tail connectors through this process. ### So I thought this was a cool new idea for your product roadmap, but you did this a year and a half ago. Another one was looking at the logs for errors. You can imagine that across 600 different connectors, you get tons of different types of error messages for all different kinds of things. And so it was really hard to surface those errors appropriately to customers within our UI when something went wrong. A lot of them were very unhelpful. A lot of them, our customers couldn't do anything about them. And so we needed to surface those errors to Fivetran internally versus externally. And this has been a hard challenge for many years. We built an AI app on top of all of our logs that then goes through, breaking them down into 51 different types. Whereas we had 350 before, but a lot of them were duplicates. It’s been hugely helpful for us to debug what's going on and make sure our support team is jumping on the right things. Humans are really good at fixing things if they know what the problem is. And machines are really good at scanning through tons of data and understanding the patterns and what's happening. I think a lot more of that will continue to happen, especially as like we add more and more data sources, and with more complexity we're really focused on making the latency as short as possible. ### You folks have been on the record over the years as being a little contrarian on streaming. Streaming often gets a lot of tension. There’s a lot of hype, that faster is always better, but there's been some scrutiny around that too. What are you folks seeing now that's making you pay more attention to latency? In general, 5-10% of organizations need streaming data where it's in real-time. There's some workload or on-the-floor dashboard that folks are looking at in their manufacturing plant or whatever. But a lot of times, executives across organizations will say they need real time. But what is the actual outcome of this? The problem with real-time streaming is that there's a really high cost. It's a ton more data. There's a lot more tooling you have to build. It's a lot more complicated. Generally what we found is that for 90 % or more of cases, micro batches work quite well down to 15 minutes, or one minute. Now we're realizing that if we can get down to five-second latencies, that may move away from needing to have streaming. Streaming may become 1% of your overall use case. Customers generally want things to be faster. We're doing the hard work to get us there. ### What’s something that you hope is true of the data ecosystem over the coming five years? I hope that the data lake ecosystem turns into the core ecosystem that people are building on top of. I think it’d be better for our customers ultimately. And I think it provides a lot of optionality and obviously the tooling has to all build around it as well. So I think, you know, in five years, that's what I’d hope for. --- --- title: "AWS re:Invent got us in the data spirit" description: "AWS re:Invent was one for the books—read all about it here." url: "https://www.getdbt.com/blog/aws-reinvent-2024-recap" date: "2024-12-06" authors: ["Jeff Mills"] categories: ["Partnerships"] --- # AWS re:Invent got us in the data spirit dbt Labs was back in Las Vegas nearly two months after [Coalesce](https://coalesce.getdbt.com/on-demand) wrapped up. This time it was for AWS re:Invent, one of the largest conferences in the desert with more than 57,000 technologists and data nerds in attendance. There was a flurry of exciting product announcements, great conversations, and an energy that’s found in few other places. We’ll get to the news that we think will have the most impact on the dbt Community in a moment, but first, we want to thank AWS for hosting, our customers for joining us and making the event special, and everyone who stopped by our booth to learn about dbt. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/466b18a75598c23427ccab134789f356444ba914-3978x4920.jpg) ## Data Mesh supporting major use cases at Moderna In October, dbt Labs announced [Cross-Platform dbt Mesh](https://www.getdbt.com/blog/introducing-cross-platform-dbt-mesh) to include cross-platform references, allowing customers to collaborate on dbt projects that are running in multiple data platforms. Fast forward just a few weeks, and this new feature is already operational at Moderna. [Sri Kamireddy](https://www.linkedin.com/in/sri-kamireddy-180b678/), Principal Cloud Architect at Moderna, joined dbt Labs co-founder Connor McArthur on stage at re:Invent to share her experiences connecting Amazon Redshift and Athena together in a single analytics workflow, governed by a unified dbt DAG. The company is entrusting dbt Cloud and AWS with ensuring vaccine supply chain resilience so that it can provide critical and timely support to its network of customers. We’re thrilled to be able to support the work that Sri and the data team are doing at Moderna. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/3977031ad990a54e654d337bc701ec2d252c23cc-5712x4284.jpg) ## AWS announces Iceberg support for Amazon S3 tables Many dbt Cloud customers use Amazon S3 for their data storage needs. [AWS announced this week](https://press.aboutamazon.com/2024/12/amazon-s3-expands-capabilities-with-managed-apache-iceberg-tables-for-faster-data-lake-analytics-and-automatic-metadata-generation-to-simplify-data-discovery-and-understanding) its support for Iceberg table formats in S3. Having the world’s largest cloud provider validate the future of Iceberg should give you the confidence you need to start deploying Iceberg. As more and more companies support Iceberg, you’ll see dbt supporting those implementations soon after. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/18258a2589d836f832cd226e4f1e7c76a0e68ab5-1280x720.jpg) ## SageMaker Lakehouse As our partners introduce more and more data structures, dbt works to support those structures. Seeing [AWS lean into the Lakehouse data structure](https://www.businesswire.com/news/home/20241203118816/en/AWS-Unveils-the-Next-Generation-of-Amazon-SageMaker-Delivering-a-Unified-Platform-for-Data-Analytics-and-AI) is exciting. Lakehouses provide cost-effective and flexible structures for a company’s data. Most importantly, we were thrilled to see our joint customer Roche highlighted in the AWS announcement. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/05b6f05bd4aedda3fb8fd826e1fa0e58b4a29e2b-1232x694.png) ## The momentum continues We can’t wait to be back next year. If you’d like to hear about our announcements, please join us for the [second webinar](https://www.getdbt.com/resources/webinars/one-dbt-accelerate-data-work-with-cross-platform-dbt-mesh) in our series digging deeper into our product announcements from Coalesce. We’re covering AWS and cross-platform mesh there. And be sure to[ join us](https://www.getdbt.com/events) all over the world as we engage with our community over the coming year. --- --- title: "Testing is not enough: Transforming data quality with Write, Audit, Publish" description: "Stop testing after the fact. Catch bad data before it hits prod." url: "https://www.getdbt.com/blog/testing-is-not-enough-transforming-data-quality-with-write-audit-publish" date: "2024-12-05" authors: ["Elias DeFaria"] categories: ["Insights"] --- # Testing is not enough: Transforming data quality with Write, Audit, Publish ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/579f960c259b04246d7cc1b2759de3cc6dbfdb65-1024x748.gif) [_This blog post originally appeared on the SDF Labs website. _](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs) Before we had frameworks like SDF and dbt, data professionals relied on ancient relics of the past known as "stored procedures" buried deep in unversioned cloud warehouses. These managed their pipeline’s SQL, but often lacked business logic enforcement or validations. The rise of data testing within frameworks like dbt brought data quality into the mainstream. However, while dbt made testing more accessible, the way tests are currently executed remains a significant weak spot in many data workflows. Most testing frameworks today only validate data _after_ it has been written to production. This might sound logical—after all, isn't that when you'd see the results of your transformations? But, consider what happens when a test fails. Downstream queries, dashboards, and models start consuming incorrect data almost immediately. In the best-case scenario, you catch the issue before it propagates too far - maybe the next query in the DAG fails or a dashboard goes down. In the worst-case scenario, a critical business decision is made based on faulty data—a costly mistake. Inevitable business logic issues combined with a reactive approach to testing erodes trust between data teams and the business consumers who rely on accurate insights. Enter the **Write, Audit, Publish** paradigm - a game-changer in ensuring data quality _before_ it hits production. ## **What is Write, Audit, Publish (WAP)?** Well, it’s certainly not a reference to a Cardi B song. Inspired by best practices in software engineering, Write, Audit, Publish (WAP) is a "blue-green deployment" for data pipelines. Instead of immediately overwriting production data, a typical WAP paradigm stages updates in a table, runs tests against this staged data, and only promotes it to production if all tests pass. The process ensures that no erroneous data ever reaches production, avoiding unnecessary downtime or downstream impact. Here’s a breakdown of how it works: 1. **Write**: The transformation runs and stages its results into a temporary non-production table. 2. **Audit**: Data quality tests are executed and the results are audited in order to identify and rectify any issues. 3. **Publish**: If tests pass and issues are resolved, the non-production table is promoted to production (ideally in a way that doesn't recompute the results). This ensures atomicity and avoids rerunning computations. By decoupling testing from production updates, WAP empowers teams to build trust in their data pipelines without the constant fear of breaking downstream systems. A traditional pipeline execution with data quality tests might look something like this, where tests are executed after the model’s data has already been updated. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/2da33d4fea614b5d205e506f7338a093f9b8fe04-4020x1488.png) Critically, this means downstream tables would still run even if the tests failed, resulting in faulty data all the way down to the data consumer be it a dashboard or embedded application. With WAP, updating _Table_A_ would look something more like this, where table A is first temporarily staged, then only updated in production if all tests pass. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/5b005457d57d09daf280ea2fa9e9eaffdd2e91ed-2524x1416.png) ## **From paradigm to reality: Meet** **_sdf build_** As of version _0.10.7-p_, SDF introduces a new command that abstracts the Write, Audit, Publish workflow: _sdf build_. With this feature, teams can seamlessly adopt WAP without worrying about the underlying mechanics. The simplicity of _sdf build_ lies in its ability to: - Automatically stage data transformations with a __draft_ suffix. - Run all configured data tests on the staged data. - Publish validated data to production _without rerunning it_. This is done through `ALTER` statements that rename the draft table and drop it in a transaction. It also works out-of-the-box with complex materialization strategies like incremental models and snapshots, so the same semantics and optimized computation are preserved between _sdf run_ and _sdf build_. Think of sdf build as a bridge between the testing revolution dbt started and the operational excellence software engineering has perfected. It brings predictability, reliability, and trust to your data workflows—all in one intuitive command. SDF build is available as of [preview release 0.10.7-p](https://github.com/sdf-labs/sdf-cli/releases/tag/v0.10.7-p) and works on native SDF workspaces in Snowflake and BigQuery. Support for running _sdf build_ on dbt projects and models is coming soon. ## **The business value of SDF Build** Implementing WAP with _sdf build_ doesn’t just improve data quality—it drives measurable business outcomes: - **Reduced Dashboard Downtime**: Dashboards reflect only accurate data, minimizing disruption to decision-making processes. - **Lower Maintenance Costs**: By catching issues before production, the need for backfilling data or debugging pipeline failures is significantly reduced. - **Increased ROI on Developer Spend**: With less time spent fixing broken pipelines and more time focused on value-add projects, developer productivity soars. - **Save Compute on Bad Data:** By short circuiting pipelines with poor data quality, unnecessary compute that would have produced unusable data is now skipped. Consider the impact on a team’s overall efficiency. With fewer firefights and improved reliability, data engineers can redirect their focus to innovation and data consumers can trust the data they’re seeing. ## **Case study: Accelerating data pipelines in health insurance** A large health insurance company had previously implemented its own version of Write, Audit, Publish using dbt. While effective, their custom solution required constant maintenance due to extra configuration in their orchestrator and came with a significant performance overhead. After migrating to _sdf build_, they achieved an **80% efficiency increase** in one of their critical DAGs, reducing execution time from **22 minutes+ to just 4 minutes and 30 seconds**. This dramatic improvement was enabled by SDF’s runtime parallelism and low overhead, thanks to its Rust-based architecture. Moreover, all of this was accomplished without additional engineering effort. The clean abstraction of _sdf build_ eliminated the need to maintain their homegrown WAP solution (thousands of error prone lines of Python), freeing up their team to focus on higher-value initiatives. ## **The end of fragile pipelines** Testing alone was a great first step, but testing in isolation isn’t enough to ensure trust in your data. Write, Audit, Publish offers a proactive, resilient approach to data quality, and with _sdf build_, it’s never been easier to integrate this paradigm into your workflows. [_Read the official Documentation_](https://docs.sdf.com/guide/basics/build_and_deployment) --- --- title: "Why analytics engineering and DevOps go hand-in-hand" description: "How empowering your analytics engineers to build data project infrastructure can help you scale." url: "https://www.getdbt.com/blog/analytics-engineering-devops-relationship" date: "2024-12-04" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Why analytics engineering and DevOps go hand-in-hand Scaling your data estate can be daunting. Organizations with many projects might find themselves in a logistics nightmare trying to maintain consistency across hundreds of data projects manually. It doesn't have to be that way. By allowing analytics engineers to leverage patterns in existing tools and embrace concepts from DevOps, you can streamline scaling your data estate on dbt Cloud while avoiding too much pressure on your DevOps engineers. Duet Technologies, an early-stage company that’s created the first provider network for Nurse Practitioner-owned practices, found they could scale more easily by bringing their analytics engineers into the DevOps world. Let’s look at how they did it. ## From analytics engineering to infrastructure Duet Technologies aims to support Nurse Practitioners (NPs) in transforming and easing access to primary care—helping people, especially in underprivileged and under-resourced areas, find NPs who can provide them with care while helping to tackle the challenges that NPs encounter while running their practices. The founding Analytics Engineer at Duet, Katie Clairborne, is an analyst turned engineer. She started her career as a business analyst working in spreadsheets before she moved on to data analysis using tools like [Tableau](https://tableau.com). From there, she was introduced to the world of engineering via the version control available in dbt Cloud. Not stopping there, she continued to learn about [DevOps and CI/CD](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud) using GitHub Actions while also becoming familiar with Infrastructure as Code with tools like Terraform. Katie believes that tools like [dbt Cloud’s Visual Code Editor](https://www.getdbt.com/blog/dbt-cloud-launch-showcase-2024#:~:text=Low%2Dcode%20development%20environment) can be very effective for learning engineering topics such as the building blocks of SQL. Further, she argues that analytics engineering isn’t just an entry point into the world of software engineering. It can even be a jumping-off point for further specialization into areas like infrastructure engineering and DevOps. Analytics engineers may not realize it, but they already have the skills they need to solve their infrastructure problems, such as scaling their dbt Cloud deployments. ## Framing the problem When you hear phrases like “multi-project deployment” or “multi-project collaboration,” [dbt Mesh](https://www.getdbt.com/product/dbt-mesh) might be the first thing to come to mind. However, multi-project collaboration requires multiple sets of cloud resource configurations. And those configurations need to be consistent. Consider a single dbt Cloud project. This project will have at minimum two [environments](https://docs.getdbt.com/docs/dbt-cloud-environments)—one for development and one for production. In practice, however, there may likely be other intermediate environments, such as a staging environment. Each of these environments will also have one or more [jobs](https://docs.getdbt.com/docs/deploy/jobs). There will probably be a CI job that runs as part of a pull request. However, there may also be a merge job that runs when the code changes are merged into main. There could even be a production deployment job that’s triggered via an API call by GitHub release. Setting up just a single project requires a significant number of steps for each environment: - Naming the project - Configuring the directory - Setting up external connections such as BigQuery - Setting up service accounts or other credentials on the external service side - Getting the authorization setup on those correctly - Connecting a git repository Each of these is also unlikely to be a single action. Each may have its own number of small steps. When Katie did this, she counted 225 steps over about 15 minutes to configure a single dbt Cloud project manually. (And she was already very familiar with dbt Cloud.) When put into the perspective of an organization that could have five, ten, 15, or even more projects, this is a _lot _of manual steps. That’s plenty of chances to make mistakes. The likelihood that an individual or team can consistently execute these correctly for every single project isn’t high. ## Solving the problem via people, processes, and technology This raises the question of how we can set up and maintain these projects in a more automated, sustainable fashion. To accomplish this, Katie relied on the three pillars of DevOps: people, processes, and technology. ### People Katie believes that organizations should give analytics engineers the opportunity to extend concepts they’ve learned from dbt. Integrating analytics engineers into a DevOps function, she says, is the most efficient use of an organization’s personnel. Those analytics engineers may not be able to solve all of those problems independently. However, they also don’t need to hire or borrow a DevOps engineer to support their analytics engineering efforts. Analytics engineers bring an innate understanding of the problem they’re attempting to solve. Their skill sets transfer well to infrastructure, even if the specific tooling may be new to them. A DevOps team provides an organization-level framework within which analytics engineers can contribute. The team can help them with guidance, best practices, and standardized workflows. ### Processes Optimization at the right time is crucial. - Optimize too early and you risk wasting time and effort building features you may end up not using. - On the other hand, not optimizing at all has a different set of downsides: teams may get better at executing all of these steps manually, but they’re simply deferring high maintenance costs. Katie recommends applying the [rule of three](https://en.wikipedia.org/wiki/Rule_of_three_(computer_programming)) and optimizing around the time of the third project. She suggests this is when teams have a decent idea of what they will benefit from. Thus, they can spend the extra time optimizing to gain dramatic improvements that’ll pay dividends in the future. ### Tools Lastly, Katie encourages using tools like Terraform to manage cloud resources programmatically. A scalable data project system needs to reduce both project setup time and maintenance costs. To this extent, ideally, we want to manage dbt Cloud resources programmatically. [dbt Cloud Terraform Provider](https://github.com/dbt-labs/terraform-provider-dbtcloud) is exactly the tool for this. It was originally developed by a community member known as Gary James, but now it’s a fully official piece of software maintained by dbt Labs. [A Terraform provider](https://developer.hashicorp.com/terraform/language/providers) is a Terraform plugin, like [adapters](https://docs.getdbt.com/docs/supported-data-platforms) in dbt. Just as you might use a [Snowflake](https://slowflake.com) or [BigQuery](https://cloud.google.com/bigquery) adapter in dbt, you can use a Google Cloud or dbt Cloud provider in Terraform. ## Using Terraform to transform dbt Cloud transformations Both dbt and Terraform need to be told how to connect to data platforms or cloud providers. In dbt, this varies depending on the product and environment being used. dbt Cloud has account-level connections for service accounts and personal credentials for developers. If developers are using the dbt Cloud CLI, they might choose to download those credentials. This is very similar to dbt Core’s [profile.yml](https://docs.getdbt.com/docs/core/connect-data-platform/profiles.yml) file. In Terraform, these connections are configured using a provider declaration block that includes an account ID, a dbt Cloud service token, and a dbt Cloud host URL: provider “dbtcloud” { account_id = var.dbtcloud_account_id token = var.dbtcloud_token_id host_url = var.dbtcloud_host_url } Once you’ve set up the provider, you need to configure your source code as a dependency. Just as before, both dbt Cloud and Terraform do this in similar ways. In dbt, you use the dbt deps command to install packages. Likewise, Terraform provides a terraform init command that installs providers and modules. If you look at the file structure of both tools, they’re very similar: a subfolder for installed dependency code, a lock file for consistency, and a configuration file. When defining resources in Terraform, you declare data sources as inputs—just as you’d use sources in dbt to create [models](https://docs.getdbt.com/docs/build/models). Similarly to previewing the model SQL in dbt before running that model, you can use the [Plan command in Terraform](https://developer.hashicorp.com/terraform/cli/commands/plan) to see the infrastructure changes that will occur when you run the Apply command. Likewise, both dbt and Terraform produce artifacts in the form of files that describe their state at a moment in time. These are the [manifest.json](https://docs.getdbt.com/reference/artifacts/manifest-json) or the [Terraform state file](https://developer.hashicorp.com/terraform/language/state), respectively. All of the resources are listed as well as various attributes about them. Both also have commands for building graphs and state comparison, as well as executing over only changes elements. ### How this helps with scaling dbt Cloud If we go back to our description of environments and jobs within a dbt Cloud project, we could have a TF file for environments —one for jobs and one for project-level resources. However, if we’ve got a large number of dbt Cloud projects, each with its own code repository, imagine how this could be handled. While you could copy these files into each repository, consider what would happen if you needed to make a broad-scale change as an organization to the way these projects are configured. You would need to go into each repository and make the same changes over and over, then validate that it was made correctly. In dbt, you can use [macros](https://docs.getdbt.com/docs/build/jinja-macros) with arguments to avoid repetition. You can do effectively the same thing with [modules](https://developer.hashicorp.com/terraform/language/modules) in Terraform. Modules provide a way to create multiple related resources. Given the example below: module “dbtcloud_project” { source = “modules/dbtcloud_project” project_name = “coalesce” } The specified source module defines all the resources a dbt Cloud project needs to get started. The project name is provided as a variable. Each team can then call this module in their repository. This creates a pre-configured general structure that the entire organization can use while retaining the flexibility to configure a subset of elements differently if needed. When Katie did this, the manual procedure that previously took 15 minutes was completed in just_ under a minute_. ## Conclusion This isn’t even everything that’s available with dbt Cloud for DevOps automation. There are things like [dbt-jobs-as-code](https://github.com/dbt-labs/dbt-jobs-as-code), which uses YAML files to define dbt Cloud Jobs. [dbt Cloud Terraforming](https://github.com/dbt-labs/dbtcloud-terraforming) generates Terraform configuration files from existing dbt Cloud configurations, making it easier to incorporate Terraform. By treating dbt Cloud infrastructure as code, we simplify both project setup and project maintenance. Creating an easy way to tear down, adjust, and recreate those projects can be incredibly valuable. Additionally, you can transparently maintain consistency across projects. Curious to see how you can scale your data operations with dbt Cloud and DevOps? Watch the full presentation to learn more. [Watch video](https://youtu.be/6les3dVoYh0?si=rFoa2IkMfxTNZEsY) --- --- title: "How SurveyMonkey sharpens dbt with data observability" description: "How reactive monitoring with Monte Carlo and proactive data management with dbt proved to be a useful combination." url: "https://www.getdbt.com/blog/surveymonkey-monte-carlo-dbt" date: "2024-12-03" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # How SurveyMonkey sharpens dbt with data observability SurveyMonkey is a company that lives and breathes data. For a while, its primary focus was getting data in front of business users. Over time, however, it realized multiple teams across the company were having similar—and, in some cases, unexpected—issues with data. Let’s take a look at how data quality and data processing both began as pain points across SurveyMonkey—and how the company’s data engineering team used a combination of reactive monitoring with Monte Carlo and proactive data management with dbt to yield measurable improvements. ## SurveyMonkey’s data landscape SurveyMonkey is the leader in global online forms and surveys. With over 25 years of experience, they’ve helped organizations of all sizes, from startups to Fortune 500 companies, answer close to 88 billion questions. Over 3,000 organizations look to SurveyMonkey for insights on their customers, employees, and products. Which means they process a _lot_ of data. To make smarter business decisions, you need to know what you don’t know. SurveyMonkey helps you **ask**, **listen**, and **act**, transforming customer insights into valuable, actionable strategies. The magic happens behind the survey, in the company’s backend data landscape. SurveyMonkey imports data from multiple data sources into Snowflake Enterprise. It runs 180+ workflows, including data transformation pipelines powered by dbt, to transform, tailor, and mine meaning from all this data. SurveyMonkey also uses Airflow for orchestration and Monte Carlo for data observability. The company’s dbt footprint is sizable. It maintains **600 models** and **2,000 test cases**, pulling data from close to **900 data sources**. ![SurveyMonkey data landscape](https://cdn.sanity.io/images/wl0ndo6t/main/8658d57926becaf71444af7723bb3adbb2997f24-1278x647.png) ## The data problems SurveyMonkey unearthed (with a survey, of course) When SurveyMonkey’s data engineering team first tackled data quality and data governance, its focus was building its dbt models. It focused more on **availability** of the data (getting it into the hands of business users) versus the **quality** of the data. Soon, however, the team began receiving pointed requests from a diverse group of stakeholders: - **Executives** wanted data pipelines to run early so they’d have scorecards in the AM - **Marketing** wanted to monitor and track the performance of their campaigns in real-time against benchmarks and goals - **Directors** requested proactive detection and notification via Slack of anomalies discovered by Monte Carlo Instead of fixing these problems immediately, the data engineering team asked itself: Are other groups in SurveyMonkey having the same issues? Data engineering team manager Samiksha Gour sought to answer that question with—what else?—a survey. The results were illuminating. When asked what the common data challenges were in their organization, 53% cited data quality issues (inconsistencies, inaccuracies, etc.). Which Gour expected. What was surprising, however, was that 50% cited data **processing** issues. Many, many teams were concerned about the performance of their data pipelines. ![Common data challenges](https://cdn.sanity.io/images/wl0ndo6t/main/b3849b3fc0c78d54bd960308a0effb8835f83e46-1250x565.png) ## How Monte Carlo + dbt enhanced data quality To tackle these issues, SurveyMonkey integrated two tools it was already using: Monte Carlo and dbt. SurveyMonkey leverages dbt for data transformation. It uses [dbt models](https://docs.getdbt.com/docs/build/models) to clean raw data from various sources to create high-quality, usable data sets. The company leverages [dbt tests](https://docs.getdbt.com/docs/build/data-tests) to verify data quality and [dbt documentation](https://docs.getdbt.com/docs/build/documentation) to create consistent, well-documented data models that serve as single sources of truth. The company employs [Monte Carlo](https://www.montecarlodata.com/) to monitor these dbt models and pipelines. Monte Carlo’s monitoring capabilities detect data anomalies, ensuring data accuracy and maintaining data governance. SurveyMonkey integrated these two tools in several different ways: ### Standardized quality checks dbt tests became mandatory for every model. ### Regular performance reviews Rather than abandon a model after pushing it to production, the team conducted regular performance reviews and optimizations for all dbt models. ### Automated monitoring and alerts for Monte Carlo-detected anomalies and failed dbt tests The team provided proactive notifications that something was wrong with data vs. waiting for the issue to show up on a scorecard. ### Converting Monte Carlo anomalies into dbt test cases For anomalies where the data team didn’t want the pipeline to proceed, they’d create a test that stopped it from running. Business users were informed of the problem and were responsible for addressing the issue in the upstream data source. That meant business users could unblock themselves instead of waiting on data engineering to unblock them. ### Leveraging dbt documentation with Monte Carlo Asset SurveyMonkey used a combination of manual documentation and AI-assisted suggestions [generated by dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot) to help document their 6,000-some tables for business users. Users would then search [Monte Carlo Asset](https://docs.getmontecarlo.com/docs/assets), the company’s single pane of glass for data assets, to filter down the available tables based on the documentation. ## How dbt and Monte Carlo improved SurveyMonkey’s business After putting these measures in place, SurveyMonkey measured their impact on the business. One of the most impressive results is the data engineering team **reduced its Snowflake credit usage by around 73%** for close to 10,000 credit jobs. A key driver was performance monitoring, which led the team to fix and merge queries, redo models, and delete unused models and tables. In particular, the team used performance monitoring to identify long-running update jobs doing cross-joins, simplifying the SQL statements that powered them. ![Data transformation efficiency over time](https://cdn.sanity.io/images/wl0ndo6t/main/43834baefde6910d77c74c93ee5ab5e406894b67-1278x576.png) A second benefit was that SurveyMonkey could scale while keeping its costs stable. Despite adding many more dbt models (e.g., by bringing a marketing analytics media platform in-house), the team’s job execution time remained stable while its cost per credit in Snowflake steadily decreased. A third benefit was a marked increase in data quality. In March 2024, when the team brought in the third-party marketing analytics platform, it saw a spike in data anomalies reported by Monte Carlo. This was because the team itself didn’t fully understand this new data. It needed time to work with it, grok it, and shape its data pipelines to produce healthy, accurate data. Remember, however—the team had a principle of converting Monte Carlo anomalies into dbt test cases. As its dbt test cases spiked, the anomalies started to drop. ![Converting frequent Monte Carlo anomalies into dbt tests](https://cdn.sanity.io/images/wl0ndo6t/main/f87f335655affb7f99ec24a4e4ed205028ec01a1-1284x570.png) Finally, SurveyMonkey saw its regular performance reviews pay dividends. In one case, as a result of obsoleting unused models and revamping inefficient queries, **it saw a 94% reduction in pipeline runtime and a 97% reduction in Snlowflake credit usage**. ## Lessons learned The data engineering team admits that their journey, at times, wasn’t a very smooth ride. However, they learned some valuable lessons along the way: ### Monte Carlo and dbt is a winning combination Using two tools that worked well together worked in the team’s favor, unlocking wins they might not have secured otherwise. ### Don’t introduce anomalies to business users at the beginning of the project It doesn’t make sense to send anomaly notifications to business users when the data engineers themselves still don’t understand the data. Take a month and wait for the data pipelines to stabilize before you involve business users. ### Stop running pipelines on bad data A stopped pipeline gets everyone’s attention. If you detect an important anomaly, make a test case for it. Then, you can stop the pipeline and restart it once your business users fix the upstream source. ### Onboard new users to the system from the get-go Specifically, SurveyMonkey now trains new business users to leverage Monte Carlo for incident routing and table research. ## What’s next for SurveyMonkey’s data journey? Using dbt and Monte Carlo, SurveyMonkey’s data engineering team successfully tackled two key data problems simultaneously: data quality and data processing times. In doing so, it decreased its total data processing time and costs—which enabled the company to handle even more data. The team isn’t stopping there. Some key improvements it plans to make include: ### Enhanced tool compatibility Exploring future updates to bridge the gap between dbt and Monte Carlo features—specifically in handling dbt snapshots effectively. ### Expanded monitoring capabilities Specifically, developing and implementing more custom monitors. ### Using data domains A [data domain](https://www.getdbt.com/blog/data-domains) is a logical grouping of data, along with all of the operations it supports. They enable domain teams to own and operate their own data while interacting with other teams via a [data mesh](https://www.getdbt.com/blog/what-is-data-mesh) architecture. SurveyMonkey aims to leverage data domains to identify further areas of data improvement. ### Stakeholder tracking Implementing a system to track stakeholder trouble tickets and monitor their progress closely. To get more insights on SurveyMonkey’s use of dbt with Monte Carlo, watch the full presentation from Coalesce 2024. [Watch video](https://youtu.be/eoqUH14sshk?si=ZzvjJWfCGmnEgI5U) --- --- title: "November dbt Community update" description: "Stay updated with the dbt Community. In November we had an AMA with Erica Louie, 10 Meetups, community awards, and more." url: "https://www.getdbt.com/blog/november-dbt-community-update" date: "2024-12-02" authors: ["Kathryn Chubb"] categories: ["Community"] --- # November dbt Community update Welcome to the November edition of the dbt Community Update, your monthly roundup of all things happening in the [dbt Community](https://www.getdbt.com/community). This month was packed with events, achievements, and incredible stories. Highlights included an insightful AMA with Erica Louie, the announcement of our [dbt Community Award](https://docs.getdbt.com/community/spotlight) recipients, five in-person [dbt Meetups](https://docs.getdbt.com/community/spotlight), and a ton of great discussions across Slack. Are you ready for the recap? Let’s get started. ## Community Slack AMA Each month, we host a live Ask Me Anything (AMA) event in the #dbt-community-merge channel on Slack, bringing together analytics engineers, data enthusiasts, and industry experts for open and honest conversations. This month, we featured Erica Louie, Senior Analytics Manager, EPD at dbt Labs. As one of dbt Labs' earliest employees (back when we were Fishtown Analytics), Erica has witnessed firsthand the evolution of dbt, and she brought her rich experience to a discussion on data careers, team-building, the dbt Semantic Layer, and scaling analytics projects. Here are five takeaways from Erica’s AMA: **1. Pathways into data careers** Erica shared her unconventional journey from art to analytics, highlighting how data appealed to her as a flexible, creative career. She emphasized that anyone can succeed in data with curiosity, a knack for storytelling, and the ability to interpret patterns. Key advice: cultivate empathy, stay transparent, and invest in relationships. **2. The power of hard and soft skills** While technical expertise is a must, Erica stressed the equal importance of soft skills in data roles. She discussed how effective communication and empathy can transform data teams from “order-takers” into trusted consultants who drive impactful business decisions. **3. Navigating leadership and burnout** Erica opened up about stepping back from her role as head of data and hiring her own manager—a bold move that underscored the value of self-awareness and prioritizing well-being. Her message? Leadership requires empathy for both your team and yourself. Check in regularly to stay aligned with your values and find joy in your work. **4. Scaling data teams and tools** From the dbt Semantic Layer to the Cloud CLI, Erica highlighted strategies for building scalable data ecosystems. She underscored the importance of measuring data team success through metrics like uptime, cost efficiency, and dashboard impact, while staying grounded in your role’s purpose. **5. Finding balance between service and strategy** Data teams often juggle executive requests with long-term projects. Erica’s advice? Build trust by delivering on immediate needs while carving out time for innovative, strategic work. And in a lighter moment, Erica shared her adventures with her converted van, Bessie—a reminder to embrace passions outside of work. ### Catch the full conversation Missed the AMA? No problem. Watch the full recording here: [Watch video](https://youtu.be/YwOxoTQGcjM?si=y-0lT277HmeRzqaU) ### Get ready for another AMA in January [Join us for the next AMA](https://www.getdbt.com/resources/webinars/community-ama) in January. Register now to get the link to watch live and join the [#dbt-community-ama channel in Slack](https://getdbt.slack.com/archives/C025ZN1L679/p1716308109072499) to participate in the conversation. ## dbt Community award recipients We are excited to share more about the dbt Community Award recipients announced at Coalesce in this quarter's dbt Community Spotlight. The [Community Spotlight](https://docs.getdbt.com/community/spotlight) is where we highlight folks from around the world who have gone above and beyond to contribute for the benefit of others. ![dbt Communty award recipients at Coalesce](https://cdn.sanity.io/images/wl0ndo6t/main/e3b4683de653a22a2b913feee45c08b61b47fb32-4000x2667.jpg) dbt Community Award recipients embody what it means to be an advocate within the dbt community, exemplifying collaboration, expertise, creativity and leadership. This round, we are featuring Christophe Oudar, Ruth Onyekwe, the original dbt-athena maintainers (Jérémy Guiselin, Mattia, Jesse Dobbelaere, Serhii Dimchenko, and Nicola Corda), Mike Stanley, Meagan Palmer, Bruno de Lima, Opeyemi Fabiyi, and Jenna Jordan. Visit the [Community Spotlight](https://docs.getdbt.com/community/spotlight) page to learn about their backgrounds, their plans to grow as leaders, and their experiences—both learning from others and sharing their own knowledge. If you’re interested in being selected for future rounds of the Community Spotlight, learn more about [becoming a contributor](https://docs.getdbt.com/community/contribute). ## November dbt Meetups In November we had 10 dbt Meetups: Taipei, Seattle, China (online), Halifax, Bratislava, Boston, New York, Taipei again, Brisbane, and Hasselt. These events continue to be a cornerstone of the dbt Community, bringing members together to share knowledge, network, and collaborate. Stay tuned for even more Meetup opportunities in December by checking out our [Meetup page](https://www.meetup.com/pro/dbt/). ![Photo from Meetup](https://cdn.sanity.io/images/wl0ndo6t/main/5336af3ca587960bac786591704faf9d8b2bda8c-2048x1313.jpg) ![Photo from Meetup](https://cdn.sanity.io/images/wl0ndo6t/main/33bde14f58dfe395d8c709e956d8c739a31a2302-2048x1536.jpg) ## How has dbt transformed your career and your life? One of the highlights of Coalesce this year was the “How has dbt transformed your career and your life?” wall. Attendees filled it with stories of growth, opportunity, and life-changing moments—so many that we had to clear the wall several times to make room for more. ![How has dbt transformed your career and your life wall](https://cdn.sanity.io/images/wl0ndo6t/main/7b569dce1e79d43433498e891d25acc1e0adcf58-2854x3461.jpg) These reflections showed that dbt is more than a tool; it’s a catalyst for transformation. From landing dream jobs to feeling empowered in their roles, the wall was a powerful reminder of the impact our community has on real lives. Seeing these stories reinforced what makes the dbt community so special—empowering each other to do meaningful work and grow together. ## Community announcements We’ll wrap up this month's update with some of the exciting announcements that are regularly posted in our [#announcements](https://getdbt.slack.com/archives/C0VLZM3U2/p1715777392876319) channel on Slack. ### 2025 State of Analytics Engineering survey The 2025 State of Analytics Engineering survey is live. As the data landscape shifts, we're at a pivotal moment of introspection. How do we ensure our work stays compelling, aligned, and impactful? How do we keep data products at the forefront of driving value for our organizations? That’s exactly what the 2025 State of Analytics Engineering Survey is here to uncover. This is your opportunity to share your team's pains, gains, and strategic investments—and help us capture the pulse of data teams worldwide. [Take the survey](https://docs.google.com/forms/u/1/d/e/1FAIpQLSe8hp6evq_Qr78b2gQZsZAxXTTI-y5sPhwtgyj4594oWESFbQ/viewform?usp=sf_link). ### Upcoming events - [Register for our next Community AMA](https://www.getdbt.com/resources/webinars/community-ama) in January - [Register for our upcoming webina](https://www.getdbt.com/resources/webinars/one-dbt-accelerate-data-work-with-cross-platform-dbt-mesh?utm_medium=social&utm_source=linkedin&utm_campaign=q4-2025_post-coalesce-recap-dec_aw&utm_content=____&utm_term=all_all__)r: ​​One dbt: Accelerate data work with hybrid deployments and cross-platform dbt Mesh - Ongoing [Cloud Demo with Experts](https://www.getdbt.com/resources/dbt-cloud-demos-with-experts/) in North America, EMEA, and APAC-friendly times - Add [dbt Events](https://www.addevent.com/calendar/Tb314369) to your calendar ### Upcoming dbt Meetups We’ve got a busy month coming up with six [in-person dbt Meetups](https://www.meetup.com/pro/dbt) scheduled. If you’re looking for opportunities to learn with fellow members of the dbt Community, and have fun while doing so, join us at one of the sessions listed below: - 🇧🇷 São Paulo | Tuesday, December 3rd, organized by Bruno Souza de Lima and Thales Donizeti - 🇳🇴 Oslo | Wednesday, December 4th, organized by Glitni - 🇩🇪 Berlin | Thursday, December 5, organized by Eva Schreyer and Victoria Perez Mola (dbt Labs) - 🇩🇪 Cologne | Thursday, December 5, organized by Hicham Babahmed and Stephan Durry (dbt Labs) - 🇨🇭 Zurich | Thursday, December 5, organized Astrafy - 🇹🇼 Taipei | Friday, December 13th, organized by Karen Hsieh, Laurence Chen, Allen Wang, and Damon Liao There are so many exciting things going on in the dbt Community, and we can’t wait to see you all there. If you haven’t yet, [join the community](https://www.getdbt.com/community) today. --- --- title: "From dbt Core to dbt Cloud: Why Warner Brothers Discovery made the switch" description: "dbt Core worked well for Warner Brothers Discovery. Moving to dbt Core brought them even better results." url: "https://www.getdbt.com/blog/warner-brothers-core-to-cloud" date: "2024-11-27" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # From dbt Core to dbt Cloud: Why Warner Brothers Discovery made the switch Warner Brothers Discovery moves data. A _lot_ of data. And they’ve used dbt Core to manage it for years. However, a sudden spike in data demand revealed multiple weak points in their architecture. Here’s why the company moved from dbt Core to dbt Cloud, how they did it, and the benefits they reaped from the move. ## An end-to-end first party data platform Warner Brothers Discovery already knew the value of its data. To this end, they created a common data platform—a place where all of their businesses warehouse their data. In addition to this common platform, some of their businesses also had their own respective data platforms. One such is the CNN data platform called Zion, lovingly named after the city in _The Matrix_, which serves as a first-party news analytics platform. Zion is a three-layered platform containing a collection layer, an information layer, and a service layer. The collection layer is effectively a complex data ingestion pipeline that runs in AWS. The service layer comprises the actual consumers of the data and associated products, built using a combination of dbt Core and Airflow. This platform is fully-featured—it sports everything from an SDK to a dashboard—and is designed for resilience and scale. And when they say scale, they mean _scale_: Zion moves around **100TB of data** in a day, can quickly **scale 15x**, and ingests about **a billion events every 24 hours**. A platform of this size presents a lot of problems. Since it’s a news analytics platform, it’s designed to handle breaking news. That means it has to support identity resolution and stitching, behavioral events, profiling, and content metadata. Plus, it has to present data in a way that makes sense to its businesses. ## Why dbt? The two main things Zion powers are WBD’s content analytics and its Machine Learning (ML) recommendation system. The company chose dbt Core to build this system because dbt is modular and automated. dbt Core enables insight into what is being processed, how, and from where. This transparency has enabled WBD to test and validate the decisions they’re making. Moreover, dbt readily integrated with common engineering practices like WBD’s [Continuous Integration and Continuous Development (CI/CD) pipelines](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud), further smoothing adoption. The magnitude of their data was also no problem, requiring no esoteric changes in [Snowflake](https://www.snowflake.com/) to dynamically change a test model’s provisioning. A simple configuration change was all the team needed. Lastly, but certainly not least, dbt Core is also very cost-effective—it basically doesn’t cost anything, but it also didn’t cause cost spikes for, e.g., the team’s SQL queries against their data warehouse. ## WBD’s journey with dbt Although they started creating Zion in 2017, WBD adopted dbt Core only in 2021. The company then quickly scaled up its use for CNN+. In 2024, it adopted dbt Cloud to help empower its data stakeholders. In 2025, it plans to leverage it to enable those stakeholders to use the tools and cloud platforms they want to use. With dbt Core, WBD is running **450 models**, with about half of these covered by **1,600 tests**. Most of those models run every hour and, even under anomalously high loads on a high-traffic day, only experience about 10 minutes of variance in model run time—instances like [when a tanker hits a bridge](https://edition.cnn.com/us/live-news/baltimore-bridge-collapse-03-26-24-intl-hnk/index.html) and people keep rewinding, then rewatching the moment the bridge was struck. WBD’s architecture is a typical [medallion-like architecture](https://docs.getdbt.com/guides/optimize-dbt-models-on-databricks?step=3), with multiple [DAGs](https://www.getdbt.com/blog/guide-to-dag) supporting multiple time aggregations. This way, each team can have data processed for them in different functional areas. ## Why dbt Cloud? The description above already seems like a big win. If WBD was already so successful with dbt Core, why move to dbt Cloud? The answer is that, over time, WBD’s data needs grew—often with sudden, huge spikes in demand from unforeseen events such as Queen Elizabeth II’s Funeral. With each of these spikes, the company realized its dbt Core architecture faced multiple challenges. These challenges included: ### Inconsistencies in job performance Sometimes jobs resulted in unpredictable, unreliable outcomes that also increased costs. ### Infrastructure management and scaling Difficulties implementing distributed processing, load balancing, or horizontal scaling. ### Direct dependency on engineering teams for business analytics WBD had no support for a [data mesh architecture](https://www.getdbt.com/blog/what-is-data-mesh), an approach to data engineering in which data domain teams can create, launch, and manage their own data sets as [data products](https://www.getdbt.com/blog/key-components-of-data-mesh-creating-and-managing-data-products). This hobbled grass-roots data efforts. Stakeholders required continuous engineering support to launch their own data products. **** ## How dbt Cloud helped Moving to dbt Cloud enabled WBD to solve these issues in a variety of ways. The monolithic nature of their 450 models in dbt Core not only impeded the independence of their stakeholders. It was also challenging to find concepts or semantics. Breaking this into smaller conceptual chunks using [dbt Mesh](https://www.getdbt.com/product/dbt-mesh) immediately improved this and also made it easier to manage dependencies between projects. Decentralizing data ownership and management allowed for better scalability by incorporating distributed processing, load balancing, and horizontal scaling. While this improved reliability under higher demand, WBD also implemented fault-tolerance mechanisms such as automated retries and error handling. All in all, the more efficient resource allocation and utilization of cloud services also resulted in lower costs. WBD faced many challenges at an organizational level with which this move helped: ### Increasing stakeholder adoption The move to dbt Cloud enabled them to provide self-managed projects that are built from a centrally managed platform without the need for constant engineering support. ### Enhancing developer onboarding WBD improved the flow of onboarding engineers and the experience of their analysts by creating robust deployment tooling and project scanning using dbt Cloud’s [scheduler](https://docs.getdbt.com/docs/deploy/job-scheduler), [environments](https://docs.getdbt.com/docs/dbt-cloud-environments), and [environment variables](https://www.getdbt.com/blog/introducing-environment-variables-in-dbt-cloud). ### Support through tooling While easy deployment is important, ongoing support is also a common source of stress for developers. By reducing the redundancy in their code, WBD could create centrally controlled projects with common macros to solve common problems. ### Standardize observability and alerting By utilizing dbt Cloud’s [webhooks](https://www.getdbt.com/blog/introducing-webhooks-in-dbt-cloud), the team could send Slack notifications directly to the engineering team to enable them to mitigate problems as soon as possible. ## How WBD migrated to dbt Cloud Zion is a large and important platform. WBD aimed to minimize the impact to the existing platform while maximizing the effectiveness of the migration result. To this effect, the team performed the migration in steps: 1. Create a POC with something small and easy to understand, with minimal impact. 2. Isolate that process or pipeline from the Zion ecosystem. 3. Examine what did and didn’t work. 4. Begin building out macros and project tooling. 5. Pick the next smallest and easiest-to-understand piece. 6. Keep some relation to the first isolated piece so the mesh can be tested, then build up and standardize for observability and maintainability. 7. Finish moving the rest of the processes and pipelines. 8. Lastly, instrument for observability and maintainability. Since simplifying the architecture of Zion was also a primary goal of this move to dbt Cloud, WBD took the opportunity to increase the code legibility. Simultaneously, it made the infrastructure easier to use and maintain by implementing [Infrastructure as Code (IaC)](https://aws.amazon.com/what-is/iac/). Using tools like Terraform, they streamlined things with single-command, push-button project deployments. Another benefit of this architecture is that the data engineering team could swap out the underlying data sources without impacting their stakeholders. They can now run cost metrics and projections and switch to whatever resources are most optimal for a domain team’s use case. In other words, the domain team doesn’t need to worry about making those decisions themselves. Domain experts can focus on launching products using their domain expertise while the engineering team optimizes their data performance under the hood. **** ## Choosing the right orchestration WBD found that there were instances where it made sense to continue to use Airflow for orchestration, while there were other instances where dbt Cloud jobs excelled. dbt Cloud jobs worked best for: ### Last mile delivery Models that are primarily concerned with refreshing dashboards or providing endpoints. ### Complex run environments In particular, projects requiring dynamic incremental loads or dynamic variable settings. ### Non-standard incremental processes Processing done over fixed incremental windows. ### Backfilling I.e., processes where data needs to be replayed or rerun for specific periods. Additionally, WBD used Terraform, Terragrunt, and [AWS CodeBuild](https://aws.amazon.com/codebuild/) to reduce the number of deployment scripts and automate tasks to eliminate manual intervention. ## Migration benefits With the move to Infrastructure as Code using dbt Cloud, Warner Brothers Discovery was able to see many gains: - Launch new projects more quickly using quick-start templates to deploy new projects on dbt Cloud - Automatically set up alerting and observability - Integrate new projects with the existing Zion ecosystem using dbt Mesh - Create fully automated deployments with Terraform, Terragrunt, and Code Build From the overall transition to dbt Cloud, WBD saw the following positive impacts: - Implemented a multi-project framework by adopting a data mesh architecture, which improved data organization and made the entire system easier to understand - Adoption of the platform by their Data Analytics Research and Testing team (DART team uses the platform to explore lineage and develop their own data products) - Improved run times and cost efficiency - in some cases, by up to 75% - Enhanced monitoring and alert systems (Slack and switchboard warnings) - Visibility into migration progress and refined roadmaps - Improved documentation and lineage maps As WBD looks to the future, it aims to support multi-compute. Whether it’s Databricks or Snowflake, they want their platform to enable platform-agnostic functionality out of the box. Using dbt Mesh, they can finally accomplish that goal. ## Empowering stakeholders, streamlining systems Moving from dbt Core to dbt Cloud and dbt Mesh enabled Warner Brothers Discovery to reduce costs while increasing the number of domain teams who could create and manage data products without engineering assistance. The result was better stakeholder adoption, higher reliability, and less time spent babysitting data deployments. Find out how dbt Cloud can accelerate your digital transformation and streamline data operations—[contact us for a demo today](https://www.getdbt.com/contact). Watch Warner Brothers Discovery’s session at Coalesce to learn more about how they moved from dbt Core to dbt Cloud. [Watch video](https://youtu.be/uUnAqHLg9IM?si=CLiRhEnGbv94pxcY) --- --- title: "How Virgin Media O2 streamlines data operations with dbt Cloud" description: "How newly-merged Virgin Media O2 rebuilt its data infrastructure in the cloud while improving customer loyalty." url: "https://www.getdbt.com/blog/virgin-media-o2-dbt" date: "2024-11-26" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # How Virgin Media O2 streamlines data operations with dbt Cloud Virgin Media O2 faced several data-driven challenges all at once—including shifting their focus from products to customers and pulling off a migration from on-prem to the cloud. The data team achieved some initial successes using dbt to migrate to Google Cloud Platform (GCP). However, a ride-along one day with customer support revealed that support technicians weren’t getting the data they needed when they needed it. Customers paid the price in terms of long wait times. Fortunately, after some experimentation, the company found that dbt Cloud made it easy to lower data processing times. The result was more loyal customers—and at a lower cost to boot. Here’s how they did it. ## Long customer support times Virgin Media is a broadband television and home phone services supplier. [O2 UK](https://www.o2.co.uk) was the largest mobile network provider in the UK. In 2021, these companies joined forces in the largest UK telecom merger to date to create a new venture, Virgin Media O2. Over time, Virgin Media O2 noticed a shift in their customers' desires. Years ago, new products drove consumers to purchase new service contracts. The company noticed that, once technology had met the needs of the consumer, they had less incentive to purchase new products. ‌This led Virgin Media O2 to shift their focus from products to customers. This was only one of the transitions the company embarked on to adapt to the needs of its customers. Another was a digital transformation of their e-commerce platform. An initiative to move their on-premise solution to Google Cloud Platform (GCP) was the perfect opportunity to prove their customer-focused direction by improving their [net promoter score](https://www.surveymonkey.com/mp/net-promoter-score-calculation/), a standard measure of customer loyalty. The company decided to take a dual approach. A group of hundreds of contractors would focus on doing a lift and shift of the existing platform into the cloud. Meanwhile, it organized an internal team of four people to rewrite the solution from the ground up to be cloud-native. The internal, cloud-native team achieved their goal within a year. Throughout the year, they worked alongside customer support to identify improvements. One key improvement came from a “go and see day,” where data engineering rode along with support. The engineers observed that, since the merger, a regular support call would require information from multiple systems and pages. This, in turn, led to longer support calls and longer wait times for the customer. The engineering team identified the opportunity to reduce the amount of time it took to support a customer by transforming the data and providing support with a single interface that included all of the data needed by the support team. To provide even more value to the customer, the interface would also include data-backed recommendations for upgrades and services that the support team could recommend to customers. ## Reducing five-hour data pipeline run times The team envisioned their work as building modular data backed products and models that service Artificial Intelligence (AI), Machine Learning (ML), and MLOps. They achieved this by consolidating data from various sources into Google Big Query, then running data transformation jobs to produce purpose-built data sets. One of these data-backed products supported the needs of the customer support team. The data transformation job to produce the data needed initially began as a daily run that took five hours to complete. They saw large spikes in usage of Google Big Query around the time the transformation job would run. That was expected. However, it also represented a lot of waste in the form of reprocessing and waiting. The team aimed to reduce the amount of data being processed to only what they needed and also to reduce the amount of idle time. The team determined that they should move to an incremental data processing system that's designed to handle only data updates, such as new or modified data. They evaluated two different approaches for this: push and pull. ### Push system A push system would kick jobs off every hour to do the work in increments. The team found some issues with processing their data in this way to achieve their goals: - **Timing**: If the hourly run processes the data in less than an hour's time, that’s wasted processing time, which the team wanted to reduce. If a run takes longer than an hour, this causes delays in processing and potentially re-processing. That, again, adds to waste. - **Manual Intervention of Issues: **When an issue in processing does arise, there’s a risk of losing data and having to make code changes to recover a point in time. Improvements in data processing would add to wasted processing time. However, running the jobs too close together would exacerbate potential issues caused by failures. ### Pull system A pull system continually pulls data to be processed. When a job has finished its processing, it triggers another job to start. This solved the timing issue the team saw in a push system, as the data flow was continuous and not wasting any processing time. This also solved creating code changes to recover lost data, as failures in the processing wouldn’t trigger the next job. **** ## How dbt Cloud enabled the new pull system To pull of re-processing of data efficiently, the team turned to incremental models in dbt Cloud. In dbt Cloud, an [incremental model](https://docs.getdbt.com/docs/build/incremental-models) is a table in the data warehouse. Upon the first model run, it transforms the entirety of the source data. On subsequent model runs, it only processes rows from the source data that have been created or updated since the prior run. This solves the problem of over-processing by only transforming changes. Circular referencing of the data model allows the data to flow continuously by batching the processing of the incremental data and then retriggering another run. This was achieved by using the [{{ this }} Jinja function](https://docs.getdbt.com/reference/dbt-jinja-functions/this) in their dbt models. This enabled: - Using the reference of the latest timestamp of processed data in starting a new job. - Using the data pipeline to keep track of what data has been processed and where to instruct the next job to begin processing. This solves the problem of wasted processing time, as data was now flowing continuously. The team used data segmentation to allow for sections of the data to follow a batch processing model. At the same time, it allowed for the data needed for just-in-time processing to flow continuously. Key stages of the data pipeline will always receive full data refreshes. However, models that don’t require this can use the incremental delta tracking process to reduce the amount of data processed and the amount of time it takes to deliver the result. ### More customer loyalty at less cost The team quantified their success in stages: 1. Initially, when the entire data set was being processed daily as a batch job, the section of data the team was optimizing processed 8TB of data per run, which took 47 minutes apiece to complete. 2. In the next phase, using incremental models in a pull system, a run that processed 130GB of data took 140 seconds to complete. 3. Finally, implementing delta tracking in conjunction with the incremental pull system, runs could process 32GB of data in 50 seconds. The result of these changes equated to £100 savings per run or £36.5K per year. The team also fulfilled its goal of increasing its net promoter score, which rose by 26%. **** ## Transforming data to transform business The Virgin Mobile and O2 merger led the data engineering team to take a huge risk. Rather than reproducing the existing infrastructure in the cloud, they aimed to rebuild the system from scratch for better performance. The team hoped this would open up support for AI, ML, ML Ops, and Large Language Models (LLMs), all of which require low-latency data model production. The goal of re-platforming the data infrastructure was to provide stability and standardization. That allowed the team to innovate quickly and focus on building in quality over the course of the journey. This journey started with following a traditional approach to data transformation by batch processing the data daily. That gave them stability and also standardized their processes. Once stable and standardized, the team could rapidly innovate to produce specific models that would serve as products to their consumers. If a team wanted a portion of their data model, such as the example given around increasing the value and speed of the support team, they could focus on reducing the latency by employing just-in-time data processing. In the end, using features built into dbt and dbt Cloud, Virgin Mobile O2 accelerated its time to market while also improving customer satisfaction and overall data quality. Find out how dbt Cloud can accelerate your digital transformation and streamline data operations—[contact us for a demo today](https://www.getdbt.com/contact). Watch Virgin Media O2's session at Coalesce to learn more about how they streamlined operations with dbt Cloud. [Watch video](https://youtu.be/iY7rbutVwm0?si=h1lKjqNCjofEnhq-) --- --- title: "Five tips and tricks for getting the most out of dbt Explorer" description: "How many of these features did you already know about?" url: "https://www.getdbt.com/blog/five-tips-and-tricks-for-dbt-explorer" date: "2024-11-25" authors: ["Roxi Pourzand", "Alexis Jones"] categories: ["Product"] --- # Five tips and tricks for getting the most out of dbt Explorer Spending your working hours putting out fires or fielding never-ending tickets is a short path to burnout. Luckily, [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) is designed to help you move beyond reactive workflows and offers the holistic context and breadcrumbs you need to build, fine-tune, and troubleshoot your pipelines in a more proactive way. **When you fire up dbt Cloud, dbt Explorer should be your starting point.** It’s a knowledge base that offers a visual representation of all of the metadata across your dbt pipeline. You can double-click into any node to get detailed context about that data asset, its dependencies, its freshness, who it's used by, how to improve it, and more. Using dbt Explorer, you have a launchpad to easily understand how your data models are interconnected, what data products they inform, and which models or nodes may need your attention, with the necessary insights on how improve them. > “dbt Explorer makes it easy for our consumers to understand the entire lineage from the source to reporting—and all of the data quality checks or issues along the way—without having to go ask a dev." _- Robert Goodman, Lead Developer of Enterprise Data Analytics at Lennar_ Since launching dbt Explorer a year ago, we’ve been steadily shipping a lot of new functionality to make it the best place to navigate, understand, and improve your dbt projects. This post will dive into five tips and tricks for getting the most out of dbt Explorer so you can build trust in your data products, ship with more confidence and velocity, and tune your projects so you can optimize costs. ## Tip #1: Find the exact resource you need by mastering search filters and selector syntax dbt Explorer has advanced search capabilities that make it easy to find the exact resource you’re looking for; saving developer time, reducing siloed workstreams, and enforcing model consistency. ### Search Let’s say I’m a data analyst at a B2C company that’s looking to understand how many of our transactions get routed through our call centers. I know we have column in our data that tracks whether a call center was part of the transaction cycle, but I don’t know exactly what the column is called or which data source it lives in. This discovery exercise is simple with dbt Explorer. When I enter “call” in the search bar: - I am rendered a long list of resource names, column names, resource descriptions, warehouse relations, and code that match my search criteria. I can easily narrow this list down by using the `Column name` filter. Now, I’m able to see all relational nodes that contain that column in their schemas. ![Explorer search detail](https://cdn.sanity.io/images/wl0ndo6t/main/76ecc5e1f7243648bb3d2af7687492384097b1fd-2272x884.png) - I find the relevant column I’m looking for, `CALL_CENTER_SK` , and see that it’s tied to our public “transaction” model that is exposed. Using the lineage lenses at my disposal (read more about these in Tip #4), I can see this is a healthy model that’s often consumed by many stakeholders, and therefore I know this has the data I’m looking for. ![CALL_CENTER_SK](https://cdn.sanity.io/images/wl0ndo6t/main/f32c8387fe27b7063742c32fa36ed0733260e575-1333x487.png) Without dbt Explorer, this exercise becomes a lot more tedious. I’d have to search and sift through potentially hundreds of tables and their schemas to confirm which is the correct column. Then I’d have to run a number of test queries to see if the column looks correct. Finally, I’d likely have to ask someone on the data team to confirm which table is good for consumption. ### Selector syntax Seemingly unruly lineage graph? Explorer has you covered with logical selector syntax based on the same selection syntax used in the dbt CLI. You can filter the lineage graph in the same way you would filter resources during a dbt invocation. Let's go through some of the basics. Just like running a single model with `dbt run -s my model`, you can search for that model using its name. ![Selector syntax](https://cdn.sanity.io/images/wl0ndo6t/main/46b71ce77f0c25f743ff4c501051ed5ef3e6827b-2694x600.png) You can also use graph operators to see the model's lineage; to see everything just one step downstream of a resource model, use “+1” after the resource name, like `transactions+1`. If you want to see additional dependencies _past_ one, you can add +2, +3, and so on. ![Graph operations](https://cdn.sanity.io/images/wl0ndo6t/main/51e72d367615044faa6ec1ce5407c355bc851523-2696x1818.png) If you want to see one step _upstream_ of a particular resource, simply append “1+” before the resource name, for example `1+transactions` . You can build on this logic to show one step upstream _and_ downstream of a resource with the syntax `1+transactions+1`, [and so on](https://docs.getdbt.com/reference/node-selection/graph-operators). If you want to see specific resource _types_ only, you can use the “resource_type” specification, for example `resource_type:model` . This goes for any resource type in the DAG (model, sources test, seed, snapshot, exposure, metric, semantic_model, macro, group). !["rescource_type"](https://cdn.sanity.io/images/wl0ndo6t/main/158e9785ae3e15c9df2f4eff16046e1f42ec8cc1-2700x1818.png) Similarly, you can use the + operator to see additional levels of lineage for the resource type. For example, you can see everything one step down from every source in your dbt project using `resource_type:source+1` ![+ operator](https://cdn.sanity.io/images/wl0ndo6t/main/9725adbcea8aa0b81e8ec6bd38da37ce083e0f65-2698x1838.png) We support a whole host of selector methods and you can find auto-suggested selectors in the lineage search bar. To read more about [search](https://docs.getdbt.com/docs/collaborate/explore-projects#example-of-keyword-search) and [syntax selector](https://docs.getdbt.com/reference/node-selection/syntax) in dbt Explorer, see our docs. > “I always have dbt Explorer up on one half of my screen to be ready to answer questions about where data is coming from or how a column is defined. It's a tremendous time saver for me.” _- Brian Gillet, Director of Data and Analytics at Hazel Health_ ## Tip 2: Discover cross-project assets and view lineage from a single pane of glass Oftentimes, organizations manage multiple projects within one dbt account. For example, there may be a project for the marketing team, another one for the central data team, and a third for the finance team. And with [dbt Mesh](https://www.getdbt.com/product/dbt-mesh), these projects can reference each other to promote better collaboration and governance while keeping code DRY. dbt Explorer offers a few ways to discover and understand cross-project assets and lineage. ### Find all public models Using dbt Explorer, it's easy to discover all of the public models across your account. Using this view, you can get an at-a-glance understanding of what projects use these models, who owns them, and then dive into lineage with a single click. ![Public models in account](https://cdn.sanity.io/images/wl0ndo6t/main/7fd3c6b3ecfbf00ec00b029d0147951c3dd66f49-3508x884.png) ### Navigate cross-project lineage It’s also easy to view lineage for more than one project side-by-side in dbt Explorer. If your data pipeline includes a reference to another project, simply double click that project’s node in your lineage graph. ![cross-project lineage](https://cdn.sanity.io/images/wl0ndo6t/main/5b08efe76728f94329fe5bd65ca0444c092de115-2964x2002.png) When you do, we’ll automatically render a new tab with _that_ project’s lineage (zoomed into what should be most relevant based on your flow) so you can visualize both graphs side-by-side. ![side-by-side graphs](https://cdn.sanity.io/images/wl0ndo6t/main/51b075eac5fb0693cb8a78a7da5bf2d9c3564b36-3832x2132.png) This feature keeps you in your flow, giving you an intuitive way to view downstream or upstream dependencies used by other projects. This keeps your DAGs manageable while still providing a useful comparative view into your dependencies ## Tip #3: Build trust with detailed context into what dashboards—and teams—your models power All of the data that you curate in dbt models is in service of your business initiatives, and now, native in dbt Explorer, you have the detailed context needed to connect the dots between a data model and the business value it drives. Two new interrelated features—auto-exposures and model query history—make this happen. ### Auto-exposures With [auto-exposures](https://docs.getdbt.com/docs/collaborate/auto-exposures) (now available for Tableau, coming soon for other BI platforms), you can have your lineage graph automatically build out downstream Tableau dashboards. This gives data teams automatic context into how and where models are used, so they can prioritize data work to promote data quality. Coming soon, you can trigger downstream dashboards to automatically refresh as soon as new data is available, giving business stakeholders confidence that they’re always making decisions from the freshest data. ### Model query history You can also easily layer in a “lens” across your DAG to build context for each node on things like materialization type, freshness, and—as shown in the image below—model consumption. By quickly grasping the relative popularity of your models, you have a data-driven “to do” list of where you should allocate your engineering resources, as well as built-in empathy for the consumers behind the models you build. ![auto exposures with model consumption lens](https://cdn.sanity.io/images/wl0ndo6t/main/3765797b835a350cd4f673626594b180378681f8-3164x1710.png) Having this context is critical to ensure that you keep your business-critical dashboards up and running. After all, no one wants to get a frantic call the morning of a board meeting that the dashboard is broken. With these details at your fingertips, it’s easy to do one better and actually _optimize_ the inputs that drive that dashboard so you can continue to build trust with your stakeholders. > “Leveraging the model query history feature in dbt Cloud has transformed our approach to optimization. It empowers us to gain deep insights into our SQL execution, identify performance bottlenecks, and enhance our data models with confidence. We're excited to see how this feature evolves in the roadmap ahead." _- Gary How, Data & Analytics Architect at Kenvue_ ### Data health tiles A bonus feature for building that trust is the turnkey ability to embed health tiles that provide trust signals like data freshness and data quality directly in those downstream dashboards. This gives your stakeholders confidence in the data they’re about to use, and empowers them with the transparency they need to trust the data you provide. ![Health tile](https://cdn.sanity.io/images/wl0ndo6t/main/2926b7aeb743d6e0b74b6d4f9d3af0386a85e440-1878x1312.png) ### In-app health signals These health signals are also accounted for throughout the in-app dbt Cloud experience, giving _developers and other dbt users_ an at-a-glance understanding of whether the model they’re about to use is fresh, error-free, tested, documented, and more. ![In-app health](https://cdn.sanity.io/images/wl0ndo6t/main/88ac65cacf1ce240c8ffa9cde3c3807728c1d549-2132x684.png) ## Tip #4: Supercharge your context and debug faster with lineage lenses Using [lineage lenses](https://docs.getdbt.com/docs/collaborate/explore-projects#lenses), you can visualize your lineage graph from a number of different parameters—beyond the default of resource type—so you can grok critical details that help you build more resilient, efficient pipelines and debug issues faster. These lens overlays include: - Model layer (staging, intermediate, marts) - Materialization type (table, view, materialized, incremental, ephemeral) - Model execution status (success, fail, error, warn, skipped) - Test status (pass, error, fail, warn, skipped) - Column-level evolution to see how your columns change across the pipeline (passthrough, transformed, renamed, etc.) - Model query history (actual metric for the last 30 days and visual representation for high, low, medium of consumption queries against the models) (highlighted in Tip #3 above!) With this layered context about your data estate at your fingertips, it becomes much simpler to understand pipeline issues (test status, column-evolution lens), know which resources to spend development time on (model query history, model execution status), and identify ways to simplify and streamline your pipeline (model query history, materialization type). ![lineage lenses](https://cdn.sanity.io/images/wl0ndo6t/main/86d6055c602b68b78bd579de12f8052c133b9faa-1336x820.png) ### Column-evolution lens You can also use lineage lenses to easily visualize how your _columns_ evolve across your DAG. When debugging a pipeline issue, it’s a huge timesaver to be able to quickly understand whether a column has simply been reused, or if it in fact has been transformed somewhere in your pipeline. Having this transparency of what transformations are happening at the column level allows analysts to easily confirm how a particular column evolves throughout the pipeline. This helps users understand how a particular data point was calculated or helps developers and analysts decide how they should extend the model for further analysis. Additionally, when columns in a source table are modified, renamed, or removed, column-level lineage shows exactly which downstream tables, views, or dashboards will be affected. This visibility allows teams to proactively anticipate and plan for downstream impacts of changes, preventing disruptions in analytics or reporting. ![Evolution lens](https://cdn.sanity.io/images/wl0ndo6t/main/ba0bb39e17f73d481f20954618305b5db7fc7f5e-3702x1930.png) Overlaying this context onto your lineage graph with lineage lenses is a powerful sidekick as you build, troubleshoot, analyze, and improve your data pipelines. Having this context at your fingertips helps you understand how data evolves from ingestion to analysis, promoting data quality without compromising velocity. Check out this demo video to learn more about lineage lenses in dbt Explorer. [Watch video](https://www.youtube.com/watch?v=wdxtzNujVx0&list=PL0QYlrC86xQlwGtQsDjWhmBIRYedXSSCT) ## Tip 5: Fine-tune and tidy up your data estate with project recommendations Another great way dbt translates your project metadata into actionable insights is with [project recommendations](https://docs.getdbt.com/docs/collaborate/project-recommendations) in dbt Explorer. No one likes an urgent fire drill alerting you to an outage or quality issue in your data pipeline. You can get ‌ahead of these potential issues with project recommendations that surface proactive ways to improve the test coverage, documentation, and overall project health of your dbt models. You can filter the list by severity (high, med, low), improvement category (documentation, performance, testing, etc.), and rule names. ![Project recommendations](https://cdn.sanity.io/images/wl0ndo6t/main/a86eaefc4763f65e85b59b18b50d8aae22538563-2170x1466.png) With these insights at the ready, data teams can proactively tackle improvements to the performance and quality of their data projects. The end result is more resilient pipelines and better trust with stakeholders…done in a way that’s proactive, data-driven, and manageable for data teams. > “dbt Explorer is an indispensable ally for any data-driven organization aiming for excellence in their analytics workflows. We gained valuable insights into project data quality and adherence to dbt best practices. It not only helped us pinpoint areas for code enhancement but also significantly improved our documentation practices. We achieved substantial enhancements in data quality percentages, effectively mitigating data errors in the bronze/silver layer and ensuring a higher standard of data quality for our end consumers.” _– Shravan Banda, Solutions Architect at World Bank_ ## Get started with dbt Explorer today dbt Explorer is generally available to all dbt Cloud customers. Given the amount of new features we’ve shipped into Explorer this past year, we'd be curious: how many of these features are you using today? Hopefully, this post inspired you to fire up Explorer and discover how it can help improve your workflow. It’s really easy to get started. Just navigate to the “Explore” tab in dbt Cloud. Maybe start by assessing which models are most popular, or seeing what additional tests or documentation you can build (or better yet, have [dbt Copilot](https://www.getdbt.com/blog/introducing-dbt-copilot) build them for you!) to better tune your projects. We also offer a [free online course](https://learn.getdbt.com/learn/course/dbt-explorer/introduction-to-dbt-explorer/introduction-to-dbt-explorer) on getting started with dbt Explorer where you’ll get hands-on instruction on how to use many of these features. --- --- title: "Maturing as an analyst alongside the ADLC" description: "Learn how analysts can mature alongside the ADLC with practices that improve accessibility, velocity, and correctness in workflows" url: "https://www.getdbt.com/blog/maturing-as-an-analyst-alongside-the-adlc" date: "2024-11-20" authors: ["Rachael Gilbert", "Alex Talbott", "Paige Berry", "Erica Louie"] categories: ["Insights"] --- # Maturing as an analyst alongside the ADLC In the recent post about the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle), we outlined a framework for evaluating the maturity of data workflows. We broke the process down with guiding principles across eight stages of the ADLC. ![ADLC loop](https://cdn.sanity.io/images/wl0ndo6t/main/1fc981ff485ca62fb80c5c9d4bde3789904514f8-2400x1260.jpg) This subject matter is vast, and we have only started to scratch the surface in aligning on what maturity means across these phases. This model is intended to serve as a springboard for debate, discussion, and deep reflection. Today, we reflect on what it means for us analysts in particular. Per the ADLC, a mature analytics workflow must manifest features such as **accessibility**, **velocity**, and **correctness**. What are the table stakes in our analyses that ensure this? How do we keep our exploration nimble but our output mature? How do we ourselves evolve as professionals alongside our systems? Tristan states that in the ADLC: > > “There is no ‘right’ or ‘wrong’ way to conduct exploratory data analysis. Rather, it specifies a set of requirements that all users should have”. Here we share some of the requirements we've found useful in our own work as analysts at dbt Labs. ## Why the analysis phase is challenging If you have ever worked as an analyst or with one, you have likely (and intimately) felt the friction in building toward the following characteristics. ### Accessibility Analysis can be extremely tough to communicate at the **right** level with the **right** context to the **right** people. Things get missed, glossed over, or misconstrued. On top of that, code complexity (or let’s be honest, plain bad code) can also outweigh any benefit of building further on the work. ### Velocity Long-term, dashboard overhead continues to bloat (and bloat!). Time is spent reactively, either fixing charts or trying to answer the "this looks wrong" Slack message, instead of pushing our organizational knowledge forward. Given the interconnected nature of the ADLC, short-term velocity gains from tech debt upstream (e.g. a feature release without proper tracking) can often hit analyst velocity too (additional time spent on logic workarounds). ### Correctness The cost of those maintenance and upstream data quality issues hurt not only velocity but also correctness. Much of the analyst's time is spent validating their work, often at odds with velocity. Do we move quickly or do we move accurately? ## Here's what we do about it In order to mature our analyses, our team continuously leverages the best practices checklist below. This isn't, in any way, exhaustive. However, we believe that when conducting an analysis, data practitioners should be able to answer the below seven questions. This checklist helps us maintain focus on the right things, making our work easier to leverage (**accessibility**), faster to turn around (**velocity**), and less prone to error (**correctness**). And those are some of the hallmarks of a mature analyst. ### 1. Do we understand the question behind the question? It feels natural to take a question at face value. We’ve learned the hard way that a question can be very subjective, often meaning something different to the asker than to the listener. Say a PM asks for a deep dive on feature X retention. Maybe this is tied to a very specific retention definition in a company OKR, or maybe they're just more broadly trying to get at feature value or a customer segment. What is at the actual heart of it, what do they want to achieve? Are we aligned on how it'll be used to make decisions? As our head of data recently wrote about, always “[start with the why](https://roundup.getdbt.com/p/a-compass-not-a-map)”. ### 2. Is that question the right starting point? The stakeholder is ready to go, and we know exactly what they want. But wait—are we positive that makes sense? We need to ensure we're confident in the landscape surrounding the ask. We have a different perspective than our stakeholders, and it never hurts to check their assumptions. This is a critical step. This may feel obvious to a seasoned analyst who does this implicitly, but explicitly making this part of your process ensures consistency across the team. What if that feature X was only rolled out to new users so far? In that case, we may want to take a step back and examine the broader funnel first. Ignoring the behaviors of existing users could skew overall retention takeaways. ### 3. What small steps can we take to break down complexity? Fight against bloat. If it sounds urgent, if it sounds complex, or if it adds to maintenance overhead, we need to make sure it is fully understood upfront. If not, rescope and reprioritize. Are we sure it can’t be addressed by existing work? What exactly needs to happen when? **Humans are wrong all the **time, whether it be on priority, premises, reasons, or solutions. Over decades of analytics work, we have learned it never hurts to **pause, simplify, and then iterate**. The more we can break down big work into small steps, the less likely we are to spin cycles accidentally going in the wrong direction. [A mature ADLC does not mean perfectionism](https://roundup.getdbt.com/p/a-compass-not-a-map). For example, maybe a quick correlation is a good sanity check before a full-fledged mixed effects model. Maybe ensure that the PM is monitoring the top metric they requested before building an entire dashboard. Our backlog is our friend. Don’t ignore all the asks and tangents, but prune and prioritize them ruthlessly. ### 4. Are we explaining our work enough? Right now, our lovely new analysis is fresh in our heads. Excessive code comments may feel…excessive. But will they feel that way when we have to update this work a year from now? What about if other people want to borrow our logic? We must be clear in our assumptions and our approach as we write our code. We must write out why we are doing things that may not be obvious to others. We can’t always assume a shared perspective. If the work isn't maintained in an ongoing manner or if it’s just scratch work, it **never** hurts to make that explicit. It's easy for us to stumble across work that looks like the logic we want, but it can be hard to know if we should trust it. ### 5. Have we established sufficient confidence in our findings? Do we feel confident that we can trust these results? Any survivorship biases we need to reflect on? Do shared averages break down in time or main subgroups in the way we’d expect? Do shared percentages have sufficient volume to back them up as trends? If there are known limitations to our confidence, **make sure **that's** conveyed**. It can be very useful to regularly communicate if this is a high, medium, or low confidence analysis; this is critical for building trust. Alongside that, flag if this isn’t maintained so the company knows how long-lived that confidence may be. ### 6. Are we communicating the takeaways in an optimal way? For sharing our work, we need to make it visible and make it easy. Some guidelines that we like to keep in mind here ([if you want to go even deeper, check out our past work](https://locallyoptimistic.com/post/share-your-data-insights-to-engage-your-colleagues/)): - We go ‌where our stakeholders are. This often means not our BI tool but rather Notion, Slack, etc. - Before we spin up yet another dashboard (YAD ™️), are we sure there isn’t an existing one that'd make more sense to add to? Fight against that sprawl. - We try to communicate regularly, even if the status update is a boring “still working on it”. We also often boost the end results more than once (e.g. in Slack, plus in our meetings, plus in our data newsletter). A million things are happening each day, and people get busy. Ensure visibility. - We get to the point. Rather than just sharing a link or a novel, we remind ourselves to frame everything with a highlighted “why is it important, and here’s what to do with it” upfront. No one else is thinking about the nuances of our work as much as we are. That being said, for those who want them, we have details at the ready in links and dropdowns. This makes it more efficient for our varied audience to get the varying levels of detail that they need. ### 7. Can we leave the campground cleaner than we found it? Keeping scope creep in mind, what did we learn from our project? Is there anything worth converting into a win around scale or efficiency? Any small tweaks we should make in favor of our vision of accessibility, velocity, and correctness? For example, were we surprised there wasn’t documentation on something? Is it ‌easy to add this back in, since we now have learned it? Did we use a cool approach or learn something insightful that’s worth sharing with our team? ## Developing people and processes alongside systems The ADLC framework outlines in very broad strokes [a path towards data maturity](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). We’ve already started to dig into [what that means in terms of pitfalls and focus areas](https://roundup.getdbt.com/p/a-compass-not-a-map). Here, we dig further into the analysis phase, focusing on a few key aspects of maturity (such as accessibility, velocity, and correctness) as they relate to analysis workflows and thought processes. More to come as we introspect further on evolving in the analysis phase of the ADLC, and you can always catch us in the [dbt Community Slack](https://www.getdbt.com/community/join-the-community) to discuss. We’re all on this journey together. --- --- title: "Deliver reliable, personalized embedded analytics with the dbt Semantic Layer" description: "Learn how the dbt Semantic Layer can help you bring trustworthy, personalized analytics to any end-user experience." url: "https://www.getdbt.com/blog/deliver-reliable-personalized-embedded-analytics" date: "2024-11-20" authors: ["Jordan Stein", "Chakshu Mehta"] categories: ["Product"] --- # Deliver reliable, personalized embedded analytics with the dbt Semantic Layer Too often, valuable data is stuck in internal dashboards or buried in databases, out of reach for customers, partners, and decision-makers. While it powers internal insights, it rarely delivers value externally. Today’s data consumers want more. They want personalized dashboards that show their specific purchasing trends, reports with individualized performance metrics, and they want these insights wherever they consume data. These experiences build trust and drive engagement, but they’re challenging for data teams to deliver. Personalization at this level often requires significant time, effort, and resources. Custom dashboards, for example, require aligning metrics, designing APIs, and coding to ensure reliable backend data for front-end visualizations. It’s a complex process, and mistakes easily slip through the cracks and impair customer experience and trust. ## The challenges of traditional BI for personalized analytics Traditional BI workflows for delivering personalized analytics often add complexity, relying on fragmented processes that slow development and increase the risk of inconsistencies. Data preparation is handled in one tool, where datasets are selected and organized to meet visualization requirements. Custom APIs are then created in another tool to retrieve the necessary data. Finally, the data is manually assembled, parsed, and bound to visualization libraries. This disjointed workflow not only slows progress but also introduces errors at every step. This disjointed approach introduces inefficiencies and risks errors at every stage. For example, changing a report from monthly to yearly aggregates requires updates to data preparation, APIs, and the front-end—delaying projects and increasing maintenance costs. Then there’s the challenge of **metric consistency.** Traditional BI tools often store metric definitions in silos, leading to discrepancies and confusion about which metrics are accurate. These inconsistencies erode trust in data, frustrate users, and complicate decision-making across the organization. ## Centralize your metrics and deliver them anywhere To unlock the full potential of your data, businesses need a centralized, governed source of truth for metrics—one that ensures consistency and can be easily embedded across multiple applications. The [**dbt Semantic Layer**](https://www.getdbt.com/product/semantic-layer) makes this possible. It simplifies and accelerates embedded analytics development by enabling data teams to define and manage complex metrics centrally. From there, teams can deliver reliable, version-controlled, and personalized analytics to downstream tools quickly and cost-effectively. > "The dbt Semantic Layer gives our data teams a scalable way to provide accurate, governed data that can be accessed in a variety of ways—an API call, a low-code query builder in a spreadsheet, or automatically embedded in a personalized in-app experience. Centralizing our metrics in dbt gives our data teams a ton of control and flexibility to define and disseminate data, and our business users and customers are happy to have the data they need, when and where they need it."** **_ > - Hans Nelsen, Chief Data Officer at Brightside Health_ ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/60e1d7a421e57cde65752331aeb30f05ada0c96e-1296x1024.png) **** ## How the dbt Semantic Layer powers embedded analytics Here’s how the [dbt Semantic Layer for embedded analytics](https://www.getdbt.com/product/embedded-analytics-use-cases) helps you overcome common challenges and deliver reliable, personalized analytics without the hassle of repeatedly reconfiguring your backend. ### 1. Centralize metric definitions [Define your metrics and logic once](https://docs.getdbt.com/docs/build/about-metricflow), including complex metrics such as advanced calculations, aggregations, or time-based logic. Reuse them across your tools or web apps. This ensures consistent and accurate data, eliminates conflicting metrics, and provides a governed foundation for reliable reports and visualizations personalized for any user. ### 2. Speed up embedded visualization development Building custom endpoints for each visualization can be a bottleneck. With the dbt Semantic Layer, standardized data models, SQL generation, and developer-friendly APIs eliminate the need for bespoke backend infrastructure. This streamlines development, enabling teams to focus on delivering engaging user experiences that delight end users. ### 3. Streamline queries in downstream tools With our [developer-friendly APIs and SDKs](https://docs.getdbt.com/docs/dbt-cloud-apis/sl-api-overview), you can dynamically generate filtered views and user-specific dashboards without creating separate endpoints for each use case. For example —using our beloved Jaffle Shop workflow— let's say you're the owner of the Jaffle Shop Pennsylvania region. Using the dbt Semantic Layer, you can effortlessly display all purchases in Pennsylvania or provide a customer with their order history over time. This approach simplifies development, reduces complexity, and ensures cost-efficiency as your data needs grow. For organizations managing complex data workflows or serving diverse user bases, this is a game-changer—it simplifies maintenance and enables seamless updates to metrics or logic without disrupting downstream applications. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/e3a595eb1f9b9e6ebefa1388c5b20d86ef0da878-2416x1506.png) ### 4. Improve analytics performance The dbt Semantic Layer uses several layers of [query caching](https://docs.getdbt.com/docs/use-dbt-semantic-layer/sl-cache) to increase performance of commonly run queries. This means your system can deliver responses faster without reprocessing the same data repeatedly. By reducing the strain on your database, query caching not only speeds up response times but also ensures scalability as user demand increases. This keeps your analytics fast and reliable. **** ## Accelerate the development of personalized analytics with the dbt Semantic Layer By adopting the dbt Semantic Layer for your embedded analytics, you can bypass cumbersome workflows and fragmented development cycles. Instead of managing changes across multiple data models, APIs, and front-end components, you gain a centralized, streamlined system that simplifies data management and accelerates development. Discover how customers are using the dbt Semantic Layer to transform their analytics. Watch our Coalesce session with Bilt Rewards [‘Making data rewarding at Bilt’](https://www.youtube.com/watch?v=6vQ_VsXDGFc) to see how they deliver efficient, trusted analytics in their customer-facing apps. Or, check out our on-demand webinar [‘Scaling embedded analytics with Brightside Health’](https://www.getdbt.com/resources/webinars/scaling-embedded-analytics-with-dbt-semantic-layer-virtual-event) to learn how they provide personalized analytics to both patients and providers, empowering better medical care. If you're looking to simplify your analytics workflow, deliver consistent metrics, and deliver personalized embedded analytics fast, try the dbt Semantic Layer for your embedded analytics strategy today. To get started building out your dbt Semantic Layer configurations, check out our [documentation](https://docs.getdbt.com/docs/use-dbt-semantic-layer/quickstart-sl), and reach out in Community Slack (#dbt-cloud-semantic-layer) if you have any questions. [Watch video](https://www.youtube.com/watch?v=6vQ_VsXDGFc) --- --- title: "1000X faster SQL linting" description: "Boost performance, consistency, and readability across your data pipelines." url: "https://www.getdbt.com/blog/1000x-faster-sql-linting" date: "2024-11-19" authors: ["Lukas Schulte"] categories: ["Insights"] --- # 1000X faster SQL linting ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/5b5a8b0f14ec08bcfef34d9301cee4ca92b312db-1138x628.gif) Hygiene, readability, consistency, and correctness of SQL code are some of the many barriers to moving forward with self-serve data. Whether it be convoluted tables with a growing number of columns, or a 1,000+ line unformatted chain of SQL CTE’s conspicuously named “dim_person”, many data teams struggle to scale while keeping engineering efficiency high. If a smart analyst who writes SQL for a living can’t understand the SQL behind their data pipelines, you’ve got a problem. Code is read more often than it is written. In software development, Linters and Formatters have been standardizing code quality, development expectations, and readability for decades. They help engineers focus on logic rather than layout. They automate consistency, and streamline collaboration. Today we’re excited to announce high performance SQL linting and formatting with SDF providing 100X - 1000X performance improvements over the existing standard. **The result: linting and formatting become virtually zero-cost operations in daily development and in CI/CD.** Both the Linter and Formatter have common sense defaults so that configuration is minimal and out of the box behavior should work for most engineering teams. SDF lint and format are supported for all SDF projects (including macros!) and non-templated general SQL. Support for dbt projects and dbt templating is coming soon. **_As of SDF release v0.10.0-p we introduce SDF native SQL linting and formatting functionality for with up to 1000X performance increases over SQLFluff in large SQL projects._** In SQL development the current de-facto standard for linting is [SQLFluff](https://sqlfluff.com/) which provides sensible rules and a high degree of configurability. Unfortunately, SQLFluff has severe performance limitations (due to its python runtime), and only provides high-level syntax reviews rather than feedback driven by a deep semantic understanding of SQL. SDF’s SQL Linter and Formatter have a high degree of compatibility with SQLFLuff but are based on SDF’s own dialect-specific all-Rust SQL parsers and highly parallelized visitor algorithms. This results in incredible performance and out-of-the box compatibility with every SDF workspace. In fact, SDF lint is so fast that it is primarily limited by the time needed read files from disk! ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/c8323cd32e34ab3ae98d9ee6db3a8833b27aad0d-1274x352.png) All tests performed on 7950X3d (16 cores, 96GB Memory) with SDF Default rules. SQL pulled from [Gitlab](https://gitlab.com/gitlab-data/analytics/-/tree/master/transform/snowflake-dbt), [Snowflake Sample TPC-DS](https://docs.snowflake.com/en/user-guide/sample-data-tpcds), and internal SDF workspaces. ## Benefits of SDF Lint Linters catch common mistakes (like syntax errors) before code is executed. They enforce coding standards and help maintain a consistent style across an organization. By increasing uniformity, code reviews and collaboration between engineers becomes easier and more fluid. SDF Lint offers major improvements to the developer experience, and to organizations keen to standardize and unify their SQL code. 1. SDF lint is remarkably fast and accurate, with massive performance improvements over the current standard SQLFluff. SDF’s linter is written in Rust, highly parallelized, and underpinned by proper ANTLR grammar definitions for supported SQL dialects. 2. SDF lint is easy to use and integrated into every release of SDF. There are no package dependencies, python virtual envs, or integrations to manage. 3. SDF Lint supports all Jinja macros and configuration (variables) within an SDF workspace. 4. SDF Lint strives to be compatible with SQLFluff for syntax rules. SQLFluff aliases are provided for SDF rules where possible. Check out the linter rules as they compare to SQLFluff [here](https://docs.sdf.com/linter/overview#rules-reference). 5. SDF Lint can be used independently of SDF’s transformation layer on raw SQL. All that’s needed is an SDF workspace specifying SQL file include paths or directories, and optionally a linting configuration. ## Easy configuration and sensible defaults Every SDF workspace now implicitly includes the linter configuration below. _`workspace: name: my_workspace ... defaults: dialect: snowflake --- # This the default lint configuration sdf-args: lint: > -w capitalization-keywords=consistent -w capitalization-literals=consistent -w capitalization-types=consistent -w capitalization-functions=consistent -w references-quoting -w structure-else-null -w structure-unused-cte -w structure-distinct -w convention-terminator`_ To modify the linter defaults, add an _sdf-args_ block in your SDF YML configuration, or just specify the rules you’d like in the command line when running _sdf lint_. Run _sdf lint —help_ to learn more about SDF’s lint rules and configuration options. ## Getting started SDF Lint is now available in preview build **_v0.10.0-p_** For more, see the [linter documentation](https://docs.sdf.com/linter/overview), or join SDF’s community [slack](https://sdf.com/join)! **To get started** 1. Download SDF Preview. SDF has 2 release channels: _stable_, and _preview_. The linter and formatter are in preview today, and will be released to stable in a future release. 1. To install SDF Preview, see the installation [documentation](https://docs.sdf.com/introduction/install) 2. If you already have SDF installed run: `sudo sdf system update -p` to join the preview channel. 2. Create or `cd` into your SDF workspace 3. That’s it - start linting! **Common commands:** - `sdf lint` → runs the linter on your whole project - `sdf lint path/to/file.sql` → runs the linter on a specifc file - `sdf lint —fix` → Auto-fixes rules where possible - `sdf format` → Format all SQL files captured in the workspace ## Limitations As of this initial release, SDF supports a limited set of dialects. 🟢 Snowflake 🟢 BigQuery 🟡 Redshift 🔴 Trino 🔴 Other SQL Dialects (not immediately planned) We look forward to supporting more dialects, and reaching higher parity with SQLFluff soon. ## Summary SDF is focused on building first-class, high performance tooling for data development. Underpinned by best-in-class semantic understanding of many SQL dialects, our mission is to provide a next generation transformation layer and dialect agnostic database engine. --- --- title: "How to build a semantic layer" description: "A semantic layer brings consistency to your most important metrics. Here’s how to set one up easily with dbt Cloud." url: "https://www.getdbt.com/blog/build-semantic-layer" date: "2024-11-18" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # How to build a semantic layer Even if you have an amazing data analytics story, your users may still tell you it’s challenging to find the data they need. Some may not be comfortable writing SQL code or well versed in creating accurate aggregations. Others may not know where to find the correct data. A semantic layer can help by providing a single source of truth for an organization’s key metrics. It brings consistency, discoverability, and democratization across teams, giving business users the data they need to self-service answers to their own data questions. Building a semantic layer from scratch isn't easy. That’s where dbt comes in. We'll look at the value of building a semantic layer and show how you can build one easily if you're already using dbt to model your data transformations. ## Why build a semantic layer No matter what business you're in, there are certain metrics that cut across teams. ‌These might include things like estimated sales per quarter, quota attainment, net versus gross profit margins, and others. The problem is that teams within a division often derive these metrics individually for their own reports, using their own tools. [A Forrester survey in 2021](https://www.forrester.com/blogs/the-bi-fabric-baby-is-slowly-but-surely-growing-up/) found that over 61% of organizations use four or more BI tools. A staggering 25% use 10 or more. This results in inconsistency across teams, which undermines trust in data. A [**semantic layer**](https://www.getdbt.com/blog/semantic-layer-introduction) acts as a translation layer between data and human language. It combines metrics along with context - such as documentation, [data lineage](https://www.getdbt.com/blog/guide-to-data-lineage), etc. This provides users with details on how a given metric was calculated, who calculated it, and on what data it is based‌ on. It creates a “hub and spoke” model for analytics so that data stakeholders can be sure they’re working off the same metric everywhere, every time. A semantic layer also provides both **discovery** and **reuse**. It enables business users who might not have the confidence to write their own queries and aggregations to find and use the metrics they need. And it enables other business users and analytics engineers to incorporate the work of others easily into their own reports and applications, rather than reinventing the wheel every time they need to leverage a standard calculation. ## The dbt Semantic Layer For years, dbt has worked to bring increasing standardization to data analytics code. Using dbt, you can [model all of your data transformations as code](https://docs.getdbt.com/docs/build/models) that can be [version-controlled](https://docs.getdbt.com/docs/collaborate/git/version-control-basics), [tested](https://docs.getdbt.com/docs/build/data-tests), and [deployed automatically](https://docs.getdbt.com/docs/deploy/continuous-integration) with every change. Over the years, however, we’ve noticed many customers struggling with common gaps in their analytics workflows. One has been the lack of a consistent approach to developing, deploying, and operationalizing analytics code. ‌That's why we've championed [the Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) as a method for standardizing and streamlining your analytics code development process. The other issue has been metrics consistency. Once companies exceed a certain size, they struggle with providing reliable metrics while also giving business users and data application developers the freedom to choose their own BI tools and create their own data workflows. This is why we've built the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) as part of dbt Cloud. ‌Using your dbt models, you can define metrics that are common to the organization. You can then grant access to this metrics layer using [role-based access control (RBAC)](https://docs.getdbt.com/docs/cloud/manage-access/about-user-access), enabling data stakeholders to access metrics via API calls and a wide variety of BI tools. ## How to build a semantic layer with dbt Once your data models are defined in dbt Cloud, it’s easy to add metrics definitions to the dbt Semantic Layer. This consists of five steps: - Define the metrics you need - Set up the environment - Create verified data sources - Create metrics derived from the data - Integrate metrics into data products Let’s look at each step in detail. ### Define the metrics you need As always, the first step in measuring is deciding what you need to measure. As part of [the Planning phase of the ADLC](https://www.getdbt.com/blog/adlc-plan), you should plan whatever metrics you want to push as part of a new analytics code deployment or change. Work with each project’s stakeholders to hammer out agreed-upon definitions for metrics, resolving any inconsistencies across teams. If you’re just getting started with building a semantic layer, don’t try and create dozens of new metrics all at once. Instead, identify three to five key metrics for the business across a couple of dbt projects. The ADLC is all about making small, right-sized improvements to your analytics codebase, rigorously testing and deploying the smallest unit of work possible with each push. Once you’ve defined and released one metric and worked out any kinks in the process, you can apply the lessons you learned when deploying the others. ### Set up the environment If you’re a dbt Cloud Team or Enterprise user, you’re ready to set up your environment to start building out your dbt Semantic Layer. If you’re only using dbt Core, [you can use our step-by-step guide](https://docs.getdbt.com/guides/core-to-cloud-1?step=1) to transition the projects containing your metrics to dbt Cloud. (You can transition as many or as few of your dbt projects as you want to dbt Cloud, migrating at your own pace.) [Setting up the dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/setup-sl) requires a previously successful dbt run. Once that completes, you can set up a connection to a Snowflake, BigQuery, Databricks, or Redshift data warehouse from one of your environments (development, staging, staging, production, etc.) to drive metrics definitions. Here, you’ll supply the credentials you use to connect to each service that contains the data required to drive your metrics. ![Setting up the dbt Semantic Layer](https://cdn.sanity.io/images/wl0ndo6t/main/cbb99b466f81e704a060c3be717df2d21d62692a-878x486.png) ### Create verified data sources To build metrics, you need a dbt model. If you don’t have one for the data that drives your metrics, define, test, and deploy those models before continuing. Next, you should familiarize yourself with [the key concepts of the dbt Semantic Layer](https://docs.getdbt.com/docs/build/about-metricflow), which is built on our own [MetricFlow project](https://github.com/dbt-labs/metricflow). In particular, you should understand [semantic models](https://docs.getdbt.com/docs/build/semantic-models). These correspond to models in your dbt project and consist of three core pieces of metadata: - **Entities** (nouns) - Your data table and their relationships - **Measures** (verbs) - The aggregation function you’re calculating (e.g., total sales). This can consist of your metric or can represent multiple metrics aggregated into a single, new metric - **Dimensions** (adjectives/adverbs) - Aspects of the entities you can use to slice and dice your data - location, time period, etc. You can then define your semantic model using YAML alongside your dbt YAML model. If you have a dbt Cloud Enterprise account, you can make this even easier by using [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot) to generate your semantic model for you. As with dbt models, you can—and should—write [thorough documentation](https://docs.getdbt.com/docs/build/documentation) for your semantic models. This should highlight everything data consumers need to understand, use, and have trust in your metrics. ### Create and deploy metrics With your semantic models defined, you’re ready to commit, build, and deploy your first metrics. Once you’ve committed your changes and another team member has [signed off on the pull request](https://www.getdbt.com/blog/how-to-review-an-analytics-pull-request), create a [deploy job](https://docs.getdbt.com/docs/deploy/deploy-jobs#create-and-schedule-jobs) and run it to create a new semantic model in your environment. dbt Cloud will create the new metrics as well as any documentation. ### Integrate metrics into data products Your data consumers can now find your business metrics using [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) and consume them in their own BI tools. Consumers will be able to see, not just the metric, but its associated documentation and a map of its data lineage. This gives users the knowledge they need about how to use the data - and confidence that it’s sourced and derived correctly. ![Integrate metrics into data products](https://cdn.sanity.io/images/wl0ndo6t/main/906b544fe46d2536f39a92c203a3b21e631da05f-1600x645.png) dbt Cloud supports [a number of out-of-the-box integrations](https://docs.getdbt.com/docs/cloud-integrations/avail-sl-integrations) for popular tools such as [Tableau](https://tableau.com), [Microsoft Excel](https://www.microsoft.com/en-us/microsoft-365/excel), [Google Sheets](https://sheets.google.com/), and others. For tools not directly supported as of this writing (such as PowerBI), you can use [exports](https://docs.getdbt.com/docs/use-dbt-semantic-layer/exports) to create custom integrations. You can use our Java, Python, and R clients to consume metrics in your applications ([check out these examples](https://github.com/dbt-labs/example-semantic-layer-clients/)). For other languages, you can leverage [the Semantic Layer REST API](https://docs.getdbt.com/docs/dbt-cloud-apis/sl-api-overview). ## Conclusion dbt models provide a common way to talk about data within your organization. A semantic layer takes this further by providing a translation layer between your data and your everyday business language. Using the dbt Semantic Layer, you can build this language directly on top of your existing dbt models. Once published, data consumers can easily find and use these standardized metrics using whatever tools they choose. This provides a single, centralized source for the data that matters to your company. To learn more about how dbt Cloud can bring consistency and simplicity to your data, [contact us for a demo today](https://www.getdbt.com/contact). --- --- title: "Data as an assembly line" description: "Cedric Chin runs Commoncog and Xmrit, a free tool to create and share XmR charts." url: "https://www.getdbt.com/blog/data-as-an-assembly-line" date: "2024-11-17" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Data as an assembly line _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/data-as-an-assembly-line-w-cedric). _ Cedric Chin runs Commoncog—a publication about accelerating business expertise. He joins Tristan to talk about the analytics development lifecycle, how organizations value (or misvalue) data, and why “data teams are not some IT helpdesk to be ignored.” **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Amazon Music](https://music.amazon.com/podcasts/333fe811-1b14-499c-b609-9bfb8f06d1ae/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways from this episode ### So the thing that made me want to reach out to you and have this conversation is that I [published a precursor](https://roundup.getdbt.com/p/the-analytics-development-lifecycle) to the full [analytics development lifecycle (ADLC) post](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) that I put out. And in it, I talk about how the role of data analysts is to do insight generation. And I made the statement that analytics isn't fundamentally an assembly line ### **And it was around, really how we should perceive the value that's created by data analysts and by data practitioners more generally, and how we should think about the value that data generates. I was representing what's probably almost conventional wisdom.** ### **But your pushback on this topic was, what if we do conceptualize data more as an assembly line? Why don't you try to frame what you see as the value that data can provide to organizations when it's living up to its fullest potential.** **Cedric Chen:** Maybe I should start with the conventional view of data in the data world. Data professionals I talk to say they answer questions, generate insights, the business person asks a question and then the data professional give them some answer after sometimes very hard work of figuring out correlations, relationships, whatever. And then the business person, and this is a common experience that I've heard from every data professional, the business person goes, okay, cool. And it does nothing with it, right? Not always, sometimes it really does lead to impactful stuff. But a lot of times it turns out that the business person sort of carelessly asks you as a data person or the data team a question that actually doesn't matter to the business. They're just curious. And then they don't know that you have to, you as the data professional have to go through hell and high water to get the answer to that question. And a data professional told me after reading a whole bunch of stuff that I covered, after I'd gone through [not just what Amazon did, but also other companies like Amazon who use these ideas and how they actually use data](https://commoncog.com/the-amazon-weekly-business-review/), that it's almost as if the problem with most businesses and most business people is that there's too many questions they can be asking. And there's no way to narrow down the set of questions that really matter in the business. And because you can't differentiate between what's a good question, what's a bad question, the data team is like this service desk that is overwhelmed with questions. It turns out that in really good data-driven companies of the type that I dug into, the term that I'm using is process control or statistical process control, which is the set of ideas that was used to ramp up production for World War II and then was used to transform industrial Japan post-World War II and led to the creation of the Toyota production system. This is a particular style of using data called statistical process control. And the basic idea is just, how does your business work? Your business is a system. It's a process. It has some inputs. It has some outputs. And you can figure out what those inputs and outputs are. Amazon has a framework that is much simpler to understand, which is controllable input metrics and output metrics. Obviously, the ultimate output metrics you care about in the firm, in the company, are things like financials, profit, revenue, free cash flow. But there's a complex set of inputs and intermediate inputs as well that lead out the other end. And there's some lag from the inputs to the outputs. Your executives running the company, running this complex machine, need to have an idea of how the inputs flow true to the outputs. Now, given that frame, how do you actually piece together the causal model of your business? And the core thing you have to grapple with is that it's not that easy to say, “OK, we've driven these inputs, and then we get some outputs out the other end.” The problem that we have to deal with as business people is variation. Most people can't deal with variation. Variation just means that something wiggles. And we know that when we step on a weighing scale that our weight wiggles, right? It doesn't just stay the same. But it's even worse in business. Metrics can wiggle, sales metrics can wiggle upwards of 40, 50%. And that's perfectly normal. Nothing else has happened. And so if you have a way of differentiating between a real wiggle that you have to worry about and you should investigate and routine variation, which is a normal wiggle you can ignore, you have a way of separating signal from noise. And you have a way of finding out things that actually move the needle on the metrics that you care about or the outcomes that you care about. Now what this does, if you're able to differentiate between routine variation and exceptional variation is that you unlock the most common of human learning loops, which is trial and error. And that data usage context results in a way of using data and asking questions of a data team that is very different from a company that doesn't have this weekly process of figuring out how the business actually works. ### **I follow along with the fundamental insight here, which is to say that, what we should be doing is building a causal model for our organization. And that requires defining key metrics and understanding the relationships between the metrics.** ### **And then having a reasonable statistical view of what's noise and what's actual variation. And then making decisions based on that. But obviously the devil is in the details. Does this model of the world always work or did Amazon happen to have a perfect business model where they could just observe the inputs and outputs and draw causal relationships really effectively?** Yes, it works. But there are certain things that are harder to measure using this methodology than others. So marketing attribution is not solved by this method. It's really, really hard. And if you look at early Amazon, they defaulted to things that they could actually tell there's an impact, even if the impact is vague. They didn't do TV ads. They did it for a while and they realized that, we can't really find controllable inputs and controllable outputs. We can do controllable inputs on the cost of the amount of money we spend on TV ads, but we can't tell the output, right? So let's stop. So what did they do? They defaulted to affiliate marketing. This was affiliate marketing in the 1990s. You can imagine how bad that was. And they could sort of model that behavior. The other thing that they could do was they could model word of mouth growth. They could say that within a certain number of months, a certain percentage of people who turn out to be new users in this new purchases in this first month, right? They didn't even have the term “cohort”. They called it “vintages”. So by the fourth vintage, would be like 70 % that will become a steady state number of customers, right? It is possible to figure out. One of the nice things about running Commoncog is that it's a dinky little business, and it doesn't really matter if I tell you about how my business works, and I can tell you my metrics. We wanted to figure out how the newsletter grows. The first step is just measure. And then every week, take a look at your WBR. We do a full Amazon WBR. You go take a look and figure out what routine variation looks like. So you develop a fingertip feel of what normal looks like. And what's some of the controllable inputs that you can think about? Well, LinkedIn posts or Twitter posts. And it turns out that there's a linear relationship between visits to the site when I post on Twitter and if I ramp up my Twitter posting, there's a relationship and increase in visits from Twitter. But there's no change in new starter signups. Similarly for LinkedIn. Maybe some of them do sign up, but it's not exceptional variation. It's not special variation. One day, October last year, we saw exceptional variation in in-depth readers, which is people who read more than one page, and unique new starter signups. were like, holy shit, what happened? And it turned out that a 20,000 substack, an investing substack, had linked to Commoncog. Now, that's interesting. And it turned out, by the way, that the author of the substack had tweeted just one week before and no exceptional variation on any of our metrics. And this person had 70,000 followers on Twitter, but zero budget on metrics. So what's the obvious thing? The obvious thing is, OK, we need to go run an experiment. What if in terms of the bang for buck for for our effort, it makes more sense instead of posting on social media or hiring someone to post on social media, to go and hunt down subsets and try to get them to link to us. It could be by buying ads and we have to experiment with that, or we could try to be friends with them so that they link to us naturally. There are a range of experiments that we could do because we now know that this is a lever that will result in a change in the output metric that we care about. ### **Nothing you're saying should be hard. And it's certainly not technically hard if you're good enough. For some reason BI tools don't generally draw XMR charts, right? You have a tool that does that, right?** All the data tools are good enough. Yeah, we have an open source tool that we made because BI tools don't generate XML charts and we want people to steal it. We want BI tool vendors to steal the code. ### **You want to drive change in the industry so that people can do this instead of BI tools.** The funny thing is that the most people who are using the tool are business people who want to get results. And they don't have data teams, so they can't bother to deal with the data team and wait for them. So just give me the CSV, and I'll paste it into this tool so that I can run the experiments, which is sad. Commoncog is for people who want to get good at business, because I want to get good business. So I bring them along for the journey. The open source thing is more like I feel for data people, I have an affinity for and I empathize with them. Plus my wife is a data analyst, so I feel her pain. And it's just I'll give you an answer to the question you did in articulating. Why is this not more widespread? ### **What I was poking at originally was that maybe it's not relevant in all cases and maybe it is to a greater or lesser degree relevant in different cases, but I think that you would probably take the belief that it's generally a useful tool in most contexts. And my guess is that it somehow has to do with organizational dynamics. Is that right?** Yes, it has to do with power. It has to do discipline. It has to do with will not skill. You have to bear in mind that Amazon did all of this in 1997 with Excel 1997, which sucks. And they could forecast within a 3% error rate their growth of their business at the time. Amazon got to a billion dollars on Excel. So if they got into a billion dollars running the WBR on Excel, what excuse do other organizations have? You have the modern data stack. You have everything that we didn't have back in the day. And literally, the way it worked was that every Sunday night in a shared folder, all the departments would drop Excel files into a shared folder. So if the tools are not the problem, what's the problem? The problem is the social technical dynamics, right? The WBR is not just a metrics review meeting. It's also a political tool. If somebody who is an executive asks one of your metrics owners under you in your department a question that you cannot answer or you don't know the answer, you'll be shamed in front of the entire organization. So what happens in that kind of political context? You really care about data and you don't ask stupid questions of your data analysts because what matters in this dynamic is you have a set of output metrics you have to hit by the end of the quarter. You have to figure out what the input metrics are. And the rule is we don't talk about the output metrics. We only talk about the controllable input metrics. So you need to go figure out what those controllable input metrics are. So now a fire is lit under your ass to work with your data team. And the data team is embedded inside your organization to figure out what those controllable input metrics are. And the way you figure it out is that you do trial and error quickly. You drive this and see if it pushes the output metric after some lag. No, doesn't work. Let's try another one. And you have to instrument, and it's OK. If you say that you’re still instrumenting, they say, “OK, it's fine; you can't push the controllable input metric. We'll give you some time to instrument.” Everybody understands that in the org, right? But once you do, you better figure out what the controllable input metrics are. And then you start setting targets. And then it becomes a very tight process where there's a fire under us because every Wednesday morning you have to present to executives. ### **It's a system that once it's working, I can see it being incredibly effective and self-perpetuating. It is also a system that I can really imagine being hard to create in the first place. Because most executives at a company do not actually want to subject themselves to that type of scrutiny in front of all of their peers**. One beautiful thing when you have a business review that measures the company end to end is that when there's a problem in one part of the company, you can say, hey, you in the different department, can you help this guy out? We're part of a team. Let's work together. Because things that change in one department, especially if you figure out controllable input metrics, output metrics. One person's output metric is sometimes another person's input metric. And so everybody has a bird's eye view of a very complex business. Amazon's WBR is designed so that you can do 500 metrics in exactly 60 minutes. You don't go over time except during the holiday season. But the point is just the practice that matters. There was a trend, over the last 12 months or so in the data community, at least the data people that I'm connected to on LinkedIn, talking about metrics trees. That's just as good. The point is that you need some kind of practice like this. And unfortunately, it requires shoving down by the CEO. At least all the successful examples I've seen has been somebody in a position of power on the executive team enforcing this. And when it works, it is a wonderful environment to work in as a data person, because everybody is motivated on the same business problems. Everybody has the same causal model of the business. And the data team is not some IT help desk to be ignored. They are critical to achieving your goals. ### **One of the reasons data people don't talk about this more on LinkedIn is that it's not something that's controllable for them. It requires the type of executive sponsorship that then gets books written about it.** ### **Probably many data people have never worked in a context like this. I think it probably was more common in an era of where they were more physical manufacturing type processes that we were trying to measure. So probably the experience set of people operating this environment is just actually not that high. Would you agree with that?** Somebody in my community, observed that this kind of operational excellence only emerges when you have a very low return on invested capital. Not very low, single digit. Because if you think about it, if your margins are super, super high, your return invested capital is super high, because like Google, you can be sloppy. It doesn't really matter. Who cares? You have a network effect. You're just generating gushes of cash. If your return on investor capital is negative, then it doesn't matter. Your operational excellence is just staving off the inevitable. You're just going to die. You're in a textile mill in America, and you're fighting against it. But if you have a 3% return on investor capital, an operational excellence can double that for free, for effectively free. These kinds of methods tend to spread in organizations or industries where you have a single digit return on investor capital, because if you don't have it, then you die. It just so happens that Amazon is a low margin business that just happened to hire somebody from manufacturing to deal with the scale of volume in their fulfillment centers. And then the ideas spread out and they use it to crush their competitors because the competitors just did not understand what was going on in their businesses. I often tell people that maybe you can get away with this in high margin businesses, right? But if you are in a low margin business and you up against Amazon, Amazon has process control and you have something naive like the North Star metric framework, you're going to get crushed because you are just working towards one metric. Amazon is working towards like 16 and they can spin off things to like, let's just undercut you in this particular way and we'll process control that and process control the cost that we're spending to undercut you. And we can do pincer movements. So it's an organizational capability that is super powerful. --- --- title: "How to build the business case for dbt Cloud" description: "Learn to build a strong case for dbt Cloud with practical steps, templates, and insights. Unlock your data team's potential." url: "https://www.getdbt.com/blog/how-to-build-the-business-case-for-dbt-cloud" date: "2024-11-15" authors: ["Alexis Jones"] categories: ["Learn"] --- # How to build the business case for dbt Cloud Data professionals know the struggle: your organization is bursting with data, but making sense of it—and convincing others of the need for better tools—feels like an uphill battle. Enter the [Building the business case for dbt Cloud](https://www.getdbt.com/resources/making-the-case-for-dbt-cloud) whitepaper. This resource is your step-by-step guide to aligning your data initiatives with your company’s strategic goals. It’s designed to help you secure the buy-in and budget you need to make dbt Cloud the backbone of your data workflows. Inside, you’ll find templates, proof-of-value strategies, and insights from dbt Cloud users who’ve seen firsthand how it transforms data practices. Here’s a sneak peek. ## Why dbt Cloud? Despite recent advancements in analytics—AI, cloud, open table formats, self-service—we’ve still yet to solve core data problems plaguing our organizations: data quality, data velocity, and cost optimization. This is in part because, despite its best efforts, the modern data stack has created data silos. These silos need to be centralized so organizations can build the holistic context needed to confidently embrace data at scale. The [data control plane](https://blogs.idc.com/2021/04/22/every-path-has-its-puddles-improving-enterprise-intelligence-with-a-data-control-plane/) is where this centralization should happen. **dbt Cloud is a data control plane that centralizes the metadata from your analytics workflow and makes it actionable, so your teams can ship and use trusted data, faster.** dbt Cloud is interoperable across data platforms and provides a universal, real-time view of what’s happening across your data estate. Its built-in user experiences are designed to accelerate data delivery and help users understand and improve data quality and compute costs along the way. dbt Cloud offers varied user interfaces and integrations so that stakeholders of all technical stripes can participate in the data workflow and have the trusted insights needed to translate data into strategic decisions. ![dbt control plane](https://cdn.sanity.io/images/wl0ndo6t/main/ec55dbb118a3d4e1ce72bdd626b77d887cde4aeb-7500x3334.png) ## Scale analytics successfully with dbt Cloud Everyone wants to “do more with data” and become a “data-driven company”, but there are common roadblocks that derail these initiatives at scale. Data teams are still grappling with issues related to data quality, data velocity, and ambiguous data ownership. This is shown in [our most recent State of Analytics Engineering report](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024). The adoption of dbt Cloud by data teams has led to significant improvements in data speed, quality, and trust. These teams have reported impressive results, including [50x multipliers](https://www.getdbt.com/case-studies/axs) in data velocity and [dramatic increases in data quality and data trust](https://www.getdbt.com/case-studies/pepperstone) as evidenced by the 30-80%+ reduction in data inconsistencies and [improved eNPS scores](https://www.getdbt.com/case-studies/safetyculture). By implementing dbt Cloud as their standard, these teams are experiencing tangible benefits. It’s time to help more teams at more organizations realize similar, outsized results. This paper includes detailed, step-by-step guidance designed to help you successfully adopt and ramp dbt Cloud within your organization. We include templates for mapping technology solutions to business goals, parameters to consider as you evaluate proposed solutions, considerations for scoping a proof of value, and more—so you can have a measured, structured approach to securing the support you need to adopt dbt Cloud at your organization. The tips, templates, and tales in these pages have been gathered via dozens of interviews with dbt Cloud customers—several of whom you’ll meet below. ## Step 1: Align with company objectives When seeking budget or headcount approval, the best place to start is where the business wants to end: your company objectives. Mapping technical or people problems to the organizational goals they jeopardize will not only help you speak the same language as non-data stakeholders, but also provide leverage in your request for resources. ![Map analytics workflow](https://cdn.sanity.io/images/wl0ndo6t/main/eb817dcc0d8e3f326acacf48c32eab0646b10822-5100x1900.png) In the above example, company upsell rates haven't progressed as expected. Polling key stakeholders on why they believe they’re behind on achieving this goal will help connect business context to data roots. Use trends identified in these interviews to fill in the “contributing factors” column. This will highlight the need for additional data resources to hit these company-wide goals. This is just the first of **eight actionable steps** for building the business case for dbt Cloud. Each step is designed to give you the tools, templates, and insights to align data priorities with strategic objectives. [Download the full guide](https://www.getdbt.com/resources/making-the-case-for-dbt-cloud) to learn how to confidently make your case and unlock the potential of your data team. **** --- --- title: "Uniting Core and Cloud with One dbt" description: "Tristan's reflections on the biggest theme of Coalesce 2024." url: "https://www.getdbt.com/blog/uniting-core-and-cloud-with-one-dbt" date: "2024-11-10" authors: ["Tristan Handy"] categories: ["Insights"] --- # Uniting Core and Cloud with One dbt _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/one-dbt). _ A few weeks ago, I wrote about the [feature announcements](https://roundup.getdbt.com/p/recovering-from-the-party) from Coalesce. This issue I want to zoom in on our One dbt theme I announced from the main stage. It’s a big deal for us, and will be important for the long-term health of the community. ## One dbt IMO, the biggest underlying stressor in the dbt Community over the last several years has been the underlying tension caused by the steward of the OSS roadmap (dbt Labs) going through the growing pains of becoming a real software business. Many, many humans use dbt Core as a central part of their jobs. It has become an important professional dependency for many of us, and so there has been some stress in the community associated with topics like: - Is the license going to change? - Will Core continue to get new good features or will dbt Labs abandon it? We’ve done our best to answer questions like this over the past couple of years. We’re not changing the license—it’s staying permissive open source, Apache 2.0. We continue to ship new and important features to Core—check out the new incremental functionality and unit tests, both launched this year, both things I’ve wanted for a long time. But despite this, many folks in the community correctly identified a fundamental point of friction between dbt Labs and the dbt Community, one that had created real underlying tension. I think that we (dbt Labs) have—_unintentionally!_—created a context that caused a bifurcation in the dbt Community into two somewhat distinct sub-communities: the Core community and the Cloud community. The Core community was much larger, full of early adopters, and loved the permissionless nature of dbt Core. They didn’t mind that running Core came with some friction, because it was such an important tool in their professional tool kit that learning to operationalize it became a point of pride, not a problem to be solved. Often, but certainly not always, folks in this community were skewed toward the Data Engineer persona. The Cloud community was smaller. It was just as passionate about dbt, its fundamental principles, and its impact. Its members often, but not always, tended to be a part of larger organizations and talked about best practices more frequently inside of their own, internally-facing analytics communities. Often, but not always, folks in this community skewed toward the data analyst persona. The product itself did not absolutely force users to make a choice—core OR cloud—but it didn’t make it easy to use the two versions of dbt together. All of this drove a gap between the two communities. And we exacerbated this gap in the way we set up dbt’s documentation, dbt’s onboarding experience, our product launches, and countless other aspects of the interactions that members of both of these communities had. We continually—although unintentionally—drove a wedge between these two groups, one tiny hammer blow at a time. Over time, identities started to form. **“I’m a Core user.”** **“I’m a Cloud user.”** Divisions in a community are bad. Divisions in a community that ossify into identities are worse. Instead of having a pragmatic conversation about feature set and pricing, you’re having a conversation about who someone is. We certainly haven’t fixed all of the underlying problems yet. But at Coalesce this year, we called it out publicly, and we committed to bringing these communities, and the products, closer together. This was an important part of our “One dbt” theme and we talked about it _everywhere_. Here’s how I framed it in my keynote: > In the data industry, we are obsessed with these zero sum debates. Snowflake versus Databricks, Data Engineer versus Analytics Engineer, data team versus the business, dbt Core versus dbt Cloud. It is incredibly easy to frame things this way, as fights, as battles. > > But none of these should be **OR** conversations. They should be **AND** conversations.  > > We need to focus on how to bring people together, not pit them against each other. Snowflake AND Databricks, Data Engineers AND Analytics Engineers, **dbt Core AND dbt Cloud**. I have always wanted dbt to be a unifier, not a divider. dbt works across Clouds, across Data Platforms, across Teams. ## Reactions at Coalesce The feedback we got from attendees throughout the conference was, quite honestly, _incredible_. Cloud users expressed an excitement about building bridges to Core users, often who worked on peer teams at the same organization. Core users expressed excitement about being able to use Core alongside Cloud-native experiences. You know how sometimes in a personal relationship there will be this underlying unsaid thing. Maybe that thing has gone unsaid for years, and maybe it’s been there for so long that the two people in the relationship have forgotten that it’s even there. But it continually saps energy and joy from the relationship. Eventually someone works up the courage to say “you know what I’m really pissed about [this thing].” And then all the sudden the floodgates open. At first the conversation is tense. It’s hard to talk about these kinds of things, after all, which is why that topic went unsaid for so long. But eventually, as both sides go deeper, the conversation develops some safety. Trust is built. Both sides display empathy for the perspective of the other that they slowly come to understand. Eventually—and this “eventually” can take hours or years—the relationship is stronger than ever. That’s what it felt like started—started!—to happen at Coalesce 2024. If I had to summarize the response from the scores of attendees I spoke to about this divide and about our One dbt message, it was: _“You’re goddamn right. Now let’s fix it.”_ And behind this response was this incredibly complex cocktail of emotions: relief, frustration, gratitude. It was a lean in, not a lean out. A “let’s do this.” ## Making it real But now we have to manifest One dbt in the product. Let me give you two examples of things we’re thinking about. Standard caveats apply: this is not an official roadmap, but these are real ideas we’re thinking through. **Example 1.** We need to create an easy way to ingest Core metadata into dbt Cloud so that Cloud users can do things like:  - Create cross-project references to upstream Core projects - View the entire DAG in dbt Explorer Right now, doing this is possible, but hard. In a One dbt world, the idea that some users (and some workflows) will be on Core and some will be on Cloud is a default assumption. **Example 2.** We currently have two CLIs: there is Core (obviously), but there is also the dbt Cloud CLI. Many folks don’t even know about the Cloud CLI, but it has many advantages over Core and is my personal favorite dbt developer experience. The problem is, moving back and forth between those two CLIs is friction-ful and users tend not to do it.  In a One dbt world, we wouldn’t force users to choose “which CLI to install.” In a One dbt world, the idea that you would have to choose wouldn’t make any sense—there should only be one. I could go on; there are a lot of examples of these types of problems. dbt Cloud has a ton of product surface area at this point, and most of these product experiences have not been designed with a One dbt mindset. There’s real work to do. ## Reflecting back There are not many people, or many companies, that have the very good fortune to steward large open source products or lead large communities. _I am consistently humbled by the trust the dbt Community has placed in me, and in us, over the years._ As I am not shy about saying, I don’t know all the answers, and I am not perfect. As I reflect on the journey that got us here, I can’t help but be self-critical. But I’m so very grateful for the opportunity to name this problem and for the opportunity to correct it. There is a tremendous amount of goodwill in the global community of dbt users, and a genuine desire to build a bright future together. Creating space for commercial innovation while keeping the community healthy and aligned is, I believe, a very achievable goal, and is absolutely central to how we need to ensure a sustainable future for dbt over the very long term. Thanks for your patience along the journey, thanks for your support, and I’m looking forward to pushing forward together. There is only one dbt. LFG. --- --- title: "What's new in dbt Cloud - November 2024" description: "Learn about the latest features and capabilities that just landed in dbt." url: "https://www.getdbt.com/blog/whats-new-in-dbt-cloud-november-2024" date: "2024-11-07" authors: ["Alexis Jones", "Sara Gawlinski"] categories: ["Product"] --- # What's new in dbt Cloud - November 2024 Welcome back to our regular installment of "What's new in dbt Cloud" where we recap all the latest innovations landing in dbt since [our last announcement](https://www.getdbt.com/blog/whats-new-dbt-cloud-august-2024) back in August. It's been a very busy few months at dbt Labs and we're still coming down from [Coalesce 2024](https://www.getdbt.com/blog/coalesce-2024-product-announcements), where over 2,000 data enthusiasts joined us in Las Vegas (with another 8,000+ tuning in online) to connect, learn, and get inspired. You can check out the sessions on-demand and be the first to hear all the need-to-know information about next year’s event [here](https://coalesce.getdbt.com/). Alright, on to everything that's new in dbt Cloud! ## ♾️ Analytics Development Lifecycle In case you missed it, in September we [published a whitepaper](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) about what we believe is the right next step forward to mature analytics practices at organizations of any size: aligning on the analytics development lifecycle (ADLC). We encourage you to read the paper to learn more about how embracing this vendor-agnostic process can help accelerate and improve analytics at your company. You can also check out our [new-and-improved website](https://www.getdbt.com/product/dbt-cloud) to learn more about how dbt helps teams embrace various stages of the ADLC. ![Visual of the 8 phases of the analytics development lifecycle (ADLC)](https://cdn.sanity.io/images/wl0ndo6t/main/7fb61879a024adcb59b5f73665f7d804dbe8d223-1554x770.png) ## 📈 dbt Semantic Layer Harness the power of the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) with new features that make it easier to build, consume, and scale your semantic layer strategy. Centralize metrics, ensure data consistency, and deliver insights more efficiently than ever before. Here’s what’s new: 🪄**Auto-generate semantic models with dbt Copilot (beta):** Building and deploying your semantic layer is now faster and simpler. With [dbt Copilot](https://docs.getdbt.com/docs/cloud/dbt-copilot), you can automatically generate semantic models, reducing manual work and allowing you to standardize metrics and logic across your organization with just a click. Contact your sales rep if you’re interested in getting involved in the beta. ![GIF showing how to use dbt Copilot to generate semantic models](https://cdn.sanity.io/images/wl0ndo6t/main/ef1baf611786d2f8a68a92a5d90832fcc8084e90-1729x1292.gif) 🌎 **Query the semantic layer within the IDE**: You can now [query](https://docs.getdbt.com/docs/build/metricflow-commands) metadata, metrics, preview compiled SQL, and run exports directly in your development environment in the Cloud IDE. This makes it easier to build and manage semantic models, speeding up governed data development. And, with parity between the IDE and Cloud CLI, you can choose the environment that works best for your team for a seamless experience. 📊 **Microsoft Excel integration**: The dbt Semantic Layer integration with [Microsoft Excel 365 and Desktop](https://docs.getdbt.com/docs/cloud-integrations/semantic-layer/excel) is now generally available! This enables business users to self-serve data from governed metric definitions through a simple drop-down interface query builder directly in Excel. Whether you're in finance, accounting, or any other department that relies on Excel, you can now easily access data directly without needing assistance from the data team. ![Screenshot of the dbt Semantic Layer integration with Microsoft Excel](https://cdn.sanity.io/images/wl0ndo6t/main/3c2b5824b7735a22ab8e75ac9800193da716d0ab-2048x1036.webp) 📅 **Custom calendar support in MetricFlow:** The dbt Semantic Layer now supports [custom calendars in MetricFlow](https://docs.getdbt.com/docs/build/metricflow-time-spine#custom-calendar-) (now available in Preview). This allows you to define and use your own business or fiscal calendars, aligning your metrics and reporting with custom timeframes like 4-4-5 retail calendars or non-standard fiscal years. With this update, your analytics will better reflect your organization’s unique structure, delivering more accurate and relevant insights. 📤 **Exports improvements:** We've [enhanced our exports experience](https://docs.getdbt.com/docs/use-dbt-semantic-layer/exports) with new features like database configurations, limit and order configurations, and tagging. In addition to export-specific settings like `export_as`, `schema`, and `alias`, you can now configure the `database` setting to select the most suitable database for each export. Limit and order settings, previously outlined in the documentation, are now configurable directly in your YAML files. Additionally, tags allow you to run exports based on specific tags, just as they work for models. 📈 **Embedded analytics repo:** Embedded analytics has emerged as a prominent use case for the dbt Semantic Layer. To make it easier to learn how to get started, we introduced a [new demo environment](https://github.com/dbt-labs/embedded-sl-demo) where you can explore integrating the dbt Semantic Layer into backend systems for an embedded analytics use case. Dive into how the Jaffle Shop uses the dbt Semantic Layer to deliver personalized sales metrics to independent merchants by dynamically filtering data by store location with the Python SDK. **** ## 🔎 dbt Explorer [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) is dbt's built-in, automated data catalog. Use it to gain the holistic context and breadcrumbs that enable you to move beyond reactive workflows so you can build, fine-tune, and troubleshoot your pipelines more proactively. Here's what's new: **📊 Auto-exposures for Tableau:** Now in Preview, you can automatically populate your dbt lineage with downstream exposures in Tableau (and Power BI to follow). This gives data teams automatic context into how and where models are used, so they can prioritize data work to promote data quality. Coming soon, you can trigger downstream dashboards to automatically refresh as soon as new data is available, giving business stakeholders confidence that they’re always making decisions from the freshest data. Auto-exposures are automatically accounted for throughout dbt Cloud, including in dbt Explorer, scheduled jobs, and CI jobs. [Read the docs](https://docs.getdbt.com/docs/cloud-integrations/configure-auto-exposures) to learn more. ![Lineage graph in dbt Explorer showing the downstream Tableau dashboard powered by dbt models](https://cdn.sanity.io/images/wl0ndo6t/main/c41ef96e0994bcca9599700e73d05b500890fd7a-2752x1538.png) **💡 Model query history: **dbt Explorer now surfaces how frequently models are queried, helping data teams focus their time and infrastructure spend on popular data products as well as easing discovery by making analysts aware of widely used data models. Model query history can be viewed in performance charts, as a lineage lens in your DAG (pictured), and as a new column in your list of models. This feature is currently in Preview for Snowflake and BigQuery, with additional platforms coming soon. [See the docs](https://docs.getdbt.com/docs/collaborate/model-query-history) to learn more. ![Model consumption lens in dbt Explorer gives users at-a-glance context into model query count](https://cdn.sanity.io/images/wl0ndo6t/main/ca5e660198502c485f80f8eb40d13a7f4055fa7a-1351x996.png) 🆗 **Data health tiles: **Now GA in dbt Cloud, you can embed health signals like data quality and freshness within any dashboard, giving your downstream stakeholders at-a-glance confirmation of whether they can trust the data they’re about to use. Users can also navigate back to dbt Explorer with a single click to investigate further. [Read the docs](https://docs.getdbt.com/docs/collaborate/data-tile) to get started. ![See data health signals like data freshness and data quality directly in BI tools like Tableau (shown)](https://cdn.sanity.io/images/wl0ndo6t/main/2926b7aeb743d6e0b74b6d4f9d3af0386a85e440-1878x1312.png) 👍 **In-app trust signals:** Data health signals aren’t just for downstream dashboards. Now in Preview, health signals are accounted for throughout the in-app dbt Cloud experience, giving users an at-a-glance understanding of whether the dbt resource they’re about to use is fresh, error-free, tested, documented, and more. [Read the docs](https://docs.getdbt.com/docs/collaborate/explore-projects#trust-signals-for-resources) for more. ![Get at-a-glance health signals(freshness, test coverage, and more) for your dbt resources directly in-app](https://cdn.sanity.io/images/wl0ndo6t/main/88ac65cacf1ce240c8ffa9cde3c3807728c1d549-2132x684.png) ## Develop dbt Cloud offers [multiple, accessible development environments](https://www.getdbt.com/product/develop) to foster organization-wide data collaboration. Here's what's new: **🖼️ Visual editing experience (beta):** With a low-code visual editing experience in dbt Cloud, users can create new or explore existing dbt models using a drag-and-drop interface that compiles directly to SQL. Users have the additional flexibility to jump back and forth between this visual representation of their models and SQL code to dive deeper. The visual editing experience is fully integrated with version control, dbt Explorer, and the added capability to leverage AI for code generation. Reach out to your account team to get involved in the ongoing beta. ![Screenshot of the visual editing experience in dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/a537de6f3ccf9fc6e63489d15b113bc4a0095ff8-1800x1142.png) ## Deploy dbt makes it easy to validate and [safely deploy models into production](https://www.getdbt.com/product/deploy). Here's what's new: 🤲 **Microbatch incremental strategy:** The new [microbatch strategy](https://docs.getdbt.com/docs/build/incremental-microbatch#what-is-microbatch-in-dbt) allows you to break up large time-series datasets and process smaller batches for faster transformations and improved performance and resiliency for runs. Microbatch is currently available in beta for dbt Cloud versionless and dbt Core v1.9 for BigQuery, Postgres, Snowflake, and Spark with Redshift, Databricks, and Athena coming soon. 📸 **Snapshots improvements:** We’ve been hard at work enhancing [snapshots](https://docs.getdbt.com/docs/build/snapshots) in dbt, and we’re excited to introduce a few new improvements. Snapshots can now be configured via YAML files for a cleaner, more consistent setup alongside your models. Additionally, you can now customize the names of meta fields, offering greater flexibility to tailor snapshot metadata to your needs. 🔄 **Advanced CI:** Continuous integration in dbt just got even smarter. With the ability to compare changes in CI (now in Preview), each CI job will include a breakdown of the columns and rows that are being added, modified, or removed in your underlying data platform as a result of executing your dbt job. Users can see a summary of these changes inside their PR in Git. This additional context allows data teams to catch any unexpected behavior before code is deployed into production, improving data quality and increasing trust among all collaborators. [Read the blog](https://www.getdbt.com/blog/announcing-advanced-ci) to learn more. ![Compare changes of your data builds with Advanced CI in dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/04ca430f6275d0aafe560e2e1a1b9595b4c436bc-1244x758.png) ### Observability Catch problems before your stakeholders notice with dbt Cloud's [built-in observability features](https://www.getdbt.com/product/test-and-observe). Here's what's new: ⚠️ **Job warn notifications.** Now you can [get notified via Slack or email](https://docs.getdbt.com/docs/deploy/job-notifications) if a job run encounters _warnings_ from tests or source freshness checks—in addition to already available options for notification on job success, failure, or cancelation. ## Platform We’re always making improvements to the dbt Cloud platform to make it more scalable, reliable, and accessible. 🧊 **Iceberg table support:** dbt [now supports the Apache Iceberg table format](https://docs.getdbt.com/blog/icebeg-is-an-implementation-detail#fixing-it). This enables data teams to work more efficiently with large-scale data lakes while maintaining the familiar dbt workflow. Support for Athena, Spark, Databricks, Starburst/Trino, and Dremio are GA and Snowflake is currently in beta. Iceberg table format support is a critical capability to enable [cross-platform dbt Mesh](https://getdbt.com/blog/introducing-cross-platform-dbt-mesh) (currently in development). ☁️ **Further support for Azure deployments:** We launched the ability to deploy dbt Cloud multi-tenant natively on Microsoft Azure in Europe a few months ago, and we’re excited to now share the US region is joining the fold. This hosting option is in addition to our support for [AWS deployments](https://docs.getdbt.com/docs/cloud/about-cloud/tenancy), bringing the same powerful dbt Cloud experience to even more data teams in more regions — regardless of your choice of cloud. [Azure multi-tenant support](https://docs.getdbt.com/docs/cloud/about-cloud/tenancy#available-features) (hosted in both America and Europe) is currently in Preview for dbt Cloud Enterprise customers. **🔐 MFA enforcement for all users:** Multi-factor authentication (MFA) is now required for all dbt Cloud users. The next time a user logs in to dbt Cloud with a username/password, they’ll be required to set up MFA—through SMS, an authenticator app, or a WebAuthn-compliant security key. This new posture will help bolster overall security of your dbt Cloud account. [Read the docs](https://docs.getdbt.com/docs/cloud/manage-access/mfa) to learn more. 🏂 **External OAuth for Okta & Entra with Snowflake:** Using External OAuth, you can federate data warehouse access for developers using an OAuth Flow with an identity provider. Now, instead of Snowflake acting as the identity provider, you can leverage Okta or Entra ID to authenticate. Available to enterprise customers with Snowflake connections. [Read the docs](https://docs.getdbt.com/docs/cloud/manage-access/external-oauth) to learn more. ✍️ **Sign commits from Cloud IDE:** You can now configure a private key from dbt Cloud so GitHub can verify your identity when committing code from the Cloud IDE. This capability is now GA for Enterprise accounts. [Read the docs](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/git-commit-signing) to learn more. 🔌 **New adapters:** dbt Cloud now integrates with [AWS Athena](https://docs.getdbt.com/docs/cloud/connect-data-platform/connect-amazon-athena) (GA) and [Teradata](https://docs.getdbt.com/guides/teradata?step=1) (Preview), enabling more organizations and teams to collaborate on data workflows. ## What's next As always, we're excited to get these new features in your hands and look forward to your feedback. Be sure to join us for our ongoing "One dbt" webinar series where we dive into the latest features powering dbt Cloud. The next one is happening in December and is all about how to enable cross-functional teams with trusted data with cross-platform dbt Mesh. [Save your spot here](https://www.getdbt.com/resources/webinars/one-dbt-accelerate-data-work-with-cross-platform-dbt-mesh)! --- --- title: "Getting the most out of your dbt Cloud deployment" description: "Leverage dbt Cloud features and workflows to maximize your team’s ability to deliver quality data products efficiently." url: "https://www.getdbt.com/blog/getting-the-most-out-of-your-dbt-cloud-deployment" date: "2024-11-06" authors: ["Neha Hystad"] categories: ["Learn"] --- # Getting the most out of your dbt Cloud deployment When it comes to any software investment, the ultimate measures of success are whether it helps you do your job better, drives your organization toward its goals, and perhaps even gives your career a boost. This year at Coalesce, I got so much energy from hearing our customers talk about the impact that dbt has had on their lives. We even had [a wall at our booth](https://www.getdbt.com/blog/coalesce-2024-recovering-from-the-party) dedicated to stories from dbt users. Certainly, the topic of how to maximize success and get to ROI quickly is resonant and timely as data organizations look to justify budgets and improve their standing as strategic partners to the business. This also happened to be the focus of my Coalesce presentation, which you can [watch in its entirety right here](https://coalesce.getdbt.com/on-demand/coalesce-2024-increasing-roi-with-dbt-cloud). [Watch video](https://www.youtube.com/watch?v=C81KkvEDZU0) Below are a few tips on how to maximize the ROI of your dbt Cloud deployment, with a focus on onboarding additional users and teams, and then enabling those teams to provide a performant, governed data delivery experience. ## Laying the foundation for success The first step is to define what successful data delivery looks like for your organization. Who are your stakeholders? How frequently does the data need to be updated? Are there certain SLAs for data freshness and quality? Defining the success criteria will help tailor your approach to using dbt Cloud effectively. Next, think about your organizational structure. Are you a smaller team managing all analytics centrally, or are you planning to scale up and empower multiple teams to build and maintain their own data models? Your deployment strategy should align with your team composition and growth plans. ## Onboarding teams efficiently Efficient onboarding is critical for a good ROI. The faster new teams or users can start using dbt Cloud to deliver data products, the quicker your organization will see value. We’ve all experienced slow, painful onboarding processes. What should take a few days often stretches into weeks because of issues with setting up access and configuring data connections. With dbt Cloud features like role-based access control, single sign-on, and global connections, you can streamline onboarding users. ### Streamlined access and connections To onboard users quickly with the right access controls, start by implementing [single sign-on (SSO)](https://docs.getdbt.com/docs/cloud/manage-access/sso-overview) and [role-based access control](https://docs.getdbt.com/docs/cloud/manage-access/enterprise-permissions). This ensures that users can quickly and securely access dbt Cloud with the appropriate permissions. #### Recommended implementation: 1. Set up groups in the identity provider that map to dbt Cloud groups. 2. Use SSO so users can log in seamlessly and group mapping to ensure users are added to the mapped group in dbt Cloud. 3. Assign groups access to projects using pre-defined permission sets. 4. Designate project admins to manage access for their teams, reducing bottlenecks and administrative overhead. ![Groups](https://cdn.sanity.io/images/wl0ndo6t/main/0a94d2519888aede8594e9d7b2e93d92206d63d3-2314x960.png) ### Efficient data platform connections Configuring data platform connections manually for each project can be time-consuming and error-prone. Instead, use [global connections](https://docs.getdbt.com/docs/cloud/connect-data-platform/about-connections#connection-management) in dbt Cloud. These connections can be reused across projects, allowing teams to self-serve their setup and reducing duplicate work. For developers, consider using [external OAuth authentication](https://docs.getdbt.com/docs/cloud/manage-access/external-oauth) or [native warehouse authentication](https://docs.getdbt.com/docs/cloud/manage-access/set-up-snowflake-oauth). This will help them easily connect to data platforms without managing credentials manually. ![Global connections](https://cdn.sanity.io/images/wl0ndo6t/main/0ac1ed2d4dc2d1db52944a15017e416960887b9e-2206x1250.png) #### Key takeaways for onboarding: - Use SSO and mapped groups for efficient and secure access. - Set up global connections at the account level to reduce redundancy. - Implement OAuth to simplify developer authentication. ## Delivering governed and efficient data products Once teams are onboarded, the next step is to enable those teams to build high-quality and up-to-date data products, while optimizing the resources required to do that efficiently. With multiple teams contributing, it's crucial to have visibility into how they’re using dbt Cloud and whether they’re following best practices. The **Data Delivery Insights dashboard **(beta) surfaces metrics like job performance, test coverage, and source freshness to give you this visibility. Here are some ways you can use these metrics to empower teams with the visibility they need to deliver consistent, high-quality data assets. - **Test coverage**: Use the dashboard to check if models have adequate tests. If coverage is low, encourage teams to add more tests to ensure data reliability. - **Models built**: Understand how active a project is based on how many models have been built in the last 90 days - **Source freshness**: Review sources and flag any stale data sources with the appropriate teams. Data Delivery Insights is currently in beta—if you’re interested in joining the beta, reach out to your account team. ### Optimizing for efficiency and quality Once you have visibility into the reliability and efficiency of your data delivery workflows, use dbt Cloud features like **Explorer** and **micro-batch incremental models** to improve those metrics. ![Explorer query history lens](https://cdn.sanity.io/images/wl0ndo6t/main/617188403f80391bedbb0e5e150d50ae6a2983e1-2994x1420.png) 1. **Explorer: **Improve data quality and accessibility by using lineage with the new query history lens to audit and document frequently queried models while removing or consolidating underused ones. 2. **Micro-batch incremental models**: Reduce average model build time with microbatch incremental models. While incremental models allow you to limit the data that needs to be transformed to only the rows that have been created or updated since the last run, microbatch allows you to break up large datasets into smaller batches to make updates faster and more manageable. **** ## Conclusion You can maximize your dbt Cloud deployment by onboarding teams more efficiently and supporting those teams to provide high-quality, governed data products. Scale onboarding and deliver data faster with functionality like SSO, role-based access control, and global connections. Get visibility into the health of data delivery across your organization, and build a strategy for optimization powered by dbt features like microbatch incremental models and Explorer. If you’re interested in being a design partner or getting access to beta features, reach out to your account team. We’d love to hear your feedback and help you make the most of dbt Cloud. --- --- title: "The data jobs to be done" description: "Erik Bernhardsson on his serverless platform for AI, data, and ML teams, and his take on the future of data engineering." url: "https://www.getdbt.com/blog/the-data-jobs-to-be-done" date: "2024-11-03" authors: ["Daniel Poppy"] categories: ["Insights"] --- # The data jobs to be done _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/the-data-jobs-to-be-done-w-erik-bernhardsson). _ Erik Bernhardsson, the CEO and co-founder of Modal Labs, joins Tristan to talk about Gen AI, the lack of GPUs, the future of cloud computing, and egress fees. They also discuss whether the job title of data engineer is something we should want more or less of in the future. Erik is not afraid of a spicy take, so this is a fun one. **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Google Podcasts](https://podcasts.google.com/feed/aHR0cHM6Ly9hbmFseXRpY3NlbmdpbmVlcmluZ3JvdW5kdXAubGlic3luLmNvbS9yc3M) - [Stitcher](https://www.stitcher.com/show/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways from this episode ### You might be only the second person who's a repeat guest. You were on during Season One, and back then you were still in stealth mode with the project that has become now Modal. What have you been up to in the last few years? Erik Bernhardsson: When we talked, it was the depths of COVID. I had just quit my job and was hacking on something that's now turned into Modal. And back then, my idea was I wanted to build a general-purpose platform for running compute in the cloud. And in particular, focus a lot on data, AI, machine learning use cases. Three years later, I'm still working on that. What we discovered along the way was AI and Gen AI is a great application, because when we started building this, I didn't have like a clear use case in mind. I just had an idea that if I build this platform, people will use it because it seems like a gap in the market. Turned out that Gen AI is a killer app for what we built. We've seen a lot of interest in using large-scale applications for audio, video, image, diffusion models, biotech models, and video processing. So we run a big compute cloud, a big pool of GPUs, CPUs in the cloud. And then the other side of that is we offer an easy-to-use Python SDK that makes it very nice, very easy to take code and deploy it into the cloud and in a way where you don't have to think about scaling and provisioning and containers and all that stuff that typically you have to do if you build your own stack. ### You are building a compute platform that's built on top of the cloud providers. Why do people need Modal? How is it different than just operating directly with the services that organizations like AWS provide? There’s more room in the cloud space. AWS, Oracle, and Azure and all of these have done a fantastic job building a foundational layer of storage and compute. I've used AWS for a good 15 years, and I love AWS for what it enables me to build. But it's still not easy to use. And it still gets in the way of iterating quickly and shipping things. AWS and others have a very solid place in that stack; they're an amazing compute and storage provider. But there’s plenty of room to innovate in the layer above, which I think of as the sort of layer that you were talking about, which is the Snowflake layer, and that's also where I would put Modal. We repackage a lot of the cloud primitives in a way that suits what data, AI, and machine learning teams want to accomplish. And by focusing on one particular use case, we can offer a much better user experience. Cloud providers are hard to use because they try to build for every user at the same time, which means no user is going to have a good user experience. At the end of the day, they're massively successful businesses generating tons of money from storage and compute. That's where I think they belong. That's the core part of the stack. Anything above that is going to try to drive the usage of storage and compute. They just want to drive demand to the underlying services. And you've seen this with the success of Snowflake. Snowflake, builds on top of AWS and the other clouds and in a way competes with them, but not really, because at the end of the day, like AWS and the other clouds, they get the money either way. ### So there's room for a layer on top of the hyperscalers and Databricks and Snowflake are both sitting in that place today, and you're a more nascent entrant there. How do I frame the thing that you're helping users with versus Databricks—and especially Databricks because they’re very focused on AI use cases and you're doing a lot of work there too? I hesitate to position ourselves against Databricks because in many ways they're a fantastic company. But aspirationally, we share the same vision of what we want to accomplish. They're obviously massively ahead of us by 10, 15 years, but their vision is the same; they want to build an end-to-end platform to serve data, AI, and machine learning needs. We come at it with a very different architectural approach. We basically said, you can't run this yourself. We're going to be not just the software layer, we're also going to be a hosted infrastructure as a service platform—building for containerization, building for cloud, building for this multi-tenant super elastic compute pool. That meant that we could make very different architectural decisions. Obviously, I have a bias, but I think we have the right tailwinds because we're thinking about architecture in a very different way. ### There are a couple of companies recently that are doing some version of helping you have an abstraction layer across the different hyperscalers and across the different availability zones to do resource pooling. Why is that happening today? I honestly think it just comes down to the fact that GPUs are expensive. In order to make the economics work, you have to run them at very high utilization. And because you run them at very high utilization, you're going to have poor availability, which means that for any single availability zone, you may actually be close to capacity most of the time. Which means in order to do this well, you need to go to different availability zones, you need to go to different regions, and you need to go to different clouds. It's a big part of what we've been spending time doing—integrating with a bunch of different cloud vendors, using all the different regions and zones, and then just getting capacity to the customer wherever we can find capacity. That's actually a fun, interesting problem in itself. We monitor its prices, which change dynamically 24/7. We solve a mixed integer programming problem to figure out the optimal placement, given the resource constraints. How do we allocate the pool of machines in the cheapest possible way? So this is a fun, interesting thing. ### There's a fixed number of GPUs in the world. There are fewer GPUs than the number of workloads in the world, I think. that a true statement? The fundamental economics of GPUs is that most of the cost goes to Nvidia. To recoup that cost, you need to run them at very high utilization. Most of the CPU cost is power, which means that it's a more variable cost, which means that you don't really care about utilization of CPUs as much. So AWS could just over-provision and run things at much lower utilization. But GPUs to make the economics work, you need to run that at high, high utilization. Hopefully GPU prices will come down. That's what I'm hoping for. But right now, it’s a supply and demand problem. ### If somehow GPU prices come down over time and they become more like x86 processors in the way that the market works, do we still care about all this hard work that you're doing to combine resource pools of GPUs across multiple availability zones? Maybe it becomes less relevant, but on the other hand, the value of the platform then becomes more important. I think a lot about this for a lot of AI startups. Are you long on GPU prices or are you short on GPU prices? If GPU prices go up, what happens to the value of your company? Does it go up? The truth is if GPU prices were to crash, It would be hard for us in the short term because we have a bunch of long-term contracts and the revenue would go down quicker. But I actually think in the long run, it would be good for us. Because having an abundance of GPUs is very good for customers. It's good for the world. But I also think for a lot of infrastructure providers, in a way, we focus on the software, not the hardware. And if the underlying cost of the hardware goes down, the relative value of the software goes up. ### Let's talk about egress fees. One of the driving forces of being cross-cloud, and cross-availability is GPUs. In the past, one of the reasons not to do that had been that it's just really expensive to move your data around. I think in certain cases that's still true, but it's starting to change. Can you say more about what's happening with the egress fees? AWS still has very high egress fees, but they're coming down. I think there's a lot of pressure coming from R2 and Cloudflare, for instance. The interesting thing about Gen AI is that egress fees don’t really matter that much. And that's been a weird re-architecture of a lot of compute. Part of why we've been able to build a multi-region, multi-cloud architecture is that if you think about it, something like Gen AI doesn't need a lot of bandwidth. It's very compute-hungry. The bandwidth data is actually very little compared to the compute. I have a feeling that over time, egress fees will come down and region distinction will matter less and less, except for latency-sensitive applications. But it turns out a lot of stuff is actually not that latency sensitive. ### One of the biggest conversations in the data industry today is what's going on in the file format catalog wars—Unity, Iceberg, Delta—and there's a lot of focus on making sure that different systems can talk to each other. But one of the things that we're not talking about yet is that inevitably if you have a global company, you're probably not using a single cloud provider in a single availability zone. So you also have to solve a fabric issue of where the data is physically located. It's not just a format thing. Yeah, that seems hard. ### Okay, let's go from the future of the cloud to the future of data engineering. You spent a long time as a data engineer or building tooling for data engineers. I was at Spotify for seven years. I built a lot of music recommendation systems. And then I was at a company called Better for many years as the CTO. I've been focused on data, AI, machine learning. The precursor to all of that, or the prerequisite is that you got to get to data. And so that sort of necessitates doing a lot of data engineering, data cleaning, building data pipelines. I ended up building my own workflow called Luigi. ### No one uses it today, but you're skipping by a time period in which a lot of people used it. A lot of people used it 10 years ago. I was deep in that swamp. Data engineering is funny because, in a way, I kind of don't want it to exist. My prediction has always been it's going to go away at some point. ### I want to explore this with you because in many ways, I agree with you. At the same time, there are a ton of humans in the world that call themselves data engineers that use dbt. And so the last thing in the world I want to do is like say that data engineers suck because I don't believe that. The funny thing, though is that technology progresses over time and the jobs to be done that humans need to do change. I've seen this so many times a team of data scientists benefiting tremendously from just injecting a lot of data engineering skills. Suddenly they can get the data, I don't really like the idea of it being someone's job to shuffle data around. I want everyone to think about what the business needs and to build business applications. I would say the same thing about any internal platform team. All these internal platform teams tend to be somewhat ephemeral and transient. All these titles too, right? There are data engineers, data scientists, and analytics engineers. To me, it doesn't matter. There are always going to be people who need to work with data, AI, machine learning, and that slice is going to grow and grow and grow. But the actual composition of those teams is going to change a lot. And so I don't really pay that much attention to titles. ### I just wrote a white paper on the analytics development life cycle. And there are three different jobs to be done in the analytics development life cycle. There's a developer—people who create reusable assets for other people. There's an analyst—people who interact with the data to try to draw conclusions about the real world. And then there's a decision maker—people who get the recommendations from the analyst and make decisions. If you start slicing it up more than that, I think you inject friction into the process as opposed to adding clarity. I have to compliment you on creating that label analytics engineer. Because at that time there were so many people out there unsure of how they fit into their organization. You brought them an identity. That was eye-opening for so many people and created a sense of belonging in a community. ### It was my favorite thing was when people told me that they got an analytics engineer title and their pay went up 50%. I was like, great, but you're doing valuable work. You should be paid for it. You gave a lot of people recognition and I think that you should get a lot of credit for that. ### We've talked about data engineering is a valuable thing to be done. There's a trajectory here where probably fewer humans need to turn knobs and dials. What about machine learning and AI? So there are ML engineers. More and more people describe themselves as AI engineers. Do we need specific titles for these things or are we all just software engineers? It's a difference between AI engineers and machine learning engineers, that AI engineers use TypeScript and machine learning engineers use Python for the large part. A lot of the recent stuff around LLMs—not to trivialize it—but it's a low-code machine learning tool for people for which machine learning felt unapproachable. Suddenly they were given a new tool and they could stitch a bunch of prompt engineering and get it to work. ### Let's wrap up with the question that I love to ask everyone. What do you hope will be true of the data AI software industry over the coming five years? I hope that people will never have to think about containers, infrastructure, provisioning, or resource management. I really hope that all of that will be abstracted away in the next five or 10 years. That people will focus on application code and business logic and building cool AI shit, and then rely on infrastructure to take care of all the other stuff. --- --- title: "Prove ROI for data analytics initiatives" description: "Discover how to quantify the ROI of analytics with metrics, case studies, and best practices to make your data investment count." url: "https://www.getdbt.com/blog/analytics-roi-best-practices" date: "2024-11-01" authors: ["Joey Gault"] categories: ["Pulse"] --- # Prove ROI for data analytics initiatives Before diving into ROI measurement, it's important to understand the specific problems that modern data analytics initiatives solve. Organizations typically struggle with several interconnected issues that create measurable inefficiencies. [Data transformation processes](https://www.getdbt.com/blog/data-transformation) often rely on outdated methods, including hand-coded stored procedures or drag-and-drop tools that lack transparency and governance. These approaches lead to inconsistent data transformation methodologies across teams, resulting in duplicated work and conflicting results. The extensive rework required to reconcile these inconsistencies translates directly into lost productivity hours and delayed project timelines. Trust in data erodes when analysts and business users cannot trace data lineage or understand how metrics are calculated. This lack of confidence forces teams to spend significant time validating results rather than generating insights. The cumulative effect is missed deadlines, frustrated stakeholders, and reduced confidence in data-driven initiatives across the organization. These challenges create quantifiable costs that can serve as the baseline for ROI calculations. Teams spend excessive time on data preparation rather than analysis, engineering resources are consumed by repetitive data requests, and business decisions are delayed while teams verify data accuracy. ## Establishing measurement frameworks The most effective approach to proving analytics ROI involves establishing clear measurement frameworks before implementation begins. This requires identifying specific metrics that can be tracked consistently over time and establishing baseline measurements that reflect current state inefficiencies. Developer productivity represents one of the most measurable areas of impact. This includes tracking the time required to complete common data transformation tasks, the frequency of data pipeline failures, and the effort required to implement new data models or metrics. By measuring these activities before and after implementing modern analytics practices, teams can demonstrate concrete productivity improvements. Data quality metrics provide another quantifiable dimension. Organizations can track the frequency of data issues, the time required to resolve data quality problems, and the number of support tickets related to data discrepancies. Improvements in these areas translate directly into cost savings and increased confidence in analytical outputs. Collaboration efficiency offers additional measurement opportunities. This includes tracking the time required for cross-team data projects, the frequency of conflicting metric definitions across departments, and the effort required to onboard new team members to existing data processes. Modern analytics approaches that emphasize standardization and documentation typically show significant improvements in these areas. ## Learning from independent research The most compelling ROI evidence comes from independent third-party research that examines real-world implementations across multiple organizations. [Forrester Consulting's Total Economic Impact](https://www.forrester.com/policies/tei/) study provides a valuable benchmark, having examined organizations across diverse industries including construction, life sciences, energy and utilities, and B2B software. The study methodology involved creating a composite organization with $2 billion in annual revenue, 25 data engineers, and 200 data analysts. This approach allows for standardized comparison while accounting for the scale effects that influence ROI calculations. According to a [Forrester Consulting Total Economic Impact study commissioned by dbt Labs](https://www.getdbt.com/resources/study-forrester-tei), the composite org achieved 194% return on investment, with breakeven achieved within the first six months of implementation. The specific benefits identified in the study provide a template for measuring ROI in other organizations. Developer productivity increased by 30% through accelerated workflows and reduced context switching. Data rework time decreased by 60% as teams moved away from manual, error-prone processes toward automated, testable data pipelines. Data analysts experienced a 20% reduction in time spent on data gathering and preparation, allowing them to focus more time on actual analysis and insight generation. Data transformation costs decreased by 20% through more efficient processes and reduced compute waste. Perhaps most importantly, every organization surveyed reported increased trust in data and faster time-to-business value. While these qualitative improvements are harder to quantify directly, they often represent the most significant long-term value creation. **** ## Calculating direct cost savings Direct cost savings provide the most straightforward component of ROI calculations. These savings typically fall into several categories that can be measured and tracked consistently. Labor cost reductions represent the largest category of direct savings. When data engineers spend less time on repetitive tasks and data analysts require less time for data preparation, organizations can either reduce headcount or redirect existing resources toward higher-value activities. The Forrester study found that organizations could avoid hiring additional data engineering resources as data volumes grew, representing significant cost avoidance. Infrastructure cost optimization provides another source of direct savings. Modern analytics approaches often reduce compute waste through more efficient query patterns and better resource utilization. Organizations can track cloud computing costs before and after implementation to quantify these savings. Reduced rework costs offer additional direct savings. When data pipelines are more reliable and data quality issues are caught earlier in the process, organizations spend less time and resources fixing downstream problems. This includes both the direct cost of engineering time and the indirect costs of delayed business decisions. ## Measuring productivity improvements Productivity improvements often represent the largest component of analytics ROI, but they require careful measurement to avoid overstating benefits. The key is focusing on activities that can be measured consistently and that represent meaningful business value. Time-to-insight metrics track how quickly teams can answer new business questions or implement new analytical capabilities. Organizations should measure the complete cycle from initial request to delivered insight, including data discovery, transformation development, testing, and deployment. Improvements in these timelines directly translate into faster business decision-making. Self-service capabilities reduce the burden on centralized data teams while enabling business users to answer their own questions. Organizations can track the number of ad-hoc data requests handled by engineering teams, the time required to fulfill these requests, and the frequency of follow-up questions. As self-service capabilities mature, these metrics typically show significant improvement. Collaboration efficiency improvements can be measured through project completion times, the frequency of cross-team data projects, and the consistency of metric definitions across departments. When teams can build on shared data models and common definitions, project timelines compress and results become more consistent. ## Quantifying risk reduction Risk reduction represents a significant but often overlooked component of analytics ROI. While these benefits are harder to quantify than direct cost savings, they often represent substantial value creation over time. Data governance improvements reduce compliance risk and increase confidence in regulatory reporting. Organizations can track the frequency of data governance issues, the time required to respond to audit requests, and the consistency of regulatory reporting across different systems. Improvements in these areas reduce both direct compliance costs and the risk of regulatory penalties. Operational risk reduction occurs when organizations can identify and respond to business issues more quickly. This includes detecting fraud, identifying operational inefficiencies, and recognizing market opportunities. While these benefits are harder to measure directly, organizations can track the frequency of issues caught through data monitoring and the speed of response to identified problems. Decision-making risk decreases when business leaders have access to more reliable, timely data. Organizations can track the frequency of decisions that need to be revised due to data quality issues and the confidence levels of executives in data-driven recommendations. Improvements in these areas reduce the risk of poor strategic decisions. ## Building the business case Creating a compelling business case requires combining quantitative measurements with qualitative benefits in a framework that resonates with executive stakeholders. The most effective approaches focus on business outcomes rather than technical capabilities. Start by establishing clear baseline measurements across the key areas identified above. This requires collecting data on current state performance before implementing new analytics capabilities. Without solid baseline measurements, it becomes impossible to demonstrate concrete improvements. Project benefits conservatively, especially in the first year of implementation. The Forrester research shows that organizations typically see results during the first year as teams adapt to new tools and processes, with ROI increasing as maturity grows. Conservative projections build credibility and create opportunities to exceed expectations. Include both direct cost savings and productivity improvements in ROI calculations, but be explicit about the assumptions underlying each category. Direct cost savings are easier to verify and should form the foundation of the business case. Productivity improvements often represent larger potential value but require more careful measurement and validation. Account for implementation costs comprehensively, including not just technology licensing but also training, change management, and the opportunity cost of team time during implementation. The Forrester study found breakeven within six months, but this timeline assumes proper planning and execution. ## Measuring long-term value creation The most significant analytics ROI often comes from long-term value creation that extends beyond immediate cost savings and productivity improvements. These benefits require longer measurement periods but often represent the most substantial business impact. Strategic decision-making improvements occur when organizations can identify market opportunities, optimize operations, and respond to competitive threats more effectively. While these benefits are harder to quantify directly, organizations can track business outcomes that correlate with improved analytics capabilities. Innovation acceleration happens when teams can experiment with new ideas more quickly and validate hypotheses with data. Organizations can track the number of new initiatives launched, the time required to test new concepts, and the success rate of data-driven experiments. Competitive advantage develops when organizations can respond to market changes more quickly than competitors or identify opportunities that others miss. While this advantage is difficult to measure directly, it often represents the most significant long-term value creation. ## Conclusion Proving the ROI of data analytics requires a systematic approach that combines direct cost measurements, productivity tracking, and long-term value assessment. The key is establishing clear baselines, measuring consistently over time, and focusing on business outcomes rather than technical capabilities. Independent research provides valuable benchmarks, with studies showing ROI of 194% and breakeven within six months for organizations that implement modern analytics practices effectively. However, these results require proper planning, execution, and measurement to achieve. The evidence is clear that modern analytics approaches, particularly those that emphasize governance, collaboration, and software engineering best practices, deliver substantial returns on investment. The challenge for data engineering leaders is not whether these investments provide value, but rather how to measure and communicate that value effectively to drive continued organizational support and investment. But tools alone don’t make the difference—architectures and practices do. That’s where dbt shines. With [dbt](https://www.getdbt.com/product/dbt) you get built‑in testing, version control, data lineage, and semantic consistency—all of which give you the guardrails and governance needed to scale analytics with trust. When changes happen, you can see what broke, why, and when. When metrics are defined, they stay consistent across teams. And when analytics ROI is under scrutiny, your pipeline becomes a transparent, auditable asset—not a black box. In short: measure everything you can, change what you must, and operate with the kind of disciplined transparency that inspires confidence. Use dbt not just to build pipelines, but to build trust. The organizations that do this well will be the ones that sustain analytics as a competitive advantage—not just a cost center. --- --- title: "Coalesce 2024 edition: What is next for data teams?" description: "Brooklyn Data Co. Founder Scott Breitenother joins Tristan Handy in conversation at Coalesce 2024 in Las Vegas" url: "https://www.getdbt.com/blog/coalesce-2024-edition-what-is-next-for-data-teams" date: "2024-10-21" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Coalesce 2024 edition: What is next for data teams? _This post first appeared in [The Analytics Engineering Roundup](https://roundup.getdbt.com/p/coalesce-2024-edition-whats-next). _ Scott Breitenother, founder of data consultancy Brooklyn Data Co., joins Tristan at Coalesce 2024 in Las Vegas to discuss the early days of dbt, the evolution of data teams, and what's next for the dbt community. **** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Google Podcasts](https://podcasts.google.com/feed/aHR0cHM6Ly9hbmFseXRpY3NlbmdpbmVlcmluZ3JvdW5kdXAubGlic3luLmNvbS9yc3M) - [Stitcher](https://www.stitcher.com/show/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways from this episode ### **So, you have some news to share that I think you've been talking with some different folks about. So, I don't think we're going to break this live or anything, but it's a little bit of new information. Do you want to get it out there? Sure.** **Scott Breitenother:** I started Brooklyn Data in the summer of 2018. We grew rapidly, I think big thanks to dbt and all the success the dbt community has had. We built something really special. And about a year and a half ago, we went through an acquisition. We were acquired by Velir, which is an amazing digital agency based out of Boston. I couldn't have asked for a better home for Brooklyn Data. For the last year and a half, I've been working hard on integrating. Any acquisition is like a marriage. I feel like I can do a whole podcast about integrations and acquisitions and everything I've learned. But the story is that tue integration has been great. We're one company, and at the end of October I'm going to be stepping back from the day-to-day. I'll still be a board member of Velir and Brooklyn Data. I'm going to be involved in the strategy, the big picture, all that kind of fun stuff. But I won't be doing any more hands-on work, and I’m not quite sure what's next. No plans. So I feel like I'm gonna have to ask you for some help on coming up with hobbies. ### **When you were at Casper, you were very involved in dbt and the dbt community, and then you started Brooklyn Data. I remember asking how you were managing to be a husband, a new father, and a founder. What did it take to get this thing off the ground?** I think the answer is I didn't sleep a lot. You hear the saying it isn't work if you love what you're doing. And I mean, it's true. The reason I have to find a hobby is because building Brooklyn Data has been a hobby. Any free processor time I had on my mental CPU, I was thinking about new and innovative ways to grow Brooklyn Data. Not because I was on the clock, not because I had to, but because I loved it. We met when you were working with us at Casper and helping us set up dbt. And I felt very privileged to, I don't know, be in the know on something like dbt. It was special and cool. I told all my friends because dbt was great and it was changing my life, my team's life, and the data scene, especially in New York. We had a lot of people asking us for advice because at that point we had probably one of the more sophisticated dbt setups in 2017, 2018. There was a point where we were like, there's a business here. There were very smart people working on very cool pieces of software—dbt, Snowflake, Looker at that time. But there was a gap in professional services to guide companies on this journey to get the most out of these tools. ### **The classic consulting companies had not really woken up to this, and you couldn't just show up to Accenture and say, “Help me build a modern data stack and move faster.” Do you think that the McKinsey of the modern data environment is just McKinsey?** Is there still a great independent consultancy to be built out of this era of data? I don't think there’ll be a pure-play large-scale data consultancy. It's too early. IYou need to be a broader agency. What's happened over the years is that the industry has just matured, and you need scale for brand awareness, to scale to invested systems and account management and sales teams and sponsoring conferences and all the things that, you know, when it was just me and my Brooklyn apartment coding I couldn't have even imagined or afforded. And so you need that scale. One of the reasons we chose to join Velier, which is a digital agency, is you need multiple complementary service lines. Ddata work by it's very nature will ebb and flow. Not every organization is going to refactor their data warehouse every single year. Don't get me wrong, we have long, durable, amazing relationships with our clients. It typically starts with kind of a bigger project, then it evolves to different projects. And for many of our clients we have a long, durable, similar-sized relationship for many years. The best way to kind of keep those long durable relationships is to take that trust that you have built with this client and help them with other stuff. The beautiful thing about data is it plugs into everything. ### **I started my career at Deloitte, and I've now gotten to see things from the other side partnering with the largest consulting organizations in the world. Trust is just not replicable without years and years of relationship building.** People buy from people. With services, the people are the product so it's doubly true. I reinforce to my colleagues at Brooklyn Data that we're somewhat in the hospitality business—we're in the surprise and delight business. Listen and read behind the lines, give people what they want, but also give them what they need. Truly successful consultancies are empathetic. At the end of the day, there is going to be a transactional relationship to any third-party consultancy, but it should feel as much like a partnership as possible, and I think we've done that. A lot of our clients call us a consultancy that doesn't feel like a consultancy, and I think that’s a compliment. ### **You are a little bit emblematic of a data practitioner of a particular era. You started as an econ person, right?** Yes, I studied business. I did strategy consulting for four years. Excel, PowerPoint. And it was very quantitative. Very quantitative. No SQL. I build economic forecast models in Excel. I remember waking up at 2 in the morning early in my career with Excel nightmares. Now I look back and I see the testing and all the version control that you have in data would have saved a lot of heartache early in my career. ### **Your Excel macros were strong?** I’m extremely good at Excel shortcuts. ### **In 2009, I remember hanging out with some friends and someone asked, “What is the thing you think you're top 10,000 in the world at?” And I said I'm really good at Excel. Everybody else had better answers, but it’s a point of pride.** I mean, it's still my happy spot. I love when I code dbt. I don't do it much very often, but when I do, it's very tangible, very fulfilling. I love when I do Excel. And no one's looking, but I'm still gonna format it really nicely. No one's gonna see it. This is just for me. ### **I really like it to always be the correct font size course across the entire sheet. And then the** One hundred percent. ### **You got a job running data at Casper. How did that happen?** Yeah, that's a head scratcher for me. I'm not quite sure how I got it. I was living in London at the time, working consulting. Moved to New York, had taken a sabbatical from my consulting job and decided I was going to work in tech. And I didn't know what that meant. I just literally went to meetups and reached out to anybody I had the most tenuous connection with on LinkedIn, went to a meetup every night for months. I do really, really recommend to everybody, networking is key. Build your network. Go to meet people. Actually meet people. Meeting people in real life builds a much stronger relationship than digitally. A friend of a friend introduced me to the founders of Casper. They brought me on as a freelancer and I remember in the interview, they asked, “why should we hire you?” And I said, “You don't want to hire someone that knows the tool and only knows the tool, you want to hire someone who knows how to think and can learn any tool.” ### **You were the first person data hire?** Yep, I was employee number 16 at Casper. I was there for four years. The data team grew to 15 or 16 people. The Casper experience was amazing. I learned so much. It set my whole career up. People talk about the PayPal mafia, I think there's a Casper mafia. Everybody is doing very cool things. Someone took them private. I mean, I still buy Casper mattresses. I don't know anybody that works there, but I buy it out of brand affinity. ### **I bought the top-of-the-line mattress.** Oh, I'm still cheap. It's supposed to keep you cooler. Does it work? ### **I love it. It's actually worth it.** You're convincing me. ### **We had a podcast about data and now we're talking mattresses.** ### **Here's the reason that I wanted to go back in time to this. Historically, data teams weren’t a thing. There were IT organizations that** did data things because executives needed dashboards. But all of a sudden, all the tools were changing, and there were startups with existing data infrastructure. You got the opportunity to take people who understood strategy and give them control of the infrastructure. And it produced very interesting things. I agree. The evolution of the Casper data team exemplified that. When we started, everybody in the Casper data team looked a lot like me, former management consultants. And actually we looked a lot like our stakeholders. We’d start to specialize, so we'd have people on the data team that’d report with the marketing team, and they’d go to marketing meetings and be embedded. But we got to the point, I’d say two years in, where the scale of the problem, of the data, of the complexity of the infrastructure we were managing got to a point that we saw a bit of a divergence and evolution of the team where we started seeing more STEM degrees. More people with engineering or tech backgrounds focused on writing good, clean pull requests. I look back and those people that I'm describing, those were early analytics engineers. We organically saw the split of the analyst and the analytics engineer. I think your vision is to get them all speaking a common language. And that was SQL. That was a big aha moment six, seven years ago, when there was a clear debate, Python versus SQL. And we all forget about that. And I'm not saying that Python lost, but SQL has very much become the language of communication between analysts and engineers. ### **What do you think has created the opportunity for change over the past decade? Because things in data have changed a lot.** Compute and storage got cheaper. If you strip away all the things, that's the thing. Cloud driving compute and storage to be cheap for you to not just save some data, just save it all. Let's not look at five days of data. Let's look at all the data. I think Redshift was very much the big unlock. It's not like, because I remember early days I was chatting with, you know, I think Fivetran made it easier to bring data in. dbt made it easier to transform and scale and use GitHub and version control. So it's just all these technologies that allow people to do something easier and easier and cheaper and cheaper. ### **What are the companies that are going to get created, like the professional services companies, what are they going to get created based on now?** The modern data stack wave was gigantic, and I feel very privileged to be at the right, exact right place at the right time when it happened. Kind of like the dbt community, I don't know if I will ever see something like that in my lifetime again. You see people talking about AI, of course. ### **Are there boutique AI consulting shops?** I haven't looked around, but I'm sure, there has to be, right? I think the only difference maybe with AI, though, is the modern data stack flew under the radar for many years. Even today, we win against the big global consultants because we know the tools better. I've been using dbt since 2016. Yeah. You know, it wasn't enterprise ready back then. It wasn't on the radar. You look at Accenture making one to two billion dollars a year on AI work, that didn't exist. The whole AI industry, we all agree that there's a there there. But I also think we all universally agree that 90 percent of the startups we see will fail. And these VCs are paying really high valuations for investments. I think a lot of people are going to get washed. But they'll also find that one next thing, too. It's going to be beautiful, creative destruction in the AI space. ### **I'm bullish on the underlying technology versus I'm bullish on the equity returns currently.** One hundred percent. I'm reading a book right now on what moves markets. Not many people made money on canals and railroads. They were hugely valuable for America. But a lot of people got wiped out. It’s all about timing. ### **Do you have any thoughts on Iceberg? Are you finding customers engaged in this?** As a conceptual concept, yes. I don't think any of the use cases are perfectly ready for primetime. I think everyone's excited about the direction. It's really exciting to see Snowflake support Iceberg, both as a native format and external tables. And you play that forward, theoretically, you know, you're abstracting storage. May the best compute platform win. Everybody I talk to is excited about this. But I think they're all in a holding pattern. They're experimenting it. They're using it for ad hoc use cases. ### **I don't know of anybody with their entire data lake on Iceberg and they use eight different compute platforms.** But I'm excited to experiment. I'm excited to flip on Iceberg as a materialization method in dbt and play around with it. The team is already starting to test into it. I think there might be a moment where we go to our clients and say it's hit the tipping point and we should convert everything to Iceberg and start thinking about it from an Iceberg first mindset when doing architecture. ### **Fivetran has gone very far at Iceberg. They're making it very easy to take data and deliver it directly to Iceberg and then you can use it from there however you want.** That aligns, because when we bring Fivetran into enterprise clients a lot of them do want some sort of landing zone in some sort of cloud storage before bringing it into their data warehouse. And I think that speaks to both an Iceberg trend and just a need in the enterprise to have a separate landing zone. ### **One of the other interesting things that I've heard this week is about dbt Semantic Layer. The thing that I'm hearing over and over again is people saying they are finally ready to make the investment to do this.** ### **Have you seen conversations with clients where there's an interest in a semantic layer that spans BI tools, or is this another conversation like Iceberg, where in theory sounds great, but we're just like not there yet.** One hundred percent there's interest. People are really excited about doing this. There are a few things that have happened. It was only GA last summer, right? There is the network effects of all the integrations. One of our clients works at a large enterprise, and when you announced the Power BI integration, people were really excited about that. That's game-changing. It does definitely feel like a moment. It’s much more tangible and less theoretical than Iceberg at the moment. ### **I'll take that as a compliment. Okay, so let's close on the community. You have seen a real journey from a couple folks in a room at Casper. Here at Coalesce there are people who this is the first time they've been a part of a community that's outside of the people that they work with at a large enterprise.** ### **What do you think this group of people needs next? Whether it's from their software or maybe more importantly, from each other. What is the dbt community?** First I’ll step back and just say how grateful I’m just to be an early member. You say the dbt community, I just say my friends. I go to Coalesce for a few things. I go to see the product announcements. I go to see which vendors are cool. I go to see people. And I think for a lot of people that's why, because it's the community. I mean we love the Slack version, but nothing beats in person. I's just been such a wonderful experience. And what I've loved is how inclusive it is. Come as you are used to mean one thing, now it also means wear a button-down and a suit to Coalesce. Or come in a t-shirt. Be an executive at a Fortune 500 or be an analyst at a startup. The dbt community is very welcoming. dbt is a collection of people that have shared interests, shared challenges, and now have a common language to express those challenges, a common community to learn and interact. The community is only going to get bigger. The personas are only going to get more varied. I really love the [One dbt aspect](https://www.getdbt.com/blog/coalesce-2024-product-announcements) because it's community but it's also the tool. There are people that are never going to learn SQL. There are people that might be able to but need an onramp that's a little less intimidating than the command line. And that's where the Cloud IDE came in. And so now the drag-and-drop is going to be really interesting, and it's a good opportunity to make the tool more welcoming. It's going to be very interesting to see what the talks look like next year. Because there's going to be talks for people that literally know dbt only as a drag-and-drop interface. It's a different community. And so we just need to keep welcoming and open to all the different personas. --- --- title: "dbt Labs’ Coalesce 2024 Unifies Data Practitioners and Data Leaders With One dbt" description: "Coalesce 2024 brought more than 11,000 data enthusiasts together to introduce One dbt, a commitment to a unified dbt experience." url: "https://www.getdbt.com/blog/dbt-labs-coalesce-2024-unifies-data-practitioners-and-data-leaders-with-one-dbt" date: "2024-10-21" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs’ Coalesce 2024 Unifies Data Practitioners and Data Leaders With One dbt _New Salesforce collaboration complements cross-platform, persona and cloud vision_ **PHILADELPHIA**, October 21, 2024: [Coalesce 2024](http://coalesce.getdbt.com) – hosted by [dbt Labs](http://getdbt.com), the pioneer in analytics engineering – brought together more than 11,000 data enthusiasts in Las Vegas and virtually to share its vision for One dbt: a commitment to a single unified dbt experience, regardless of persona, data platform, or cloud an organization uses. Three days of engaging keynotes, interactive training, breakout sessions, and industry executive roundtables focused on realizing the business value of the [analytics development lifecycle](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) (ADLC) and the opportunities for AI. During the event, dbt Labs announced [significant product enhancements](https://www.getdbt.com/blog/coalesce-2024-product-announcements) to expand dbt Cloud to become a data control plane. As a data control plane, dbt Cloud works across platforms and supports users across every stage of the ADLC—regardless of their title, technical aptitude, chosen data platform, or where they build and consume data. Data leaders from [Amazon Web Services](https://aws.amazon.com/), [Bilt Rewards](https://www.biltrewards.com/), [Fifth Third Bank](https://www.53.com/content/fifth-third/en.html), [Roche](https://www.roche.com/), and [Siemens](https://www.siemens.com/global/en.html) took the stage during the event’s keynotes to share stories of dbt’s transformative impact on their organizations’ data practices and their careers. Product announcement highlights include: - [**dbt Copilot**](https://www.getdbt.com/blog/introducing-dbt-copilot), the AI engine in dbt Cloud that helps users accelerate their analytics workflows. dbt Copilot is designed to automate tasks that previously required repetitive manual work, significantly improving productivity, data quality, and stakeholder trust. - A new **[visual editing experience](https://www.getdbt.com/product/develop), **a low-code, drag-and-drop environment for building and exploring dbt models, designed to democratize the ADLC to more types of users and enhance collaboration. - [**Cross-platform dbt Mesh**](http://getdbt.com/blog/introducing-cross-platform-dbt-mesh), built on the emerging standard in open table formats, Iceberg, allows dbt to act as the lingua franca for how data pipelines are defined enterprise-wide regardless of the underlying data platform. dbt Mesh’s new cross-platform capabilities eliminate data silos while maintaining data governance in the increasingly complex, multi-platform environments that are dominant in enterprise settings. - [**Advanced CI**](http://getdbt.com/blog/announcing-advanced-ci), which allows users to compare code changes as part of the CI process to catch unexpected behavior before new code is merged into production. This improves organizational trust in data and helps optimize compute spend by only materializing correct models. In addition, dbt Labs unveiled [a new collaboration with Salesforce](https://www.getdbt.com/blog/dbt-labs-and-salesforce-announce-strategic-partnership) to unite Salesforce Data Cloud AI, automation and analytics solutions with dbt Cloud. According to Ali Tore, Senior Vice President of Advanced Analytics at Salesforce, the relationship is “... helping organizations make faster, more informed and trusted decisions by providing an integrated, end-to-end solution for data transformation, metrics management, and analytics—unlocking the full potential of their data to drive impactful outcomes.” Leaders from [Salesforce](http://salesforce.com) and [Tableau](http://tableau.com) reinforced the potential of these new integrations to provide a streamlined, trustworthy, end-to-end data experience for users. The dbt Labs Partner Program also made headlines during Coalesce, with the [announcement](https://www.getdbt.com/blog/dbt-labs-names-seasoned-channel-leader-shawn-toldo-vice-president-worldwide-partner-ecosystem) of seasoned channel leader Shawn Toldo as Vice President, Worldwide Partner Ecosystem. Toldo is joining dbt Labs on the heels of momentous program growth, as dbt Labs has tripled the number of implementation partners and doubled its technology partners since the Technology Partner Program’s official launch in 2022. At the event’s Partner Summit, dbt Labs honored nine partner organizations for their exceptional collaboration, go-to-market strategy, technical expertise, and deep commitment to customer success. Winners of the 2024 Partner Awards are: - [phData](https://www.phdata.io/), Global Services Partner of the Year; - [Aimpoint Digital](https://www.aimpointdigital.com/), Innovation Partner of the Year – Americas; - [Spaulding Ridge](https://www.spauldingridge.com/), Customer Impact Partner of the Year – Americas; - [Indicium](https://www.indicium.tech/), Emerging Partner of the Year – Americas; - [Xebia](https://xebia.com/), Services Partner of the Year – EMEA; - [Devoteam](https://www.devoteam.com/), Emerging Partner of the Year – EMEA; - [Deloitte](https://www.deloitte.com/au/en.html), Services Partner of the Year – ANZ; - [Ippon Technologies](https://au.ippon.tech/), Emerging Partner of the Year – ANZ; and - [Classmethod](https://classmethod.jp/english/), Services Partner of the Year – Asia/Japan. dbt Labs [celebrated](https://www.getdbt.com/blog/coalesce-2024-highlights) dbt Community members during Coalesce, presenting eight awards to honor their unique contributions to the dbt ecosystem. dbt users also weighed in on how dbt has transformed their careers and lives, leading to more than 200 handwritten responses pinned to the dbt Labs booth wall, inside the event’s Discovery Hall. “Coalesce is the manifestation of One dbt. It creates space to bring together the diverse types of people, conversations, and innovations that make dbt a unifier,” said Tristan Handy, founder and CEO at dbt Labs. “The thousands of people who attended in person and online are all part of it – one community, one movement, all aligned around the idea that analytics can be done better. I’m inspired by the incredible work at dbt Labs and in our vibrant partner and customer communities, all working toward this transformation.” The sixth annual installment of the event, Coalesce 2025, will take place in Las Vegas in fall 2025. To stay up to date on calls for speakers, ticket availability and next year’s agenda, visit [https://coalesce.getdbt.com/coalesce-2025](https://coalesce.getdbt.com/coalesce-2025). For a deep dive on the big announcements and themes unveiled at Coalesce 2024, visit [https://www.getdbt.com/resources/webinars/one-dbt-the-control-plane-for-data-collaboration-at-scale-virtual-event](https://www.getdbt.com/resources/webinars/one-dbt-the-control-plane-for-data-collaboration-at-scale-virtual-event). **About dbt Labs** Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 50,000 teams using dbt every week. To learn more about dbt Labs, visit [getdbt.com](http://getdbt.com) and follow us on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/). --- --- title: "Common challenges to scale data operations" description: "phData's Dakota Kelley provides practical advice for data leaders scaling their data operations." url: "https://www.getdbt.com/blog/common-challenges-to-scale-data-operations" date: "2024-10-16" authors: ["Dakota Kelley"] categories: ["Insights"] --- # Common challenges to scale data operations _This is a guest post by Dakota Kelley, senior solutions architect at [phData](https://www.phdata.io/)_ Scaling data operations within an organization is no small feat. As teams grow, processes become more complex, and the volume of data will expand exponentially. This often leads to several challenges that will slow down your progress and create massive challenges. Here are some of the most common roadblocks that companies face when trying to scale. ## Lack of standardization One of the most significant hurdles to scaling is the absence of consistent standards across modeling, development, and processes. When each team or individual operates under a different set of rules or lacks clear guidelines, workflows can vary wildly across the organization. This lack of standardization makes it difficult to collaborate, track progress, and ensure the quality of data outputs. This results in disjointed efforts, duplicative work, and errors that could easily be avoided with a standardized framework. Often leading to fatigue and burnout in your data team. ## Unclear ownership Another common issue is the absence of a clear owner for data initiatives. In many organizations, overwhelmed data teams are tasked with handling an influx of requests and projects. However, with no designated responsibility for overseeing specific datasets or processes, the data produced may not fully meet the needs of the requesters. This often leads to a cycle of finger-pointing between the data team and business stakeholders, as both groups avoid taking responsibility for the outcome. The lack of accountability stalls progress and can undermine ‌trust between teams. ## Inefficient workflows Without streamlined workflows, productivity can take a major hit. Inefficiencies in the way tasks are handled lead to delays in decision-making and increase the likelihood of errors. Moreover, these slowdowns reduce organizational agility, making it harder to respond to market changes or customer needs in real-time. Inefficient workflows also contribute to employee frustration, which in turn impacts team morale and performance. ## Minimal operational oversight Operational oversight is crucial for identifying issues and addressing root causes, yet many organizations struggle to implement it effectively. When metrics aren't tracked, or issues go unmonitored, it becomes nearly impossible to diagnose problems, let alone prevent them in the future. This lack of insight hinders the ability to perform meaningful analysis, which is critical for driving continuous improvement. Without operational oversight, organizations miss out on valuable opportunities to optimize processes and enhance performance. ## Embrace the ADLC Successfully scaling data operations requires addressing these common challenges head-on. Which is done best by embracing the Analytics Development Lifecycle (ADLC). **** ## Set and enforce standards > “You should shoot for high standards and believe they’re obtainable.” - Buster Posey ### Business impact Establishing clear standards and codifying processes can significantly improve operational efficiency and drive meaningful business impact. By creating a unified framework, organizations can reduce redundancy, streamline workflows, and eliminate unnecessary rework. This structure not only enhances productivity but also makes onboarding and task transitions smoother for employees. This means that teams can quickly adapt and contribute to new projects. Moreover, the collaborative nature of developing standardized solutions fosters cross-team alignment, accelerates development cycles, and often results in higher-quality outcomes. In short, standardization is a key enabler of scalability, innovation, and long-term success. ### Technical best practices Implementing standards in a repeatable way that can drive true business impact requires a variety of technologies. Tools like **SqlFluff** help enforce SQL style and coding standards, ensuring that queries across teams follow consistent, maintainable formats. **Meta-testing** further enhances this by setting conventions for naming, testing, modeling, and documentation, making it easier to collaborate and review work across the organization. A well-structured **Git workflow**, paired with **pull request templates**, streamlines the development process by encouraging thorough code reviews and reducing the chances of errors. Finally, by standardizing the **CI/CD pipeline**, teams can automate deployment and testing, reducing manual intervention and improving overall efficiency. Together, these best practices not only improve code quality but also accelerate development and ensure scalability. ## Provide clear ownership > “No one can come and claim ownership of my work. I am the creator of it, and it lives within me." - Prince ### Business Impact Clearly defining roles within data operations enhances governance by minimizing the risk of unauthorized changes and ensuring that the right people have control over critical data processes. By assigning subject matter experts (SMEs) with approval authority and ownership, organizations can foster workflow stability and robustness, ensuring that decisions are made by those with the most expertise. Furthermore, adopting a [**data mesh**](https://www.getdbt.com/product/dbt-mesh) approach empowers individual teams to own their data domains, promoting greater collaboration and efficiency. This decentralization allows teams to work autonomously, improving productivity while maintaining alignment across the organization utilizing cross-project references. Together, role clarity and a data mesh strategy drive both operational resilience and business agility. **** ### Technical best practices Establishing the appropriate processes, ownership boundaries, and hand-offs requires us to utilize our technology beyond the basic features. Using **Codeowners** within a Git repository ensures that specific teams or individuals are assigned ownership, making it clear who is responsible for approving changes. A well-defined **organizational structure** for Git workflows and approval processes further enhances accountability and streamlines development. This same principle should be applied to **dbt projects.** Establishing governance and ownership within teams ensures that models, tests, and documentation are consistently maintained and improved in a way that makes it easy to enforce with code ownership. By embracing data mesh—scaling horizontally with clear roles, structure, and standards—organizations can expand their data capabilities efficiently without sacrificing quality or collaboration. This creates a strong foundation for long-term scalability and success. ## Establish operational excellence > “Watch the little things; a small leak will sink a great ship” - Benjamin Franklin ### Business impact Improving operational efficiency requires a proactive approach to addressing recurring issues at their source, preventing them from resurfacing and disrupting workflows. By continuously tracking key metrics such as warehouse sizes, model run-times, and compute usage, organizations can optimize costs, ensuring resources are used efficiently without unnecessary overspend. Additionally, identifying patterns in data operations allows teams to update processes and standards regularly, ensuring high-quality output is maintained while mitigating potential issues before they escalate. This ongoing refinement not only drives cost savings but also boosts productivity and overall business performance. ### Technical best practices dbt artifacts, which drive the Discovery API and [dbt Explorer](https://www.getdbt.com/product/dbt-explorer), offer invaluable observability across all dbt projects, allowing teams to monitor performance, spot patterns, and detect problems across the organization. By leveraging these insights, teams can conduct rigorous root-cause analysis to swiftly identify and resolve recurring issues. This level of visibility is crucial for understanding the **financial impact** of technical decisions, particularly when it comes to **concurrency** and **scheduling**—factors that can unknowingly drive up cloud spend. By proactively addressing these inefficiencies, organizations can not only maintain operational stability but also optimize their cloud usage and reduce costs, ensuring smarter and more scalable data operations. dbt has a new Advanced CI feature that lets you see how code changes affect your data better. It does this by comparing the last production state to the latest pull request commit. It allows teams to see changes in primary keys, rows, and columns, either directly in dbt Cloud or through Git comments. Helping teams avoid introducing changes that result in major breakages or loss of data. Ensuring trustworthy data products with efficient operations. **** ## Excellence is a journey It’s time to collaborate, empower individuals, and focus on continuous improvement through innovation. --- --- title: "Coalesce 2024: Recovering from the party" description: "Coalesce 2024 was our biggest, announc-iest show yet. Here are my notes from the afterglow." url: "https://www.getdbt.com/blog/coalesce-2024-recovering-from-the-party" date: "2024-10-14" authors: ["Tristan Handy"] categories: ["Company"] --- # Coalesce 2024: Recovering from the party Coalesce 2024 was the fifth Coalesce. Every year gets a little tighter, a little bigger. And as we build our innovation engine and mature the underlying platform, the product announcements get ever-more-significant. We’re only a couple of days out from the event so I don’t have stats to share here. But if you’ve been to a few Coalesces, I will say that I think [Benn does a really fantastic job of summarizing](https://benn.substack.com/p/something-lost-something-found) the journey that the event has taken over the last five years. I don’t want to endorse every single thing in his post, but his core point—the transition from purely-community-vibes to yes-AND-also-with-ROI—is spot on. Travel budgets in 2022 were bountiful; if we still wanted people to show up in 2023 and 2024 we needed to demonstrate real ROI from the event. Fortunately, ROI doesn’t need to crowd out the quirky, the playful, the unique. We can build relationships, safeguard the dbt Community’s special place in the ecosystem, and still get shit done. Speaking of special, the photo below is almost too much for me. We had a spot in the Discovery Hall where we asked users how dbt has changed their careers and their lives. The responses were 😭😭😭 I realize that you cannot actually read the notes that folks left in this picture, but if you took the time to stop and write, please know that this is the stuff that keeps me and everyone at dbt Labs showing up at work every day. Thank you. So much 💜 ![How has dbt transformed your career and your life?](https://cdn.sanity.io/images/wl0ndo6t/main/7b569dce1e79d43433498e891d25acc1e0adcf58-2854x3461.jpg) Ok, on to the announcements. ## What Shipped at Coalesce 2024? This year was by far our biggest year for product announcements. If you want the entire laundry list, find it [here](https://www.getdbt.com/blog/coalesce-2024-product-announcements)—there’s way too much for me to hit in this newsletter. There were four big product themes: trust, collaboration, cross-platform, and AI. We expanded on our vision for the **Data Control Plane**: ![Data control plane](https://cdn.sanity.io/images/wl0ndo6t/main/dd8db82ba281b83cb762629f66e779c8bca50097-1771x780.png) And we wrapped all of this up into the singular theme of the entire conference: **One dbt**. But I’ll need to say more about One another day because…there is just a lot to say about it. For today I want to focus on the product announcements. Here’s the summary slide of the product keynote, hitting the various product announcements by theme: Not to pick favorite children, but here are the things I think are the most critical. ### Visual editor We welcomed new users into a visual editing experience that allows them to write dbt code using a drag and drop interface. This interface both reads and writes native dbt code, and therefore brings a brand new set of personas—less technical data practitioners—into the mature, governed, [ADLC](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) in a first-class way. This announcement got the biggest response of any announcement at the show. Almost everyone I spoke to throughout the week said “I have people at my company that I want to give that to.” I have never understood the “data prep” space as a standalone category. It seems like a space that has been stuck in time roughly a decade ago. Over the past year, many of our biggest customers have been asking us to help them solve their “data prep problem” (their words!). It shows up for them in two specific ways. First, license costs are high and only going up. Second, governance characteristics of these systems are incredibly low as they don’t conform to ADLC best practices AND they often need to pull data (including SPI) locally in order to operate on it. Having dozens or thousands of these licenses floating around inside your company is a real risk and we developed the dbt Cloud visual editing experience in direct response to customers asking for help mitigating this risk. This will be a multi-year investment for us; expect to continue to see innovations from us in this space. ### Cross-platform dbt Mesh [dbt Mesh](https://www.getdbt.com/product/dbt-mesh) was the biggest announcement of Coalesce 23, and this year we gave it a major upgrade. Downstream nodes in a mesh can now [be built on upstream nodes from a different data platform](https://www.getdbt.com/blog/introducing-cross-platform-dbt-mesh). Redshift can `ref()` Athena (etc). All of the implementation details are abstracted away; just use a [two-argument ref statement](https://docs.getdbt.com/reference/dbt-jinja-functions/ref#ref-project-specific-models) and it all just works. As we shared on stage, over half of our enterprise customers have dbt running on at least two data platforms. I fundamentally do not believe we are going to see one, or even two, winners in the data platform space. This is not Windows in the ‘90s, or even iOS and Android in 2012: the data platform ecosystem is not a monopoly or a duopoly; at best it is an oligopoly with 6-10 real players. But in reality I think it is better to just think about it as a competitive market. This is good for users—no one needs the Oracle-vs-Microsoft dynamic that existed in 2003 at the start of my career. But it also creates complexity and bifurcation. Because today, different teams that use different data platforms inside the same company typically do not know about or have any access to the data assets that live inside the other platform. This leads to duplication, inefficiency, and inaccuracy. Under the hood, dbt’s new cross-platform ref capabilities are powered by [its support for Iceberg](https://docs.getdbt.com/blog/icebeg-is-an-implementation-detail). Iceberg without dbt can be a real pain to use, but I am a _huge believer_ in its ability to move the market in practitioner-favorable ways. I’m delighted by our ability to abstract away the complexity behind a perfectly dbtonic interface. dbt’s cross-platform capabilities, and its support for open table formats, is still in its early days. I fully expect this to a major area of innovation for us in the coming years as the entire ecosystem is reshaped around this new set of standards. ### Two smaller, but lovely, features The above are what I believe to be the two biggest announcements coming from our biggest innovation themes. But I have two other product announcements that I just love so much that I just have to sneak them in. They are both real, important additions to dbt: - [Microbatching](https://docs.getdbt.com/docs/build/incremental-microbatch#what-is-microbatch-in-dbt) allows incremental model authors to load, and backfill, large models in date-based batches. It has been obvious that we would need to support this since all the way back in 2017 when [Max B](https://preset.io/about/) scolded me about dbt’s lack of support for this. All I can say is: we pushed ‘traditional’ incremental models a long way! It was time though. You should definitely implement this ASAP for your very largest tables. - [Advanced CI](https://www.getdbt.com/blog/announcing-advanced-ci) enables users far greater safety when merging a PR. The core of this feature is data diffing—being able to examine the differences of datasets produced in a development schema and the production schema. Zero differences, lots of safety in hitting ‘merge’. ![Coalesce party](https://cdn.sanity.io/images/wl0ndo6t/main/d50971bf69d054216aec03779a24648c25e32e4d-4000x2667.jpg) --- --- title: "Coalesce 2024 highlights" description: "The vision for One dbt and the entire dbt community." url: "https://www.getdbt.com/blog/coalesce-2024-highlights" date: "2024-10-11" authors: ["Daniel Poppy"] categories: ["Company"] --- # Coalesce 2024 highlights Coalesce 2024 is a wrap. Whether you were with us in person or online or catching up now, here’s what happened in Vegas at Coalesce 2024. Here’s hoping it stays with you wherever you are. A quick aside, I attended Coalesce for the first time last year. The part that really stood out for me was that people _really_ wanted to be at Coalesce. People were excited to be at Coalesce. How many conferences have you attended out of obligation? Each year the dbt Community gathers at Coalesce to be together with peers and friends. Yes, folks get practical advice for scaling their data organization, yes training, yes networking, all the the things. But there’s something incredible that this community just wants to be together. [Watch video](https://youtu.be/6b5dcns8IDM?si=eo_FXJX2l-B9ZXxR) My single favorite thing at Coalesce this year was the wall for the dbt community to share how dbt has transformed their lives. People shared truly life-changing stories. We had to clear the wall several times for folks to add new stories. This is real and personal to everyone, and such a delight to see. ![Note cards pinned to a wall with the text, "How has dbt transformed your career and your life?"](https://cdn.sanity.io/images/wl0ndo6t/main/f65f0d1032c132251f9b8061b0288aca5083bbc0-2831x1909.jpg) ## dbt Cloud product announcements [So. many. announcements](https://www.getdbt.com/blog/coalesce-2024-product-announcements). The overarching theme at Coalesce is One dbt, because dbt works across clouds, across data platforms, across teams, and across personas. This is a framework for everyone in the data ecosystem to solve analytics problems together. You can see this embedded in all the features for dbt Cloud—the data control plane for enterprise analytics—to orchestrate and observe your entire data ecosystem. ![a visual depiction of dbt Cloud's data control plane](https://cdn.sanity.io/images/wl0ndo6t/main/a0de6c37b3365c84936bed52331432645ec12f0f-1771x780.jpg) > “One dbt gave me goosebumps.” - Joseph Aranez, Principal Solutions Engineer at Trust & Will. dbt Cloud centralizes metadata across your business, and provides an interface for every data stakeholder to contribute. Here’s just some of what we announced: - [dbt Copilot](http://getdbt.com/blog/introducing-dbt-copilot) is an AI engine for data teams to accelerate their analytics workflows. dbt Copilot helps users get more done using the power of generative AI. It makes everyone interacting with data more productive. Automate away the mundane so you and your data team can focus on projects that matter. - [Cross-platform dbt Mesh](http://getdbt.com/blog/introducing-cross-platform-dbt-mesh): Support for the Apache Iceberg table format allows us to extend [dbt Mesh’s](https://www.getdbt.com/product/dbt-mesh) cross-_project_ references to now support cross-_platform_ references. This is crucial for maintaining data governance in the complex, multi-platform environments we’re seeing more and more of. Snowflake support is currently in beta, and Athena, Spark, Databricks, Starburst/Trino, and Dremio are GA. - dbt Cloud’s new visual editing experience is a drag-and-drop environment to build and explore dbt models. This new interface makes it easy for less technical people to participate in the data development process, BUT it’s also still dbt. That means all models are tested, documented, and version controlled, and you can see the SQL as the projects build. Your data team can see what your business partners are creating and work with them more closely on projects. - [dbt Core 1.9](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.9): You can now use the new [microbatch](https://docs.getdbt.com/docs/build/incremental-microbatch#what-is-microbatch-in-dbt) strategy to optimize your largest datasets, and a streamlined [dbt snapshot](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.9#snapshots-improvements) configuration. - For a complete list of our product announcements, [check out our product launch post](https://www.getdbt.com/blog/coalesce-2024-product-announcements) for everything from [Advanced CI](https://www.getdbt.com/blog/announcing-advanced-ci), adapters for Teradata and AWS Athena, and auto-exposures with Tableau. ## Salesforce Data Cloud and dbt Speaking of Tableau, we announced that [we are teaming up with Salesforce](https://www.getdbt.com/blog/dbt-labs-and-salesforce-announce-strategic-partnership) to unite dbt and Salesforce Data Cloud AI. Salesforce Data Cloud, Tableau, and Agentforce customers now gain access to dbt Semantic Layer, data modeling layer integrations. > "Through our new partnership with dbt, we are aiming to broaden the trust, extensibility, and value of Tableau by incorporating dbt models and metrics directly into the product. This will bring together two industry-leading solutions to accelerate time to value, foster seamless collaboration, and provide a world-class analytics experience to our customers." - Ali Tore, SVP of Advanced Analytics, Tableau ## dbt executive summit and customer advisory board dbt grows with the data leaders in the community. At Coalesce, we were thrilled to host our first executive summit—a forum for 72 executives to share, through peer exchanges, how data and AI are shaping the future of customer experiences. Tristan Handy, the CEO and founder of dbt Labs, and Brandon Sweeney, the President and COO of dbt Labs, hosted the summit. The executive summit included industry experts, including AI expert Allie Miller, who also gave our Day 2 keynote. We also had a chance to host our Customer Advisory Board at Coalese. This was a chance for top customer executives to talk about the future of data engineering, analytics, and AI. The goal is to engage some of our most innovative customers to surface best practices and insights to inform and validate our product plans. **** ## dbt community award winners ![dbt community award winners participating in a panel on a stage at Coalesce 2024](https://cdn.sanity.io/images/wl0ndo6t/main/d12b60ae2ae32c539b901f31fe655deaeac6ceca-4032x3024.jpg) Coalesce is a community event, so it was a real joy to celebrate dbt community members who embody what it means to be an advocate within the dbt community. This year’s winners included a data manager who helped people in the dbt Slack with 2,375 posts(!) in the past year and the founder of the Lagos dbt meetup. [Please make sure to join the dbt community if you haven’t already](https://www.getdbt.com/community/join-the-community). ## dbt partner award winners Coalesce is a great chance for dbt partners to come together. The dbt community is wide, and when we talk about One dbt, we’re also talking about all those who contribute to the dbt partner ecosystem. We have tripled the number of implementation partners and doubled our technology partners since the Technology Partner Program launched in 2022. We were thrilled to announce our 2024 consulting partner awards at Coalesce, including our Global Services Partner of the Year phData. [We also announced at Coalesce the appointment of Shawn Toldo as Vice President, Worldwide Partner Ecosystem.](https://www.getdbt.com/blog/dbt-labs-names-seasoned-channel-leader-shawn-toldo-vice-president-worldwide-partner-ecosystem) ![Photo of phData employees awarded Global Services Partner of the Year along with dbt Labs' Brandon Sweeney and Amy Deora ](https://cdn.sanity.io/images/wl0ndo6t/main/0cf3228fdbcf6ae11ae9d3b2ed6847cd911f2ba3-4000x2667.jpg) ## Coalesce sponsors A key part of making this Coalesce happen is our sponsors. A big shout out to all of our partners, including our platinum sponsors AWS, Snowflake, and Tableau. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a90a2b4f7d747f554a8675606e00241546319094-2400x1254.png) ## Coalesce 2025 Sessions will be published on YouTube, and conversations and connections will continue on dbt Slack. [Join us there.](https://www.getdbt.com/community/join-the-community) Coalesce is our highlight of the year, and it’s not an exaggeration that this was the best one yet. It was so much fun that we’re running it back next year. We’ll see you back in Las Vegas for Coalesce 2025. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/400195912dd31f4ad489051d6ef70749c9522f4d-720x405.png) --- --- title: "dbt Labs Names Seasoned Channel Leader Shawn Toldo Vice President, Worldwide Partner Ecosystem" description: "Significant partner program momentum continues with new executive hires and a focus on global expansion." url: "https://www.getdbt.com/blog/dbt-labs-names-seasoned-channel-leader-shawn-toldo-vice-president-worldwide-partner-ecosystem" date: "2024-10-09" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Names Seasoned Channel Leader Shawn Toldo Vice President, Worldwide Partner Ecosystem _Toldo joins alongside Raymond Wong, former Head of Partner Programs at Snowflake, to focus on global program expansion, building on rapid growth since 2022_ **PHILADELPHIA** - October 9, 2024 - [dbt Labs](https://www.getdbt.com/), the pioneer in analytics engineering, today announced the appointment of seasoned channel leader Shawn Toldo as Vice President, Worldwide Partner Ecosystem. Toldo’s appointment comes during a period of significant momentum, as dbt Labs has tripled the number of implementation partners and doubled its technology partners since the Technology Partner Program’s [official launch](https://www.prnewswire.com/news-releases/dbt-labs-announces-formal-launch-of-its-technology-partner-program-301597851.html) in 2022. With hundreds of partners now in the program including Snowflake, Databricks, AWS, Microsoft, and Salesforce, the dbt Labs partner ecosystem has more than doubled in size since 2022. Toldo will focus on continuing to expand the partner ecosystem’s geographic footprint and increasing the program’s number of technology and services partners. He also will lead, mentor, develop, and grow the existing partner team, while driving cross-functional collaboration needed to ensure effective integration and execution of partnership initiatives. As a career sales and channel leader, Toldo’s expertise is in transforming go-to-market strategies and operating models to execute at scale. His leadership will guide the company’s strategy as its partner program matures. “The data ecosystem is _heavily_ partnership-driven, and our relationships with our partners have been a key part of dbt’s success to-date,” said Tristan Handy, founder and CEO of dbt Labs. “The goal of our partner network has always been to connect the market with partners whose solutions complement and extend dbt’s capabilities. Shawn’s expertise is going to help further expand our footprint to make even more of an impact.” Toldo has more than two decades of experience leading global partner ecosystem teams at companies such as Appian, VMware, and Microsoft. In his most recent executive leadership role at OneTrust, he built partner capabilities, accelerated annual contract value growth, and raised customer satisfaction through its partners. In addition to Toldo, Raymond Wong, who previously led the Global Partner Programs and Investments at Snowflake, recently joined dbt Labs as Sr. Director of Worldwide Partner Programs, Strategy, and Operations. Prior to Snowflake, he led the Partner Programs and Investments at VMware, accelerating partner-influenced bookings and services across their product portfolio. He also held leadership roles at TD Synnex Corporation, overseeing product management and operations. Wong will focus on building and scaling dbt Labs’ partner programs to drive global impact across the services, technology, and cloud platform partners. “Over the past two years, dbt Labs’ partner ecosystem has solidified itself as a crucial part of delivering value to dbt users,” said Toldo. “I’m impressed by the program’s rapid growth and am energized by the opportunity to lead it into its next phase.” dbt Labs’ partner network offers a global reach that allows data practitioners to extend dbt’s capabilities with seamless integrations, while providing partners with increased connection to its user base and guidance on integrating with dbt to meet user needs. The company has [implemented several recent product enhancements](https://www.prnewswire.com/news-releases/new-dbt-cloud-enhancements-empower-organizations-with-trustworthy-data-at-scale-302144421.html) aimed at solving the partner ecosystem’s most pressing challenges, including expanding to a multi-cloud approach with support for dbt on Azure. dbt Labs also announced several additional enhancements at its annual [Coalesce](http://coalesce.getdbt.com) event earlier this week. To learn more about dbt Labs’ partner program, visit [https://www.getdbt.com/partners](https://www.getdbt.com/partners). **About dbt Labs** Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 50,000 companies using dbt every week. To learn more about dbt Labs, visit [https://www.getdbt.com/](https://www.getdbt.com/) and follow us on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/). --- --- title: "Introducing cross-platform dbt Mesh" description: "Half of enterprises that use dbt Cloud work across multiple data platforms. Cross-platform dbt Mesh addresses that complexity." url: "https://www.getdbt.com/blog/introducing-cross-platform-dbt-mesh" date: "2024-10-08" authors: ["Connor McArthur"] categories: ["Product"] --- # Introducing cross-platform dbt Mesh At Coalesce 2024, we announced that we are building a new, multi-platform data governance and sharing capability in dbt Cloud called **cross-platform dbt Mesh**. This is a natural evolution of [dbt Mesh](https://www.getdbt.com/product/dbt-mesh): whereas dbt Mesh originally supported references within a single data platform, cross-platform dbt Mesh allows you to coordinate dbt _across _data platforms. This capability is made possible by rapid adoption of Apache Iceberg in the ecosystem, and pairs with [our launch of Iceberg support in dbt](https://docs.getdbt.com/blog/icebeg-is-an-implementation-detail). Our team is actively working with design partners to develop this capability, and we're iterating toward a beta that will include support for Athena, Databricks, Redshift, and Snowflake. As we move towards GA, we will continue to add support for any Iceberg-compatible platform. **** We're excited to build on the momentum behind dbt Mesh as the means for managing data complexity at scale. And, further, we're excited to empower various teams within enterprises to collaborate on data, regardless of which data platform they choose. But, before we talk about the _what _and _how_ of cross-platform dbt Mesh, I want to share a reflection on how we got here, and where we're going. ## The need for multi-platform flexibility dbt adoption and data complexity have exploded over the past eight years. In the first few years, we considered a project with a few hundred models to be very complex; today, 5% of dbt projects (>2,000 of them) have _over 5,000 models_. Over the same time horizon, we've observed [analytics development becoming more like software development](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle): as the dbt workflow drives more value, more people are engaging with it than ever, more assets are under management than ever, and the systems we are building encapsulate more complexity than ever. Practitioners need better approaches and tools to tame that complexity, because, just like in software engineering, building and managing complex systems is _hard_. We also think it's one of the most interesting and important challenges in data today. ![Graph showing the growing number of dbt projects with over 500 models](https://cdn.sanity.io/images/wl0ndo6t/main/bc0f0bcfeaaffd5f0a2ebc01cc041d96acc446b7-2526x1196.png) There is one dimension of complexity that we have not addressed until now. In the early days of the modern data stack, we envisioned organizations choosing a single best-in-class tool for each step in the process: one tool for extraction, one tool for warehousing, one tool for transformation, one tool for BI, etc. That is not the reality we see in most enterprises today: data stacks are less like a single thread, and more like a patchwork quilt. Individual buyers buy the tools that suit their team, but in an enterprise context, that means many buyers are buying on behalf of many teams. It’s become clear that the question leaders are asking isn’t: “Will we adopt platform A _or_ B?” Instead, it is: “Our organization has clear use cases for platform A _and _B. How can we embrace _both_ platforms and still foster governed collaboration, data velocity, and data trust?” To make this concrete: **half of enterprises that use dbt Cloud are working across multiple different data platforms**. This would have surprised me in 2016, but over the last few years I have come to believe that it's a _wonderful _thing. Practitioners love the tools that they love, and teams purchase the tools that suit their needs. From there, at an _organization _level, the tools and workflow should stitch these various systems together to create a seamless experience. Just like software engineering, I don't care if you use emacs or vim, git CLI or GUI. As long as you are a good citizen of the software engineering workflow, you get your code reviewed, you follow the style guide, and you don't break anything; then teams happily coexist even when using very different tools. Today, working _within_ a single data platform works well, even with multiple projects. But when it comes to working _across_ platforms, the seams in the quilt are more like chasms. Practitioners across different business units aren't able to discover or re-use the contents of the "other" platforms; architects aren't able to govern how data is exchanged. Work ends up happening in siloes, and we hear things like, "I have no idea what the other team is doing – they don’t use the same platform that we use." At best, we see practitioners deploying duct tape, glue, and Python notebooks to hold these stacks together (if they are held together at all). That may be necessary today, but it looks to me like a step backward. We should aspire to avoid using "hacky" or heavyweight tools to discover and share data, unless we _really _need them. These breaks across tool boundaries lead to breaks in the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). We aspire to treat analytic code as an asset, rather than making transactional requests for data ("can you re-run your notebook to refresh the data in my system?"). We aspire to empower practitioners to "[put on the hat](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle#hats-not-badges)" and participate in the process, regardless of which tools exist in their organizational ecosystem. We aspire to empower administrators to govern data and the process by which it is produced. And, we aspire to enable discovery and re-use and avoid duplicative work. If we could make data discoverable and governable across these data platforms, then we'd get the power of the workflow _across an enterprise_, rather than just within a team or business unit. It's time to stitch these seams between data stacks together across data platforms. ## Introducing cross-platform dbt Mesh dbt Mesh lets you break down a monolithic application into constituent parts, govern how different datasets can be used downstream, and discover lineage across multiple projects. But it breaks down at the data platform boundary, which is also frequently a team boundary. And, so, we've asked our customers: what if getting unified data workflows was as easy as applying dbt Mesh across projects sitting on different platforms? If a cross-project ref "just worked" across those projects, regardless of platform? Is that something you would want? When we talk to enterprises using dbt, the answer we resoundingly get is “YES.” More than a simple “yes,” it’s almost always a gigantic-sigh-of-relief-yes. So, that's exactly what we are building. **Today, we’re excited to announce that we are adding cross-platform ref support to dbt Mesh.** > _"Cross-platform dbt Mesh makes the promise of data mesh an actual reality for us. Now, it will be possible to work on an organization-wide data model—one that all teams can contribute to and consume from—regardless of what that team's tech stack looks like. Cross-platform dbt Mesh gives the technology diversity of our data ecosystem a common denominator that we can all build around." > - Ulrik Svanborg Møller, Lead Data Engineer, Vestas Wind Systems_ ## How does cross-platform dbt Mesh work? Cross-platform mesh leverages open table formats (in particular, [Apache Iceberg](https://iceberg.apache.org/)) to interchange data. Today, Iceberg support exists, but is limited: each platform only supports a subset of the Iceberg spec. That support is becoming more complete and widespread with each passing quarter. The trend is a rapid movement toward complete support, and we're putting significant effort at dbt Labs into participating in this evolution. The end result of open table format support will be that data platforms can seamlessly interchange data, at least at the edges. And that seamless interchange is what enables cross-platform mesh. Adopting Iceberg is a prerequisite to using cross-platform mesh. That could be an intimidating prospect. For now, just know that this does _not_ require migrating every single table you manage to an Iceberg catalog. You do need access to an Iceberg catalog where you can "stage" public models so that they can be referenced by downstream projects. Once you have that set up, you can "share" public models across warehouses without copying data. > _"By truly separating storage and compute, the technological barriers that contribute to siloed thinking will be eradicated. With cross-platform dbt Mesh, self-service business users can participate in the data workflow, without having to worry about managing refreshes or data synchronization. This levels the playing field so more people can generate value from our organizational information." > - Ulrik Svanborg Møller, Lead Data Engineer, Vestas Wind Systems_ The cross-platform dbt Mesh beta will include support for Athena, Databricks, Redshift, and Snowflake. As we move toward GA, we plan to add support for any Iceberg-compatible platform. In the example below, you can see that an Athena upstream can be referenced by a Redshift downstream. ![Athena to Redshift cross-platform dbt Mesh](https://cdn.sanity.io/images/wl0ndo6t/main/7945eb908bc6423d28ca1bf836c4e6a67dee6b5d-1226x362.png) No data is copied or duplicated as part of this process. Everything just works. Your lineage will render in [dbt Explorer](https://www.getdbt.com/product/dbt-explorer), and dbt builds will immediately pick up any changes to upstream models computed in a different warehouse. The benefits you get from dbt Cloud and dbt Mesh translate to this cross-platform mesh: it unlocks governance and re-use at scale. ![Cross-platform mesh in dbt Explorer](https://cdn.sanity.io/images/wl0ndo6t/main/50dd17f357fa5c84baadb079390b72c3cf8d42a5-1999x1218.png) Zooming into a model level, it works like this: - In the upstream and downstream warehouses, integrate both warehouses with the same Iceberg catalog. - In the upstream project, classify your model as public via dbt Mesh, and configure it to be written into your Iceberg catalog. - In the downstream project, `ref` the upstream model. - Under the hood, dbt Cloud translates the `ref` to point to the right place in the Iceberg catalog. - When dbt Cloud executes the upstream, it writes data into the Iceberg catalog. - When dbt Cloud executes the downstream, it looks up the table in the catalog, **and then loads data directly from the Iceberg store**. The result: you can build a model in an upstream project, `ref` it from a model in a downstream project, and the newly-built model will be _immediately _available for consumption by the downstream model. ## What's next for cross-platform dbt Mesh? We're actively iterating towards a beta of cross-platform mesh with a few select design partners. We're focused on building complete Iceberg compatibility, supporting a broad set of platforms, and sanding off the rough edges of this new capability. If your company uses more than one of Athena, Redshift, Databricks, or Snowflake, and you are interested in learning more about how to use this, contact your account team—we'll keep you informed when we go into beta in the coming months. --- --- title: "Introducing dbt Copilot: The future of AI-accelerated analytics" description: "Accelerate every stage of the ADLC with dbt Copilot." url: "https://www.getdbt.com/blog/introducing-dbt-copilot" date: "2024-10-08" authors: ["Drew Banin"] categories: ["Product"] --- # Introducing dbt Copilot: The future of AI-accelerated analytics Today at [Coalesce 2024](https://coalesce.getdbt.com/), we proudly unveiled dbt Copilot—the AI engine embedded within dbt Cloud to accelerate your analytics workflows. dbt Copilot automatically generates documentation, semantic models, and data tests, while also offering powerful natural language chat, enabling any stakeholder to easily interact with your data. And this is just the beginning. Starting today, dbt Copilot is available in beta. ## The future of AI-accelerated analytics As a data practitioner, you've likely noticed the growing trend—analytics teams are increasingly adopting software development best practices to ensure their data is reliable, scalable, governed, and future-proof. We call this the [analytics development lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle), a scalable process designed to streamline and optimize data operations. By incorporating these best practices into your analytics workflows, you can deliver high-quality, trusted data at speed and scale. But, you shouldn’t have to do it all on your own. We believe practitioners should be focusing on higher-impact projects rather than getting bogged down by repetitive tasks like debugging, figuring out the right syntax for SQL functions, writing tests, and answering the same data questions repeatedly. While these tasks are necessary, they shouldn’t dominate your time. That’s where AI can act as a powerful assistant, automating routine processes and freeing you to tackle more complex, strategic challenges. That’s why we are excited to introduce [**dbt Copilot**](https://docs.getdbt.com/docs/cloud/dbt-copilot), the AI engine embedded within dbt Cloud designed to accelerate your analytics workflows. Now available in beta, dbt Copilot seamlessly integrates AI-powered assistance throughout your dbt Cloud experience, empowering you to ship data products faster, and deliver higher data quality with confidence. It takes care of the tedious tasks—like drafting documentation, building semantic models, authoring data tests, and answering well-understood data questions—within a strong best-practice framework. Most importantly, dbt Copilot keeps you in control, with a human always in the loop. The AI isn't replacing you; it’s enhancing your workflow, allowing you to use your expertise to oversee and guide the process, ensuring quality and context are preserved. > > > — Josh Carlson ## Accelerating DataOps across the ADLC with dbt Copilot Our goal with dbt Copilot is to deliver AI to make data practitioners more productive by accelerating every stage of the ADLC. Here’s how dbt Copilot is helping today and what’s coming soon: ### Develop: Automatically generate documentation and code You’re likely already managing documentation and building semantic models, or planning to implement these analytics best practices. But tasks like documenting 100 columns or defining dozens of metrics manually can be repetitive and time-consuming. #### Auto-generate documentation Today, dbt Copilot leverages generative AI to automatically draft your documentation, speeding up the process and allowing you to quickly review and refine the content to meet your standards. This enables you uphold best practices with just a click. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a888c629af75c3f011d7b94ace329a8da1e9275c-2143x1293.gif) #### Auto-generate semantic models When it comes to semantic models, defining key business metrics with precision is essential for your business, but it doesn’t need to be a long, manual task. dbt Copilot builds an initial draft of your semantic models and the metrics that power it, to help you get up and running with the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) faster. With dbt Copilot handing the foundational code, you can focus on refining and aligning your team’s efforts. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/ef1baf611786d2f8a68a92a5d90832fcc8084e90-1729x1292.gif) ### Test: Auto-generate tests We know that many dbt models have low test coverage, and a big reason for that is the tedious nature of writing tests. Defining tests can feel like manual labor—identifying primary and foreign keys, scaffolding YAML code, and profiling datasets can be time-consuming. But without proper testing, you risk data quality issues that can erode trust. Now, dbt Copilot automatically generates a baseline set of data tests, helping you catch potential errors in real-time across your entire DAG. This way, you can ensure data quality without getting bogged down by the manual work, allowing you to focus on building while staying ahead of issues before your stakeholders even notice them. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/48383d626ad6343a8f4c2cb4f0330bf029d6e961-1734x1294.gif) Looking ahead, dbt Copilot will soon help you create [unit tests](https://docs.getdbt.com/docs/build/unit-tests) in the dbt Cloud IDE. Using generative AI, dbt Copilot will scaffold out mock input data and expected output data that will stress-tests tricky transformation logic like date math, regular expressions, and long case-when statements. These unit tests will ensure your logic is valid, cutting down on troubleshooting time and ensuring the correctness of your transformations. ### Analyze: Chat with your data Answering data questions from stakeholders often requires frequent context-switching and tedious work. Centralizing business metrics in the dbt Semantic Layer has simplified this by allowing users to rigorously define key metrics like ARR, churn rate, and WAUs, making it easier to provide stakeholders with self-service access to the data they need. Now, we’re taking it a step further with dbt Copilot. In addition to querying governed metrics from a BI tool, stakeholder can now ask their questions in conversational language, turning the process into a seamless "chat with your data" experience. This not only delivers instant insights to your stakeholders but also frees up your time to focus on higher-impact work. ![Chat with your data with dbt Copilot](https://cdn.sanity.io/images/wl0ndo6t/main/0e1669cddbb5d53953c00dfbe1fe52096ebdccb5-3200x2018.png) Looking ahead, you can use dbt Copilot to discover trusted data sets more efficiently with AI-powered semantic search in dbt Cloud. **** ## The future of analytics starts now with AI-driven workflows By integrating AI into every stage of the ADLC, dbt Copilot automates the tedious tasks, freeing you to focus on strategic, high-impact work. Currently in beta, dbt Copilot offers key features including: - [Auto-generated documentation](https://docs.getdbt.com/docs/cloud/use-dbt-assist) to speed up documentation creation and review. - [Auto-generated semantic models and metrics](https://docs.getdbt.com/docs/cloud/dbt-assist) for faster adoption of the Semantic Layer. - [Auto-generated data tests](https://docs.getdbt.com/docs/cloud/use-dbt-assist) to ensure data quality in real-time. - “Chat with your data” for natural language querying of well-defined metrics to provide instant insights to any downstream data stakeholder. In the future, dbt Copilot will include even more powerful tools, such as natural language SQL generation to accelerate development, AI-powered unit tests, and much more. We're excited about how these features will help you achieve more with your data, and we're committed to expanding AI-powered support across all stages of the ADLC. If you’re interested in a participating in our beta please sign up [here](https://docs.google.com/forms/d/1B8txoOrJlfbjmCHTxtmRtNOXKQd86LVz1ajM_tHkRl8/viewform?edit_requested=true). To learn more, be sure to join our upcoming virtual event to dive deeper into the new dbt features announced at Coalesce 2024, including dbt Copilot. We'll have live demos and experts ready to help you make the most of these new features, so bring your questions and join the conversation. [Register here](https://www.getdbt.com/resources/webinars/one-dbt-the-control-plane-for-data-collaboration-at-scale-virtual-event) now to watch live or receive the recording. --- --- title: "Announcing advanced CI" description: "Advanced CI offers deeper insights into development and production builds, ensuring data integrity and accuracy." url: "https://www.getdbt.com/blog/announcing-advanced-ci" date: "2024-10-08" authors: ["Reuben McCreanor", "Sara Gawlinski"] categories: ["Product"] --- # Announcing advanced CI Today, at [Coalesce 2024](https://coalesce.getdbt.com/), we announced advanced CI: a powerful enhancement to our existing continuous integration capabilities that allows data teams to validate and compare changes _before_ merging them into production. Advanced CI builds upon our [slim CI](https://docs.getdbt.com/docs/deploy/continuous-integration) feature by offering deeper insights into the differences between development and production builds, ensuring data integrity and accuracy without driving up unnecessary compute. Advanced CI is generally available today to all dbt Cloud Enterprise customers. ## Improve data quality while promoting velocity with CI In data-driven organizations, business-critical decisions and insights rely on data products built on high-quality data. To deliver data products at scale, having a robust and automated CI/CD pipeline is not simply a nice-to-have, but a must-have. CI/CD pipelines allow teams to integrate and deploy changes more frequently, while ensuring that the code and data models are continuously tested, validated, and safely deployed into production environments. Simply put: by embracing CI/CD, data teams are able to ship trusted data products faster. Without a proper CI/CD pipeline, even small changes to data sets or transformations could introduce undetected errors that lead to costly downstream failures in reports, dashboards, or anything that consumes data from the warehouse. For data teams, it isn’t just about faster releases; it’s about trust. Trust that every new commit will integrate seamlessly with the current state of production, and trust that changes won’t break critical data workflows and downstream tools. Historically, dbt Cloud has contributed to this trust through [slim CI](https://docs.getdbt.com/docs/deploy/continuous-integration), a feature designed to catch errors early in the process by running only the models that have changed in a PR. This optimizes compute costs while ensuring that models build correctly. Slim CI doesn't just check for broken code; it verifies that your changed models will build correctly, providing you with the confidence that your models will always build in production without having to risk breaking production to find that out. With these CI jobs, users can accelerate development velocity with build and test automation and create a standardized and governed way to deliver code. While slim CI ensures your code integrates smoothly with production, simply knowing your code changes _will build_ doesn’t mean that your code changes are actually _correct**.**_ One small change can alter the value of an entire column, even if all models build and all user-defined tests pass. Users need a way to validate that the code they’re merging is doing what they expect it to do. This is exactly the problem that dbt Cloud’s new advanced CI feature addresses. ## What is advanced CI? Advanced CI builds on top of slim CI with a "compare changes" feature to provide a deeper level of insight into the changes introduced with each pull request. By surfacing the differences between what is being built and what is already in production, users can ensure that they only merge accurate models into production. For each CI run, dbt Cloud compares the models built from your development branch against the latest production build, surfacing differences such as: - Rows or columns that have been added, modified or removed - Changes in the values within a column - Changes to column data types or column orders - Changes or duplicates in any of the primary keys - What percent of rows have been changed, added, or removed relative to the entire data model ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f63b47698d27116619cf938645ae4f238c8ce942-2480x1410.png) This gives engineers the ability to pinpoint exactly how their changes will impact data models and reports before changes are merged into production. If a column’s values have shifted, or if there are unexpected nulls, users will know before that data becomes accessible to end users or downstream systems. This granular view of changes builds confidence that the data is accurate and trustworthy before it's exposed to end users. By temporarily [caching](https://docs.getdbt.com/docs/deploy/advanced-ci#about-the-cached-data) this changed snapshot within dbt Cloud, advanced CI executes these comparisons seamlessly as part of the CI job workflow, where they can be reviewed as part of the CI/CD process without rerunning every time. This can be especially useful when working with complex transformations, where even small changes can have a domino effect on downstream data quality. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/e9c19012b75bce700fd90679f699c66a9f6d70ce-1870x1366.png) ## Why advanced CI matters for data practitioners As a data practitioner, advanced CI gives you the confidence that every PR you merge into production will not only build, but that it will generate the correct changes you intended for the business. By utilizing advanced CI as part of your data quality workflow, you get: ### Enhanced data quality Say you have just modified a model that feeds into a downstream reporting dashboard. Even if the model builds successfully, what if a join condition introduces nulls into a critical column? Advanced CI surfaces this change in both a git comment and the run details, reducing the risk of bad data reaching production. ### Greater developer velocity Advanced CI eliminates the need for manual testing where teams might have previously reviewed row-level differences themselves. This enables your team to focus on building models and implementing new features, rather than spending that time debugging or resolving production issues after the fact. ### Cost efficiency Data issues caught early in the development cycle are far cheaper to resolve that firefighting in production. With advanced CI, engineers can proactively catch breaking changes, reducing the cost of reactive troubleshooting, reruns, or (gasp!) data downtime. This is particularly useful as your data ecosystem continues to grow, making it increasingly challenging to catch every potential error manually. ## The future of CI in dbt Cloud The question that guides our evolving product strategy is always: How can we give our users more confidence in their data quality? Future improvements we are currently investigating include: - Downstream dashboard impact analysis via auto-exposures (_How can I know what's at stake by better understanding where the data is used downstream?_) - AI-powered PR reviews (_How can we do a better job _of empowering_ users to improve both their code and data development process at scale?)_ - Data quality monitoring (_How can we detect and alert data practitioners to issues with the data, so that they are always the first ones to know?)_ ### Getting started Advanced CI is now generally available for all dbt Cloud Enterprise customers as an opt-in feature. To learn more check out the [docs](https://docs.getdbt.com/docs/deploy/advanced-ci). --- --- title: "dbt Labs Unveils AI Innovations and New Features to Improve Collaboration and Multi-platform Analytics" description: "New features and enhancements, announced at Coalesce 2024, cement dbt Cloud as the data control plane for enterprise analytics." url: "https://www.getdbt.com/blog/dbt-labs-unveils-ai-innovations-and-new-features-to-improve-collaboration-and-multi-platform" date: "2024-10-08" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Unveils AI Innovations and New Features to Improve Collaboration and Multi-platform Analytics _One dbt unlocks the potential of the Analytics Development Lifecycle_ **PHILADELPHIA**, October 8, 2024: [dbt Labs](https://www.getdbt.com/), the pioneer in analytics engineering, today announced new features and enhancements to [dbt Cloud](https://www.getdbt.com/product/dbt-cloud), cementing the platform as the data control plane for enterprise analytics. Announced during the keynote of dbt Labs’ annual conference, [Coalesce 2024](https://coalesce.getdbt.com/), the new innovations support users across various stages of the analytics development lifecycle by providing unprecedented cross-platform flexibility, empowering more people to contribute to the analytics workflow while accelerating speed and productivity, and improving organizational trust in data. This includes the launch of dbt Copilot, an AI engine embedded across dbt Cloud to bolster productivity while improving data quality. **Unifying analytics workflows with One dbt** A central theme to Coalesce 2024 is One dbt, a commitment to create a single unified dbt experience, regardless of infrastructure, data platform or cloud an organization is using. Through One dbt, capabilities will be unified for data transformation, observability, orchestration, cataloging, and semantics, regardless of data platform or cloud an organization chooses to use – providing a universal view of what’s happening within the data estate, accelerating data delivery, tuning data quality, and optimizing compute costs. According to dbt Labs’ 2024 [State of Analytics Engineering](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024) report, decade-long challenges like building data transformations and improving access to compute have largely been solved, but major hurdles related to data quality, data literacy, and data ownership continue to plague the industry. By embracing the[ Analytics Development Lifecycle](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) (ADLC), an integrated, mature analytics workflow, organizations can now benefit from a standardized, scalable way to overcome these issues and move faster with trusted data. “The data industry has made real progress towards maturity over the past decade,” said Tristan Handy, founder and CEO of dbt Labs. “But real problems persist. Siloed data. Lack of trust. Too much ‘duct tape’ in our operational systems. Our announcements from this week go a long way toward fixing these gaps: One dbt experience that is cross-platform, multi-persona, trusted, and infused with AI. All facilitating a single mature workflow: the Analytics Development Lifecycle.” **Delivering enhanced flexibility, collaboration, and trust with the dbt Cloud data control plane** dbt Cloud is a data control plane that supports users across every stage of the ADLC—regardless of their title, technical aptitude, chosen data platform, or where they build and consume data— and further accelerates the workflow with AI. dbt Cloud centralizes metadata and makes it actionable across the ADLC workflow. A suite of new features and capabilities across the dbt Cloud data control plane, introduced at Coalesce, have been designed to scale adoption across a more diverse set of data practitioners, make data development more accessible, streamlined, and governed, and build and automate high quality data pipelines. These include: - **[dbt Copilot](http://getdbt.com/blog/introducing-dbt-copilot),** the AI engine in dbt Cloud that helps users accelerate their analytics workflows. dbt Copilot is designed to automate tasks that previously required repetitive manual work, significantly improving productivity, data quality, and stakeholder trust. Today, this includes the ability to auto-generate tests, documentation, and semantic models (all in beta), an AI-chatbot that allows business stakeholders to ask natural language questions of their data (in beta as part of the dbt native app in Snowflake), and the ability to bring your own OpenAI API key (GA). In the coming months, dbt Copilot will extend to help automate model code generation. - **[Cross-platform dbt Mesh](http://getdbt.com/blog/introducing-cross-platform-dbt-mesh) **will build on [dbt Mesh’s](https://www.getdbt.com/product/dbt-mesh) existing support for cross-_project_ references and allow for cross-_platform_ references using the Iceberg table format as the underlying transport layer. This lets users eliminate silos while maintaining data governance, even in increasingly complex, multi-platform environments. Using cross-platform references in dbt Mesh, data teams will be able to centrally define and maintain data governance standards, see end-to-end lineage across various data platforms, and easily find, reference, and re-use existing data assets instead of rebuilding. - **Support for Apache Iceberg™** enables users to create tables in the Iceberg format and benefit from Iceberg's first-class performance and portability. Apache Iceberg support makes cross-platform dbt Mesh possible. Snowflake support is currently in beta, and Athena, Spark, Databricks, Starburst/Trino, and Dremio are GA. - **A new visual editing experience**,** **currently in beta,** **is a low-code, drag-and-drop environment for building and exploring dbt models, designed to democratize the ADLC to more types of users. Just like everything else in dbt, under-the-hood these visual models compile down to SQL and new code must be version controlled before being deployed into production. This new development interface gives downstream users (who already have the most business context) the ability to accessibly—and safely—author analytics code. Users who are more familiar with SQL can opt to use the visual editing experience to check their work and explore a visual representation of their models. - [**Advanced CI**](http://getdbt.com/blog/announcing-advanced-ci), now GA, allows users to compare code changes as part of the CI process to catch unexpected behavior before new code is merged into production. This improves code quality and helps organizations optimize compute spend by only materializing correct models. Users can see a summary of their changes in their Git pull request and dive into modified, added, and removed rows and columns within dbt Cloud. - **Data health tiles**, now GA, can be embedded into any downstream app so data consumers can have real-time context into critical trust signals like data freshness and data quality directly in the tools where they work. - **Auto-exposures with Tableau**,** **now in Preview, automatically incorporates Tableau dashboards into dbt lineage. This allows data practitioners to optimize, expedite, and automate the orchestration of end-to-end pipelines—from source to dashboard. Business users can be confident that they always have the freshest data powering their decisions. - **An upcoming Power BI integration for the dbt Semantic Layer** will enable business users who have standardized on the Microsoft ecosystem to query and analyze consistent metrics. - **Teradata and Athena are supported adapters.** dbt Cloud now integrates with Teradata (in Preview) and AWS Athena (GA), enabling more organizations and teams to collaborate on data workflows. These new features and others are enabling dbt customers like [Roche](https://docs.getdbt.com/blog/dbt-squared) to unify their data, standardize and accelerate their workflows, and maximize their investment in dbt. “dbt Core jump-started our data platform’s growth, and dbt Cloud allowed us to spread it across the globe,” said Yannick Misteli, Head of Engineering, Global Product Strategy at Roche. “Today, we are able to power our platform in 70 countries and run over 15,000 models and 40,000 tests every day. We can support our core and country teams with the workflows that best suit them and promote code to production in two week cycles instead of the previous quarter or semester-long cycles.” For more information on dbt Cloud and how it helps users across every stage of the ADLC, visit [https://www.getdbt.com/product/dbt-cloud](https://www.getdbt.com/product/dbt-cloud). **About dbt Labs** Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 50,000 teams using dbt every week. To learn more about dbt Labs, visit [getdbt.com](http://getdbt.com) and follow us on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/). --- --- title: "dbt Labs Teams Up with Salesforce to Enhance Data Transformation and Metrics Management" description: "Together, Salesforce and dbt Labs are bringing customers a streamlined, trustworthy, end-to-end data experience." url: "https://www.getdbt.com/blog/dbt-labs-and-salesforce-announce-strategic-partnership" date: "2024-10-08" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Teams Up with Salesforce to Enhance Data Transformation and Metrics Management _dbt Semantic Layer, Data Modeling Layer integrations now accessible to Salesforce Data Cloud and Tableau customers_ **PHILADELPHIA**, October 8, 2024 – [dbt Labs](http://getdbt.com), the pioneer in analytics engineering, announced at [Coalesce 2024](http://coalesce.getdbt.com) that it is teaming up with Salesforce, the world’s #1 AI CRM, to unite Salesforce Data Cloud AI, automation and analytics solutions with dbt Labs' renowned expertise in data transformation and metrics management, offering customers a streamlined, trustworthy, end-to-end data experience. “Together, Salesforce and dbt Labs are redefining what’s possible with data,” said Ryan Segar, Chief Customer Officer at dbt Labs. “We are focused on delivering tightly integrated solutions that enable customers to accelerate their journeys along the analytics development lifecycle, harnessing trusted, flexible, and powerful data insights that drive better business outcomes.” Salesforce Data Cloud, Tableau, and Agentforce customers will gain access to dbt Labs' trusted data transformation pipeline. This end-to-end capability ensures that data is seamlessly prepared, transformed, and delivered to Tableau’s analytics platform, enhancing data accuracy, quality and reliability. dbt Labs also will provide an independent metrics layer, enabling Tableau and Salesforce customers to define, manage, and standardize key business metrics across all platforms. This flexibility supports consistent and comparable insights, empowering data-driven decision-making with confidence in the flow of work. New integrations available to users include the ability to connect the dbt Semantic Layer and Tableau Pulse, export metrics from dbt Cloud to Tableau Cloud, and export dbt models to Tableau and Einstein. Additional integrations, such as alignment with Tableau Semantics to enable “bring-your-own-Semantic” use cases and enabling Tableau instant analytics from the dbt Cloud console, will be explored to help users maximize the collective potential of dbt, Salesforce and Tableau platforms. “This integration builds deeper and more seamless product integrations, bringing the best of dbt into Salesforce Data Cloud and enhancing Tableau’s data modeling, governance, and semantic layer capabilities with trusted dbt experiences,” said Ali Tore, Senior Vice President of Advanced Analytics at Salesforce. “Our collaboration with dbt Labs enables our thousands of customers to harness the power of AI-powered insights on a foundation of reliable, trusted data seamlessly in the flow of their work. We’re helping organizations make faster, more informed and trusted decisions by providing an integrated, end-to-end solution for data transformation, metrics management, and analytics—unlocking the full potential of their data to drive impactful outcomes.” More than 50,000 teams use dbt, and Salesforce customers can now leverage the same advanced data modeling techniques trusted by some of the world’s most innovative organizations. This integration allows for scalable, robust data modeling directly within Salesforce Data Cloud, optimized for both technical and non-technical users. For more on recent dbt enhancements designed to benefit Tableau users, visit [https://www.getdbt.com/blog/coalesce-2024-product-announcements](https://www.getdbt.com/blog/coalesce-2024-product-announcements). Salesforce, Data Cloud, Tableau, Agentforce and others are among the trademarks of Salesforce, Inc. **About dbt Labs** Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 50,000 teams using dbt every week. --- --- title: "One dbt: the biggest features we announced at Coalesce 2024" description: "Learn about the latest dbt Cloud features designed to help organizations embrace analytics best practices at scale." url: "https://www.getdbt.com/blog/coalesce-2024-product-announcements" date: "2024-10-08" authors: ["James Mayfield"] categories: ["Product"] --- # One dbt: the biggest features we announced at Coalesce 2024 [Coalesce 2024](https://coalesce.getdbt.com/) kicked off this morning in Las Vegas. In front of 1,800+ data practitioners, team leaders, and executives at Resorts World and thousands more watching online, we shared our vision for the future of the analytics workflow. A time where zero-sum choices about cloud data platforms, infrastructure, and even dbt Core vs. dbt Cloud become unnecessary, because everything works together in support of the analytics development lifecycle. We call this, “One dbt.” [Watch video](https://youtu.be/ynE9H9Ya2UQ) ## One dbt One dbt isn’t a specific feature. It’s an ethos that’s influencing all aspects of how we operate at dbt Labs. It represents our commitment to an integrated, governed, and scalable approach to data analytics. We are building towards a future where everyone—from data engineers to analysts to business decision-makers—has a **common unifying framework** for solving analytics problems. This framework should allow all of these humans to collaborate on any data platform, on any cloud. Their work should be accelerated by AI. And they should be able to work in the tools that make the most sense for them. You’ll see this reflected in the announcements below and certainly in our investments in the future. dbt has always been a force for unification in the data industry, bringing together people, platforms, and workflows. But this focus has ramped up over the past year for a few reasons: 1. **Market demand for platform flexibility:** Most companies run multiple data platforms. Iceberg is quickly becoming the de facto standard for how companies embrace multiple cloud data platforms and compute engines. This opens up a future where teams have greater flexibility around how and where they access data and reduces dependence on a single data platform. 2. **The need to empower more data collaborators:** The rise in self-service BI interfaces and new generative AI tools have brought business users much closer to the analytics development workflow. These users bring unique context about the business that typical data practitioners may not otherwise have. Empowering business people and data teams to collaborate creates more avenues for developing trustworthy and useful data products. 3. **The trust imperative:** Our most recent [State of Analytics Engineering](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024) report found that poor data quality, ownership, and stakeholder literacy are among the top challenges in analytics today. Consider what happens when a business user spots an error in a sales dashboard. Can they troubleshoot it themselves? Would they know where to go for support? This gets at the heart of the trust issue, and a lack of trust between business and data teams can stifle company growth. The [Analytics Development Lifecycle](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) (ADLC) is a framework for maturing the analytics workflow and improving data quality and trust, tying together the technology stack with the multiple personas who contribute to generating and disseminating organizational knowledge. ![Image of the analytics development lifecycle (ADLC) infinity loop](https://cdn.sanity.io/images/wl0ndo6t/main/4a5682d44d9f8b63ca3f5da115bcb613c3e7dc61-1618x854.png) Our newest features are designed to help our customers take advantage of these trends, helping them improve cross-platform flexibility, empower more people to safely contribute to analytics workflows, and improve organizational trust in data. **** ## New dbt features announced at Coalesce 2024 [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) is a data control plane that centralizes metadata across the ADLC—orchestration, observability, cataloging, and more. Organizations can move faster with trusted data and dbt is the lynchpin for helping data teams build, deploy, monitor, and discover data assets. Let’s dive in to the newest features powering the dbt Cloud data control plane. ![Image of the dbt Cloud Data Control Plane architecture](https://cdn.sanity.io/images/wl0ndo6t/main/e4299c0adfdce3582030abf9019b51c5f83c0a50-1771x780.png) ### Cross-platform dbt Mesh [dbt Mesh](https://www.getdbt.com/product/dbt-mesh) will soon support cross-_platform_ references, allowing organizations who rely on various data platforms to have a centralized, governed approach to multi-platform collaboration. This new capability is in direct response to the fact that, more often than not, large enterprises rely on multiple data platforms (for example Snowflake and Databricks, or Redshift and Athena) for various use cases—and they need to maintain a unified workflow regardless of what data platform a particular team uses. > _"Cross-platform dbt Mesh makes the promise of data mesh an actual reality for us. Now, it will be possible to work on an organization-wide data model—one that all teams can contribute to and consume from—regardless of what that team's tech stack looks like. Cross-platform dbt Mesh gives the technology diversity of our data ecosystem a common denominator that we can all build around." > - Ulrik Svanborg Møller, Lead Data Engineer, Vestas Wind Systems_ With cross-platform dbt Mesh, practitioners will be able to reference and re-use data assets from not only different projects, but also different _data_ _platforms_, breaking down silos and fostering collaboration across an organization’s full data estate. We are actively working with a small handful of design partners to develop this capability, and we're iterating towards a beta which will include support for Snowflake, Databricks, Redshift, and Athena. Learn more [here](https://www.getdbt.com/blog/introducing-cross-platform-dbt-mesh). ![Architecture image of cross-platform dbt Mesh](https://cdn.sanity.io/images/wl0ndo6t/main/540d0edcccc78346110e38fe86f70e86dc4abb68-1496x728.png) ### Iceberg table support dbt Cloud now supports the Apache Iceberg table format, which is a key capability that enables cross-platform dbt Mesh. This allows users to leverage the benefits of Iceberg's table format, including improved query performance and schema evolution. By supporting Iceberg, dbt Cloud enables data teams to work more efficiently with large-scale data lakes while maintaining the familiar dbt workflow. Support for Athena, Spark, Databricks, Starburst/Trino, and Dremio are GA and Snowflake is currently in beta. ### Visual editing experience We are adding an intuitive, visual drag-and-drop interface (currently in beta) to our suite of data development environments. Why? The short answer is to empower more types of users with governed workflows for data collaboration. The long answer is that the folks within an organization who are closest to the business context typically have two courses of action for getting the data they need: they can lob a ticket over to their (overworked) central data team and hope that they get the right data back in a timely manner or, they could go-it-alone with a CSV pull and a spreadsheet. Neither of these approaches is scalable. ![UI of dbt Cloud visual editing experience](https://cdn.sanity.io/images/wl0ndo6t/main/a537de6f3ccf9fc6e63489d15b113bc4a0095ff8-1800x1142.png) With the new visual editing interface, downstream teams have a governed inroad to building data products—even if they don’t know how to write SQL. Under-the-hood, it’s the same familiar dbt workflow governed by the same principles: it generates SQL in your dbt project, the code is version-controlled, and the code can be tested and documented so models are trustworthy. This accessible visual editing experience is how data teams can take themselves off the critical path for ad hoc requests, and empower their business colleagues to translate _domain knowledge_ into _real analytics code._ Even folks well-versed in SQL will find it useful to distill unwieldy lines of code into a visual representation of their models so they can better understand, optimize, and explore their data pipelines. Register your interest in joining the beta [here](https://docs.google.com/forms/d/1B8txoOrJlfbjmCHTxtmRtNOXKQd86LVz1ajM_tHkRl8/edit). ### dbt Copilot dbt Copilot (currently in beta) is dbt Cloud’s embedded AI engine with various user experiences across the dbt workflow designed to help users accelerate and automate analytics. Using dbt Copilot, users can perform tasks that previously required repetitive manual work and significantly improve productivity, data quality, and stakeholder trust. Today, this includes the ability to auto-generate tests, documentation, and semantic models (all in beta), an AI-chatbot powered by the dbt Semantic Layer (in beta as part of the dbt native app in Snowflake), and the ability to bring your own OpenAI API key (GA). In the coming months, dbt Copilot will extend to help automate code generation and workflows across all of dbt Cloud. ### Compare changes with advanced CI Continuous integration in dbt just got even smarter. While users have [long been able to validate that a pull request wouldn’t break something in production](https://docs.getdbt.com/docs/deploy/continuous-integration), with the new compare changes feature inside of CI jobs (now generally available for dbt Cloud customers on the Enterprise plan), they’ll be able to ensure that _what_ is being built meets their expectations. When enabled, each CI job will include a breakdown of the columns and rows that are being added, modified, or removed in your underlying data platform as a result of executing your dbt job. Additionally, users can see a summary of these changes inside the actual PR in their Git provider interface. This additional context allows data teams to catch any unexpected behavior before code is deployed into production, improving data quality and increasing trust among all collaborators. Learn more [here](https://www.getdbt.com/blog/announcing-advanced-ci). ![Advanced CI compare changes feature in dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/04ca430f6275d0aafe560e2e1a1b9595b4c436bc-1244x758.png) ### Auto-exposures with Tableau Now in Preview, the dbt DAG can be automatically populated with downstream exposures in Tableau (and Power BI to follow). This gives data teams automatic context into how and where models are used, so they can prioritize data work to promote data quality. Meanwhile, by triggering downstream dashboards to automatically refresh when new data is available, business stakeholders can be confident that they’re always making decisions from the freshest data. These exposures are automatically accounted for throughout dbt Cloud, including in dbt Explorer, scheduled jobs, and CI jobs. ![dbt lineage with auto-exposures for Tableau](https://cdn.sanity.io/images/wl0ndo6t/main/c41ef96e0994bcca9599700e73d05b500890fd7a-2752x1538.png) > _“With auto-exposures in dbt Explorer, it’s like going from a treasure hunt to having a treasure map. The native Tableau integration and auto-generated lineage help us see everything clearly, making impact analysis across hundreds of dashboards a breeze!” > —Rahavan Raman, Director - Data Engineering & Analytics, Zscaler_ ### Data health tiles What good is a dashboard if you can’t trust the freshness and veracity of its supporting data? With dbt Cloud, you can now embed health signals like data quality and freshness within any dashboard, giving your downstream stakeholders at-a-glance confirmation of whether they can trust the data they’re about to use. Users can also navigate back to dbt Explorer with a single click to investigate further. These trust signals provide users with a quick and easy way to assess the reliability of their data assets. By offering clear indicators of data health and freshness, teams can make more informed decisions and have greater confidence in their analytics processes. This feature aligns with our commitment to enhancing data quality and fostering trust across the entire data lifecycle. ![Embed tile with data health signals in downstream dashboards](https://cdn.sanity.io/images/wl0ndo6t/main/0a5d303438d629ac604292417d6e3c2708cddc08-1806x854.png) ### New integrations dbt Cloud now integrates with [AWS Athena](https://docs.getdbt.com/docs/cloud/connect-data-platform/connect-amazon-athena) (GA) and [Teradata](https://docs.getdbt.com/guides/teradata?step=1) (Preview), enabling more organizations and teams to collaborate on data workflows. Athena will also be one of the first adapters compatible with cross-platform dbt Mesh. Additionally, we are expanding our BI tool integrations with a new dbt Semantic Layer connection to Power BI, coming soon. ### dbt Core v1.9 dbt remains committed to advancing our open source offering for data practitioners, and starting in Core 1.9, we have two big improvements, available now in beta. First, you can use the new [microbatch](https://docs.getdbt.com/docs/build/incremental-microbatch#what-is-microbatch-in-dbt) strategy to optimize your largest datasets. This lets you process your event data in discrete periods with their own SQL queries, rather than all at once. Second, we've streamlined [snapshot](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.9#snapshots-improvements) configuration to make dbt snapshots easier to configure, run, and customize. In addition, we are solving for a list of smaller “paper cuts” upvoted by the community that you can read about in full in our [docs here](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.9). ## Embrace analytics best practices at scale From cross-platform dbt Mesh to the new visual editing experience, the latest platform features exemplify the "One dbt" ethos. We are empowering data teams to work seamlessly across data ecosystems, enabling more users to contribute to analytics workflows, and improving trust in data with a single, integrated set of governance features we call the data control plane. Want to learn more? Join our upcoming webinar, _One dbt: The Control Plane for data collaboration at scale,_ to see these features in action. [You can register here](https://www.getdbt.com/resources/webinars/one-dbt-the-control-plane-for-data-collaboration-at-scale-virtual-event). And as a reminder, if you miss any of the action this week, you can catch a recording of our opening keynote or other sessions from Coalesce [online at any time](https://coalesce.getdbt.com/). --- --- title: "Unlocking new possibilities with dbt Cloud on Azure Databricks" description: "Unlock AI and data innovation with dbt Cloud on Azure Databricks—scalable pipelines, optimized performance, and unified governance" url: "https://www.getdbt.com/blog/unlocking-new-possibilities-with-dbt-cloud-on-azure-databricks" date: "2024-10-08" authors: ["Jeff Mills", "Clarke Patterson", "Hiral Jasani", "Ken Wong", "Can Efeoglu"] categories: ["Product"] --- # Unlocking new possibilities with dbt Cloud on Azure Databricks The [rapid pace of AI adoption](https://www.gartner.com/doc/reprints?id=1-2IAZ5C96&ct=240807&st=sb) requires a strong foundation of data accuracy and governance. At a recent summit, Gartner stated that [at least 30% of genAI projects will be abandoned by the end of 2025](https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025) as organizations fail to realize value due to poor data quality, inadequate risk controls, and rising costs. The outcome of AI applications can only be as good as the underlying data. That’s where dbt and Azure Databricks have proven to deliver a powerful combination to scale the use of data and build new data products, including genAI applications. With dbt Cloud on Azure, this combination now extends to Azure data stores and services. dbt Cloud on Azure streamlines data transformation workflows in the Azure ecosystem. Along with [Azure Databricks](https://www.databricks.com/product/azure), customers now have access to a simple, open platform to meet their analytics development needs using services native to the Azure cloud. Providing access to data and transformations in the cloud of the customer’s choosing continues to deliver on the promise of One dbt. One dbt is our strategy of unifying our experiences, audiences, and communities together. **Customers especially benefit from**: ### Manageability With dbt Cloud on Azure Databricks, data teams can seamlessly manage their data pipelines in a few clicks, all while leveraging Azure’s integrated environment for managed access control and authentication (Microsoft Entra ID) and services (Power BI, Azure Data Factory and Azure OpenAI). This out-of-the-box, first-party integration between Microsoft and Databricks provides a cohesive experience, reducing the complexity of managing separate tools and platforms. ### Data access While building their data pipelines on Azure Databricks, dbt users can take advantage of optimized reads and writes directly from Azure Data Lake Storage (ADLS Gen2) and other Azure storage systems (e.g. Blob storage) for a streamlined data workload. ### Governance and security Azure Databricks offers unified governance and lineage via Unity Catalog to analysts, data engineers, and scientists to securely discover, access, and collaborate on trusted data and AI with end-to-end lifecycle visibility. Data teams using dbt on Databricks get the added protection and monitoring capabilities of the Azure Databricks environment with Azure Security Center. ### Lower cost of compute Write the most efficient dbt code by taking advantage of Azure Databricks’ flexible pricing, which charges only for the compute you use. This flexibility helps data teams of all sizes manage budgets and grow efficiently without the risk of unexpected cost spikes. ## Performance, power, and simplicity for Analytics Engineers [Databricks SQL](https://www.databricks.com/product/databricks-sql) warehouse is an optimal engine for building and running dbt projects. Built with DatabricksIQ, the Data Intelligence Engine that understands the uniqueness of your data, Databricks SQL helps dbt users become more efficient and build faster pipelines with an intelligent and auto-optimizing platform that also benefits from the simplicity, unified governance, and openness of lakehouse architecture. In 2024, Databricks SQL has undergone significant advancements, leveraging AI to automatically improve performance and efficiency, resulting in a [4x improvement in query performance over the past two years](https://www.databricks.com/blog/whats-new-with-databricks-sql). Key enhancements include: - [Intelligent Workload Management](https://docs.databricks.com/en/compute/sql-warehouse/warehouse-behavior.html#serverless-autoscaling): Optimize resources for high-concurrency BI workloads - [Liquid Clustering](https://docs.databricks.com/en/delta/clustering.html): Automatically managing data layout without manual fine-tuning - [Predictive I/O](https://docs.databricks.com/en/optimizations/predictive-io.html): Index-like performance without the need for index creation or maintenance. This means faster query execution times for dbt models, especially for large-scale dbt transformations. [Retool](https://www.getdbt.com/case-studies/retool), a joint customer of Databricks and dbt Labs, uses the combination to power most of their analytics workloads. “With dbt Cloud on Databricks SQL, we [decreased the actual spend on our daily dbt production jobs by 50% and decreased the runtime by 25%](https://www.youtube.com/watch?v=3PtU_owmgFg&t=2s&ab_channel=dbt) at the same time”, said Samuel Garfield, Analytics Engineer at Retool. “This is truly a no-compromise situation.” ### Simplified user experience By combining Databricks SQL with dbt Cloud’s intuitive modeling layer, business analysts can work more effectively. Expanding who can be successful with dbt and Databricks allows more people to participate in the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). Databricks SQL offers AI-assisted tools to simplify data analysis for the broader organization. The Databricks AI Assistant provides a context-aware tool to help dbt users create, edit, and debug SQL queries. To prevent ADLC implementations from getting siloed within centralized teams, cross-functional teams can improve collaboration with Databricks AI/BI, a new business intelligence product that allows quick visualizations based on business context, and Genie, a conversational tool answers business questions knowing the context of your own data. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/7b1e2e48896b1194fd351906bb75fe90ae93e095-964x480.gif) _Databricks Assistant to create, debug and explain code in SQL Editor_ ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/d38734dd34396b7dedacec9564171cb15da749ef-1345x971.gif) _Analysts can generate visualizations using natural language_ ### Core SQL warehouse capabilities Databricks SQL has several core features including: - dbt + [Materialized ](https://docs.databricks.com/en/views/materialized.html)Views (MV): Building efficient pipelines becomes easier with dbt, leveraging Databricks' powerful incremental refresh capabilities. Users can use dbt to build and run pipelines backed by MVs, reducing infrastructure costs with efficient, incremental computation. - dbt + [Streaming Tables](https://docs.databricks.com/en/delta-live-tables/index.html#streaming-table): Streaming ingestion from any source is now built-in to dbt projects. Using SQL, analytics engineers can define and ingest cloud/streaming data directly within their dbt pipelines. ## Looking ahead As we continue to evolve together, expect deeper integration between dbt andUnity Catalog, enhanced support for real-time data transformations using dbt Cloud and Databricks' streaming capabilities, and more seamless workflows between data engineering, data science, and machine learning teams. By leveraging dbt Cloud on Azure Databricks, organizations can build more robust, scalable, and maintainable data pipelines while empowering a wider range of users to work effectively with data. This powerful combination is set to drive the next wave of innovation in the world of data analytics and engineering. To learn more, check out this guide on [setting up your dbt project on Databricks](https://docs.getdbt.com/guides/set-up-your-databricks-dbt-project?step=1) or [take it for a spin on the Databricks Platform](https://www.databricks.com/try-databricks#account). --- --- title: "The current state of the AI ecosystem" description: "Former Analytics Engineering Podcast co-host Julia Schottenstein, now at LangChain, returns to the show." url: "https://www.getdbt.com/blog/the-current-state-of-the-ai-ecosystem" date: "2024-10-06" authors: ["Daniel Poppy"] categories: ["Insights"] --- # The current state of the AI ecosystem _This post first appeared in_ _[The Analytics Engineering Roundup](https://roundup.getdbt.com/p/the-current-state-of-the-ai-ecosystem)._ Host now guest! Former co-host Julia Schottenstein returns to the show to go deep into the world of LLMs. Julia joined [LangChain](https://www.langchain.com/) as an early employee, in Tristan’s words, to “Basically solve all of the problems that aren't specifically in product and engineering.” LangChain has become one of, if not the primary frameworks for developing applications using large language models. There are over a million developers using LangChain today, building everything from prototypes to production AI applications. If you're looking for a comprehensive overview on the state of the AI ecosystem today, this is the episode for you. **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Google Podcasts](https://podcasts.google.com/feed/aHR0cHM6Ly9hbmFseXRpY3NlbmdpbmVlcmluZ3JvdW5kdXAubGlic3luLmNvbS9yc3M) - [Stitcher](https://www.stitcher.com/show/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways from this episode ### What is LangChain and what kind of impact is it having? **Julia Schottenstein:** The two co-founders, Harrison and Ankush, got started on LangChain, an open-source project, back at the end of 2022, right before ChatGPT became an enormous household name. And what LangChain does is give you building blocks that you can build your own applications much more easily. It's a framework that can be used in Python or TypeScript. It helps you bring your company or your personal information, your APIs, or up-to-date information to the context of the LLM reasoning API, so it can do more useful things. As listeners probably know, LLMs can do a lot. They've been trained on an enormous corpus of information, but they have no idea about your private data or your private information. To do useful things for your business, you need to prompt the LLM in very specific ways, and what LangChain does is it helps bring that pipeline of data and computation to the LLM, so you can build apps and do things like customer support bots, or personal assistants, co-pilots, things of that nature. ### This is a lossy analogy, but I often think about LangChain as the dbt for AI. I think about it the same. ### With dbt, we said there's this fundamentally new compute construct—the cloud data platform—and it needs a programming framework on top of it to enable users to do the things that they want to do. ### I think LangChain came out of that same kind of need. There's a new kind of fundamental compute construct—the large language model—so what are the things that you need to do to make it useful? Is that right? Obviously I have similar biases to you having been a part of the dbt journey, but I think it's pretty similar in the sense that LangChain has become a way for people to learn about generative AI and how to build applications. There are these common building blocks. With dbt you have your models, tests, documentation, all of these components as a recipe to do work. It’s pretty similar to LangChain. We have the components you need to build an LLM application. And they're pretty simple. A lot of times it's, prompts plus a chat model, plus an input parser that, that take the raw text and makes it structured, so it can be useful in downstream workflows. It's a framework. It's open source. There's a big ecosystem around it. I think we have a million developers now who use Langchain, and our partner ecosystem is really important for making LangChain as special as it is. There are the vector databases—dozens of those—and hundreds of different models, and people will swap in and out these components as they need. You can define your logic once for your application logic, and then we've abstracted away a lot of the commonality among the underlying technologies. LangSmith is our commercial offering, and it's a companion product. The reality with all LLM applications is that the hardest part is quality. It's easy to build a proof of concept in an afternoon. But it's hard to make it work consistently well at large scale and not have embarrassing moments when it's in production. A lot of engineering work goes into polishing those rough edges, and making it worthy for end users to interact with it. And so Langsmith helps solve that problem with testing and observability. It gives you tools so that when you're making a prototype and moving to production, you can experiment with changing logic and see how it affects metrics you care about. Once it's live, it's a monitoring and observability platform so you know what people are asking of your application and how well is it responding to questions. LangSmith supports you in making your application higher quality, getting to production faster, and giving you the visibility you need once it's live. ### I wonder if the industry knows what it means to test or observe an LLM application. Software engineers have been testing and observing classic software applications for forever. But I feel like we're trying to invent the tools and processes that we need to make our AI applications high-quality. ### Do you know enough about what your users need to build what they need? We're definitely learning as we go. The big advantage of being open source is that we get to learn from teams around the world who are cutting-edge, and we can build alongside them. We try our best to move quickly, to build product, to, to help solve it, but it is an evolving practice. We do wonky things in the LLM app world where you use an LLM as a judge to run your tests. You just ask another LLM. That feels really awkward and clunky at first, until you realize that LLMs are fantastic at grading the response of other LLMs. You get this new way of bringing engineering best practices adapted to this non-deterministic, highly variable, type of application. ### [I interviewed Yohei Nakajima](https://roundup.getdbt.com/p/the-rapid-experimentation-of-ai-agents), the creator of Baby AGI and he talked about an agent that he built with a prompt to create a business—it decomposed the steps needed to start a business, and then it decomposed those steps and then it started doing the thing. And you could just leave it running for an arbitrarily long time. ### When people build applications with Langchain, are they generally trying to build things where there's a request response? Or is it more building an agent that's behind the scenes? There's a reality of what the technology can do today versus what we want it to eventually do. The majority of use cases in Langchain, just because of where people are, tend to be chat applications and RAG, as we discussed. What you're describing is what we call agentic applications, where people have different definitions for it. Some people have the definition that agents are LLM apps that act on your behalf, so things get automated. Our definition is more that you use the LLM to decide which steps to take. In this example of “build me a business” you don’t have to predefine all of the steps of, first incorporate, then hire your employees, then, then, then. You let the LLM decide the routes to take and you can let it run. Both of those worlds are hard as you can imagine. I don't think the business that this LLM built is going to be very successful. What we've been trying to do is find a nice balance where you give the LLM the ability to route and make decisions. But if you fall into certain states, we define it in code, which actions they’re allowed to take. And so that's where the conditional edges come into place. People end up building things that fall into a few categories. There's concierge search, a new way to discover information instead of clicking on a bunch of filters. You can chat and have the experience of, “I'd like to go on a cruise with my family of 10, can you send me some options?” There are also a lot of copilots built into applications to help with tasks for operational efficiency—things that are repetitive or time-consuming—to automate with an LLM. ### Anytime there's a new technology, it tends to be adopted first by a cohort of folks who work at small, digital -native companies. ### There has been a tremendous amount of attention in the enterprise of the AI trend. I am not clear if the enterprise is AI curious or is really putting applications in production. Do you have insight into the relative stages of digital native versus traditional enterprise? We've seen a lot of activity in the enterprise, a lot of investment. They usually start with an internal application. The first thing that people will build is productivity for their employees. We've seen really large organizations with a mandate for their employees to get trained and learn about AI. And it could be as simple as “how do you interact with an LLM?” Rakuten, which is one of our customers, built an internal GPT, where all of their employees can build mini apps. The most common app that people built or use at Rakuten is just for practicing their English. We see frequently mandated from the top for employees to have access to an internal chat app, which is generative AI, so people can learn to prompt it and to get their work done faster. ### I'm curious about the industry structure of the AI space. The cloud data space evolved in the way that it did almost completely because there was a standard to unite around, which is SQL. ### Do you have beliefs about what the emerging categories are and how the lines are being drawn? Obviously, the model layer is the one with tremendous amount of interest. You're not an LLM application without an LLM. It's the one part of the stack that you don't get to avoid. There's a big question about whether there'll be a lot of value accrues to the application layer or to the picks and shovels layer. Picks and shovels would be Langchain, because we're an orchestrator. We build tools for developers to build end applications. Then there's the application layer, where you see a tremendous amount of interest. AI SDRs, AI marketing assistance, customer support. Customer support is an enormous category. These companies are doing really, really well overnight. You'll hear some of the numbers of going from zero to 20 million ARR in one year. It's an open question of where all the value will accrue or how sustainable those revenue streams are. You always will have both. You'll have fewer on the picks and shovels side that accrue a ton of value, and a lot of people will try to go after that to be the currency of the new industry, and then you'll have many, many more successes with the application layer. I think you'll have companies that will rise and fall pretty quickly there. ### One of the criticisms of the modern data stack has been that if you want all of the appropriate picks and shovels, you have to go to 12 different hardware stores to buy them all. There are players in the AI space that are a little bit more narrowly scoped than the stuff that you folks are building. ### But my understanding is that your goal is to provide all of the tooling required to put applications in production. Is that right? It's part of our philosophy that the data you collect in production is so important to help you improve your application that it needs to be all in one. To put a finer point on that, it's really hard to come up with realistic examples or edge cases that you need to test against to improve your application. And so it's important to be able to observe production traffic, understand where your app is falling short, and collect those examples of where you need to improve, and then feed that back in to your testing and design process so you can iterate faster. And the other cool thing that we're, we're starting to in and experiment with is feedback from production. People liked your responses; can we use that to influence and improve your prompting strategy or improve the way that you actually can design your application. All in one is our approach. Some people are more narrowly scoped and they do testing for financial data really, really well, or they provide an LLM that's really good at SQL or data analytics. It's just where people are picking their lanes. We have aspirations to support you throughout your entire development life cycle. ### Let’s talk about design. We all have built this idea of what good design looks like over the last 30 years of the internet, but we just don't know a lot about what good design looks like for these systems yet, do we? Like, are there people who consider themselves like quote unquote AI designers and what does that mean? There are two types of interesting design happening. There's the user experience—how an end user interacts with this application. Oftentimes it's chat. We have companies that are shopping online where you're now chatting instead of browsing. It's a totally new user interface. There are other more creative user experiences that we're seeing as well that are ambient where there's just an agent listening and observing and then knows to take action. As you start to dissect and think about the different types of user experiences that are now possible, how do we design for them? It's totally new worlds that I took for granted. It is these small new ways to interact with your application and software. ### In the run up to Snowflake Summit, we released a native app, and it had a a feature in it called ask dbt, powered by the dbt Semantic Layer. You can ask natural language questions and it compiles them into semantic layer requests and it gives you governed answers back in text. ### But the interesting thing was not the user interface layer that you are talking about. The user experience that we were trying to think about was the types of questions will people ask. That's the second design pattern that I was going to go into. Your application logic has to reflect how people are going to use your app. And so if they're coming in with certain types of problems that they need to get solved, you're going to design your app in a different way. ### How much do you think the creativity being unlocked in the LangChain ecosystem is still modest relative to what will happen with the next major innovations in the underlying model architecture? Are we seeing the thing or are we just seeing a tiny little precursor to the thing? It has to be the latter. I've been in this space now for a year, and the types of applications people are building today versus last year are very different. ### Because the models have gotten better or because people have gotten better at using it? Skillset mostly. People have figured it out. They've gotten more of the basics down. They've figured out how to push the technology further. Yes, the models are getting better as well and that helps. But as people learn best practices for prompting, for fine tuning, for testing, we can just go a lot further. --- --- title: "Why Flybuys migrated to dbt Cloud" description: "Flybuys needed a single data control plane to increase data quality and productivity. Here’s why they went with dbt Cloud." url: "https://www.getdbt.com/blog/flybuys-dbt-cloud-migration" date: "2024-10-04" authors: ["Mark Wan"] categories: ["Partnerships"] --- # Why Flybuys migrated to dbt Cloud Like many companies, Australia’s Flybuys found itself in an unenviable position with its data. It had multiple, redundant methods for slicing and dicing data—and none of them were easy to use. Flybuys knew they needed a single data control plane. The question was which tools and approach would yield the greatest return on investment. Recently, I sat down with Milos Zikic, Flybuys’ Lead Enterprise Data Architect, to discuss the issues that his company faced in taming data anarchy. We discussed why they ultimately chose dbt Cloud as their data control plane, how they did it, and how it paid off for them. ## Four different data architectures Established in 1994, Flybuys is one of the biggest rewards programs in Australia, with 9.5 million members. On top of offering a traditional loyalty program, they also run campaigns and perform partner offer management through both traditional and digital media channels. This means Flybuys processes data. A **lot** of data. According to Zikic, the platform churns through 1.4 billion transactions every month. Up until 2022, Flybuys had four different data teams with four different approaches to managing data: - Their data platform team was using dbt Core to perform data transformations, supported by a home-grown infrastructure platform that ran on AWS - An analytics team was with [Snowflake](https://www.snowflake.com/) stored procedures, with compute spun up in AWS - Another analytics team used a combination of [Amazon Sagemaker](https://aws.amazon.com/sagemaker/) and [R](https://www.r-project.org/) scripts to do its data transformation - Finally, a fourth team used the Sagemaker/R platform along with a separate machine learning platform While these systems worked for each team, maintaining four separate environments brought a number of problems: ### Waste Every team had to stand up—and support—its own infrastructure for running data transformation pipelines. ### Inconsistency There was no uniformity or standardization in how teams modeled data. Not everyone was testing data (or testing it in the same way). There was also little opportunity for code reuse. ### Complexity All of these systems were hard to understand and use. Data analysts who needed a new type of data product had to rely on data engineers to make it happen, rather than self-servicing their own solutions. ## 192% ROI: Why Flybuys made the leap Flybuys knew it wanted to move to a single, standardized system for creating, testing, and managing data transformations. Zikic and his team thought that using a single system would eliminate duplicate effort, boost productivity, and enable greater self-service for analysts and end users. The team saw two possibilities: - Self-host dbt Core as a centralized platform; or - Move to dbt Cloud Self-hosting dbt Core seemed attractive because it meant not paying for dbt licenses. However, it also meant standing up and maintaining a centralized infrastructure—which would consume an enormous amount of time and resources. By contrast, dbt Cloud would give them this centralized platform out of the box, with no need to maintain separate infra. It'd also get a lot of features—such as auditing and role-based access control—that would be expensive for them to get right. The team proposed both future states. It also ran experiments to gauge how each one would work and which one provided the best improvement at the lowest cost. It then built a business case for each tool. According to Zikic, Flybuys’ analysis revealed that the company could earn a 192% Return on Investment (ROI) by going with dbt Cloud. The move would save 143.6 hours—a full 17 days—per analyst every year. It would save even more time—152.6 hours—per engineer per year. Most of this cost savings would come from the added features that dbt Cloud provided. With dbt Cloud, [testing](https://docs.getdbt.com/docs/build/data-tests), [version control](https://docs.getdbt.com/docs/collaborate/git/version-control-basics), and [documentation](https://docs.getdbt.com/docs/build/documentation) all come out of the box, without the overhead of manual coordination and infrastructure management. Zikic says that dbt Cloud’s support for [Continuous Integration/Continuous Deployment (CI/CD)](https://docs.getdbt.com/guides/custom-cicd-pipelines?step=1) was also a driving factor. “Our CI/CD was very convoluted,” he said. “Only data engineers were doing it.” dbt Cloud would enable Flybuys’ analysts to set up data pipelines that they could use to test, verify, deploy, and monitor changes to data models. ## How dbt Cloud aligned with Flybuys’ values Flybuys also had to make the case that using dbt Cloud aligned with the company’s core values. ### Be financially responsible Zikic’s team made an effective case that using dbt Cloud would reduce the costs of running data transformations by reducing both complexity as well as duplication of effort. The company also aimed to save money by increasing work velocity for engineering and analytics teams. In terms of concrete numbers, Flybuys aimed to reduce spend on AWS services for data transformation by 15%. It also predicted a whopping 70% reduction in data transformation infrastructure costs. ### Build secure and trustworthy systems Flybuys needed to verify that their dbt Cloud implementation met the company’s standards for security and trust. More than that, however, Zikic argued that dbt Cloud would improve both the company’s data quality as well as its overall data security posture. By acting as a data control plane for all data, dbt Cloud would bring improved quality and heightened consistency to its data, as well as help the company standardize on metrics. In addition, dbt Cloud’s embedded documentation support would lead to better clarity and transparency of data analytics code. “Using a Software as a Service (SaaS) system like dbt Cloud provides better out-of-the-box security than we could provide in our bespoke environment,” Zikic said. **** ### Strengthen company culture Before dbt Cloud, data analysts were heavily reliant on data engineers to produce new data products. Flybuys wanted to empower its analysts to create more [data products](https://www.getdbt.com/blog/data-product-examples) on their own. That meant providing a reliable CI/CD deployment process with the proper safeguards in place to ensure no unauthorized changes shipped to production. Zikic’s team set a goal of increasing the number of quarterly data product deployments by 100%. It also aimed for zero unauthorized production pushes. Hitting these goals would increase the value of data and reduce the time it takes to provide insights to end users. ## How Flybuys managed the transition to dbt Cloud To make this transition smooth, Flybuys took it slowly. It started with a proof of concept to suss out what worked and what didn’t about its planned migration. Flybuys found immediate value in several dbt Cloud features: - dbt Cloud’s infrastructure management and focus on application development—particularly, easy-to-use Integrated Development Environment (IDE)—made it easier for anyone to set up a data pipeline. Environment management took away the pain of managing infrastructure and environmental conflicts. - Easy security set-up using Role-Based Access Controls (RBAC). - Access controls and audit logs made it easier to verify protection for workloads with sensitive data, such as personally identifiable information (PII). That was a lot more complicated to achieve with dbt Core, as workloads ran on people’s local machines. - It’s easy to manage dbt Core versions and release in dbt Cloud and keep everyone on the same version. Flybuys also identified some points where they'd still need to spend some additional time and work. Integration with GitLab required some additional footwork in terms of security configuration and repo management. Additionally, Zikic’s team kept Airflow in the mix to enable integration with non-AWS resources. After identifying and addressing the gaps, Flybuys started transitioning to dbt Cloud. It started small—just two use cases supported by ten licenses across the data engineering and analytics teams. The team focused its efforts primarily on member segmentation and member engagement scores. Flybuys ran the pilot across Q3 and Q4 of 2022. It then evaluated the project based on the success metrics defined above. That would determine whether it moved forward with migrating other use cases onto the new system. ## The results So how’s it going two years later? Zikic says that dbt Cloud is now the official platform for productionizing enterprise-grade data transformation pipelines at Flybuys. The company has 100+ projects on dbt Cloud, with dbt Core only managing a handful of legacy jobs. It’s also integrated dbt Cloud with a number of its other systems, such as [FiveTran](https://www.fivetran.com/) and [Hightouch](https://hightouch.com/). Analysts and data engineers report a high level of satisfaction with the new toolset. The main benefit is that dbt Cloud gives them a single, simple way to build new data projects. Using dbt Cloud, both data engineers and analysts can focus on delivering pipelines instead of managing environments. Zikic said work remains to be done. In particular, Flybuys is still running some of the legacy tooling it had hoped to replace. The team plans to continue reducing this, introducing additional vendors that easily integrate with dbt Cloud as part of its modern data stack. ## Conclusion Every team’s data journey will be different. No matter what data challenges your team faces, dbt Cloud can help by serving as your data control plane—a single, comprehensive platform that all data stakeholders can use across your organization. **** Find out more about how dbt Cloud can address your company’s specific needs—[schedule a demo today](https://www.getdbt.com/contact). --- --- title: "The ADLC: Your compass for navigating modern data analytics" description: "Discover how the ADLC can serve as your compass in navigating modern data analytics." url: "https://www.getdbt.com/blog/adlc-compass-modern-data-analytics" date: "2024-09-29" authors: ["Alex Welch"] categories: ["Learn"] --- # The ADLC: Your compass for navigating modern data analytics _This post first appeared in_ _[The Analytics Engineering Roundup](https://roundup.getdbt.com/p/a-compass-not-a-map)._ It’s 2003, and you’ve just watched Captain Jack Sparrow reclaim his beloved Black Pearl. You rode the edge of your seat as he evaded the English navy, outsmarted an undead pirate crew, and navigated treacherous waters with a compass that doesn’t point north. Fast forward to 2024. You are navigating the equally unpredictable seas of modern analytics. Your treasure? It's not buried gold on a remote island, but the invaluable insights hidden within your data. Your crew? A talented team of data professionals. And your compass? That's where the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) comes in, ready to guide you through the challenges ahead. But here's the catch: just as Jack's compass behaves differently based on his deepest wants, the ADLC's guidance changes with your team's unique circumstances, skills, and business needs. It's not about following a predetermined route; it's about navigating the waters of data analytics using eight interconnected stages as your guide: Plan, Develop, Test, Deploy, Operate, Observe, Discover, and Analyze. ## Understanding your compass: ADLC basics Imagine you’re holding Captain Jack Sparrow’s compass. It’s not your ordinary navigational tool—it’s a bit eccentric, incredibly powerful, and, most importantly, personalized. That’s the ADLC in a nutshell. The ADLC isn’t just another rigid framework to implement and forget. [It’s a versatile guide comprising of eight interconnected stages](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle#the-adlc-model): Plan, Develop, Test, Deploy, Operate, Observe, Discover and Analyze. But here’s the kicker—unlike a traditional compass, the ADLC doesn’t always point in the same direction for everyone. Your team’s skills, your organization’s needs, and the maturity of your data practice all influence which way the needle swings. For a scrappy startup, the compass might pull strongly towards the [Develop](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle#develop) and [Deploy](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle#deploy) stages, emphasizing quick iterations and rapid value delivery. A heavily regulated organization might find the needle frequently pointing towards [Test](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle#test) and [Observe](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle#operate-and-observe), prioritizing data governance and reliability. The ADLC's strength lies in its adaptability. It’s not about rigidly following a predefined path, as you would a map, but about using these stages as guideposts to chart your own course. Sometimes you’ll move sequentially through the stages. Other times you’ll jump back and forth as needed. The key is to let your compass guide you based on what your organization truly needs at any given moment. Remember, just as Jack's compass led him to what he desired most, the ADLC is designed to guide you towards your most valuable treasure—a mature, efficient, and impactful data practice. ## Assessing your current position: data practice maturity Before setting sail with the ADLC, it’s crucial to understand where your data practice currently stands. You need to know your starting point to chart an effective course. Data practice maturity isn’t just about having the latest tools or the biggest team. It’s about how effectively you use what you have, how robust your processes are, how well you can adapt to changing business needs, and, most importantly, what value to drive. Your data journey might begin with scattered information across systems, manual processes, and a small team. At this point, your ADLC compass likely points towards [Discover](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle#discover-and-analyze) and Develop, focusing on understanding your data landscape and building foundational pipelines. As you progress, you’ll establish more robust pipelines and begin to deliver regular insights. Here, your focus might shift to Test and Deploy, improving reliability and speed. Further along, your processes become largely automated, delivering significant business value. You might concentrate on Operate and Observe, fine-tuning systems and implementing proactive problem-solving. At later, more advanced stages, data becomes core to decision-making and you begin to explore cutting techniques. Your compass might guide you towards [Analyze](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle#discover-and-analyze) and [Plan](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle#plan) while you navigate complex cross-functional challenges. These aren’t distinct categories, but a continuum. Your position on this journey isn’t a judgment—it’s a reality check. The key is honesty about your current state to guide your ADLC implementation effectively. **Some questions to ask to help triangulate:** 1. How is your data currently stored and managed? Is it scattered across systems? Is it centralized but with limited accessibility? Or is it well-organized and layered with governance? 2. What’s the state of your data process? Is it a mix of automated and manual processes? No automation? Or is it highly automated with reliable workflows? 3. How does your organization use data in decision-making? Is it being used to drive the strategic plan or only for operational decisions? Do they trust it at all? 4. What’s your team’s data expertise level? Are you learning as you go? Do you have a solid foundation with some specialized skills? Or do you have advanced skills across multiple domains? 5. How do you approach data quality and reliability? How reactive are you to issues? Do you have adequate test coverage? Before you plan your next move, take a good look at your data practice. How reliable are your data pipelines? How comprehensive is your documentation? How aligned is your team with business objectives? Your answers will help you understand which way your ADLC compass should be pointing next. ## Charting your course: aligning with business needs Now that you’ve got your ADLC compass in hand and you’ve assessed your current position, it’s time to chart your course. But here’s the catch: your destination isn’t a fixed point on a map—it’s the ever-shifting landscape of your organization’s needs and priorities. The ADLC isn’t just a technical framework—it is a business tool. Each step should solve a stakeholder’s pain point or support a key initiative. Your compass might be pointing in one direction, but if there’s no business value in that direction it’s time to recalibrate your understanding of data’s place within the organization. To keep your ADLC journey on course: ### Start with the “why” Ask yourself, “_Why does this matter to the business?”_ No clear benefit? Recalibrate. ### Engage your stakeholders Regular check-ins with business leaders help adjust your course. Their perspectives, goals, and pain points can help you adjust your course. Sometimes the most valuable data projects aren’t the most technically exciting ones. ### Prioritize flexibility Business needs can change overnight. Maybe a competitor launches a new product or a [global event disrupts your supply chain](https://www.weforum.org/agenda/2021/07/covid-19-pandemic-global-supply-chains/). Ensure your ADLC implementation can pivot when needed. ### Balance short-term wins and long-term goals Deliver quick wins to build trust, but don’t lose sight of the long-term strategy. Your ADLC compass should help you navigate both. ### Practical tip Use a “Business Value Checklist” for each ADLC stage. Ask: Which pain point does this address? How does it align with top priorities? Can we quantify the impact? What’s the cost of not doing this now? This checklist is just one tool in your toolkit. We’ll explore more later, including regular health checks and feedback loops. Your ADLC compass is pointing towards value, not just completion. If you can’t articulate the business value of a step, it might be time to change course. ## Navigating the winds of change: adapting the ADLC Implementing the ADLC isn’t always smooth sailing. Your team’s skills will evolve, business priorities will shift, and new technologies will emerge. The key to success? Adaptability. ### Adjusting based on team skills: - **Skill gap analysis:** Regularly assess your team’s strengths and weaknesses. Maybe you’ve got a strong analytics engineer team but lack data visualization experts. Your ADLC compass might point towards up-skilling in the Analyze stage. - **Cross-training**: Encourage versatility. A data engineer dabbling in analysis can bring fresh perspectives to both roles. - **Strategic hiring:** Use the ADLC to guide recruitment. Identified a bottleneck in the Test stage? It might be time to bring a QA specialist on board. - **Leverage external expertise:** Don’t be afraid to call in reinforcements. Consultants or temporary hires can help you navigate the tricky parts of your ADLC journey. ### Adapting to changing business priorities: - **Agile ADLC sprints:** Break your ADLC journey into short sprints. This allows you to reassess and pivot more frequently based on business needs. - **Priority mapping:** Regularly map your ADLC activities the current business focus. If the business suddenly pivots to customer retention, your Discover and Analyze shift to churn prediction. - **Feedback loops:** Establish regular check-ins with key stakeholders. Their input can help you adjust your ADLC compass before you veer off course. - **Modular implementation:** Treat the ADLC stages as building blocks, not a fixed sequence. You might need to jump back to Plan midway through Develop if business requirements change. The ADLC is meant to be a flexible guide. For teams early in their evolution, this may mean focusing on basic skill development across all ADLC stages. A team further along might emphasize cross-training in specific areas. Your ability to adapt to your team’s evolving skills and your organization’s changing needs is what will ultimately determine your success. ## Avoiding the sirens: common pitfalls Even with the best compass, treacherous waters remain. Beware of these common challenges that can derail your efforts: ### Perfectionism Striving to perfect each stage can stall progress, especially as you begin to scale. This is when the desire to create robust processes can overshadow the need for quick wins and iterative improvement. _Solution:_ Embrace iterative improvement. Start with a minimum-viable process and refine as you go. ### Technology fixation Implementing new tools without clear business justification can lead to unnecessary complexity and cost. Teams early in their journey are particularly susceptible to this, often believing that the latest tool will solve all their data challenges. Be wary! Even teams on the bleeding edge can get distracted by the shiny object. _Solution:_ Tie tool adoption to specific business needs or ADLC improvement. Prioritize solving real problems over acquiring new technology. ### Departmental silos Isolated ADLC implementation leads to fragmentation. This becomes particularly problematic once your team has grown beyond the initial centralized team. The increased specialization may inadvertently lead to isolation between teams. _Solution:_ Foster cross-functional collaboration. Ensure the Plan stage involves stakeholders from across the data team and the broader organization. ### Misaligned metrics Focusing on vanity metrics can lead to misguided efforts. This is a pitfall that teams of all shapes and sizes run the risk of falling victim to. Don’t overemphasize complex metrics that fail to tie directly to business outcomes. _Solution:_ Align your ADLC metrics directly with business impact. Prioritize quality and value over quantity. ### Lack of continuous improvement Treating the ADLC implementation as a one-time project can quickly lead to outdated and ineffective processes. This is particularly dangerous in more mature practices, which might become complacent with their existing processes. _Solution:_ Schedule regular ADLC reviews. Continuously assess if all stages are serving their purpose and incorporate new best practices. ### Over-engineering Building complex systems for hypothetical future needs wastes resources and creates unnecessary complexity. This is particularly dangerous for early and young data teams. _Solution:_ Start simple. Let actual business needs drive increased complexity over time. ### Insufficient documentation Neglecting documentation can turn your ADLC into tribal knowledge and hinder your ability to scale. This is something all stages are likely to encounter. _Solution:_ Make comprehensive documentation a required part of each ADLC stage. This supports consistency, scalability, and knowledge transfer. By being aware of these potential pitfalls, you can proactively address them in your ADLC implementation. Remember, encountering challenges is part of the process. The key is to learn from them and continuously improve your approach. ## Calibrating your compass: continuous improvement Implementing the ADLC isn’t a one-and-done process. Just like how Jack Sparrow’s compass changes course depending on what he wanted most at the moment, so too will your ADLC compass shift. The key is understanding why that happened. Implementing the ADLC isn't a one-time process. Understanding why is key to keeping your implementation relevant. ### Regular health checks Schedule quarterly reviews of your ADLC implementation. Ask key questions like: - Are all stages still aligned with our business objectives? - Are we collecting the right data? - Which stages create the most value? Which create bottlenecks? - How has our data maturity evolved, and does our ADLC approach reflect that? ### Feedback loops Establish mechanisms to gather ongoing feedback and use it to make incremental improvements to your process: - Data team members: Frontline insights can identify practical issues quickly. - Business users: Are they getting timely, needed insights? - Executive leadership: Is the ADLC supporting strategic decisions? ### Metrics that matter Develop and track metrics that reflect the effectiveness of your ADLC implementation. Regularly review these and adjust your approach accordingly. Below are some metrics to get started with: - Cycle time from data ingestion to insight delivery - Number and severity of data quality issues - Team productivity and satisfaction. ### Stay informed The data world evolves rapidly. Stay current with the following and consider how these developments might enhance or impact your ADLC implementation: - Industry best practices - New tools and technologies - Regulatory changes - Emerging data ethics considerations. ### Cross-pollination Encourage knowledge sharing within your organization. - Rotate team members across different ADLC stages - Host internal workshops to share learnings and challenges - Create a best practices knowledge base. ### Flexibility in implementation Be prepared to adapt your ADLC approach as circumstances change. Your ADLC should be robust enough to handle change, yet flexible enough to embrace it. - Business priorities shift - Team composition evolves - New data sources become available - Technological capabilities advance The goal isn’t perfection—it’s progress. Each small adjustment to your ADLC implementation can lead to significant gains in efficiency, effectiveness, and business impact over time. ## Your true north The ADLC isn’t a rigid map with a predetermined route. Instead, it guides you toward what you truly desire: a data practice that delivers real, measurable value to your organization. And just as Jack’s compass behaved differently based on his changing needs, your ADLC implementation will be unique to your organization’s needs, maturity, and aspirations. - Your current position (data maturity) influences how you interpret the compass readings. - Your desired destination (business needs) determines which direction you should sail. - The winds of change (evolving team skills, new technologies, and shifting priorities) require constant course adjustments. - Treacherous waters (common pitfalls) demand vigilance and strategic navigation. - Your unique voyage (company stage, needs, and goals) shapes how you apply the ADLC principles. - Regular calibration (continuous improvement) ensures your compass remains accurate over time. The ADLC’s power lies not in blindly following its stages, but in thoughtfully applying its principles to guide your data journey. It’s about striking the balance between structure and flexibility, between best practices and practical realities. As you embark on your own ADLC journey, remember that there’s no “perfect” implementation. What matters is that you’re moving in the right direction, continuously learning and improving along the way. Your North Star **_isn’t_** a flawless data practice—it’s one that consistently delivers value to your organization while adapting to new challenges and opportunities. It’s time to grab your ADLC compass and chart your course. Where will your data journey take you? --- --- title: "Peak data performance: Moving from dbt Core to dbt Cloud" description: "Learn how WHOOP improved data governance, transparency, and collaboration by transitioning from dbt Core to dbt Cloud." url: "https://www.getdbt.com/blog/moving-dbt-core-dbt-cloud" date: "2024-09-27" authors: ["Sara Gawlinski"] categories: ["Learn"] --- # Peak data performance: Moving from dbt Core to dbt Cloud Many teams start off on dbt Core to give them a common model for managing their data transformation code. However, at some point, they need a more centralized, governable solution. That’s where dbt Cloud comes in. dbt Cloud serves as an organization-wide hub for data transformation, metrics, governance, and security. It unlocks a host of features—from data discovery to release management—that teams using dbt Core would otherwise have to build themselves. Recently, I had the pleasure of speaking with Matt Luzzi, the Director of Analytics for human performance data company WHOOP, about how they managed the transition from dbt Core to dbt Cloud. Learn when and why WHOOP made the move, how teams managed the migration, and the additional business value that WHOOP unlocked by making the move to the cloud. **** ## What brought WHOOP to dbt Core WHOOP seeks to unlock human performance in a data-driven manner by developing an internal algorithm to analyze biometric data around sleep recovery and strain. The company, which has been around for 12 years, bases everything it does in data. The analytics team at WHOOP is responsible for helping the company become and remain data-driven. It partners with everyone from marketing to product growth to engineering to run experiments to think about new ways to market or build solutions. When WHOOP started, it didn’t have dbt or, really, any central data orchestration. It was all SQL scripts created to run ad hoc. That meant analysts had little visibility into how data was created. It also meant the company had no data for success metrics (e.g., data test pass/fail rates, data incidents/table, etc.) on getting new data into production on a daily basis. Initially, the analytics team used dbt Core as a free solution to create a single, centralized layer for orchestration and transformation. As everyone on the team was SQL-literate, this was a pretty easy transition. It gave WHOOP a single, systematic approach to data transformation that offered greater visibility and code reuse than the previous ad-hoc approach. ## The motivation to move to dbt Cloud However, WHOOP soon realized that dbt Core didn’t solve other lingering problems. One key problem was that there wasn’t ‌centralized governance or a single source of truth. Two different analysts could be creating more or less the same dbt model and have no visibility into one another’s work. Another missing piece was scheduling and orchestration. Since dbt Core isn’t a centralized service, the team had to rely on external tooling to enable orchestration. That didn’t afford them real visibility of what failed or which downstream models were skipped. The team knew it needed to switch to a more shared, centralized solution. At the same time, the team was planning a transition from AWS Redshift to Snowflake. With this migration, they wanted to do more than just “lift and shift” the data. They wanted to ensure that the data was cleaned and well-governed. “We wanted to be really deliberate about what we brought into Snowflake,” Matt said. ## Moving to dbt Cloud incrementally That’s why the analytics team decided to shift to dbt Cloud. Instead of porting over their original dbt Core code, they started from scratch, reassessing their first principles. The team (around 12 people at a time) identified all the metrics they wanted to track across the company and worked backward from there. This “backwards” approach worked well for the analytics team. First, it guaranteed that everything that ended up in Snowflake was data that they needed. Second, it meant there was a single source of truth for all data sets (as opposed to say, one table for sales and a separate table with similar data for product growth). WHOOP was also happy to realize that moving to dbt Cloud didn’t mean scrapping the infrastructure it built around dbt Core—CI/CD processes, dev/test environment, etc. The team brought new workloads onto dbt Cloud but left existing workloads in the Core architecture they’d built. As new analysts onboard to WHOOP, they go straight to dbt Cloud, which unlocks a number of new use cases for the company. ## Structuring cross-team work Using dbt Cloud also allowed different data teams within WHOOP to collaborate more closely together and avoid data silos. Other teams—such as data engineering and data science—noticed the work the analytics team was doing and how easy it was to work with dbt Cloud. These teams could onboard themselves easily by creating their own [dbt Cloud projects](https://docs.getdbt.com/docs/build/projects) and Git repositories. That gave each team its own separate workspace and version history. To facilitate working with core data assets, the analytics team created a WHOOP Commons dbt project. This project contains company-wide data as well as reusable code, such as macros, that are useful across data projects. Using [dbt Mesh](https://www.getdbt.com/product/dbt-mesh), each team can find these common assets via [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) and reference them from their own projects. “We don’t have to create workarounds to bring models in from different projects like we would in dbt Core,” Matt said. “It’s all native within the cloud platform and all of our models are transparent. It gives us a lot of visibility into where things are and how they’re being used.” ## Managing production releases with dbt Cloud WHOOP says that another area where dbt Cloud brought added value was by enabling greater deliberation and quality in its production release processes. At WHOOP, software development teams don’t push changes to production on Fridays. If something breaks, that means someone’s on the hook to work throughout the weekend to fix it. In this same vein, the analytics team wanted a deliberate release process where it could understand the full impact of a change before they pushed it live. To accomplish this, Matt and his team created a process where they fork their dbt code off of production every Friday. Engineers accumulate changes in a QA branch. Later in the week, an analytics engineer compiles these changes into a release and obtains code owner review. All changes also have [accompanying unit tests](https://docs.getdbt.com/docs/build/data-tests) that run against both test and, eventually, production data. This ensures that metrics that shouldn’t be changing aren’t changing, and that code changes are doing what they were intended to. Communication is also a critical part of the release strategy. Matt and the team publish release notes in a Slack channel—what metrics might be changing as a result of code changes, which tables are being added or removed, etc. “This holds us accountable while also giving visibility to the organization,” Matt said. “This is all possible because of the dbt ecosystem and the developer framework it provides.” As a result, Matt said, the analytics team has reduced the number of accidental errors pushed to production to essential zero. “We’re really proud that we don’t push bad code out. It helps me sleep better at night.” ## Conclusion As WHOOP continues its data journey, Matt said his team is looking at leveraging other critical features of dbt Cloud. In particular, they’re looking at how [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) can better prepare them to support AI use cases. “I foresee that being a huge player in the world of AI and natural language chatbots. If you’ve thought out the semantics of these models really thoroughly and accurately in a way that a machine can understand them, it's only a matter of time before we can have conversations with this level of data. By having all of the metadata in dbt in a way that's discoverable and consistently laid out, it sets us up for a world where that's a very real near-term possibility.” Learn more about how dbt Cloud can improve data discoverability, enable data governance, and prepare you for an AI future—[contact us today for a demo](https://www.getdbt.com/contact). --- --- title: "Creating value from GenAI in the enterprise" description: "Capital One's head of enterprise data tech on building a strong data culture for what comes next." url: "https://www.getdbt.com/blog/creating-value-from-genai-enterprise" date: "2024-09-22" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Creating value from GenAI in the enterprise _This post first appeared in_ _[The Analytics Engineering Roundup](https://roundup.getdbt.com/p/creating-value-from-genai-in-the)._ Nisha Paliwal is the managing vice president of enterprise data tech at Capital One and the co-author of the book [Secrets of AI Value Creation](https://docs.google.com/document/d/106Uo0RtzwSxnZzVoPg4vpA_1uyOqasDS/edit). Nisha and Tristan discuss everything from how to get children excited about coding, to heterogeneous data environments, to building strong data culture. _Join data practitioners and data leaders in Las Vegas this October at_ _[Coalesce](http://coalesce.getdbt.com/)—the analytics engineering conference built by data people, for data people._ **_Use the code podcast20 for a 20% discount. _**[Register now](https://coalesce.getdbt.com/). **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Google Podcasts](https://podcasts.google.com/feed/aHR0cHM6Ly9hbmFseXRpY3NlbmdpbmVlcmluZ3JvdW5kdXAubGlic3luLmNvbS9yc3M) - [Stitcher](https://www.stitcher.com/show/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways from this episode ### I feel like Capital One almost invented the modern credit card industry, the data-driven practices of segmenting customers based on traits and then giving them specific credit card offers based on those traits. This is all stuff that Capital One pioneered, and it's all driven by data. Is that fair to say? **Nisha Paliwal:** There’s a very old paper—you can Google “IBS Information Based Strategy”—on which Capital One was formed. It's a public paper. [You can download it and read it.](https://www.computer.org/csdl/proceedings-article/hicss/1998/82480311/12OmNyQpgVH) It has a lot of history of Capital One. Our founder, who's now our CEO, Rich Fairbank, started the company 30 years ago on this IBS strategy. Information equals data, right? So data is everything to this company, and has been everything to this company for many, many years. Everything we do in this company, there’s a prefix and suffix of data. Where the world is now swirling on many things, like AI, it starts with data. It starts with a lot of these practices, which are engraved in the company. ### It's fairly common for a digital native company founded from 2008 onwards to think of itself as a data company, but it's a bit more unusual for a company founded in 1994 to think of itself in that way. Yeah, so think of all the aspects of roles that we have to perform, whether it's risk or legal, or HR, or even our tech roles for that matter, right? If your data thinking is backing all those roles, then you suddenly have a different culture in the company. A lot of things you hear at Capital One come from empirical evidence. And where does empirical evidence come from? Data, of course. So whether it's our job as technologists or HR or legal, what is that empirical evidence? What is the data telling us? We are guided by data in every strategy conversation, which is very different than many companies. And then the other point I'll make is about risk management. For all 50,000 people in the company, all our jobs are risk management. And how does one do good risk management? There are a lot of people who go from Capital One and come back, and when we see it outside, we’re startled on how these companies do risk management. We deal with people's money, and so we need to have a risk management angle. Where does good risk management come from? Knowing your data. Some of this is engraved in the culture. That's how we train everybody. That's how everybody's coached. That's our common vocabulary. ### How does that impact the type of humans that you bring onto the team? Capital One is a very quantitative and data-driven financial institution. Does this make it hard to hire people from other financial institutions who might have different practices? Is there culture shock? That answer will depend on the teams and level, of course, but the ingredients to be successful I can talk about. One is learning. Your appetite for learning has to be insurmountable. I can't imagine the type of learning I've had in my years at this company. Second, which is unique in this company, is a very intense strategy process. And that strategy process is happening every year. It's not two to three years or five years. We get to hear from our CEO every year on the strategy. Everybody coming in gets a lot of training, coaching, onboarding to understand how this company operates. Our onboarding is so unique because we put a lot of rigor around it. I have 30 or 35 interns, some are interns for 10 weeks. We make sure we surround them with the ecosystem that’s needed to thrive in a place like this. Data doesn't know which role you are in. Data is everything to a company. And when something that important is everything to the company, we are talking of culture. You have to make sure anybody walking in the door, regardless of job, sees the importance of what we are doing. Data is how you make decisions, how you close books, how you do anything. Don't we all want to have trust in the data that is in our hands? Data doesn't lie. But for data not to lie, it has to be correct, so I can trust in that data. And this is where the collaboration point comes in, because no particular job will ensure that it is right. But collectively, we can make sure it is right all the time. ### You recently co-wrote a book called [Secrets of AI Value Creation](https://aivaluesecrets.com/). Can you tell us a little bit about that book and what you cover in it? The beautiful part of the book is it has the framework for how we can think about AI. A common question that I hear from many CEOs when I go to any conferences is, “Okay, where do I start?” So the book does a great job of giving that framework. I’d say it's more for business CEOs who want to start somewhere. It's not a heavy tech read. It doesn't go into a lot of tech but it goes into this framework, which gives them all the components that we need to think about. ### Let's imagine that you're at some conference and a CEO of a large organization approaches you and they ask you the question, “I haven't yet tried to figure out my AI strategy. I'm worried that this is all a bunch of hype. Do you think that I should roll up my sleeves and dig in, or do I need to let this thing cook for another couple years before it's really worth my organization's time?” What do you say? I’d say start with the four-step framework in the book: vision, strategy, architecture, and execution. I think on tech topics, we often start execution too quickly. The first question I’d ask is, where are you taking your company? And does any tech matter for the direction that you are taking? And then get into these steps of strategy and architecture. Strategy guides you in that direction of where to start. Do I start small? Do I let it bake? The biggest mistake with waves of technology I have often seen is we go to execution first. And then because we haven't set up that vision and strategy, we struggle to find the ROI. The book explains this framework with tons of great examples from great companies on how they have approached all of this. ### I feel like one of the things that's so hard for leaders to reason about with AI is that it's hard to know what’ll work. It feels like hard to engage with AI today, because you can't really ask somebody unless they're one of the very few leaders in the field and say “Hey, I want to do X. Is that even possible? How do you escape this circular logic path? That's the R&D research part you're talking about or the innovation part. There are endless resources; I don't think this topic is new anymore. You should see the AI hackathon that my daughter does in high school. The question is, where do you want to invest in that technology? You need to invest very carefully because these things are ’t cheap. These are places where we put a lot of time and energy without an ROI. That's the value creation part of it. So how do you start small enough? And of course, this is a data podcast. You can’t do this without data. ### The book starts out with a bit of a cautionary statistic. I think it was something like 85 percent of organizations are failing to live up to their AI goals. That clearly creates urgency around having a framework, having best practices, et cetera. What does a failed AI project look like? What we were hearing when we wrote the book was often a lack of clarity on the areas where they want to use AI. We talked to a lot of CFOs who were struggling with whether it was an investment they wanted to pursue. We saw data quality issues. Do I trust the data? I don't think the problems are unique between industries. ### Are there the dead-obvious use cases for AI yet that everybody has decided is a clear win? A hundred percent of everything is never a true thing. But I’d say what I’ve read is everybody is doing some sort of customer interaction. I think for AI to get mainstream, we still have to get very comfortable with the usage and the result of it. We’re still putting a lot of humans in the loop to validate, because we need that validation. --- --- title: "dbt Cloud on Microsoft Fabric: The recipe for high-quality AI data" description: "Learn how you can use dbt Cloud and Microsoft Fabric to create high-quality data to power the next generation of AI apps." url: "https://www.getdbt.com/blog/dbt-cloud-on-microsoft-fabric" date: "2024-09-20" authors: ["Jeff Mills"] categories: ["Product"] --- # dbt Cloud on Microsoft Fabric: The recipe for high-quality AI data Earlier this year, our partners at Tableau and Salesforce found that 86% of data leaders in a survey agreed that AI outputs are only as good as their inputs. Unfortunately, creating a petabyte data architecture at scale with high-quality data is a challenge that many companies are still struggling to crack. One way to overcome this challenge is through the one-two punch of Microsoft Fabric and dbt Cloud. In this article, I’ll show how you can use both systems together to deliver consistently high-quality data with minimal architectural overhead. ## Microsoft Fabric: Data architecture as a service First, let’s look at the current landscape for ML, AI, and data. This graphic from Firstmark Capital shows, at a glance, the challenges that data engineers are up against. ![MAD Landscape](https://cdn.sanity.io/images/wl0ndo6t/main/9a8701dffe30c19d092365214a78b81d907a1fdd-2190x1134.png) Source: [The 2024 MAD Landscape](https://mattturck.com/landscape/mad2024.pdf) It’s super complex. It’s overwhelming. This array of different technologies leaves engineers and other data users at a loss. Which of these hundreds of technologies should I use? Which ones integrate well with one another? Microsoft Azure faces some of the same challenges. It offers a family of individual Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and SaaS features to store, secure, and analyze data at a petabyte scale. Many of these are great services. However, they’re piecemeal. You have to stitch them together to create an enterprise-grade solution. Enter [Microsoft Fabric](https://www.microsoft.com/en-us/microsoft-fabric). Fabric combines new and existing capabilities into a single, unified SaaS platform that serves as a one-stop shop for data analytics. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0df150c0dca925e388b35f3dcff9837db3887a3a-1209x434.png) Fabric’s components all work in concert to enable full, end-to-end data management: - [Data Factory](https://learn.microsoft.com/en-us/fabric/data-factory/data-factory-overview) provides the capability to move and transform data, while Data Science enables your ML flows and modeling. - [Real Time Intelligence](https://learn.microsoft.com/en-us/fabric/real-time-intelligence/) provides access to low-latency data for IoT and real-time visualization use cases. - [Power BI](https://www.microsoft.com/en-us/power-platform/products/power-bi) enables visualizing and reporting on data, no matter where it’s stored. - [Data Activator](https://learn.microsoft.com/en-us/fabric/data-activator/data-activator-introduction) adds observability to your data stack. You can configure real-time alerts—emails, Teams messages, etc. - on actions such as when a lead or a buyer requires follow-up. Across all of these features, [OneLake](https://learn.microsoft.com/en-us/fabric/onelake/onelake-overview) provides data storage for any data object type. Meanwhile, [Microsoft Purview](https://learn.microsoft.com/en-us/purview/purview) for data cataloging solves longstanding challenges that data engineers and analytics engineers face in finding and using data. Additionally, Artificial Intelligence (AI), made available through Microsoft Copilot, cuts across all of these services, differentiating Fabric from competitors. For example, data engineers can leverage Copilot for everything from creating complex SQL joins to choosing data model types or even writing Spark code. Meanwhile, business analytics or Power BI developers can leverage AI to develop new insights or visualizations on data. ## dbt Cloud: A data control plane for AI Microsoft Fabric is a great architecture for delivering the next generation of AI-powered solutions. However, when talking with dbt customers, it becomes clear a lot of companies aren’t enabled for AI at scale. The problem isn’t the architecture but the data. Most people’s greatest challenge is determining how much they can trust their data. When they look at a dashboard, their first thoughts are: what am I looking at? And where did it come from? Unfortunately, with the advent of AI, many companies have just implemented a Large Language Model (LLM) or other AI solution on top of that data. Without verifying the veracity of their underlying data, the answers they get from the resulting LLM may be inaccurate. They may be hallucinogenic. This leap—from faulty data straight to AI solutions—undermines trust in AI projects, just as delivering a faulty analytics dashboard would. High-quality data requires documenting what your data does and where it comes from. There are only two places where this documentation happens. One is in a data catalog or other centralized data repository. The other is during the transformation phase—which is where dbt lives. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/3f0d1917512ba26c21e554714e38878daed93c26-861x405.png) [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) helps companies transform their data via declarative statements about how that data ought to be organized. It provides a data control plane that not only documents these transformations but makes them available through orchestration, observability, a data catalog, and a semantic layer. **** With a data control plane, you do more than point AI at your data. You point it at the transformation logic so that it now knows how the data ought to be structured and how you think about that data. These layers are all powered by the work that data teams perform in transforming and shaping the data. The transformations, in essence, represent metadata. That metadata powers both the analytics and the AI apps that organizations aim to build. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/907f881058220a980237c943c3d5b63c1126f6f2-1177x499.png) ## A cheesy example Using dbt Cloud, you can bring the industry standard for data transformations to Microsoft Fabric. dbt Cloud makes it easier for people who work with data to build data transformations. ‌It brings more people across the company into the data fold by lowering the technical bar. Additionally, like Fabric, dbt Cloud works with a variety of data sources and destinations across the data industry. Customers use dbt Cloud to help them connect the dots between Fabric and the other touchpoints of their data ecosystem. That breaks down siloes, making previously isolated data part of a single, cohesive data estate. Let’s see how this works in practice with a cheesy (literally) example. Boss Big Slice, the boss of pizza chain Big Pizza, wants to know which province in Canada consumes the most pizza. So they call in their data experts, Mr. Ketchup and Miss Pepper. Mr. Ketchup is the company’s seasoned data analyst who powers all of the company’s Power BI dashboards. The problem is, he’s overworked. He’s the company’s one-stop shop for analytics and he’s just drowning in CSV files. He’s burnt on the bottom and ready to deliver himself to a vacation. In comes Miss Pepper, the organization’s new data engineer. She was hired to help the company build an enterprise-grade data infrastructure. Together with Mr. Ketchup, they’re ready to help take Big Pizza’s data operations to the next level. However, before they can get started, Boss Big Slice discovers during a meeting that their dashboards contain data from restaurants that are closed. In other words, their data quality is in shambles. They need to fix it—and fast. Miss Pepper steps in. She verifies the problem—but, since there’s no documentation for this data outside of the column names, it takes her a while to figure out a solution (filtering by companies she knows to be closed). She knows there’s gotta be a better way. ## How dbt Cloud works with Microsoft Fabric Miss Pepper’s grand plan is a new data architecture powered by Microsoft Fabric and dbt Cloud. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/5a4a50c2a81f9a9b4b4c3b123bee4992666c15e4-1164x713.png) In this new architecture, the team moves data from sources such as AWS, Snowflake, and Azure SQL via an ingestion process. The transformed data is stored centrally in OneLake. At the heart of this architecture within Microsoft Fabric is dbt Cloud. dbt Cloud supplies [modeling and data transformation](https://docs.getdbt.com/docs/build/models), [testing](https://docs.getdbt.com/docs/build/data-tests), [version control](https://docs.getdbt.com/docs/collaborate/git/version-control-basics), and a [CI/CD release management system](https://docs.getdbt.com/docs/deploy/continuous-integration) for data code changes. This ensures that data is always high-quality. Finally, data access is provided through tools such as Power BI, which analysts can use to power reports, and Data Activator, which provides data analytics and data event notifications via real-time analytics. This guarantees that every slice of data is fresh and actionable. Miss Pepper creates the data pipeline in dbt Cloud that powers this transformation. This pipeline orchestrates the movement and transformation of data from its source to the reports. She finds the data she needs from multiple sources to capture information on the restaurant’s current operational status. She sets this transformation to run on a schedule so the data updates periodically. After the transformation is completed, she runs a Fabric notebook that provides additional insights by performing sentiment analysis on the reviews. Because she’s using dbt Cloud, Miss Pepper can easily create a test to verify that she’s only ingesting data from currently open restaurants. She tests this by deleting the restaurant status filter from her code and running the pipeline in a staging environment. Sure enough, the data pipeline test fails and alerts her to the issue ASAP. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/acd758ae72ea5acf09b15cf4c0c4acb40cf30c71-948x529.png) She can then look at the failure in the dbt Cloud dashboard and perform root cause analysis to find where the failure occurred. ‌Using dbt Cloud, she can see a full data lineage map. This shows her data, its sources, and all of the tables it impacts downstream. Clicking through this table will give her a description of the table, the tests that failed, and the table’s relationships. Using this, Miss Pepper can find the specific column that caused the failure. Miss Pepper fixes the code, which triggers dbt Cloud to perform an automatic rebuild and run all of her tests. Since this code change caused the test to pass, she can now push this code change live and run the production Fabric pipeline. After the change, the pipeline runs successfully, creating all the required tables in Fabric Synapse Data Warehouse. Miss Pepper was even able to do some advanced sentiment analysis on reviews, classifying them as positive, mixed, neutral, and negative—all thanks to the built-in AI services offered by Fabric. Since Fabric is SaaS, Miss Pepper didn’t have to provision any infrastructure to build this new architecture. Creating a data pipeline in Fabric also gives her access to all of its data connectors. That means she could improve the report over time by including additional data sources. Boss Big Slice sees the final report and is very pleased. They can finally answer the question: which province in Canada has the best pizza? According to this 2018 pizza data set, the answer is Quebec, which had an impressive 81% average pizza rating score in 2018. But that's not all. This map was updated to show only the pizza restaurants with a perfect rating score. Miss Pepper was also able to use sentiment analysis to score the reviews of the restaurants. Out of the pizza restaurants with perfect ratings, Stripes Pizza and Kitchener had the highest number of positive reviews. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/7fef9140c07e649a447e4f2459dd089403ccbc48-949x539.png) ## Conclusion dbt Cloud brings an enhanced developer and user experience, powering a [DataOps](https://www.getdbt.com/blog/what-is-dataops) approach to data management. Its SQL-based framework makes it accessible to a wide range of users. dbt Cloud’s low-code approach to testing and documentation streamlines the development process. Additionally, dbt automatically calculates dependencies and lineage, reducing developer workload and accelerating development. It also allows business users to access ‌documentation, making data more transparent and accessible. Microsoft Fabric’s SaaS platform means that companies can manage data without managing infrastructure. Its large number of data connectors and options like database mirroring and shortcuts help manage and maintain [data gravity](https://www.getdbt.com/analytics-engineering/case-for-elt-workflow). Fabric supports AI and analytics across the platform, providing Copilot tools to speed up daily data operations, along with options for advanced machine learning. Used together, dbt Cloud and Microsoft Fabric provide a robust, end-to-end strategy for delivering high-quality data to power the next generation of AI applications. --- --- title: "dbt Labs Names Enterprise Software Veteran Sally Jenkins as Chief Marketing Officer" description: "Jenkins’ expertise in strategic marketing will guide and help scale the company as it continues its growth momentum." url: "https://www.getdbt.com/blog/dbt-labs-names-enterprise-software-veteran-sally-jenkins-as-chief-marketing-officer" date: "2024-09-18" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Names Enterprise Software Veteran Sally Jenkins as Chief Marketing Officer _Appointment comes as dbt Labs strengthens its leading position in the enterprise data market_ **PHILADELPHIA – **September 18, 2024 – [dbt Labs](https://www.getdbt.com/), the pioneer in analytics engineering, today announced the appointment of Sally Jenkins as Chief Marketing Officer. Jenkins takes on this new role at a critical juncture for dbt Labs, as the importance of trustworthy data is emphasized by the rapid growth of artificial intelligence, especially for the enterprise. Jenkins’ expertise in strategic marketing will guide and help scale the company as it continues its momentum and solidifies itself as the leading data partner for enterprise organizations. In the last year, dbt Labs has seen strong customer and community growth – reaching more than 4,600 global dbt Cloud customers and 100,000 dbt Community members. Jenkins will focus on supercharging this trajectory, leading brand and demand strategies that differentiate the organization, generate pipeline and grow dbt Labs’ leadership position in the enterprise data market. She will also be responsible for formulating a long-term go-to-market strategy that delivers a positive customer experience and supports industry-leading growth. “Sally has deep experience in data, open source, and enterprise software,” said Tristan Handy, founder and CEO of dbt Labs. “I could not be more excited about her combination of expertise, because I am a deep believer in the power of marketing to educate and advance an entire ecosystem. That’s what we’ve been doing for our entire journey at dbt Labs, and Sally is going to help us scale that to the next level.” Jenkins has more than 25 years of experience spanning hyper-growth startups to Fortune 500 companies and prominent brands, with an extensive background as CMO at several leading data and enterprise technology companies including Elastic, Informatica, and, most recently, SentinelOne. Throughout her career, she has focused on delivering brand affinity, building loyal communities and increasing customer acquisition to support revenue goals. She also brings extensive international experience and cross-functional expertise to her role, spanning enterprise and consumer SaaS/cloud environments. “The world is hyper focused on AI, but the success of these applications is dependent on high-quality, trustworthy data,” said Jenkins. “The foundational role dbt Labs plays in ensuring data quality for mission-critical company functions is crucial. dbt Labs has been a pioneer in the data space since its inception, and I’m eager to help continue to elevate the brand and focus on supporting the company’s unparalleled growth at such a vital time for the industry.” Jenkins joins dbt Labs as the company gears up for its annual user conference, [Coalesce](https://coalesce.getdbt.com/), on October 7-10 in Las Vegas. The annual event provides opportunities for customers to build connections, experience new dbt Labs products and brings industry leaders together to build the future of data. Her new role also comes on the heels of several impactful additions to the dbt Labs leadership team, with [Austin Stefani](https://www.prnewswire.com/news-releases/dbt-labs-names-former-rubrik-exec-austin-stefani-as-chief-revenue-officer-302073528.html), [Mark Porter](https://www.prnewswire.com/news-releases/dbt-labs-names-data-industry-veteran-mark-porter-as-chief-technology-officer-302050434.html), and [Brandon Sweeney](https://www.prnewswire.com/news-releases/dbt-labs-appoints-tech-veteran-brandon-sweeney-as-president-and-chief-operating-officer-301983208.html) joining the team as CRO, CTO, and President/COO, respectively. To learn more about open positions at dbt Labs, visit: [https://www.getdbt.com/dbt-labs/open-roles](https://www.getdbt.com/dbt-labs/open-roles) **About dbt Labs** Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 40,000 companies using dbt every week and over 4,600 dbt Cloud customers. To learn more about dbt Labs, visit [https://www.getdbt.com/](https://www.getdbt.com/) and follow us on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/). **Press Contact** Method Communications for dbt Labs - [dbtlabs@methodcommunications.com](mailto:dbtlabs@methodcommunications.com) --- --- title: "Developer productivity on GitHub Copilot" description: "GitHub Next's Dr. Eirini Kalliamvakou on making sure tracking productivity reflects reality." url: "https://www.getdbt.com/blog/developer-productivity-github-copilot" date: "2024-09-08" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Developer productivity on GitHub Copilot _This post first appeared in_ _[The Analytics Engineering Roundup](https://roundup.getdbt.com/p/developer-productivity-on-github)._ Dr. Eirini Kalliamvakou is a senior researcher at [GitHub Next](https://githubnext.com/). Eirini has built a career on studying software engineers, how to measure their productivity, and how developer experience impacts productivity. Recently, Eirini has been working on quantifying the impacts of GitHub Copilot. Does it help software engineers be more productive? Tristan and Eirini explore how to quantify developer productivity in the first place, and finally, whether Copilot‌ makes a difference. In the search for real business value, this research is a real bellwether for things to come. **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Google Podcasts](https://podcasts.google.com/feed/aHR0cHM6Ly9hbmFseXRpY3NlbmdpbmVlcmluZ3JvdW5kdXAubGlic3luLmNvbS9yc3M) - [Stitcher](https://www.stitcher.com/show/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways from this episode ### Can you give us a little bit of background on Copilot? I think Copilot came out ahead of ChatGPT, which means that it clearly predates the groundswell of attention that AI has gotten. **Eirini Kalliamvakou:** Fundamentally, Copilot is an AI-powered code completion tool. So as developers are in their developer environment, in their editor and they're typing, they're writing their code, they get suggestions from Copilot. Sometimes that means completing the line of code that they are writing. Sometimes it means that it gives them much longer snippets. Sometimes it means that they write a comment and then it suggests code. The concept of auto-completion was not new. The AI-powered completions and the way that Copilot suggests them, that's where the interface comes in. It doesn't have a special interface per se, but the experience of it appears as Italicized text, which we call ghost text, makes it very easy to just ignore it and keep typing if it's not something that you're interested in or it's not a suggestion that fits exactly what you were going to write. But it also makes it very helpful when you do accept it because it‌ gives you this boost of progress. ### Developer productivity has historically had a bit of a bad rap. Can you introduce us to what you believe is a good way to measure developer productivity? So you said a word earlier about developer productivity, uh, you said obsessed, and I think that definitely captures it. I think the tech industry is kind of obsessed, which sometimes seems unfair because I think productivity generally is something that other industries and other sectors are interested in. I don't find myself talking about the productivity of researchers, for example, or doctors and lawyers, but we talk a lot about developer productivity. What I think is best is if the metrics and the measurements and the models that we're using to track productivity, reflect reality. One of the problems with how traditionally we've been thinking about developer productivity is it does not match the reality of software engineers, because it comes from archaic concepts of productivity from more industrial settings, but now we're talking about more knowledge work. And those two don't match. It's an error of judgment when we always only go for things that we can measure or even more so things that we can measure easily. It's very easy to count lines of code or numbers of PRs or all of these things. And they do have their place, right? They show a pace of activity for the developer. But then the conclusions that we draw from that, we have to be careful. If we see a slower pace, is it the developer that is at fault? Is it the system that they're working with? Is it the process that development is happening in their team or in their organization that is the problem? Most of the time we don't do that diagnosis afterwards. I see it in my work a lot of the time where people ask for a lot of metrics and a pluralistic view of productivity, but they actually say, “Can you please narrow it down to this one metric that I can put on my dashboard and I can look at every day?” I think the correct thing is to have a view of productivity that matches developers’ reality, some of which is perceptual and some of it is observed. We need to have metrics and ways to measure that capture both of those. This is something that I have tried to practice whenever I'm called to do, or I'm interested in doing measurements of productivity. We know, and, and we definitely live and breathe that at GitHub too, that it's not just a single thing. So when the time came to do productivity evaluations for Copilot, yes, we looked at the acceleration of developers as one aspect of productivity, but we also looked at satisfaction. We looked at the state of flow. We looked at cognitive load. ### Can you talk a little bit about the SPACE framework? [The SPACE framework](https://github.blog/news-insights/research/research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness/#figure-summary-of-the-experiment-process-and-results) and the metrics that you mentioned are examples of metrics that you see. You could use the interesting parts and the impactful part about the SPACE framework is that it says you need to look at more than one thing. S stands for satisfaction and wellbeing. P is performance. A is Activity, which is usually when we think about automated measurements. C is Communication and Collaboration, and E is for Efficiency and Flow. So there's a lot more there than just Activity. And the Space Framework is a great reminder to look at more than just Activity. Don't just only count how much of things are happening or how fast. Think about some of these other dimensions. ### When you talk about developer experience or developer productivity to senior executives, CIOs, CDOs, et cetera, do they get it? Do they care? Do they consider this a strategic priority? It's definitely a very different conversation to have that with, as you say, a CTO, a CIO versus a head of engineering versus a developer. We're talking to completely different audiences and the conversation ends up very different. When I talk with executives about this, there is some reluctance. And there are a lot of questions. The first wave of questions is always, “But what is the developer experience?” And I think sometimes it comes with some preconceived notions that developer experience is all about. I don't know, like having perks at work and so on. ### Ping pong tables. The beers on a Friday and the ping pong tables and so on. At the very beginning, that's all people thought developer experience was. Now, fast forward a few years and a few studies, and we realize that there's a lot more to that. So the first wave of questions is always, what is it? And what we talk about is the developer experience, which is essentially how satisfied or hampered developers feel with the tools and the processes that they use, so that we start developing an understanding of what is it that boosts them. It starts with how satisfied people are with the tools and the processes that they're using today. We have narrowed it down to three main factors, and when I say we, I mean as a research community. One, it has to do with the flow states. I could be talking about the flow state for, for hours if, if someone lets me, but essentially is, you measure that as ‌satisfaction, which is very perceptual. With the amount of deep work that developers are able to do, the frequency of interruptions and surprise, surprise, whether they find their work engaging versus boring and repetitive or non-impactful. Then there's feedback loops, which is the time to get. And that's something that you can get from a perceptual angle. You can ask people approximately how much time does it take you or how slow or fast is a particular feedback loop. You can also add data to this because you can see in your systems for at least a few things, from the start of a flow to the end. You can also divide it into value-added work and not. Then there's cognitive load, which is how easy it is for people to understand and read code that already exists or artifacts that they're dealing with. It also depends on how easy it is to deploy code, not just understand it, but also be able to to work with it. And how intuitive are the tools and the processes that people have to work with daily. So these three pillars, the moment we start talking about those, they start making a lot more sense and become a lot more relatable. And then the next wave usually of questions is, why does that matter? Why, of all the things that I can use my finite resources for why should I put it towards improving developer experience and how do I go about it? And that's why [we did a study a few months ago that was specifically focused on measuring the tangible impact of introducing improvements](https://github.blog/news-insights/research/good-devex-increases-productivity/). ### And you showed meaningful results. What was it? 50 percent? It's a 50 percent productivity boost for developers to have deep work that is dedicated, like it's blocked off and it's a significant amount of deep work. There's so much piling evidence in terms of the state of flow and the state of deep work for developers, and how it's not only fundamental to their being able to make progress, which is a more traditional view of productivity, but it's also fundamental to them feeling good about what they're doing. We did find significant results. And I mean that both in the statistical sense and that we found significant results and impactful ways that improvements on developer experience can affect the bottom line for enterprises. That helps move the conversation a little bit further than before, because now we have evidence, we have data that we can show about what you can expect to get in return. --- --- title: "Understanding AI data engineering" description: "How AI is changing data engineering—and the groundwork you need to lay for it to be a success." url: "https://www.getdbt.com/blog/ai-data-engineering" date: "2024-09-07" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Understanding AI data engineering The work of data engineers has changed radically in the past 10 to 15 years. The advent of the cloud and cloud-based data warehouses such as [Amazon Redshift](https://aws.amazon.com/redshift/), the shift [from ETL to ELT](https://www.getdbt.com/blog/etl-vs-elt), and the birth of data lakes and data warehouses mean that data engineering teams can process and handle larger volumes of data faster than ever before. On the other hand, data engineers still struggle to meet the demand for large-volume, high-quality data. This pressure has only grown stronger with the emergence of Generative AI (GenAI) use cases, which require large and accurate data sets to produce useful results. The good news is that we can use GenAI itself to help meet this demand. GenAI is poised to disrupt how data engineers work by automating many routine and even complex analytics workflow tasks. Implemented correctly, engineers can use AI-assisted workflows to ship more data quality products, more quickly, and with higher quality. We’ll look at what AI data engineering is, how it can improve every stage of your data workflows, and the processes and systems you need to make it successful. ## What is AI data engineering? AI data engineering leverages GenAI technology to assist in the creation or modification of various assets associated with your data workflows. It combines the power of [Large Language Models (LLMs)](https://www.cloudflare.com/learning/ai/what-is-large-language-model/), AI models trained on massive amounts of data, with data from your existing pipelines—database schemas, data models, tests, documentation, metrics—to produce a first draft of artifacts based on natural language descriptions that your engineers can refine, test, and deploy. Raw data is never ready out-of-the-box for use in analytics or AI workloads. It requires [transformation](https://www.getdbt.com/blog/data-transformation), cleaning, testing, and documenting to mold it into a format suitable for asking questions and driving business decisions. Creating these data pipelines and the infrastructure that supports them is the heart of [data engineering](https://www.getdbt.com/blog/what-is-data-engineering). This work can frequently become a chokepoint for creating new production-ready data sets, as it takes time to create and test a high-quality data pipeline. AI data engineering cuts this workload by automating the creation of assets—often a tedious task—that constitute a data pipeline. It doesn’t replace a data engineer; it augments them. Think of AI data engineering less as a robot and more as a cyberskeleton or a mecha suit that enhances the powers of its wearer. Similar uses of GenAI in other areas of software engineering have yielded amazing results. For example, [Github found that 55% of engineers who used its Copilot feature](https://resources.github.com/learn/pathways/copilot/essentials/measuring-the-impact-of-github-copilot/) reported faster task completion, with 50% reporting faster time-to-merge their changes to ready them for deployment. ## How AI data engineering assists the analytics code process AI data engineering can provide a boost to data engineers, analytics engineers, analysts, and business users in five critical areas: - Coding - Testing - Documentation - Metric and semantic models - Discovering data Let’s look at each in detail. ### Coding Data transformations require selecting data from one or more sources and reshaping it into a format suitable for querying for a specific business use case. This involves writing code—usually in SQL or Python—that integrates data from multiple tables into fewer tables, as well as corrects any underlying issues - malformed fields, missing values, etc. As in software engineering, AI data engineering can provide a boost by generating the base SQL statements a data engineer needs for their transformations. Engineers can generate simple or complex SQL statements, complex regular expression patterns, and even perform bulk edits on existing SQL or Python code. This assistance can be particularly helpful for junior engineers. But it’s beneficial for even senior engineers who need to perform complex queries. By issuing instructions in natural language and allowing AI to write the code, engineers don't have to waste time looking up and reminding themselves of the peculiarities of SQL syntax. ### Testing Sadly, code doesn't always work the way we intend it to. Additionally, it may work in normal conditions but encounter issues when dealing with edge cases—for example, values outside of expected ranges, malformed values, etc. [Building tests](https://docs.getdbt.com/docs/build/data-tests) for data transformation code that make assertions about the resulting data provides confidence that our transformations work as expected under a variety of circumstances. [These tests can take several forms](https://www.getdbt.com/blog/adlc-test): - Unit tests (Validate small portions of data model logic) - Data tests (Ensure generated data is sound - e.g., all required fields have values, no value is malformed, etc.) - Integration tests (Test the entire project end-to-end) Testing is one of those areas in software engineering that gets the short shrift come crunch time. Everyone knows they _should_ be doing it, but the overhead involved means it sometimes gets left out in the rush to ship. AI data engineering can generate basic tasks for a new or revised data model, eliminating much of this upfront coding. That reduces the psychological barriers to creating an adequate test bed. It also frees engineers to focus on refining or adding elements to the tests that bring true value to the quality of each data set. ### Documentation [Documentation](https://docs.getdbt.com/docs/build/documentation) is another one of those assets that everyone knows they should write but sometimes don't. And that’s a shame because good documentation is critical for the discoverability and usability of datasets. Documentation tells downstream consumers where data comes from, how to use it, and how various calculations were derived. This makes data more usable and provides confidence in its validity and accuracy. If you use dbt for your data transformations, you already get some documentation for free in the form of automated [data lineage](https://www.getdbt.com/blog/guide-to-data-lineage) generation. This is useful for tracing the origin of data back to its source, which increases data confidence and also assists in troubleshooting issues. AI data engineering goes further, creating descriptions for your tables in their field based on their name, context, and similar data assets in your projects. This is extremely useful when you have hundreds of fields to document. GenAI-generated documentation can provide an initial first cut of descriptions for all tables and fields. Engineers can then check these into source control, where they and other team members can gradually improve the documentation over time. ### Metrics and semantic models Another area where AI can help is in defining consistent metrics. Providing global metrics available across your organization for key values—e.g., revenue—makes it easier for everyone to access key data without re-writing the wheel or introducing subtle errors. A [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) is a framework for metrics that defines a common representation of your data using standard business terminology, translating it from SQL or Python into common business language. Besides ensuring consistency, this democratizes access to data by making key values available to all data stakeholders. Defining a semantic layer requires tools for defining and exposing new metrics globally. [dbt Cloud’s Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) is one implementation of this concept that supports defining both metrics and semantic models that ship with your data transformations. With AI data engineering, you can not only generate these models automatically—you can ask the GenAI engine to recommend useful metrics based on your data transformation definitions. ### Discovering data Not all business stakeholders have an advanced or even basic knowledge of languages like SQL. That can prevent them from self-servicing answers to questions they have about data, or limit what they can do in visual BI reporting tools. AI data engineering can help here as well. Using AI, any stakeholder can query their data, not with a programming language, but with simple natural language expressions. The AI engine handles converting these into the required code under the hood. This further democratizes data and reduces the support requests that stakeholders make to the engineering team. ## Benefits and challenges of AI data engineering AI data engineering can markedly decrease the time required to create new data assets. By using natural language queries, engineers of all levels can spend less time looking up programming syntax and hunting down the source of minor formatting issues in their code. AI data engineering also breaks down barriers to implementing artifacts—such as tests and documentation—that are an essential part of data quality. GenAI can take on the grunt work involved in these tasks, leaving human engineers to focus on those parts where they can truly add value. Additionally, an AI data engineering solution that uses your organization’s data as input can produce more consistent output. A GenAI engine can be instructed, for example, to follow certain best practices when generating SQL code. The challenge is that AI isn’t a silver bullet. AI-generated code isn’t always accurate. Sometimes, in fact, it’s dead wrong. One developer, for example, found [only one LLM could successfully generate working code for a battery of tasks](https://www.zdnet.com/article/yikes-microsoft-copilot-failed-every-single-one-of-my-coding-tests/), such as creating a Wordpress plugin or finding a bug. Another study found that, while AI can help enforce best practices in some areas, [it might encourage engineers to skip following them in other areas](https://arxiv.org/abs/2211.03622), such as security. ## An AI data engineering assistant (with guardrails) An AI copilot can be a revolutionary addition to data engineering as part of an end-to-end, mature analytics workflow process, such as the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). Processes like the ADLC ensure alignment with business objectives and establish checkpoints to ensure code quality and conformance to best practices. dbt Cloud acts as your data control plane, providing a consistent approach to developing, testing, deploying, and documenting data in line with the ADLC, no matter where your data lives. Using dbt Cloud, you can: - Model your data using a combination of SQL and YAML code - Promote it to production with a rigorous testing and verification process supported by the platform - Monitor the performance and usage of your data [dbt Copilot](https://www.getdbt.com/blog/dbt-copilot-is-ga) integrates with every step of your data engineering workflow, using your own data—its relationships, metadata, and lineage—to automate routine tasks and implement routine tasks like testing, documentation, and SQL formatting that are essential for delivering high-quality data products. Besides generating artifacts for your data pipelines, dbt Copilot can enforce code consistency using a custom style guide. You can use Copilot out of the box with OpenAI or [bring your own OpenAI key](https://docs.getdbt.com/docs/cloud/enable-dbt-copilot#bringing-your-own-openai-api-key-byok). [Ask us for a demo](https://www.getdbt.com/contact) to learn more about how dbt Cloud with dbt Copilot can transform your data engineering practice. --- --- title: "Data leadership in the age of AI - Part 2" description: "In Part 2, AI leaders from dbt, M&T Bank, and Microsoft unpack multimodal data, governance, talent, and scaling AI responsibly." url: "https://www.getdbt.com/blog/data-leadership-ai-2" date: "2024-08-28" authors: ["Drew Banin"] categories: ["Insights"] --- # Data leadership in the age of AI - Part 2 Recently, I sat down for a chat with Andrew Foster, Chief Data Officer at M&T Bank, and Karthik Ravindran, who leads Data Governance and Enterprise Data at Microsoft. [In the first part of our talk](https://www.getdbt.com/blog/data-leadership-ai), we focused on the importance of a human-centered approach to AI. In the second part, we dove deep into questions such as how to handle multi-modal data, challenges with AI, and how to leap the hurdles that prevent us from getting AI apps into production. ## Handling multi-modal data **Q**: We’ve focused on structured data quality for 20 or 30 years. And now GenAI shows up and suddenly there’s a renewed focus on unstructured data. How are your organizations dealing with the new multi-modal data model that AI is bringing to the fore? **A**: There are many businesses in the world powered by unstructured data. Think of banks and the paperwork that goes into securing a mortgage. Many other businesses run almost solely off of structured data. And then you have semi-structured data — e.g., a table or a form in the middle of a PDF. Being a SQL person myself, I tend to think of the world in structured data sets. My implicit desire with unstructured data is to structure it—find features you want to extract from emails or PDFs and put them into a well-understood schema with defined constraints. Once that’s done, you can query, aggregate, represent, and operationalize that data. This process will likely involve GenAI, but it'll also require the more deterministic or classic Machine Language tooling we’ve developed, which still has a ton of utility. **Andrew Foster**: There are two words here I think about. One is scale. The sheer volume and linear growth of data means you can’t just scale up with people. The other is tiering. There’s risk tiering—the risks of the model you’re using, like Karthik discussed, and how you’re mitigating those risks. But there’s also tiering of the data itself. Do you have an intended use for it? What’s your quality bar for it? Are you keeping a human in the loop or fully automating your decision processes? You can use these two criteria—scaling and tiering—to look at information sets across an organization and size them up based on predicted use cases. That gives you a more flexible approach to how you process data than a “one size fits all” rule. That approach will never work. It won’t scale. ## AI challenges (and successes) **Q**: Can you share an AI challenge with us? Any success stories? **A**: A couple of months ago, at Snowflake Summit, [we unveiled Ask dbt](https://www.getdbt.com/blog/introducing-dbt-for-snowflake). It’s a chatbot that plugs into your data warehouse so you can ask questions about your data using natural language. The problem we ran into was that business metrics are defined precisely. However, people often don’t know the exact wording to solicit the answer that they care about. The solution came about from another project that we’ve been working on, the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl), for the past three years or so. With a Semantic Layer, we can define standardized metrics and also tag different business-centric metadata to them. We add things such as, what is the standard definition of the metric? What are the descriptions of its dimensions? With this information, we can create the opportunity for an LLM interface into business metrics and their underlying data. This is really compelling. It lets users intuitively answer questions like, “Who’s the owner of the Salesforce account for Acme Company?” We think this is the future of chat, where your data experiences flow through something like a Semantic Layer. **Karthik Ravindran**: As physical data estates explode, the shape of the data continues to evolve. One really powerful technique is focusing [on what Gartner calls Active Metadata](https://www.forbes.com/councils/forbestechcouncil/2022/06/23/active-metadata-what-it-is-and-why-it-matters/). If we can extract technical, business, and functional metadata agnostic from the modality of our physical data estates, we can elevate data management and governance to this metadata layer. We can then set policies for data management and scale our data governance across the underlying physical data estates. Getting there was more of a human problem to solve than a technical problem, though. We got good at collecting technical metadata. And we’d generate graphs that were nice to stare at. But we didn’t have a way to connect those to business outcomes. So what we’ve developed is a split between my team, which provides the systems and solutions, and our stakeholders who take accountability for the logical data estate and business metadata. We can’t do that for them—we don’t understand the details of every single data domain. We need stakeholders to jump in and bring that expertise. The outcome has been profound. Because now, we've got a sharp sense of accountability for the domain teams and experts. **** ## Skills and talent needed for AI **Q**: What skills and talent are most critical for data teams working with AI? How are you addressing ‌talent gaps? **A**: We’re currently seeing software engineers digging in and building these LLM-based systems. ‌I hope this doesn’t come across as offensive or sound like a secret, but not many software engineers are very good at data. Some are actually shockingly bad at data. Software engineers will need to learn a lot more about how data works. By that I mean, not just the mechanics of where it lives and how to access it. I also mean how to think about data. I think creativity is an important skill and orientation for people to have when dealing with data. Most of the outsized benefits we’ve discussed today will accrue to those who find a shorter path to existing business outcomes. Slapping an LLM onto an existing workflow may get you budget for the next quarter. But it won’t result in significant efficiency gains. **Andrew**: One thing I generally look for in my team is high emotional intelligence. My team sits in the middle of a large organization with a lot of business engagement. We have to work with traditional teams of modelers—which is highly technical work in nature—as well as with various technology and business partners. And we have to find a way to bring everyone’s expertise into this ever-increasing and important space. No one operates in a silo. When you hire, it’s less about how you’re hiring for an individual and more about how you’re structuring your teams. You have to find a good balance between hiring a generalist and hiring a specialist. **Karthik**: I think adaptive leadership and change management are going to be crucial. And not just from executives and senior leadership—it’s something we need from every team member who interacts with data. Andrew made great points earlier about data literacy—that’s so critical. Tian Kai Feng published a great book called [_Humanizing Data Strategy_](https://www.amazon.com/Humanizing-Data-Strategy-Leading-Heart-ebook/dp/B0D9TSVYR3). It shows how the crux of our AI data transformation journeys are very human-centered. Tech isn’t the solution—it’s the enabler. Change management, adaptive leadership, the emotional intelligence that Andrew spoke of—those are super important things to get right. Engineers and data scientists can’t just look at a problem and say, that’s not my problem. Heck no—it **is** your problem. You can’t just assume you can build something, toss it over the fence, and it’ll work like magic. You need to get in the weeds and lead with the outcome, with the business benefit. ## Conclusion: Key takeaways **Q**: What are two or three takeaways people should take from this discussion? **A**: Like I said at the beginning, five percent of AI projects are in production today. There are a lot of interesting things we can do better and at scale. Step one is getting started. For example, with data quality governance, it’s important not to do anything outside of the risk profile of your organization. But also, don’t be afraid to start somewhere. **Karthik**: Pay attention to keeping humans in the loop. Despite all of the innovation coming out of AI, humans in the loop and “Humans + AI” is still the name of the game. It’s not “humans minus AI.” I encourage people to think about how to institutionalize this, both as a strategy and in practice. **Andrew**: Challenge your assumptions. AI isn’t incredibly different than what preceded it. But aspects of it are different enough that they could impact every part of how your organizations are run. Look at the components that interact with AI. Because some of them will need to change—and you need to recognize that early. --- --- title: "Transforming data with AI: use cases, examples, and challenges" description: "How to use AI for faster and more accurate data transformation—and how dbt Cloud can help." url: "https://www.getdbt.com/blog/ai-data-transformation" date: "2024-08-28" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Transforming data with AI: use cases, examples, and challenges Generative AI (GenAI) is reshaping how teams work—automating not just repetitive tasks, but also increasingly complex workflows. For data teams, it’s poised to transform one of the most time-intensive parts of the job: [data transformations](https://www.getdbt.com/blog/data-transformation-examples). Transformations are the backbone of data quality. They convert raw inputs into clean, trusted data that powers reporting, analytics, and decision-making. But building that transformation code takes time—writing SQL, testing logic, documenting models, and optimizing performance all require significant effort. Used thoughtfully, GenAI can streamline this work. As part of a broader [analytics development lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle), it helps teams produce high-quality transformations faster and with less manual lift. In this post, we’ll explore how AI fits into the modern analytics workflow—and how [dbt Copilot](https://www.getdbt.com/product/dbt-copilot) brings context-aware AI into the hands of data developers. ### Overview: AI's role in the data transformation process When used with the right guardrails, AI can accelerate transformation workflows across your entire data team—from engineers and analysts to business stakeholders. Here’s what an AI-powered transformation assistant, like [dbt Copilot](https://www.getdbt.com/product/dbt-copilot), can help you do: - Generate and optimize transformation code - Create data quality tests - Write documentation for models - Build semantic models and define metrics In the sections that follow, we’ll explore how each of these tasks works in practice—and what it means for your analytics workflow. **** ## What makes data transformation challenging? Raw data is rarely usable out of the box. It’s often spread across multiple systems and riddled with issues like malformed fields, missing values, duplicate entries, and inconsistent formats. Data transformation addresses these problems by extracting data from various sources, loading it into a central destination, and reshaping it into a format that data consumers can query and trust. This [ELT (Extract, Load, Transform)](https://www.getdbt.com/blog/extract-load-transform) process resolves common issues—unclear column names, incorrect data types, mismatched table relationships, overly granular timestamps—and prepares data to support diverse analytical use cases. The full transformation workflow typically runs on a schedule or on demand, continuously processing new data as it becomes available. Every data-driven organization depends on reliable, high-quality transformation processes. But getting there isn’t easy: - Writing and debugging transformation logic takes time—even for experienced engineers. - Ensuring accuracy requires robust test coverage, which adds another layer of work. - Even correct transformations can be inefficient, requiring performance tuning to scale. - Most transformation code is written in SQL (or Python), limiting contributions to technical users. - Without clear documentation, even well-modeled datasets may go unused because business teams don’t trust or understand them. These challenges make transformation a key bottleneck and a strong candidate for AI support. ## How AI can help—and how to introduce it The good news: AI can help eliminate much of the manual effort involved in building data transformation pipelines. [Large language models (LLMs)](https://aws.amazon.com/what-is/large-language-model/)—like [GPT](https://openai.com/index/gpt-4/) and [Claude](https://www.anthropic.com/claude)—are trained on vast datasets and have proven adept at generating base code that experienced engineers can refine, test, and deploy faster than writing from scratch. LLM-powered copilots are already boosting productivity across software teams. When [Accenture integrated GitHub Copilot into their workflows](https://github.blog/news-insights/research/research-quantifying-github-copilots-impact-in-the-enterprise-with-accenture/), developer satisfaction rose 90%, and 67% of participants reported using it five days a week. A context-aware copilot can bring similar value to data engineering. That’s why we built [dbt Copilot](https://www.getdbt.com/blog/dbt-copilot-is-ga)—to integrate directly into your analytics workflows and support every step of the process. That said, AI copilots aren’t a silver bullet. ### The Analytics Development Lifecycle (ADLC) AI copilots work best when embedded within a mature, collaborative analytics process. At dbt Labs, we call this the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle)—a framework modeled after the [Software Development Lifecycle (SDLC)](https://www.browserstack.com/guide/devops-lifecycle) that helps teams build and manage analytics code at scale with speed and quality. The ADLC includes structured processes and checkpoints to ensure all data transformations shipped to production are trustworthy and aligned with business needs: - **On the business side**, it ensures analytics requirements are clearly defined, mapped to KPIs, and scoped to the smallest impactful unit—so teams can ship well-tested, high-value changes quickly. - **On the technical side**, it ensures every data transformation is defined as code, version-controlled, peer-reviewed, tested before deployment, and monitored in production. Adding an AI copilot outside of this framework won’t improve code quality—and might even increase risk. To deliver real value, AI needs to be integrated into a structured, collaborative workflow like the ADLC. ### Review and testing GenAI Generative AI can accelerate development—but it’s not infallible. Every piece of AI-generated code must be reviewed and tested thoroughly before it reaches production. That’s why dbt Copilot is tightly integrated into [dbt](https://www.getdbt.com/product/dbt). It operates within the structure of the Analytics Development Lifecycle with built-in features like [version control](https://docs.getdbt.com/docs/collaborate/git/version-control-basics), [automated testing](https://docs.getdbt.com/docs/build/data-tests), [deployment pipelines](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud), automatically generated [documentation](https://docs.getdbt.com/docs/build/documentation), [data lineage](https://www.getdbt.com/blog/guide-to-data-lineage), and [data discovery and management](https://www.getdbt.com/product/dbt-catalog). This integration creates strong guardrails around AI-assisted development, allowing teams to boost productivity without compromising code quality, security, or performance. Because Copilot works directly within your dbt project, it understands your models, relationships, metadata, and lineage—generating code that’s tailored to your team’s unique context. The result: refined, governed transformations that support high-quality analytics and AI workloads. ## Generate and optimize data transformation code Rather than write SQL code from scratch, data producers can create inline SQL using a natural language description. This can shorten the time required to write code, eliminating common inaccuracies. It also ensures that generated code follows your organization’s naming conventions and best practices. ### Example: refining code with dbt Copilot dbt Copilot helps users at every experience level write better transformation code. New contributors can use natural language to generate working SQL, while seasoned engineers can rely on Copilot to refine complex logic or apply bulk edits across a project. By expanding who can confidently contribute to data pipelines, Copilot supports broader [data democratization](https://www.getdbt.com/blog/managing-data-democratization) across the organization. Working directly within the dbt Studio IDE, Copilot helps you: - Write advanced SQL transformations with ease - Apply project-wide edits or refactorings - Generate complex regex patterns - Enforce a custom SQL style guide ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0947b31f8a9387bc92a12ed8583844d0ba9ec502-1600x878.png) This helps reduce review time, increase consistency, and maintain high-quality code across your entire analytics project. ## Generate data tests [Testing is foundational](https://www.getdbt.com/blog/adlc-test) to reliable analytics engineering. That’s why [dbt includes the ability to write tests alongside your data transformation code](https://docs.getdbt.com/docs/build/data-tests)—and run them throughout your workflow. Tests can be triggered during development, on [pull request checks](https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/proposing-changes-to-your-work-with-pull-requests/about-pull-requests), and before changes are promoted to production. This ensures code is accurate and safe before reaching end users. ### Example: creating a testbed With dbt Copilot, you can automatically generate a suite of tests based on your dbt models and their schema relationships. Copilot adds these tests directly to your project, enabling validation at every stage of the Analytics Development Lifecycle. By automating test generation, Copilot helps you catch issues earlier, boost confidence in your models, and ship changes faster—with fewer surprises downstream. ## Generate documentation Even the best data transformation code falls short if data consumers can’t understand or trust the outputs. Documentation bridges this gap—providing shared context, improving discoverability, and increasing confidence in how data is used. Well-documented models help analysts, stakeholders, and new team members quickly understand where data comes from and how key fields are calculated. But for large or legacy projects, creating that documentation manually can be a daunting lift. ### Example: generating docs with SQL logic, past queries, and metadata dbt Copilot can help scale documentation by analyzing SQL logic, historical query patterns, and model metadata to generate descriptions automatically. It surfaces plain-language explanations for complex logic or obscure field names—giving teams a head start. From there, users can refine and expand documentation organically over time as they work with the data. It’s a faster, more sustainable path to building a discoverable, trustworthy data foundation. ## Generate semantic models for metrics Even with centralized data models, teams can still end up with inconsistent definitions of key metrics. One team’s “revenue” might differ from another’s based on filters, timeframes, or underlying logic. A [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) solves this by defining metrics in a consistent, centrally governed way—outside of individual BI tools. This ensures everyone in the organization is speaking the same data language, with one source of truth for business-critical metrics. ### Example: generating models with dbt Semantic Layer With the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) you can define metrics in code using familiar, model-like syntax. dbt Copilot can generate this scaffolding automatically—and even recommend common metrics based on your data models. By shifting metric definitions upstream into your modeling layer, dbt helps eliminate inconsistencies and builds trust in every dashboard and report. **** ## Accelerate your data transformations with AI today Used wisely, AI-enhanced data transformation workflows can dramatically reduce the time spent writing, documenting, and reviewing code. That frees your team to take on higher-impact projects and deliver trusted, analysis-ready data faster. With dbt as the control plane for your analytics workflows, you can build, test, and ship high-quality transformations at scale. Want to see how dbt Copilot fits in? [Book a demo](https://www.getdbt.com/contact) and explore how AI can supercharge your data workflows. ## FAQs about AI for data transformation **What is data transformation in the context of analytics?** Data transformation converts raw data into structured, analysis-ready formats. It typically follows an [ELT (Extract, Load, Transform) process](https://www.getdbt.com/blog/extract-load-transform)—pulling data from multiple sources, loading it into a warehouse, and applying transformations to resolve issues like missing values, unclear column names, or inconsistent formats. These transformations make data trustworthy and usable for downstream teams. **How can AI help with data transformation challenges?** AI, especially Large Language Models (LLMs), reduces the manual work involved in data transformations. It can generate base SQL, optimize existing code, create tests, produce documentation, and build semantic models. Tools like dbt Copilot enable users to describe transformations in natural language and receive code that’s production-ready—speeding up development without sacrificing quality. **What is the Analytics Development Lifecycle (ADLC)?** The [Analytics Development Lifecycle, or ADLC](https://www.getdbt.com/resources/the-analytics-development-lifecycle), is a structured process for developing, reviewing, and deploying analytics code at scale. Inspired by the Software Development Lifecycle, it includes best practices like version control, scoped changes, peer review, testing, and monitoring. AI copilots like dbt Copilot are most effective when integrated into this lifecycle, where guardrails support both speed and reliability. ![Infinity loop diagram illustrating the Analytics Development Lifecycle (ADLC), showing key stages from develop and test to deploy, plan, analyze, operate, observe, and discover.](https://cdn.sanity.io/images/wl0ndo6t/main/034af006313d8a5c5c1149a039fe49a270e1a863-1198x672.png) **How does AI assist with documentation in data projects?** AI can auto-generate documentation by analyzing SQL logic, historical queries, and model metadata. This provides much-needed context for otherwise opaque models—clarifying definitions, tracing lineage, and fostering trust. In legacy projects with hundreds of models, AI significantly cuts down the time needed to document and maintain clarity across teams. **How does dbt Copilot improve data transformation efficiency?** dbt Copilot accelerates development by generating SQL code, data tests, and documentation directly within dbt. It applies best practices, enforces naming conventions, and reduces the need for manual reviews. With less time spent on repetitive tasks, teams can focus on delivering business value through analytics. **What are the key features of dbt Copilot for data engineers?** dbt Copilot helps engineers generate and refine SQL code using natural language. It auto-creates test coverage, generates documentation, and scaffolds metrics for semantic modeling. These capabilities streamline workflows and make it easier to scale high-quality data development across teams. **How does dbt Copilot fit into the Analytics Development Lifecycle?** dbt Copilot is built into dbt, which supports version control, testing, automated deployments, and documentation. This integration ensures AI-generated code fits within existing workflows, undergoes proper review, and aligns with governance standards—helping teams move faster without increasing risk. **How does dbt Copilot help non-technical users contribute to data projects?** dbt Copilot empowers users with limited SQL experience to generate and modify transformation code using plain language. This widens the pool of contributors to include analysts, product managers, and other business stakeholders—unlocking collaboration and promoting data literacy across teams. **What security considerations should teams be aware of when using AI for data transformation?** AI-generated code should always be governed by existing review, testing, and deployment practices. Tools like dbt Copilot are most effective when integrated with secure version control, access policies, and QA checks—ensuring AI doesn’t bypass your organization’s data governance and security standards. --- --- title: "What is enterprise data governance?" description: "Data governance ensures integrity, security, and quality, helping enterprises turn data from a liability to a valuable asset." url: "https://www.getdbt.com/blog/what-is-enterprise-data-governance" date: "2024-08-27" authors: ["Daniel Poppy"] categories: ["Learn"] --- # What is enterprise data governance? By now, it’s common knowledge that data is a vital competitive asset for almost every company. There is, however, a flip side to this undeniable truth: data can also become a liability. Gathering, storing, and using data isn’t limited to experts on data or technical teams. It flows through the entire organization. There are countless touchpoints where people in sales, product, marketing, and other diverse teams make decisions about and take action on the information your company collects. Without cohesive structure and guidance for handling data, stakeholders may unintentionally corrupt the quality or security of your company’s data. That can spark a cascade of problems that lead to lasting damage. You can employ a data governance framework to provide clear and consistent policies and standards to ensure the integrity, security, accessibility, and overall quality of your data assets. In this article, we’ll cover what that looks like in an enterprise organization. ## What is Enterprise Data Governance? The [Data Governance Institute defines data governance](https://datagovernance.com/defining-data-governance/) as _a system of decision rights and accountabilities for information-related processes, executed according to agreed-upon models, which describe who can take what actions with what information, and when, under what circumstances, and using what methods._ That’s quite a mouthful. Exactly what this looks like will vary from org to org, depending on your specific product and priorities. However, any data governance program has the same three goals: - Establish and align rules for the entire company around how data is collected, stored, and used (and also when and how it is to be discarded) - Monitoring the global data estate to make sure governance standards and policies are being respected - Resolving any issues or conflicts affecting data in your company while supporting your data end users These are good tenets to follow for businesses of any size. But they’re critical for enterprise organizations. Smaller orgs and startups may be able to muddle through with informal data governance practices, addressing data-related cross-functional activities in meetings and emails. The larger the company, the more complex and distributed the data assets. This is where the four pillars of enterprise data governance come in. ## The four pillars of enterprise data governance Data governance plays a crucial role in ensuring the effective management, quality, and security of your data assets. To achieve this, an effective data governance program needs a solid foundation based on the four pillars of data governance: data quality, data stewardship, data protection and compliance, and data management. ### Data quality Ensuring the accuracy, completeness, and consistency of all data assets across your entire organization, no matter how large or distributed your company may be. ### Data stewardship Quality doesn’t happen by itself. Another cornerstone of data governance is creating and assigning clear and dedicated roles and/or responsibilities for **the people in charge of managing, monitoring, and ensuring data within your organization — i.e., data stewards. ** Data stewards serve as the front line for a governance program because they are collectively in charge of defining and documenting your company’s data assets, ensuring data quality, and promoting effective data sharing and usage by every shareholder across the organization. They also act as liaisons between different teams to resolve any data issues that arise and are responsible for ensuring that governance policies and procedures are implemented correctly. ### Data protection and compliance The third pillar of data governance focuses on protecting data from unauthorized access, ensuring compliance with privacy laws and regulations, and managing user access and authentication. There are three main components here: - Security measures that prevent unauthorized access to your company’s data overall. - Privacy measures to protect the privacy of sensitive and personal data within your company. - Processes and policies to ensure compliance with all applicable regulations and contractual requirements. ### Data management This pillar encompasses the processes and procedures for storing, accessing, and manipulating data effectively. It includes metadata management, data lifecycle management, and data integration (i.e., how data is structured, stored, and linked across different systems and databases in your organization). ## When does a company need enterprise data governance? Enterprise data governance becomes necessary when your organization becomes large enough that casual, arms-length management can no longer be counted on for controlling data-related activities across the entire company. As your company grows, so does the number and complexity of your data systems. Without some sort of formal data governance framework, your data estate quickly [becomes siloed](https://www.getdbt.com/blog/what-are-the-four-principles-of-data-mesh) as teams naturally diverge in their priorities and choices made around data. It quickly becomes impossible to gain informed, higher-level horizontal visibility around the masses of data pooled and handled around your organization. Likewise, many (if not most) enterprise-level companies eventually become liable to regulatory compliance and/or contractual requirements that call for a formal data program. For example, 6.3 billion people—or 79.3% of the world's population—is covered by one or more [national data privacy regulations](https://iapp.org/news/a/identifying-global-privacy-laws-relevant-dpas) like the EU’s GDPR or China’s PIPL. Certain heavily regulated industries, like finance and healthcare, are subject to even more laws like [FINRA](https://www.finra.org/) (the Financial Industry Regulatory Authority) and [HIPAA](https://www.hhs.gov/hipaa/index.html) (the Health Insurance Portability and Accountability Act). ## Enterprise data governance frameworks Enterprise data governance frameworks are platforms designed to streamline and automate the complex and multifaceted process of managing, organizing, and protecting data across a large, and even globally distributed, enterprise organization. It’s certainly possible to draw up your own data governance framework from scratch. There are many governance program templates out there, both open source and for-profit, ranging from the nonprofit Data Governance Institute to elite consulting companies like McKinsey and Gartner. The “why build when you can buy?” rule absolutely applies to data governance frameworks. Starting with a pre-built framework based on tried-and-true principles that you can adjust to your organization cuts down the time required to design and roll out a new program. Data governance tools can also cut down the time required to implement a data governance framework. With tools like [dbt Cloud](https://www.getdbt.com/blog/dbt-explained), you can create a data control plane to manage your data uniformly in a standardized framework. That means everyone in the organization—whether a data engineer, business analyst, or executive leader—can move fast with trusted data in a scalable, cost-effective way. An enterprise-quality framework standardizes your teams on the terminology and concepts most important to your company while it builds collaborative data bridges across the full organization. Business, technical, and compliance stakeholders can now communicate easily with each other, exchanging data-driven information and ideas. In short, an enterprise data governance framework empowers every person in your company to extract value from your data assets while managing cost and complexity. ## Choosing the right enterprise data governance framework There are many data governance framework tools out there. Critical components of a mature, enterprise-ready platform include data stewardship, quality control, cataloging, lineage tracking, security, compliance, and data visualization. Not every solution, however, offers all of these in a single platform that can help you understand your operations, improve your performance, and achieve your goals. Here’s what to look for in a scalable and enterprise-ready platform: - It manages data complexity in a way that’s modular, scalable, repeatable, and governed — directly inside of your data platform instead of scattered across all the different business intelligence and technical platforms used by all the different teams and departments within your organization. - It’s vendor-agnostic. It can integrate with the major data cloud platforms and data tools that are important to you. It makes it straightforward to connect your data directly to Snowflake, Databricks, BigQuery, and all other leading data cloud platforms. - It utilizes centralized, reusable models, fosters collaboration, reduces duplication, and ensures consistent data definitions across teams. - It has robust audit logging and access control features to help safeguard data integrity. - It allows your data teams to build, test, and deploy analytics code using software development best practices (like portability, CI/CD, observability, documentation, etc.) to create production-grade analytics pipelines that scale along with your workloads. - It delivers accessible and easy-to-understand data models that can be delivered into your BI tools, LLMs, and APIs so that stakeholders have the accurate data they need and when and where they need it. The [[[dbt framework](https://www.getdbt.com/product/dbt-cloud)](https://www.getdbt.com/product/dbt-cloud)](https://www.getdbt.com/) is a scalable and enterprise-ready platform that fulfills all of these requirements plus many more. dbt can help your company standardize data transformation processes, increase data quality and transparency through lineage tracking, and automate documentation while providing horizontal visibility across all your company’s data assets. The right data governance framework makes sure that you can be confident in your data—and everyone in your company, no matter what their role, can use it for making accurate and informed decisions. Your enterprise gets enhanced data quality and consistency, reduced data management costs, and accelerated insights from trusted data. When your enterprise data governance framework automates making your data flows traceable and your data-related processes transparent, you can focus on optimizing your operations, improving performance, and achieving your goals—while simultaneously minimizing any downside data security and privacy risks. --- --- title: "Data leadership in the age of AI" description: "Explore how industry leaders are tackling the challenges of AI adoption, from data quality to human-centric approaches." url: "https://www.getdbt.com/blog/data-leadership-ai" date: "2024-08-26" authors: ["Drew Banin"] categories: ["Insights"] --- # Data leadership in the age of AI According to a [2024 CDOIQ conference report](https://cdoiq2024.org/wp-content/uploads/2024/07/2024-CDOIQ-Pre-Symposium-Proceedings-as-of-July-11-2024.pdf), only 6% of planned AI applications have made it into production. It’s a sign that many companies are still struggling with fundamental issues around Gen AI. Successful Gen AI applications require a strong foundation that enables high data quality and strong governance. Fortunately, we already have many of the tools and processes to enable this. Recently, I participated in a [panel for CDO Magazine](https://us02web.zoom.us/webinar/register/WN_uFWpYuPPSFWs14OZV9LDKA?utm_content=303038624&utm_medium=social&utm_source=linkedin&hss_channel=lcp-40830869#/registration) with Andrew Foster, Chief Data Officer of M&T Bank, and Karthik Ravindran, who leads Data Governance and Enterprise Data at Microsoft. They asked us how our businesses are crafting our approach to AI, including our approaches to governance, skills and team structure, and driving business value. The first part of our discussion focused on our respective company’s overall approach and strategy. Here’s what I said about how dbt Labs is taking a human-centric approach to AI adoption while keeping business values as well as compliance front-and-center. ‌I’ve also included (my own) summaries of Andrew and Karthik’s insightful takes on how their companies are facing the same challenges. ## Human-centric AI **Q**: Some of the top challenges companies have cited with AI are data availability and data quality, business outcomes, culture and adoption, and change management. Where's your organization ‌currently with rolling out AI-powered solutions? **A**: I spoke to a Chief Data Officer [CDO] recently who called himself the CSNO, or Chief Saying No Officer. When someone told him they needed Gen AI, he’d respond, “I don’t think you do - but tell me what business problem you want to solve with it?” Finding the right applications for a new technology is key. At dbt Labs, we think of Large Language Models (LLMs) as great co-pilots for development. If they can put context front-and-center for developers, that’s helpful in tightening development loops. **Andrew Foster**: Every vendor is embedding Gen AI into their capabilities—whether the tools need it or not. For our part, we have to make sure we’re approaching it with the right discipline, taking a “safety first” approach. So our focus is boosting internal productivity and vetting tools through internal testing, incorporating all the existing controls we have in place. **Karthik Ravindran**: I think at times that, as technologists, we lose sight of the human component. With Gen AI, we need to really resonate the story of the why. Why is AI going to help humans be more impactful and scale better? What change management processes do organizations need to embrace the shift that’ll be foundational to AI success? ‌We can’t ignore the human component if we want to transcend the hype and get to true value outcomes. ## AI use cases, known and unknown **Q**: What are some of the most compelling use cases for AI today? **A**: I try and take a longer-term view, thinking about what the next five years or so will look like. Obviously, we don’t know what’s going to happen with tomorrow’s OpenAI models, let alone five years from now. Take the original iPhone apps—they felt more like web pages. As the industry evolved and people became more comfortable with them, the apps did too. We see a similar thing with Gen AI apps now. They’ve started as chat interfaces to LLMs—which is a natural progression. But I think, over the next few years, we’ll see different modes of interaction. **** Agents will be a part of that. But I think we’ll also see the use cases we discuss today obviated by new approaches that free up humans to do more high-leverage work. I also expect we’ll see something like interrogating dashboards with natural language or even bypassing parts of the entire BI workflow. **Andrew**: When email came along, lots of people got fired. They got fired because they didn’t know what to say or not to say in corporate email. You saw something similar with social media. People had to ingrain newly learned behaviors. I’m less concerned with how a tool works and whether you should adopt it. I’m more concerned about how we ensure people understand and use the tool in ways that are additive—to themselves and the organization—instead of detrimental. As an example, say you’re using something like Microsoft Teams and recording a call. It’s recorded and summarized by AI. Are you comfortable with what everyone in the company might say on that call? What’s your retention period for that data? This is less a question about picking the perfect tools and more about building an organization that’s ready for an era where these new capabilities will exist everywhere. **Karthik**: I think we’re all familiar with the use cases for Machine Learning and traditional AI. Generative AI brings about a bunch of new opportunities. At Microsoft, we think of it as a layer cake. At the bottom, you have things like Q&A bots and natural language agents that can provide basic information and assistance. The next layer would be productivity use cases like content summarization or content tuning. And then from there, it goes further northy—e.g., advanced use cases around bootstrapping content, and then agents that can help take over and automate more of our manual tasks. As you go further up these layers, however, the level of change management and risk also rises. At Microsoft, we take this framework and use it to make conscious decisions about the maturity model of an overall team. We use this to create a graduated framework to help guide us in how to introduce people to these opportunities, letting humans turn the knob in terms of how sophisticated they want to get. ## Balancing AI with data accuracy **Q**: How should organizations balance the rapid pace of AI adoption with the need to ensure data accuracy and integrity? And what are some of the techniques companies can use to build trust in AI-powered insights? **A**: I don’t think we need to reinvent the wheel here. We know that if you put garbage into an AI or ML system, you’ll get garbage out. We also know how to govern data and manage access to different classes of data for different groups. That’s a lot of what we build dbt for. It enables you to model your data, making sure it’s in the right format for analysis or operational use cases, like feeding it to an LLM. You can test data to ensure it conforms to your requirements, and even detect anomalies and issue alerts before they impact downstream data consumers. That requires, of course, understanding what level of data quality is “good.” Data is never perfect. But if someone can aggregate and zoom in on the data and say, yes, this is useful, this maps to my understanding of how the business works, then that builds trust in the data. Data quality also requires gathering feedback and staying attuned to the needs of our users. It’s hard for the data team sometimes to understand if a given set of data looks off. But it’s easy for members of our field teams to look and say, “Hey this user is on 10 different accounts and they’re the admin everywhere—what’s going on?” There’s a smell test where the people closest to the data can tell if it’s suspect. All this is to say that I don’t think we need to reinvent the wheel here. We need to take these tried-and-true principles and apply them in this new domain. **Andrew**: I see data quality as an opportunity to strengthen the effectiveness of my team and our federated roles within the organization. Typically with data quality, a technologist and an analyst sit in a room somewhere and hammer out some rules. It’s not particularly efficient. And it doesn’t scale. With AI, we can use things like anomaly detection to compare data across time and detect things such as a spike or reduction in volume. I see that as a massive enabler for ongoing day-to-day improvements, particularly because it still keeps humans in the loop. It can be self-learning, taking the inputs from humans, and learning how to point the human decision-maker to anomalous activity in a more effective manner. Last year, there was a very technology-centric, “shiny new toy” view of Generative AI. Now, based on conversations I’m having, companies are looking more at how adoption is led through human behavior. **Karthik**: We faced this challenge at Microsoft, where we had to traverse quality for our internal data estate. At first, we did what Andrew said, setting up a bunch of rules, and dashboards with green/yellow/red statuses. It was an interesting scalability challenge. But what we lacked was how to tie those dashboard statuses to business outcomes and why teams should care. We realized it’s hard to excite your stakeholders when you cast everything through a purely technical IT lens. So instead, we changed our focus on the problem to business outcomes and aligned them to OKRs. This eventually became the mainstream way of approaching data quality—top-down versus bottom-up. Generative AI brings a whole set of different quality dimensions—bias detection, ethical compliance, infringement protection, etc. But at the end of the day, it’s all anchored in data. And not just in data coming in, but in data going _out_. You can’t just unleash a foundation model in its raw form. There are tougher processes to go through first. Are you doing continuous training? Are you domain-contextualizing and training the model with inputs that are specific to your context? Are you fine-tuning? Are you prompt engineering? All of this eventually accrues to the quality of the user’s experiences and the output the models can generate. There are a lot of thoughtful investments needed on both sides of the fence. **** ## Conclusion The use cases for AI are evolving and will continue to change. However, some basic principles will remain unchanged. Data quality, strong governance, and a focus on people and processes over tools will continue to be the firm bedrock upon which production-ready AI applications are built. [In the second part of our discussion](https://www.getdbt.com/blog/data-leadership-ai-2), we touch on the more practical aspects of managing AI initiatives, such as dealing with multi-modal data and measuring success. --- --- title: "Product spotlight: Four new dbt Cloud features you need to know about" description: "Learn more about the four dbt Cloud features highlighted in our Product Spotlight Series." url: "https://www.getdbt.com/blog/dbt-cloud-product-spotlight" date: "2024-08-23" authors: ["Sara Gawlinski"] categories: ["Product"] --- # Product spotlight: Four new dbt Cloud features you need to know about We have been shipping a ton of new features in dbt Cloud. So, we decided to kick off a new Product Spotlight Series to show how they can improve your analytics workflows. We'll highlight column-level lineage and lineage lenses, our new chatbot Ask dbt, unit tests, and job chaining for easier pipeline orchestration...complete with demo videos, which you'll find below 👇 ## Column-level lineage and lineage lenses First up is some cool capabilities in [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) to help you navigate, understand, and troubleshoot your entire data estate more seamlessly. Think of dbt Explorer as the data catalog for all your dbt assets. You can instantly visualize your project's DAG and understand resource dependencies. Impact analysis and pipeline troubleshooting are no longer stressful, mind-numbing exercises. With dbt Explorer and [column-level lineage](https://docs.getdbt.com/docs/collaborate/column-level-lineage), you can visualize relationships between sources and models and where they're used downstream in other projects, metrics, and dashboards. Using lineage lenses, it's easy to dive deeper to grok critical details like model execution status, materialization type, test status, and query count metrics. Let's say you get a request from a stakeholder to add a new parameter to their data pull. You no longer have to worry that you'll break something downstream. With [column-evolution lens](https://docs.getdbt.com/docs/collaborate/column-level-lineage#column-lens), you can track exactly how columns flow, transform, and are renamed across your pipeline. dbt Explorer and these new capabilities help you build, troubleshoot, and analyze workflows more efficiently. They also keep your stakeholders happy since they can quickly get the trusted data they need. [Watch video](https://www.youtube.com/watch?v=wdxtzNujVx0&feature=youtu.be) ## Ask dbt You didn't think we'd have a product spotlight series without talking about AI, did you? ICYMI, dbt Cloud now offers a [native in-app chatbot called Ask dbt](https://www.getdbt.com/blog/introducing-dbt-for-snowflake) as part of our Snowflake native app. Democratize analytics by allowing your stakeholders to ask questions like “What is ARR growth over time?” or “What is the count of customers by plan type?”. They'll get (governed, consistent, accurate) answers without writing a single line of SQL. What makes it different than other chatbots? Context. Ask dbt is powered by the metrics defined in the dbt Semantic Layer, improving accuracy by 3x as observed in our [benchmark](https://roundup.getdbt.com/p/semantic-layer-as-the-data-interface). With Ask dbt, users can ask questions in natural language and receive insights in an understandable format. It can significantly speed up business processes and decision-making. The best news? The dbt Semantic Layer also powers your insights for _other_ analytics endpoints. Whether it's a BI interface like Tableau or Google Sheets or a personalized embedded analytics experience, you can be confident that metrics are consistent across your varied user experiences. [Watch video](https://youtu.be/pKtkWQ2b56Q) ## Unit testing You asked, we listened. Unit testing is now live in dbt. And by "we" I really do mean the "royal we". Thank you to everyone in the dbt community who helped make unit testing a reality. 🧡 While dbt users have long been able to build assertions (or tests) about their dbt models (for example: is `not null`, is `unique`, etc.), [unit tests](https://docs.getdbt.com/docs/build/unit-tests) allow you to validate the _behavior_ of model logic _before_ the model is materialized with real data. Not just "Will this model build?" but "Will it build what I expect it to build?" If a unit test fails, the model won’t build. This saves you from unnecessary data platform spend while improving data product reliability and mitigating the risk of introducing breaking changes into your pipeline. **** Lean on unit tests in dbt to: - **Save costs**: Validate logic before transforming a full production dataset. - **Improve code reliabilit**y: Reduce risk of breaking changes in production. - **Collaborate at scale**: Create stable and reliable interfaces for cross-team collaboration. [Watch video](https://youtu.be/5byGFDDj9pg) ## Job chaining Automation is the name of the game, and dbt Cloud just got an upgrade when it comes to orchestrating your jobs. Sure, cron jobs and manual configs _work, _but as your dbt projects and dependencies grow, that approach becomes unwieldy and untenable. With job chaining,** **you can trigger jobs to run automatically as soon as an upstream job successfully completes, both _within_ and _across_ projects. All you have to do is configure the jobs you want to chain and flip a toggle. Using job chaining, you can: - **Optimize compute spend:** Only run a downstream model once an upstream job has successfully completed. - **Reduce undifferentiated heavy lifting:** Don't manually trigger jobs, let the computers do it for you. - **Reclaim control**: Build flexibility into how you orchestrate your jobs, so your stakeholders are always working from the freshest data. [Watch video](https://youtu.be/iZ8_M90i1uE) ## That's a wrap Thanks for joining us for our first-ever Product Spotlight Series! We hope you learned something new about how dbt Cloud's latest features can improve your analytics workflows and foster data collaboration within your organization. As always, we love your feedback. Please continue to send it along via your rep, our community Slack, or in-app. We hope to see you IRL soon at [Coalesce 2024](https://coalesce.getdbt.com/register), which is happening October 7 - 10 in Las Vegas (or online. We'll geek out over features like these that are helping us transform data analytics together. ## --- --- title: "Demystifying AI: 7 simple tips to get you started" description: "Demystify AI with 7 simple tips to kickstart your journey. Learn how to integrate AI seamlessly into your data strategy." url: "https://www.getdbt.com/blog/demystifying-ai-7-tips-to-get-started" date: "2024-08-20" authors: ["Drew Banin"] categories: ["Learn"] --- # Demystifying AI: 7 simple tips to get you started An interesting thing is happening with AI today. As a co-founder at dbt Labs, I have the good fortune of getting to speak with dbt Cloud customers and dbt Community members frequently. And here’s something I’ve noticed: even as some organizations are already seeing real value from AI initiatives, _many_ _more_ are struggling with how to get started. The era of AI was thrust upon us quite suddenly, and the transition to leveraging AI can be daunting. There’s a lot to consider: Where does using AI even make sense? What does good governance look like? What changes might you need to make to your data architecture? **** But I’m here as the bearer of some good news: it doesn’t have to be this complicated! My goal with this post is to demystify the process of rolling out reliable AI, in the form of some **actionable tips **gleaned from working with customers who are already seeing AI success. **… But first, let’s get specific: what are we talking about here?** There are two fundamental ways that AI is changing data workflows today: 1. By making it easier and faster to **build new data products. **With the help of [copilot-like experiences](https://www.getdbt.com/blog/introducing-dbt-assist), it’s now simpler than ever to model, test, and document data with speed and precision. 2. By making it easier and faster to **consume or analyze data. **That is, making organizational data accessible to more people in more scenarios through the use of natural language interfaces. These are both tremendously exciting, and we’re doing a lot of thinking about both here at dbt Labs. However, I want to focus today on the latter one, as it arguably represents a larger paradigm shift for data teams. There’s also a ton of demand for it: I’m seeing overwhelming, almost universal interest **in unlocking self-serve data exploration powered by AI**. Business users want to be able to self-serve their own data insights. Now with the help of LLMs, they're on the cusp of being able to ask questions about their data as easily as they might search for a recipe online. If you’re seeking to make that possible at your organization, here's my advice to you. ## **7 simple tips for building a reliable AI interface to your data** ### **1. Remember: LLMs aren’t magic** You should generally approach LLMs like you would any other tool in your technology stack. There are some tasks that LLMs are well-suited for (eg. generating or summarizing text, following text-based instructions) and some tasks that LLMs are less well-suited for (eg. forecasting or anomaly detection at scale). In this way, LLMs are just tools, and _you_ get to choose how and when you wield them. **LLMs are (generally) good at:** - Generating text - Summarizing text - Extracting sentiment from text - Following text-based instructions **LLMs are (generally) bad at:** - Classification - Forecasting - Anomaly detection - "Classical ML"-style problems This doesn’t mean you can’t solve the above problems with LLMs today! They just do not perform quite as effectively as “classical ML” approaches like regression and clustering. It’s always a good idea to be problem-obsessed, and it’s never a good idea to fall in love with a particular solution. LLMs are very good at powering conversational interfaces (eg. generating and summarizing text), so a chatbot would be a great application for an LLM. For other types of problems, you shouldn't assume that an LLM is the right shape of solution to solve the task at hand. Just as you would anywhere else, remain focused on selecting the right tool for the job. ### **2. Supply relevant context** LLMs are trained, _basically_, on the Internet. LLMs _aren't_ trained on your organization’s data (with some possible exceptions). Fortunately, you can supply additional information to an LLM in the form of a well-crafted _prompt_. Generally, you can think of an LLM as a reasonably smart person that you just hired off the street. They can read a lot and work very quickly, but they don’t know anything about your organization in particular. The instructions and supplemental documentation that you'd give to such a new employee should also be present in the prompt that you supply to an LLM. **Example:** A business user at your organization might ask: “How many customers did we add last month?” If your LLM prompt is simply “You are a helpful assistant,” you're likely not going to get a great response. It may just make up a number. **This is bad:** ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/2c8361f423f26b1f6e8c0a53e77a7c4904f7ff20-1569x565.png) However, if you instead supplement the context in the prompt itself, you're likely to get a more helpful response. This prompt might look like “You are a helpful assistant. **Use the context below to answer the user’s question.” **This technique of providing additional context is called RAG (Retrieval Augmented Generation) – it’s a very important and popular technique for such use cases. **This is good:** ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/d7076f9737f9eb048442ef45bbb81facf7b163f9-1550x803.png) ### **3. Take governance seriously** LLMs are sponges for context. The more context you give them, the more accurate and helpful their responses will be. And they're very good at synthesizing this available context into their responses. While this is often a strength, it also presents as a weakness (and a risk). You _must_ make sure that the data you provide to an LLM is appropriate for your given use-case. If you are building a customer support chatbot, build your application so that each ticket gets its own fresh context to avoid leaking information across customer chats. Additionally, make sure that the context made available to the LLM doesn't include sensitive or private or otherwise confidential information that shouldn't be communicated publicly. That last point might sound obvious, but keep in mind that _LLMs require a lot of context to generate high-quality responses_. Make sure that you understand your data pipelines well enough to guarantee that the information made available to an LLM is appropriate for your use-case. **Example:** Maybe you supply information on customer support plans to your RAG question-answering pipeline. If a business user asks, “What customer support plan is my account on?”, they'll get ‌a good answer. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/fb51c6382907d2c97d67585e8d1423de8f35c9ee-1613x523.png) You can also imagine that if we for some reason supplied payment info as part of the available context for the LLM, very bad things could happen. A user could ask, “What are the last 4 digits of the credit card on file?” The LLM would happily answer that question, so you need to be sure that the information you supply is correct. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0f337b7676017a1852eedf860797a5dffb73cfea-1647x748.png) ### **4. Don’t forget about data quality** Data quality is key. You can use data quality information to contextualize the answers returned by LLMs. You want to avoid using significantly delayed or stale data to answer user questions. You also want to avoid using datasets with known data quality issues like duplicated values, null values where you don't expect them, or problematic referential integrity. **Example:** A business user at your organization might ask how revenue has changed week over week. Back in the world of BI tools, you might have seen a line chart that looks like the one below. Any discerning, analytics-minded person would look at this chart and say, “Hey, that doesn't look quite right to me. There's probably some missing data and it's not an up-to-date dataset.” ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/752a22ec94f3fa942c6045f827497c95061c9cb1-1209x534.png) If you imagine that the interface you’re using is conversational and a user has asked your LLM, “How has our revenue changed week over week?” The answer from the available data might be that revenue is down 85 percent week over week. And that's an _alarming_ answer! An _alternative_ answer that you could imagine LLM responding with is, “It looks like payment data is delayed by six days, so current data may not be accurate. At present, our reported revenue is down 85% since last week” That's a much more contextual and helpful answer. A business user could look at that, see there's likely a data quality problem, and perhaps even check to ensure the data teams are working on it. In that way, we can use metadata signals for data quality or data freshness to contextualize the responses of an LLM. Given the right metadata cues, you might prefer it to pick a different dataset to use that has similar information but is of higher quality. Or you might rather it simply say it can’t answer the given question right now, but that the user can follow along in the right channels to be alerted when the issue is fixed. ### **5. Document datasets thoughtfully** Utilizing documentation is essential for providing accurate LLM responses. To achieve this, it's important to include relevant information alongside your data to help LLMs comprehend its meaning. Keep in mind that as a documentation author, you have two types of readers: individuals such as data analysts and business analysts who will use the datasets and descriptions, as well as LLMs and systems that rely on this information. When writing documentation, you should use the language of the business. That means spelling out what acronyms stand for and providing abbreviations or colloquial terms in your descriptions. Similarly, you should denote the definitions of common synonyms. The term “rep” might mean sales director, account owner, or account executive at different organizations. By supplying synonyms and acronyms to contextualize the description of a dataset, you can help an LLM find the right data and answer the right questions. If you don’t do that, an LLM is going to have a hard time mapping arbitrary questions from business users to the right datasets. **Example:** Below is a screenshot from [dbt Explorer](https://www.getdbt.com/product/dbt-explorer). Here we have a table called dim_customers and a column called OWNER_NAME. The description of the column right now is “the user's first and last name,” which is a poor description for this column. We haven't‌ enriched the owner name column with additional metadata to describe exactly what that means. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/94e41119056fae4d04124d2e79be55edf197fba3-1660x407.png) Below is a much better description. The second column has _“the full name of the sales director (SD) assigned to this customer”_ as the description. With this type of description, you can ask questions like, “Who is the sales director for Acme Corp?” The LLM will have a much better chance of answering the question correctly. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/3a0a06684bc54ece3f8ee169b06a136f57039c65-1662x490.png) ### **6. Define your key metrics explicitly** An LLM _could_ theoretically define revenue for your organization on your behalf. But you _probably don’t want it to do that_. It might very well come up with a new definition every time you ask. That’s why for KPIs with exactly one definition, you should carefully _document_ that definition, _version control_ it, and _code review_ any changes to it or to upstream tables that impact it. If you do this using the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), you'll get structured information about the key metrics that power your business, along with the ability to access that consistently-defined metric from any connected BI or AI tool. **Example:** Below is another screenshot from dbt Explorer. Here you can see how a jaffle business might calculate what percentage of their revenue comes from food orders. There's probably some sort of metric tree like this applicable to your organization. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/d5fdec4a46d1e56eb0a2db3e6db2aa02c1b68d3a-1600x694.webp) In this example, each metric is defined explicitly, in code inside of dbt, building off of existing transformation logic. You might notice we have a formula to very precisely define food_revenue_pct. This makes it so that if you're asking a LLM to tell you how the percent of revenue coming from food orders has changed over time, you're not asking the LLM to define that metric for you. You're asking it to take well-understood, precisely defined metrics and format them in the way that you want to see them. ### **7. Build high-quality feedback loops** You’ve probably established some feedback loops with your stakeholders, even before LLMs are introduced to the picture. But with the addition of LLMs, it becomes even more critical to have strong feedback loops. Feedback loops allow you to constantly improve your LLM’s usefulness over time. You can add logging and monitoring to understand end-user satisfaction or understand if an answer was helpful or unhelpful. Then you can use that information to prioritize new data products, improve documentation for your datasets, or iterate on your LLM’s prompt. **Example:** A user might ask, “How have our new escalation policies impacted our customer satisfaction scores quarter over quarter?” However, perhaps this information isn't available to the LLM – perhaps because you haven’t loaded the data into your data platform just yet. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/5977528a5cc2e9889c8dae44ed03399e511755ce-1665x713.png) A good answer from the LLM would be, “Sorry, I don't see any information about CSAT in the database, but I can open a ticket for the data team to look into it. Is there any other context that you want me to add to the ticket?” The business might say, “Sure, I'm looking into the customer satisfaction score because it's part of our OKRs for next year.” In the background, your data team can‌ then open a ticket, and log that a person had a question that the LLM wasn’t able to answer. Now you can make an informed prioritization decision about if it's a new dataset you want to load into your data platform. Or perhaps you learn you need to update your documentation because customer satisfaction score isn't labeled correctly in the database. This approach will help you iteratively improve over time. As more people use this chat-style experience for data, your team will quickly learn more about the types of questions that they’re asking, and there the experience is or isn’t meeting expectations. ## **Seamlessly integrate LLMs into your data strategy** To summarize, here are the seven things you need to remember to get started: 1. LLMs aren’t magic 2. Supply relevant context 3. Take governance seriously 4. Don’t forget about data quality 5. Document datasets thoughtfully 6. Define your key metrics explicitly 7. Build high-quality feedback loops That sounds like a lot, but the good news is that **nothing on that list is‌ actually new**. These are things your data team has probably already been doing for some time._ The best practices required for a successful AI initiative are not actually that different from the best practices required for a successful BI initiative._ You already know the drill—document your data, build a semantic model, measure and alert on data quality, govern access to sensitive information, and prioritize feedback from stakeholders. Chances are, your team is already using tools like dbt to do this, which means you're well-positioned to leverage LLMs. By building on your existing foundation, you can share controlled, trusted, and precise datasets across your organization without having to start from scratch. Ultimately, the most important takeaway is to **focus on work that drives your business forward**—whether that means delighting customers or boosting efficiency. If LLMs can help you achieve these goals, you’re on the right track. Hopefully, these seven tips will help you make that happen. To explore how leading companies are using AI to enhance productivity and decision-making, [check out this whitepaper on dbt Cloud](https://www.getdbt.com/resources/unlocking-the-power-of-AI-with-dbt-Cloud). --- --- title: "Understanding the 3 key analytics personas" description: "What everyone gets wrong about the humans who do data." url: "https://www.getdbt.com/blog/understanding-analytics-personas" date: "2024-08-18" authors: ["Tristan Handy"] categories: ["Insights"] --- # Understanding the 3 key analytics personas _This post first appeared in_ _[The Analytics Engineering Roundup](https://roundup.getdbt.com/p/analytics-personas)._ One of my favorite stories about how dbt came to be, and why I think it has been successful in a counterintuitive way, centers on our understanding of analytics personas. In this issue I want to tell a bit of that story and then link it back to the writing / thinking I’m doing about the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) these days. First, the story. ## The origins of a counterintuitive success: Fishtown Analytics and early challenges Ten years ago, I was a bit of an unusual creature. Today, there are a lot of people with my particular mix of professional skills, but a decade ago there weren’t. Specifically, I had three specific things: - Strong analytical skills - Deep business knowledge in a few domains - Just enough software development skills to be dangerous A decade ago, it was not unusual to have one of those. Two of them could give you an unfair advantage. But if you had all three, you could do some pretty neat things. So when I went to start a data consultancy called Fishtown Analytics, I started with slightly different priors. All of the work I was going to do was going to adhere to software development best practices. My clients didn’t really care about this at the outset, but I knew it was going to help me deliver a ton of business value per hour (ultimately the most important metric for any consultant). This ended up being true, and it drove our success as a consulting business. And when I needed to build some tooling to help me implement software engineering best practices in data, I built it assuming that the user was modestly technical. Not _highly_ technical, not primarily interested in thinking about software architecture all day, but _technical enough_ to do some historically-atypical things. Such as: use git. Such as: use the command line. Such as: write Jinja macros. If you’ve only been in data for <5 years, theses behaviors may seem normal to you. In 2015 they were not. And I know this for two specific reasons: 1. When I showed data practitioners dbt back in 2016, almost none of them wanted to use it. Data engineers didn’t want to use it because they preferred to just write Spark jobs. Data analysts didn’t want to use it because they found it to be too technical, too hard to use. 2. When I showed dbt to VCs back in 2016 (I pitched maybe half a dozen folks just out of curiosity), they all told me that they weren’t confident there was a large market for a product like this. Products built for data analysts were typically visual, and products built for data engineers didn’t use boring languages like SQL. I was totally fine with this. I didn’t intend for dbt to be a widely-used product, I just intended it to be the way I created exceptional client outcomes and supported an awesome hourly rate. I thought it was neat that a few other weird folks like me also found dbt to be useful, but that was never really the point. The first people that joined the dbt community were risk takers, confident in their own abilities, willing to say ‘fuck it, I’ll learn the CLI’ even if that was a totally new skill. And when this paid dividends for them, they told others. The rest of the story is just an exponential function playing out over time. Over time, people are like water. Water always travels to the lowest point if it has a path to get there. _People always develop the skills that help them be maximally effective…_if they have a path to get there! Betting against people, their initiative, and their intelligence is always a bad bet. _If a certain group of people seems like they’re underachieving, it’s not because they’re stupid and unmotivated, it’s because they’re blocked._ dbt allowed data analysts to build mature data pipelines and significantly magnify their impact as professionals. It got them noticed, got them promoted. It unblocked them. It _never_ underestimated them. I always think about those VCs who, in 2016, told me that there was no market for dbt. They weren’t wrong—there was no market for a product like dbt in 2016. The market had to be _created_. And that required two specific things: 1. The actual result had to be 10x better. No one is going to learn challenging new skills to be 30% better; that type of hard work has to be rewarded with an order of magnitude improvement. The only way to know whether something is an order of magnitude better at the outset is to experience it for yourself. 2. There had to be a culture in at least a part of the data practitioner community that valued technical skills. In 2016 there were a LOT of data analysts who were envious of their data scientist peers and wanted to “get more technical” so that they could increase their earning potential. This meme was already present. Both of these things ended up being true; it was just really hard for VCs to observe this from the outside back in 2016. From the inside, it was obvious. ## The 3 analytics personas: The engineer, analyst, and decision maker Why do I share this story now? Is senility setting in? Am I longing for the good old days? You’ve stuck with me this long; indulge me for just a bit longer.I mentioned in [my last issue](https://roundup.getdbt.com/p/the-analytics-development-lifecycle) that I’m writing a rather lengthy whitepaper right now about the ADLC. **** In it, I write a bit about personas, and I split basically all data work into three of them: ### The engineer The engineer creates reusable data assets: pipelines, models, metrics, etc. The engineer is primarily focused on creating data assets that will be _used by others_ to create business value. ### The analyst The analyst performs in-depth analysis that drives decision making. The analyst does not make decisions, their role is quantitative investigation and presenting of analysis and/or recommendations to the decision maker. ### The decision maker The decision maker is responsible for taking quantitative outputs and translating them into action for the business. This is inclusive of everyone from a campaign manager optimizing segmentation to a CEO directing the resources of an entire company. These are not job titles, they are roles that humans play in the ADLC. They’re also not revolutionary. You could maybe pick at the details, but likely this seems fairly uncontroversial. Here’s the important point, though: **Analytics is not an assembly line.** You cannot disassemble an analytical problem and hand it out to a set of different humans and have them all come back together with an answer. Well, you can—but you can’t expect this to get you good outcomes. Analytics cannot be effectively treated as an assembly line because analytics is an iterative process that involves asking and answering questions, gathering data, poking at it, getting curious, getting stuck in dead ends, and realizing that the fact you learned way over _here_ is actually the answer to this question way over _there_. Analytics requires a neural network—currently, a human!—interacting with a really good computer. And neural networks do not cleanly submit to industrial logic, to mechanization. When you try to treat analytics like an assembly line you do get predictable outcomes (likely dashboards!), but not insights that drive ROI. You certainly don’t get agility or velocity. Insight, agility, and velocity in analytics require curiosity, flexibility, integration. The best data teams allow talented people to flex between these different roles. They allow them to take an idea and get curious about it, to explore it without needing to file a ticket or wait for anyone else. Getting stuck in someone else’s queue is where non-linear ideas goes to die. Here’s my mental model for how personas work in data. The above three personas are like hats you put on. You have your primary hat, the one you like wearing best. But over the course of the day as you get pulled into solving real problems, you will be best served by putting on different hats. The most effective data practitioners can wear all three hats. And the best data tooling enables as many people as possible to wear all three hats. Even with great tooling, you will still have a hat you prefer. But the ability to wear all of them as the situation demands is an absolute superpower—it allows you to complete a single thought yourself, without getting stuck behind someone else’s priority list. ## How focusing on non-linear growth drives innovation I believe most data tools get this wrong. I believe this is because classic product management teaches us to focus on a single persona and then ask how that persona wants to get a particular job done. This mindset assumes people are fixed. I prefer to ask the question: how do we create product experiences that have a gentle learning curve, that take every user on a journey that leads them to non-linear outcomes? How do we make it possible for many people, each on their own journeys, to collaborate together and achieve something incredible? This mindset assumes people can and will grow. **People are smart and motivated. Underestimate them at your peril.** --- --- title: "dbt explained" description: "You've heard of dbt - but what, exactly, does it do? An overview of dbt’s role in helping teams build and deliver quality data." url: "https://www.getdbt.com/blog/dbt-explained" date: "2024-08-16" authors: ["Drew Banin"] categories: ["Product"] --- # dbt explained Your business runs on data. You’ve got lots of it coming in from lots of different places. To build a competitive advantage, you need to find ways to integrate these datasets, transform them into actionable insights, and deliver them to the relevant stakeholders as quickly and accurately as possible. The challenge is how to do this without the system becoming a giant, untraceable, out-of-date mess that doesn’t scale. This is exactly the problem that dbt solves. In this article, we’ll look at what dbt does, how it works at a high level, and how it can standardize the way your organization builds, tests, delivers, and catalogs clean, accurate data. Check out this 2-minute explainer video where I walk you through it all, or keep reading to dive deeper. [Watch video](https://www.youtube.com/watch?v=rItxGK0cYj8) ## Data, data everywhere When companies begin mining their data for insights—think: clickstream data, advertising data, customer payment data, product usage data, CRM and ERP data, etc.— the initial efforts are usually ad hoc. Anxious to investigate the data they need to make a decision, you may already have some teams that are going-it-alone and downloading CSVs from different data sources and manually merging them together in a spreadsheet. Such efforts are highly manual, error-prone, and impossible to scale. Different teams trying to get to the bottom of the same insight, for example “What is our return on ad spend? (“ROAS”)”, wind up with different answers to the same question. When merging columns from Facebook Ads and Stripe, inevitably, columns don’t match, primary IDs are tagged differently, and you’re building complicated join logic to get the data you need. And the worst part? That data is stale the moment you’ve downloaded that first CSV. With an analytics stack, it’s easy to extract raw data from your data sources and load it into a cloud data platform. From there, you can integrate with BI tools to drive varied operational use cases. But what happens when your business grows? You launch your service in Europe and Asia, and now you’re allocating ad dollars to new platforms. To get to the bottom of the same insight (“What’s our return on ad spend in Europe? In Asia? Globally?”), you need to be able to dynamically splice and dice your ROAS metric across various dimensions. There will always be new data sets, new business logic, new use cases, and new stakeholders demanding new ways of looking at the data. You _could_ go back to your BI tool(s) to update your query logic to build new region-specific dashboards. But here’s the problem: when you’re managing business logic across a bunch of different (and always growing) dashboards, it’s only a matter of time before metric inconsistencies pop up. Each region sells different products, or categorizes the same product differently, or generates revenue in different currencies. The complexity is simply too vast and dynamic to manage with the current approach. It soon becomes very clear, very fast, that the one-off SQL queries you wrote are actually load-bearing, and change management involves manually updating dozens of dashboards three times per week. And you’re still constantly worried that your exec is going to call out a discrepancy in an important meeting and derail any of the trust you’ve built in your data processes. In this case, you need to standardize your data and abstract the complexity of our business logic so that everyone in the organization—whether a data engineer, BI analyst, or your executive stakeholder—can move fast with trusted data in a scalable, cost-effective way. This standardization needs to happen inside of the data platform—not across a bunch of BI dashboards. **** ## dbt and the analytics stack This is where [dbt](https://www.getdbt.com/product/what-is-dbt) comes in. dbt helps manage this complexity—in a way that’s modular, scalable, repeatable, and governed—all directly inside of your data platform. dbt is the standard for data transformation in modern environments. Using dbt, data teams can build, test, and deploy analytics code using software development best practices (like portability, CI/CD, observability, documentation, etc.) to create production-grade analytics pipelines that scale. The output is modular, clean data models that can be delivered into our BI tools, LLMs, and APIs to ensure that stakeholders have the accurate data they need, when and where they need it. dbt is foundational to the modern data integration workflow known as “ELT” which stands for “extract, load, transform.” dbt is the “T” in ELT, helping teams transform data _after_ it has landed in the data platform and shapes it into the format that analysts, managers, and executives need to drive decision-making. With dbt, you can be confident in your data and that the decisions they inform are accurate, governed, consistent, and shipped to downstream teams with agility. ## How dbt works We built dbt with two core beliefs: 1. Transformation logic should be defined in code; and 2. Data teams should have the tools to be able to work like software engineers so they can treat their data assets like a product. To accomplish this, dbt enables building robust **data pipelines**—data transformation processes that are version controlled, tested, documented, and shipped incrementally in a secure and governed way. Engineers use dbt to define a **model** using either SQL or Python. Your dbt models specify how to transform data, shaping it to normalize any dimension that matters to your business—campaign names, currencies, product categories—all according to the business logic that you define. Engineers also write tests to validate their code, and dbt automatically documents every code build to describe the data and what it represents so that future collaborators have the context they need to understand and build on existing data assets. All models are version controlled with Git and data teams create create pull requests (PRs) for other engineers to review before pushing code into production. Once approved, dbt then runs your models, **materializing** the data into a view or table. This process of merging data models into production creates a [continuous integration](https://docs.getdbt.com/docs/deploy/continuous-integration) (CI) data pipeline. dbt enables CI pipelines that materialize and test your data in different [deployment environments](https://docs.getdbt.com/docs/deploy/deploy-environments)—e.g., by creating it in a development environment and testing it in a staging environment prior to running it in production. This helps your data teams identify and resolve any issues with a data model change before it negatively impacts users. Once shipped, anyone with the appropriate access can use this normalized data to build reports, data-driven applications, and other data assets. With these normalized data sets at their fingertips, regardless of the use case, users can be confident that the analytics they’re using to make decisions are accurate, consistent, and frequently updated to keep pace with changing conditions. dbt also automatically builds and publishes documentation about every model, enabling anyone in the data workflow to navigate and understand critical context about a data asset (lineage, dependencies, freshness, materialization, exposures, etc.). This improves data velocity and collaboration, as data teams can work more efficiently and data consumers have the trust signals they need to use the data with confidence. **** ## The value of dbt By standardizing on dbt, companies can: ### Ship data products faster With dbt at the center of your data workflows, your data teams are empowered with a governed and scalable approach to not only develop and test their data models, but also explore, improve, and deploy them at scale. Data teams finally have the ability to work more efficiently, remove bottlenecks, and ultimately ship data products faster. ### Build trust in data and data teams Meanwhile, downstream teams are empowered to make more decisions, faster decisions, better decisions because they have a way to interact with that data, to understand its lineage, and to contribute to it. And if they’re less technical and are using a BI tool to tap into that data, they actually trust the metrics they’re receiving because they are consistent, regardless of where they’re queried. ### Reduce the cost of producing insights Having this centralized framework for collaboration feeds a self-fulfilling flywheel that fosters a data-driven culture that data teams can actually keep up with. All of a sudden, your data platform goes from being a cost center to a profit center. dbt also offers built-in user experiences designed to optimize data platform compute and help teams pinpoint and address inefficient, unused, or long-running models. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/42d7bc068c6ce35c63ad918338fe9e0aca3a3266-1244x754.png) Managing data at scale is full of pitfalls, including incorrect, out-of-date, or hard-to-find data. Using dbt, you can build standardized, tested, and documented data sets with high quality that anyone can confidently use to drive your business forward. --- --- title: "Analytics engineer 🤝 data analyst: How dbt Cloud enables our workflow" description: "Learn how dbt Cloud empowers collaboration between analytics engineers and data analysts at dbt Labs." url: "https://www.getdbt.com/blog/analytics-engineer-data-analyst-dbt-cloud" date: "2024-08-12" authors: ["Lauren Benezra", "Paige Berry"] categories: ["Product"] --- # Analytics engineer 🤝 data analyst: How dbt Cloud enables our workflow At dbt Labs, our data team is a centralized unit that collaborates with various departments through specialized pods. We're part of the corporate pod, which supports finance, G&A functions, and organization-wide initiatives. Our work is all about empowering our internal customers with the data they need while also maintaining our systems like Snowflake and Git repos. As a data analyst, Paige is all about that "last mile"—transforming organized data into actionable insights through queries and visualizations. She also loves playing data detective, solving the mystery when data looks off. Lauren, on the other hand, spends most of her day in dbt Cloud, fixing jobs, setting up source connections, and occasionally diving into BI tasks. Her work ensures that our data pipeline runs smoothly, and when things break, she’s the one who gets them back on track. The overlap in our roles is where the magic happens. We both have access to the entire data stack, which means we can troubleshoot and solve problems together, ensuring that nothing slows down our projects. Our shared knowledge and collaboration make us a powerful duo, capable of tackling anything that comes our way. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/6d45a13d6a295689e2d139b4279344350b9b4343-3024x4032.jpg) We'll dive into how dbt Cloud has changed our workflow, the game-changing features we can't live without, and how we've used dbt Mesh to streamline complex projects. We'll also share our tips for debugging data issues and why we believe in the power of collaboration and self-service in data teams. So, whether you're an analytics engineer, a data analyst, or just curious about how we work, stick around. We have plenty to share, and we hope you'll find our insights helpful as you navigate your own data journey. ## What are the data team pods? The data team at dbt Labs is a centralized data team of analytics engineers and data analysts. Together with folks from finance and revenue operations, we form pods that support different parts of the business. - **The product pod**: supports engineering, product, and design. - **The go-to-market pod**: supports marketing and sales. - **The corporate pod**: supports finance and G&A functions, plus organization-wide initiatives. Since our team is centralized, we support our own data team projects, like maintaining Snowflake and Git repos. We also help enable self-service, so our internal customers can answer their questions about our data. Speaking of self-service, we know it’s a spicy topic and we have a LOT to say on this, but we’re going to put a pin in it for now. ## What is your job and what parts of the data stack do you cover? **Paige:** As a data analyst, I think of it as the last mile. I take data that represents business entities and processes. After the data is beautifully organized by the analytics engineers, I combine the data, adjust the shape, run queries, and make charts to answer business questions, or even identify what business questions to ask. I also love to be a data detective. If somebody comes to me and says that their data looks weird, I get to dig through the logic, the lineage, and the sources to figure out if it’s a pipeline issue, or something changed in the behavior of our customers that we need to address. **Lauren:** Don’t forget about Git! You approve all of my PRs. The “enabling Lauren” layer is super important. I spend most of my day in dbt Cloud, where I’m fixing broken jobs and figuring out where our problems are, as well as building net new pipelines. There’s some cross-over into data engineering land, where I’ll set up new source connections. Sometimes I hang out in business intelligence (BI) land and do some basic reporting in our BI layer, but charts aren't my friends, so I try to stay away from that. ## What's your day-to-day like? **Lauren:** The first thing I do when I wake up in the morning is check Slack and look at urgent things happening in my DMs. This usually means that some job is broken, and that's where I really start my day. We have a massive job that has existed since the beginning of our project. It runs most of our models, and as you can imagine when you're running 3,000 models, it becomes very cumbersome. It takes forever and when it breaks, you have to restart it. It's the worst moment of the day. I've been working hard to break that up into more digestible chunks so that we can curb some of our Snowflake spend and make some of our data run once a week, or every 12 hours, instead of every four hours, like our production job. **Paige: **I also tend to start in Slack. Not so much to learn about what’s broken, but to make sure I see any urgent needs from my executive-level stakeholders. When they need data, they often need it right away. I also have a set of Slack channels I check to look at discussions. I like to add anything I can to help people and share insights. I’ve been at dbt Labs for almost two years, which is like a lifetime at a startup, so I have a lot of institutional knowledge. One of our core values at dbt Labs is to contribute to the knowledge loop, which is why the dbt community exists. **Lauren**: And you review my PRs! **Paige:** Yes! That’s my favorite part. You can learn so much about SQL and the data by reviewing code. Another fun thing about discussing code together is that we also get to talk about where logic lives. Does it live in a BI tool, or does it belong in dbt, where it’s version controlled, or the Semantic Layer? Should it be a metric? This is where the blurriness and overlap between our jobs is awesome. ## What benefits do you see from the blurriness between your positions? **Lauren:** While I may spend most of my time in dbt Cloud, and Paige spends most of her time in Snowflake and Hex, we both have access to the entire stack. That means we can look into any part of the pipeline and address issues, all while conferring with each other. We also share so much context, because we are working on projects together. So for PR reviews, or pairing on solutions, we have a built-in partner with much-needed context for top-notch code reviews. On top of that, we can easily stand in for each other as needed, so the progress and maintenance of our projects doesn’t need to pause. Build in redundancy! ## What did your workflow look like before and after today’s dbt Cloud? **Paige: **I started working with data in 2004. The organizations I worked for used Oracle databases, and there were many times when our things were SQL scripts in random places, like on someone's laptop, or a shared drive. We'd have command line scripts that'd run the SQL files one after another very carefully. If anything got out of order, or something got deleted or changed, it was a disaster. I remember building a data pipeline in PERL once. That was a fun one. I remember a whole day getting derailed because somebody accidentally dropped the zip code column out of the database. It's always addresses, right? I did a short stint as a software engineer and learned about engineering principles, and it all made so much sense to me. When I returned to working on a data team, I tried implementing version control in Git for our most complex SQL queries. It was better than nothing, but far from what it could be. When I discovered dbt, I immediately knew how valuable it was. Folks had finally figured out how to apply software engineering principles and best practices to data work. The company I worked at used dbt Core, so I spent a lot of my time wrestling with Python environments, dealing with Git, and making sure I was in the right repo state. I had to rebuild my local repo many times, and I had a giant list of Git commands I’d follow based on how I’d screwed up the repo. It was much better than we had before, but it wasn’t a smooth process. It was time-consuming. After dbt Cloud, the workflow is seamless. As soon as I think about what I want to know or do, before the thought even feels complete, I'm already in dbt Cloud Explorer. I look at CLL, hit the "open in the IDE" button, refresh from main, create a new branch, and run queries or make changes so fast I feel like a superhero. Working in dbt Cloud makes me feel like the computer is an extension of my brain instead of something I have to sweet talk for an hour and a half before I can get to work. These features are game-changers, and I can’t imagine working without them. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/04a5bc56e26a9661984416742478702b5508e6da-1804x454.png) ## What are the game-changer features of dbt Cloud? ### dbt Explorer dbt Explorer has changed the way that we investigate. If there is a bug in our data or there's a funky piece of data, we can go into dbt Explorer and figure out exactly what's happening. The image below shows days in the fiscal year being transformed in that particular mode. It's going downstream and there's an error. We already know that in the lineage, which is so rad. I don't need to look in any of these other pieces because those are just pass-throughs. It's using the column, but I don't need to investigate those particular models because those aren't transforming that particular column. Column-level lineage (CLL) has changed our lives. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/53d33ab5dd9d7df00011319f5c305d67c9f1614b-1631x632.png) ### Rerun prod job from failure Rerun prod job from failure is also a game-changer. Our prod job runs every four hours and takes about two and a half hours to run. When we had to rerun it from the start, we had to tell people that they’d need to wait three hours. Now we can just rerun from failure, and it might take an hour or thirty minutes. It's changed the way that we communicate with our internal customers. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/c4bd33d32c17af4a9d5883cf1874520733878b88-1708x399.png) ### Compiled tab and copy code button The compile button is great because sometimes we just want to look at the SQL in the model and then take it somewhere in the BI tool where we can make small changes and put it in a chart right away. It's great because we can hit the compile tab and see the raw SQL, copy it, and then we’re back in the BI tool with all the code. It's so much faster than what we used to do. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/e74867651c18efb5d46b080a35154358461b3ea9-1473x888.png) ### Rerun CI job from failure I also love the rerun from failure for the CI job. When we make a change and submit a PR, we run the CI jobs to make sure it won't break everything if we merge that code. As the CI job runs, maybe there's something else in the project that isn’t working but isn’t related. It'll fail, but it won't fail because of my code. In the past, I'd have to go in and put a space in a document somewhere so I could commit it and have the CI job run again. Now I just have to click a button. It's so much faster to resolve that kind of problem. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/eceed3b27c0dc7bed11541fbfb4ee3fbed04f667-1485x891.png) ### Defer to Production Deferred to Prod means that if you don't have the production, the models that are upstream of what you're working on in your development schema, then it'll look to the production schema to get the data from there. This is a great time saver because every week we used to have to remind everybody who was going to do any development that they had to run this command. Otherwise the data would be stale in their schemas or we'd have dropped them if they were stale. Instead, now we just have a toggle. As soon as we toggle it on, it’s already doing all that work and we never have to think about it. And you don't have to run the upstream models. You can just turn on Defer to Prod and already have those upstream models built for you on top of production data. They already exist and you're just using what's already built. It's so quick. ## How has dbt Mesh changed your workflow? **Paige: **We’ve meshified revenue recognition and accounting and put it into a new downstream project. Our main project has public models that we can build upon. But with dbt Mesh, it's focused. It is for specific customers and business users, so its impact is significant but narrow. We won’t break marketing data while fixing accounting models. I’m used to immediate gratification being a data analyst. That can feel challenging when I’m making a change to the repo and dbt, but dbt Mesh makes it super easy. **Lauren: **We can use versions, which creates a new version of your model. With dbt Mesh, you can point to either version, so we have parallel pipelines that are new code and old code. They're both running and we're able to see the differences between the two without creating a new version. We have cross-project dependencies, and since we can’t change the big project, we can make versions and reflect them in the downstream Mesh project. We’re making incremental changes, but they don’t have to affect every single model. Instead of being overwhelmed by 2,500 models, accounting folks can go into dbt Explore in the Mesh project and see a subset of models relative to them and their business needs. ## How does dbt Cloud enable stakeholders and self-service? **Paige:** On the data team, we have weekly office hours called self-serve data team office hours. It’s for folks at dbt Labs who are working on something and get stuck or have a question. They're often trying to figure out how to analyze something but they don't know exactly where the data is. I'll share my screen in dbt Explorer and show them how to search for that data and how to find the models. They can then do it for themselves next time. **Lauren:** One of my favorite things is helping somebody figure out how to do a PR to get the data they need to get their work done. It still goes through our PR review and we have CI jobs that are in place so that everything's protected. This process allows people to go in and answer their questions if they have the skills and the time. ## What's been your favorite moment as a team of two? **Lauren:** A few months ago we took over an accounting project, which was a beast of a project. I didn’t know anything about accounting, but we went from zero to one hundred fast. The other day, Paige was talking to the lead accountant and had memorized all of these very specific codes, and I knew exactly what she was talking about. We have so much shared knowledge, that our sleuthing now has the power of two. **Paige:** One of my favorite things is when I can’t figure something out and I'm deep in a rabbit hole, I can call Lauren and ask if we can get on a Zoom and see if she can help me figure it out. When we get on the call, the first thing she’ll say is, “You are doing a great job. We are doing a great job.” It's always so reassuring. ## Data Busters **Lauren:** Paige, you modified this list of Rules for Debugging and it's amazing. Can you walk through what they are? ### Rules of debugging 1. Understand system 2. Make it fail 3. Quit thinking and look 4. Divide and conquer 5. Change one thing at a time 6. Keep an audit trail 7. Check the plug 8. Get a fresh view 9. If you didn’t fix it… it ain’t fixed! **Paige: **The rules of debugging are from a book called Debugging by David J. Agans. It's about how to debug systems in software and hardware. I was reading the book one day and realized that I do all these things in data work. But it's a little different from what he was talking about in the book and what they mean for us as data professionals. I took the rules and I unpacked each one. Understanding the system is about figuring out the whole pipeline end-to-end. Where is the data coming from? Who's entering the data? What's it for? What's the business process it's for? Through all the transformations, all the way up to the other end. What's cool is that I shared this with Lauren, and she figured out we use dbt Cloud to do all these things. ### Rules of finding a data ghoul 1. **Use dbt Explore to understand the system**: Familiarize yourself with the end-to-end data flow, from data sources to your reporting tools, and everything in between. 2. **Use dbt Cloud scheduler to make it fail**: Create a controlled, replicable scenario where the problem appears, leveraging test data sets or simulated queries if necessary. 3. **Use the dbt Cloud debugging logs to quit thinking and look**: Analyze error messages, logs, and any other system feedback carefully. This could reveal patterns or anomalies related to the problem. 4. **Use the dbt Cloud IDE to divide and conquer**: Break down the problem into smaller parts. Segment your data, modularize your code, and isolate your variables to help identify where the issue lies. 5. **Use the dbt Cloud IDE to change one thing at a time**: When altering code, queries, or data processing methods, change one element at a time to understand the effect of each change. 6. **Use the dbt Cloud IDE to keep an audit trail**: Document your observations, hypotheses, actions taken, and their results. This facilitates collaborative debugging and prevents the same ground from being covered twice. 7. **Use dbt Explore to check the plug**: Verify the basics. Are the data sources available? Is the data being loaded correctly? Are the correct versions of software/tools being used? 8. **Use dbt Cloud Git integration to get a fresh view**: When you hit a wall, seek external perspectives. Colleagues may spot something you've missed, or suggest a new approach. 9. **Use dbt Cloud scheduler and tests to monitor your jobs. If you didn't fix it, it ain't fixed:** After implementing a solution, ensure the issue is genuinely resolved. Validate your solution under different scenarios and monitor performance over time. ## If there’s one thing you want people to know, from a data analyst and analytics engineer perspective, what would it be? **Paige: **There’s a super important path to go down before you decide to automate something. Is the investment worth it? ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/ad81f12712acd7a8d7bd25a3ba99f8c924872821-512x416.png) **Lauren:** Why are you SO passionate about this concept?? **Paige:** Our stakeholders don’t know how long things take to build. We’ll get requests to automate things that take someone 10 minutes every month. We should be proactive about protecting our time and think about what the value is of accomplishing that task. Many times people are just curious, but don’t need all that automation. **Lauren:** My advice is to Google things. It's so important, especially in this industry. I think people get intimidated. When I was in grad school, I thought that everyone knew how to code. I thought it was perfect all the time and everyone knew all of these functions. I had no idea how I was going to remember all this stuff. The answer is you just get good at Googling. We're all Googling things! It's our job as developers to Google and be better at Googling than non-developers. ## Join us at a dbt Meetup dbt Meetups are more than just events—they're opportunities for the dbt Community to come together, share knowledge, spark new ideas, and refine the craft of analytics engineering. Our conversation highlights the power of collaboration and underscores how essential tools like dbt Cloud are for helping teams work smarter and more efficiently. Much of what we've shared in this blog was part of a recent presentation we gave at a recent dbt Meetup in San Francisco. If you're interested in learning more about analytics engineering, exchanging experiences with fellow practitioners, and having some fun along the way, we highly recommend checking out an [upcoming Meetup](https://www.meetup.com/pro/dbt/). --- --- title: "What's new in dbt Cloud - August 2024" description: "Read about the latest and greatest features that recently landed in dbt Cloud." url: "https://www.getdbt.com/blog/whats-new-dbt-cloud-august-2024" date: "2024-08-06" authors: ["Alexis Jones", "Azzam Aijazi"] categories: ["Product"] --- # What's new in dbt Cloud - August 2024 Believe it or not, it's time for our regular product update launch post! Feels like we were [just announcing a bunch of new features in dbt Cloud](https://www.getdbt.com/blog/whats-new-in-dbt-cloud-june-2024)... but our product teams are 👏 shipping 👏, and so here we are again. Before we dive in, we’d be remiss if we didn’t plug [Coalesce](https://coalesce.getdbt.com/register)—the annual Analytics Engineering conference where thousands of data professionals and thought leaders like yourselves will come together in beautiful Las Vegas, Nevada to learn, share, and grow. 🌱 We have an [incredible agenda](https://coalesce.getdbt.com/agenda) that’s quickly coming together, and we hope to see you there starting October 7. **** Alright, on to the announcements! ## 📈 dbt Semantic Layer Learn about the latest features powering the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer). If you’re not sure where to start with your semantic strategy, check out our recent blog post on the [Five use cases for the dbt Semantic Layer](https://www.getdbt.com/blog/five-use-cases-for-the-dbt-semantic-layer). 🔐 **Access controls:** Admins can now federate access based on specific teams or groups of users by configuring multiple underlying data platform credentials and mapping them back to service tokens used for authentication in downstream apps. This means finance teams and data teams, for example, can have partitioned and granular access to underlying data. [Group-level access controls are now GA and available for Enterprise customers](https://docs.getdbt.com/docs/use-dbt-semantic-layer/setup-sl#4-add-more-credentials). 🌎 **Integrations:** Our Microsoft Excel integration is now in Preview! You can find it in the [AppSource Marketplace](https://appsource.microsoft.com/en-us/product/office/WA200007100?src=office&corrid=ff9d118a-e0f8-8727-807b-5e39348fdeb4&omexanonuid=&referralurl=), or get started with our [Docs](https://docs.getdbt.com/docs/cloud-integrations/semantic-layer/excel). This integration is available for both Excel Online (in-browser via Microsoft 365) and Excel Desktop. ![Find dbt Semantic Layer for Excel in the Microsoft Marketplace.](https://cdn.sanity.io/images/wl0ndo6t/main/75b408ced1b11dd80793e1f307c81e0276d9b665-1932x1396.png) 🗄️ **Declarative caching improvements:** Declarative caching now supports query time filters. Now, if you filter on a dimension that’s already included in the cache at query time, we automatically add that filter to the cached data set. This means faster delivery times, regardless of how you continue to splice your data. [Declarative caching](https://docs.getdbt.com/docs/use-dbt-semantic-layer/sl-cache#declarative-caching) is GA for all Semantic Layer customers. **🔢 Cumulative metrics now support granularity options:** Cumulative metrics are non-additive, making them tricky to calculate flexibly. Previously, if you wanted to roll up a cumulative metric to a different grain (e.g.: by week or by month), you’d need to create an entirely new metric for each grain. Now, you can keep your code DRY and streamline workflows by setting your own aggregation behavior when you request a cumulative metric by a new granularity. [Check out the docs](https://docs.getdbt.com/docs/build/cumulative#granularity-options) to learn more. 🐍 **Python SDK:** With the [Python software development kit (SDK)](https://github.com/dbt-labs/semantic-layer-sdk-python) for the dbt Semantic Layer, you can interact with semantic layer APIs and query metrics and dimensions more easily using Python. This allows you to build data products to serve internal stakeholders, or build experiences for external users, like your customers, more quickly. [Read the docs](https://docs.getdbt.com/docs/dbt-cloud-apis/sl-python) to learn more. **🔀 Metrics in CI:** You can now automatically test your semantic nodes (metrics, models, and saved queries) during your code review process! Just add validation checks in your CI job using the `dbt sl validate` command. You can also validate modified semantic nodes to guarantee the code changes you make to dbt models don't break your metrics. [Check out the docs](https://docs.getdbt.com/docs/deploy/ci-jobs#semantic-validations-in-ci) to learn more. ## 🔎 dbt Explorer [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) is your command center for navigating, understanding, and improving your dbt projects. Dive into what’s new: 🗺️ **New landing page:** dbt Explorer has a new-and-improved landing page designed to be the perfect jumping off point for diving deeper into your dbt Cloud projects. Get an at-a-glance view of important context about your projects including descriptions, latest changes, and latest issues, or scan the mart and public models tables to understand relative usage and lineage to keep pipelines healthy. You’ll also notice improved filtering and sorting of models in the model filter view. [Read the docs](https://docs.getdbt.com/docs/collaborate/explore-projects#overview-page) to learn more. ![Updated landing page for dbt Explorer in dbt Cloud.](https://cdn.sanity.io/images/wl0ndo6t/main/e46ee965f946eea7405dd47f98bc6ffe1959bc3c-1784x1828.png) 🦘 **Navigate from the Cloud IDE into Explorer:** Jump straight from the Cloud IDE into dbt Explorer to quickly visualize lineage and dependencies for the model, seed file, or snapshot you’re developing. (You can also navigate to the Cloud IDE _from_ dbt Explorer in one click 🙌.) [Check out the docs](https://docs.getdbt.com/docs/collaborate/access-from-dbt-cloud) for more info. ![GIF showing navigating between dbt Explorer and the dbt Cloud IDE.](https://cdn.sanity.io/images/wl0ndo6t/main/5a4db45749d02e41baff02a4d8c9b1b296d11b6f-1920x1080.gif) ## 🏃‍♂️Developer productivity Check out these new features designed to help supercharge developer productivity: 😎 **Environment-aware dbt snapshots:** Snapshots in dbt were originally created before we introduced the concept of deferral, and so that meant that, previously, regardless of what environment (dev, staging, prod) you were in, if you ran dbt snapshots, it would build to the same `target_schema`. Now, dbt snapshots can now be configured to be environment-aware, making testing and developing snapshots much safer and easier. Now available in dbt Cloud for customers who are running on “[Versionless](https://docs.getdbt.com/docs/dbt-versions/versionless-cloud).” [Read the docs](https://docs.getdbt.com/docs/build/snapshots#snapshot-configurations) to learn more. 🔀 The dbt Cloud CLI now supports **parallel command execution,** just like the dbt Cloud IDE. This means you can run multiple different dbt commands such as `dbt build` and `dbt compile` at the same time. Read more about parallel execution in [our docs](https://docs.getdbt.com/reference/dbt-commands#parallel-execution). 🧺 The dbt Cloud CLI now also supports **SQLFluff commands**, allowing you to easily lint your SQL files for improved consistency and readability. More info in the [docs](https://docs.getdbt.com/docs/cloud/configure-cloud-cli#lint-sql-files). ⚠️** Notifications for warnings**: You can now choose to get notified via Slack or email if a job run encounters warnings from tests or source freshness checks — in addition to already available options for notification on job success, failure, or cancelation. See [the docs](https://docs.getdbt.com/docs/deploy/job-notifications) for more. ## Platform & Partner We have lots of exciting announcements about platform improvements and partner milestones to share: ☁️ **Support for Azure deployments:** dbt Cloud multi-tenant can now be deployed natively on Microsoft Azure! This is in addition to the current support for [AWS deployments](https://docs.getdbt.com/docs/cloud/about-cloud/tenancy), bringing the same powerful dbt Cloud experience to even more data teams — regardless of your choice of cloud. Azure multi-tenant support is currently in Preview, and available to dbt Cloud Enterprise customers. [Azure deployments](https://docs.getdbt.com/docs/cloud/about-cloud/tenancy#available-features) can be hosted in Europe, with more deployment regions to follow in the coming months. 💫 **New dbt Cloud UX:** You’ve likely already noticed, but dbt Cloud has gotten a glow-up, including an updated navigation panel and account home page (now in Preview). These updates make it more intuitive to navigate between different areas of dbt Cloud. For instance, you can now quickly configure and find your favorite projects in the project switcher, while the new account homepage gives enterprise customers an at-a-glance view of each of the projects they have access to. ![New dbt Cloud homepage with improved navigation.](https://cdn.sanity.io/images/wl0ndo6t/main/885cb17e9746966a9f395e1b9ad1f8327e21fc30-1168x679.png) ⚙️** Environment-level permissions: **You can now enable developers to edit and trigger jobs in specific environments types (such as development or staging), while restricting them from accessing others (production) — giving you more granular control over permissions. More [info in the docs](https://docs.getdbt.com/docs/cloud/manage-access/environment-permissions). ☑️ **Multi-factor Authentication (MFA):** We've rolled out access to multi-factor Authentication to all dbt Cloud accounts, making it easier to enhance the security of your accounts with an extra layer of protection. See [our docs for more](https://docs.getdbt.com/docs/cloud/manage-access/mfa). 👩‍💻 **IAM User authentication** is now generally available for Amazon Redshift, giving you the option to further enhance account security. Check out the [docs](https://docs.getdbt.com/docs/cloud/connect-data-platform/connect-redshift-postgresql-alloydb#authentication-parameters) for more; also note that this will require you to be using “[Versionless](https://docs.getdbt.com/docs/dbt-versions/versionless-cloud)” in dbt Cloud. 👾 **PATs** (Personal Access Tokens) used for dbt Cloud’s Azure DevOps integration have been **reduced in scope**. The integration uses the service user to to generate a **limited PAT,** that now has only the following scopes: `code_full`, `project` and `build_execute`. These PATs are valid only for 5 minutes, and become invalid after a single API call. See our [ADO docs](https://docs.getdbt.com/docs/cloud/git/setup-azure#service-users-permissions) for more. ✨ **Polaris Catalog integration partner:** dbt is proud to be an integration partner for the Polaris Catalog launch! Using the catalog, teams can read and write from engines to a single catalog with centralized access controls. Read [Snowflake’s launch blog](https://www.snowflake.com/blog/polaris-catalog-open-source/) to learn more. ## See you in Vegas! As always, we’re looking forward to your feedback and hope that these features help improve your analytics workflows and collaboration across your company. We can't believe we're T minus ~60 days until Coalesce 2024 (😱!) and we CANNOT WAIT to be together with you all in Vegas (or online!) and share some of the exciting new features we’ve been cooking up! If you haven’t registered yet ~~(what are you even doing)~~, you can do that [here](https://coalesce.getdbt.com/register). Hope to see you all real soon! --- --- title: "dbt Labs on dbt: Streamlining KPI dashboards with the dbt Semantic Layer" description: "Learn how dbt Labs uses the dbt Semantic Layer to streamline KPI dashboards, ensuring real-time, accurate metrics across BI tools." url: "https://www.getdbt.com/blog/streamlining-kpi-dashboards-dbt-semantic-layer" date: "2024-08-02" authors: ["Paige Berry"] categories: ["Product"] --- # dbt Labs on dbt: Streamlining KPI dashboards with the dbt Semantic Layer Last year, I was Sad Paige, adrift in the Sea of WTF. I was constantly fielding questions about discrepancies in our fundamental business metrics. Each source had different numbers, causing endless confusion and stress, especially when high-level stakeholders relied on me for accurate data. Sleepless nights and tears (not the good kind) were common. Then, in August, we implemented the dbt Semantic Layer and transformed our company scorecard metrics. This change led to a single source of truth and made me Happy Paige. My goal is to help you achieve this “after” experience by sharing our learnings from the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) implementation. ## Our data team Our centralized data team at dbt Labs includes data engineers, analytics engineers, and data analysts. Together with folks from Finance and Revenue Operations, we form pods that support different business areas such as Product, Go-to-Market, and Corporate and G&A. Our complex data stack has 50 different active sources feeding into our data warehouse. Last year, the lack of a single source of truth for our most important metrics made our work challenging. ## The ARR reporting challenge One major challenge was reporting our performance on company-wide OKRs at each month’s All Hands meeting. Several OKRs were tied directly to Annual Recurring Revenue (ARR). Preparing these reports meant dealing with multiple sources: - A Looker dashboard, which was outdated - A Salesforce report from the Revenue Org - A Google sheet maintained by Finance - A Hex dashboard using a dbt model developed by Data and Finance These sources rarely agreed, leading to a frustrating and time-consuming process of reconciling numbers. Preparing for these meetings took days, and the discussions during the meetings often re-litigated the ARR numbers. ## The solution: dbt Semantic Layer Working with my Finance stakeholder, we developed a source-of-truth dashboard called the Business Overview. With executive buy-in, we put the metrics from this dashboard into the dbt Semantic Layer, providing consistent, reliable metrics from a single source of truth. Implementing the dbt Semantic Layer was a game-changer. With dbt Semantic Layer, users define metrics and dimensions in a YAML configuration file, which is then interpreted by dbt to generate SQL code that can be executed against a data warehouse. This approach decouples the definition of metrics from their implementation, allowing changes to be made in one place and propagated throughout all consuming systems automatically. dbt Semantic Layer integrates seamlessly with dbt's existing model system, allowing metrics to be defined on top of transformed data models. This ensures that all metrics are based on clean, well-modeled data, reducing the risk of errors and inconsistencies. **** ## Lessons learned ### Project background We have a large, mature dbt project with many denormalized marts. If you're earlier in your dbt journey, you might skip building wide marts and use normalized models in dbt Semantic Layer for more flexibility. ### Naming and organization - **Naming semantic models:** Prefix with sem_ to differentiate from dbt models. - **Naming metrics:** Categorize by entity/noun for better organization and understanding. - **Organizing models:** Follow a directory structure similar to existing dbt models for discoverability. ### Granularity and flexibility - Avoid over-aggregating data to maintain flexibility for adding dimensions later. - Use dbt Semantic Layer features like cumulative metrics, percentiles, and non-additive dimensions for complex logic. ## The results After launching the Semantic Layer, creating OKR performance slides became a delight. I had more time for data exploration and insights. The discussions in meetings shifted from questioning data accuracy to focusing on strategic actions. ### Quantitative benefits - **Time savings:** 20 hours per month creating OKR slides and 12 hours answering ad-hoc ARR questions. - **Automation:** Changes in the SL propagate automatically, saving time and reducing errors. - **Quality and trust:** Improved focus on data model issues and better ROI from observability and alerting. ### Qualitative benefits - **Trust and motivation:** Effort is spent on the right metrics, protecting motivation. - **Commitment to code:** Updates in the SL are visible and impactful, encouraging careful work. - **Alignment and legitimacy:** Facilitates agreement on metric definitions and emphasizes the importance of data work. ## Roadmap and future possibilities Implementing the dbt Semantic Layer brought significant improvements in our data workflow, transforming frustrating tasks into productive and enjoyable activities. As we look to the future, the applications of the dbt Semantic Layer extend beyond simple reporting. dbt Semantic Layer's API, backed by GraphQL, opens up possibilities for building custom data products and operationalizing analytics, integrating seamlessly with various systems, including CRM platforms like Salesforce. This flexibility makes it an invaluable tool for supporting machine learning processes and enhancing AI workflows by providing clean, well-defined data. We're excited about the possibilities that lie ahead and invite you to join us in exploring the full capabilities of dbt Labs and the dbt Semantic Layer. Whether you're a data professional looking to streamline your workflows or a business leader seeking to harness the power of reliable data, dbt Labs has the tools and expertise to support your goals. Transform your data workflow and become your company's version of Happy Paige. Discover the power of the dbt Semantic Layer and see how it can streamline your metrics, save you time, and reduce stress. **[Watch our on-demand webinar](https://www.getdbt.com/resources/webinars/dbt-labs-on-dbt-streamlining-kpi-dashboards-with-the-dbt-semantic-layer) and take the first step towards making your data team happy and efficient**. --- --- title: "Developing a modern data strategy for AI" description: "Discover how Generative AI is reshaping data engineering roles and the key principles for a successful GenAI data strategy." url: "https://www.getdbt.com/blog/modern-data-strategy-ai" date: "2024-08-01" authors: ["Drew Banin"] categories: ["Insights"] --- # Developing a modern data strategy for AI More companies than ever have big plans for Generative AI (GenAI) projects. However, the success of these projects depends on high-quality and well-governed data. That’s leading to a shift in how data engineers, analytics engineers, and others across the enterprise engage with data. Let's talk about current GenAI efforts, how it's changing the jobs of data and analytics engineers, and what I think are some of the main principles behind a successful GenAI data effort. ## The evolving role of the data engineer Data from TDWI provides some insight into who’s using GenAI currently - and how. Around 50% of the companies TDWI surveyed aren’t using GenAI currently - they’re still focused on self-service reporting and analytics. Of the other 50%, 27% are applying Machine Learning and Natural Language Processing to enhance their analytics. Another 23% are on the cusp of [predictive analytics](https://cloud.google.com/learn/what-is-predictive-analytics?hl=en) - i.e., using current data to forecast future outcomes. TDWI’s data also shows a lot of excitement around GenAI - i.e., using training data from foundation models (like ChatGPT) combined with a company’s own data sets to create new output. Teams are looking to GenAI for a range of use cases. Top examples include customer support chatbots, generating marketing content, writing code (a la [GitHub Copilot](https://github.com/features/copilot)), and onboarding new employees. To achieve this, companies are using more than one data platform. Many are combining traditional relational DBMSes and data pipelines with data warehouses (cloud and on-prem), [data lakes](https://docs.getdbt.com/terms/data-lake), and [data lakehouses](https://www.getdbt.com/blog/analyticsengineeringwithdbtanddatabricks). This push to GenAI, combined with the growth of data in general, has led to a lot of changes in the data-facing engineering roles. Along with data engineers, data scientists and MLOps engineers have joined the fray to assist in building out AI frameworks and the associated data pipelines. Meanwhile, [the analytics engineer has risen](https://www.getdbt.com/what-is-analytics-engineering) to assist in creating clean data sets for end users, combining data transformation and coding skills with domain-level knowledge of a business’s data. ## The role of the data engineer - now and in the future In recent years—thanks largely to increased specialization in the data space—data engineers have been able to shift from focusing on cranking out data transformation pipelines. Instead, they’ve focused more on creating the self-service infrastructure that other roles - analytics engineers, MLOps engineers, business analysts, etc. - can use to source data, create reports, and answer their own data-related questions. However, with the explosion in GenAI, we’re seeing a bit of a reversion. Many data engineers surveyed by TDWI report say they spend a lot of their time getting data into date warehouses, migrating data into data lakes, integrating with vendor data, or working on BI and analytics projects for their end users. I think we’re in a transition phase where data engineers are being asked to get data into places where it can be used, for example, as context retrieved through mechanisms such as [Retrieval Augmented Generation (RAG)](https://aws.amazon.com/what-is/retrieval-augmented-generation/). Over time, we’ll see data engineers more focused on providing AI-oriented data in a self-service fashion to data consumers, rather than being focused on building data pipelines. A lot of work in RAG right now involves curation of structured data. Most of this data transformation work continues to happen in SQL and Python. Given the ubiquity of SQL as a data access and transformation language, there’s a lot you can do to expose this data to data consumers in a way that’s easier for those without a deep knowledge of GenAI to access. For example, say that the business has a need to extract sentiment from a set of survey responses. Data engineers can build SQL user-defined functions in platforms like Databricks and Snowflake that expose sentiment analysis functions via a SQL query. Analytics engineers can then use these functions in their queries to drive reports and apps that are useful to the business. Another way I see the data engineering role evolving is in enabling the move of GenAI apps from prototypes to production. With the currently evolving crop of tools, it’s pretty easy to throw together a toy demo of how you expect a GenAI app to work. But there’s a huge gulf between that and getting the app in front of millions of users with the performance and reliability you need. I don’t see this as a major shift for data engineers in terms of their existing skillset, though. There’s a lot we can do as an industry to de-mystify some of the concepts around GenAI. People use concepts like “in-context learning” which sound ominous and foreboding at first - but just end up meaning something simple (i.e., “add the context you retrieved as text to your LLM prompt”). A lot of what might first sound complex and hard to grok in GenAI just ends up being machine learning in a trenchcoat. More than ever, we’ll need the data engineer’s grounding in fundamental information retrieval techniques to help architect our GenAI apps so they’re production-ready. ## The role of data products in enabling collaboration and governance When we talk to customers about GenAI, we get one of two reactions. Half of our customers ask us pointedly, “Where is GenAI on your roadmap?” The other half want us to sign a contract promising we will never, under any circumstances, turn on any GenAI for their datasets. This is a phenomenon we’ve seen time and again with data. We saw it when cloud data storage first became popular and companies worried about shared compute and storage models. It’ll take some time for some companies—and their security teams—to become okay with the notion of GenAI features operating over their data. That leads us to the larger questions of data governance and data collaboration with GenAI. Having data engineers, data scientists, MLOps engineers, and analytics engineers all in the mix creates several challenges, including: - Enabling these roles to collaborate seamlessly—particularly, enabling downstream engineers to consume the work of the data team - Providing a way for data consumers (MLOps engineers, data scientists, etc.) to find existing, high-quality data sets so they don’t reinvent the wheel - Preventing compliance issues unique to AI—e.g., enabling an LLM to make inferences based on non-deterministic attributes such as race, gender, or age I don’t feel, however, that most of these challenges are particularly unique to GenAI. To be sure, there are some issues of which we need to be acutely aware. Data engineers need to remain very careful about what data they supply to GenAI engines or RAG databases. In this sense, good data governance techniques - anonymization, strong and consistent data labeling and classification, etc.—are more important than ever. It’s hard work. It’s time-consuming. And it’s absolutely essential. However, in another sense, there’s nothing new here. The best way to get great results from GenAI is the same way you get great results from any data-driven project: produce high-quality data. This means going back to your data sources and ensuring that your data is sound. Do you know where it’s coming from? Is it reliable? It also means having robust data quality tests and metrics to ensure your data meets your criteria for accuracy, completeness, consistency, and timeliness. One way to ensure both quality and collaboration is through [data products](https://www.getdbt.com/blog/build-data-product). With data products, data engineers can treat a data set like a software engineering team would a software release, defining versioned, documented data contracts for each new iteration. Teams can then discover and use these data sets - e.g., one team can use the data set for BI or analytics, while another uses it for Machine Learning. I see data products playing an essential role in enabling data producers and consumers to work together. That’s why we’ve built support for data products directly into dbt via features like [data contracts](https://docs.getdbt.com/docs/collaborate/govern/model-contracts) and why we support data discovery and lineage via features such as [dbt Explorer](https://www.getdbt.com/product/dbt-explorer). Like I said, there’s nothing new here. These are solved problems. We don’t need to resolve them for GenAI—we just have to apply the lessons we’ve already learned. GenAI is changing the way that we interact with data and, along with it, the role of the data engineer. The good news is that most data engineers already have the skills and tools to help companies manage this evolution. With high-quality data, a focus on governance, and a collaboration mindset, companies can create a strong data foundation on which they can build the next generation of GenAI apps. **Get started on developing a modern data strategy for AI with this [on-demand webinar](https://tdwi.org/Webcasts/2024/07/ARCH-ALL-Developing-a-Modern-Data-Strategy-for-AI-Evolving-Roles-and-Practices-Quick.aspx).** --- --- title: "July dbt Community update" description: "Stay updated with the dbt Community. In July we had an AMA with Tristan Handy, 11 Meetups, a community spotlight, and more." url: "https://www.getdbt.com/blog/july-dbt-community-update" date: "2024-08-01" authors: ["Kathryn Chubb"] categories: ["Community"] --- # July dbt Community update Welcome to the dbt Community Update, a monthly blog about everything happening in the [dbt Community](https://www.getdbt.com/community). This month we hosted an AMA with our founder and CEO Tristan Handy. We also presented the Summer 2024 [dbt Community Spotlight](https://docs.getdbt.com/community/spotlight), had 11 in-person dbt [Meetups,](https://www.meetup.com/pro/dbt/) and facilitated a ton of great discussions on our Slack channel. Are you ready for the recap? Let’s get started. ## Community Slack AMA Each month we host a live Ask Me Anything event. This month, Tristan Handy, Founder & CEO of dbt Labs, hosted the AMA. He discussed the future of analytics engineering, the impact of AI, and the evolving role of dbt Labs in the data transformation space. Here’s a recap if you missed it. You can also check out the [full video recording](https://www.getdbt.com/resources/community-slack-ama/confirmation). ### The future of analytics and AI integration Tristan dived into how AI is expected to play a bigger role in analytics, but he was clear that AI isn’t going to replace us anytime soon. He said, “I believe that the Semantic Layer is an interactive governed analytic experience. As long as the flow is large language model to Semantic Layer to data, you can trust that the data you're getting back out is correct.” In other words, AI can help us work smarter, but it’s still up to us to make the big decisions. He also talked about how AI could be a game-changer in sorting through hypotheses, helping us zero in on the most promising ones. “What I would love is not just something that could help me move faster, but that could evaluate a hundred different hypotheses and say, here are the five that I think you should probably spend your time looking into,” he added. ### The growth and evolution of dbt Labs Tristan gave us a behind-the-scenes look at how dbt Labs has evolved from a small side project to a powerhouse in the data transformation world. Today, dbt Labs is used by over 25,000 companies, including big names like JetBlue and HubSpot. Tristan shared, “dbt Labs started out as an organization building a project on the side, it was open-sourced. And then there were so many people that used it. They were like, we literally cannot pay for enough software engineers to work on this thing to support all these users.” This overwhelming demand led them to create a commercial business to keep up with the growth and continue investing in the open-source community. His vision is clear: make dbt Labs profitable so they can keep pouring resources into the project and the community. ### Embracing competition With the rise of new competitors in the data transformation space, dbt Labs remains focused on maintaining its position as the standard tool in the modern data stack. He explained, “There have been a lot of competing takes on dbt over the past eight years since we've been building this thing. We're incredibly mindful of when new approaches to data transformation come out, being analytical about what they bring that's advantageous over dbt." It’s all about continuously improving their tools and investing in the dbt community to ensure they remain the go-to solution in the industry. ### Additional insights Tristan also highlighted some interesting trends during the session. One of them is the growing number of people within organizations who are getting involved in the dbt workflow, which is pretty exciting. Additionally, there’s been a noticeable increase in demand for Power BI integration among customers, reflecting the changing needs of the industry. As Tristan put it, “I think that customers do want to buy fewer things. And they want those things to integrate together.” This sentiment captures the importance of seamless integration in today’s data-driven world. ### Get ready for another AMA in August [Join us for the next AMA](https://www.getdbt.com/resources/webinars/community-ama) on August 28th at 5pm EDT with Joel Labes, a Senior Developer Experience Advocate on the DX team. We’ll be tackling topics such as starting a data team and dbt project from scratch, exploring the new compare changes functionality in dbt Cloud, and more. Register now to get the link to watch live and join the [#dbt-community-merge channel in Slack](https://getdbt.slack.com/archives/C025ZN1L679/p1716308109072499) to participate in the conversation! ## Summer 2024 dbt Community Spotlight Every quarter, we highlight community members in the dbt Community Spotlight. These are individuals who have gone above and beyond to contribute to the community in a variety of ways. We're excited to present the Summer 2024 dbt Community Spotlight! This round, we are featuring [Meagan Palmer](https://docs.getdbt.com/community/spotlight/mikko-sulonen), [Mikko Sulonen](https://docs.getdbt.com/community/spotlight/mikko-sulonen), and [Radovan Baćović](https://docs.getdbt.com/community/spotlight/radovan-bacovic). Visit the [Community Spotlight](https://docs.getdbt.com/community/spotlight) page to learn about their backgrounds, their plans to grow as leaders, and their experiences—both learning from others and sharing their own knowledge. ### Meagan Palmer ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/ace2b45cd2660390078f20de1ef9d6c0c7faa064-1736x1284.png) ### Mikko Sulonen ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/ac1d10468931e897b9461a4c1167059d2279defe-1736x1284.png) ### Radovan Baćović ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/4e8136dce179b54a6d9275849795594e29eddd8b-1736x1284.png) If you’re interested in being selected for future rounds of the Community Spotlight, learn more about [becoming a contributor](https://docs.getdbt.com/community/contribute). ## July dbt Meetups In July we had 11 dbt Meetups: Tokyo, Ho Chi Minh City, Taipei, San Francisco, Minneapolis, Boston, Seoul, Lagos, London, Dallas, and Dubai. The dbt Meetup in our San Francisco office was part of a new Community Event Series, dbt Labs on dbt. Two dbt Labs employees, Lauren Benezra, a senior analytics engineer, and Paige Berry, a senior data analyst, discussed how they use dbt to collaborate. They talked about how dbt Cloud has changed their workflow, the game-changing features they can't live without, and how they use dbt Mesh to streamline complex projects. They also shared their tips for debugging data issues and why they believe in the power of collaboration and self-service in data teams. [Check out the blog](https://www.getdbt.com/blog/analytics-engineer-data-analyst-dbt-cloud) to see the full discussion. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/6d45a13d6a295689e2d139b4279344350b9b4343-3024x4032.jpg) The July 4th dbt Labs on dbt Meetup took place in Tokyo. Andrew Escay, a Lead Analytics Engineer at dbt Labs tuned it remotely. Shinya Takimoto hosted a live audience from the dbt Community in Tokyo at DATUM STUDIO. They met to stream the presentation, ask questions, and mingle and network with each other. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/62dada17a451d9870c85229e6f64165657da8d87-4080x3072.jpg) ## Community announcements We’ll wrap up this month's update with some of the exciting announcements that are regularly posted in our [#announcements](https://getdbt.slack.com/archives/C0VLZM3U2/p1715777392876319) channel on Slack. ### The dbt Mesh course has been released You’ll learn how to boost data product reliability and development speed, at any scale, without sacrificing governance. This is a great chance to learn more about multi-project enablement and model governance features like model contracts, model versions, groups, and access modifiers. The course takes approximately 2.5 hours to complete. [Take the free course now.](https://learn.getdbt.com/courses/dbt-mesh) ### Upcoming events - August 28th - [Register for our next Community AMA](https://www.getdbt.com/resources/webinars/community-ama) featuring Joel Labes, Senior Developer Experience Advocate at dbt Labs - Ongoing [Cloud Demo with Experts](https://www.getdbt.com/resources/dbt-cloud-demos-with-experts/) in North America, EMEA, and APAC-friendly times - October 7-10th, 2024, [Coalesce](https://coalesce.getdbt.com/register-2024), by dbt Labs in Las Vegas and Online - Add [dbt Events](https://www.addevent.com/calendar/Tb314369) to your calendar! ### Upcoming dbt Meetups We’ve got a busy couple of months coming up with 8 [in-person dbt Meetups](https://www.meetup.com/pro/dbt) scheduled. If you’re looking for opportunities to learn with fellow members of the dbt Community, and have fun while doing so, join us at one of the sessions listed below: - 🇹🇼 Taipei | Wednesday, August 21st, organized by community members [Karen Hsieh](https://www.linkedin.com/in/karenhsieh/), [Laurence Chen](https://www.linkedin.com/in/humorless/), [Allen Wang](https://www.linkedin.com/in/allenwangs/) - 🇨🇦 Halifax | Wednesday, August 21st, organized by community members - 🇹🇼 Taipei | Wednesday, August 28th, organized by community members [Karen Hsieh](https://www.linkedin.com/in/karenhsieh/), [Laurence Chen](https://www.linkedin.com/in/humorless/), [Allen Wang](https://www.linkedin.com/in/allenwangs/) - 🇦🇺 Brisbane | Thursday, August 29th, organized by [INTELLIGEN](https://www.linkedin.com/company/intelligenau/) - 🇸🇬 Singapore | Thursday, August 29th, organized by [Jing Yu Lim](https://www.linkedin.com/in/limjingyu/), [Michael Han](https://www.linkedin.com/in/michael-ichin-han/) and [Albert Cheng](https://www.linkedin.com/in/albert-cheng-31711828/) (dbt Labs) - 🇨🇴 Medellín | Wednesday, September 4th, organized by [Factored](https://www.linkedin.com/company/factoredai/) - 🇧🇪 Belgium | Thursday, September 5th, organized by [dataroots](https://www.linkedin.com/company/dataroots/) - 🗽 New York | Tuesday, September 10th, organized by [Velir](https://www.linkedin.com/company/velir/) - 🇺🇸 Chicago | Thursday, September 12th, organized by [Analytics8 | Data & Analytics Consultancy](https://www.linkedin.com/company/analytics8/) - 🇺🇸 Phoenix | Thursday, September 19th, organized by [Slalom Consulting](https://www.linkedin.com/company/slalom-consulting/) - 🇳🇴 Oslo | Tuesday, September 24th, organized by community member Anders Elton There are so many exciting things going on in the dbt Community, and we can’t wait to see you all there. If you haven’t yet, [join the community](https://www.getdbt.com/community) today. --- --- title: "The dbt Cloud CLI is now generally available, with support for Power User for dbt" description: "The dbt Cloud CLI, the best way to develop locally with dbt, is now generally available." url: "https://www.getdbt.com/blog/dbt-cloud-cli-generally-available" date: "2024-07-26" authors: ["Azzam Aijazi", "Greg McKeon"] categories: ["Product"] --- # The dbt Cloud CLI is now generally available, with support for Power User for dbt _At our recent [Product Launch Showcase](https://www.getdbt.com/blog/dbt-cloud-launch-showcase-2024), we announced that the [dbt Cloud CLI](https://docs.getdbt.com/docs/cloud/cloud-cli-installation) is now **generally available** for all users of dbt Cloud, and now includes support for the popular VS Code extension [Power User for dbt](https://marketplace.visualstudio.com/items?itemName=innoverio.vscode-dbt-power-user), built by Altimate AI._ Here’s the thing with data transformation: it’s most effective when it’s not siloed. At successful data organizations, it’s a collaborative, transparent process: any team member with enough _business context_ can get directly involved. That helps eliminate bottlenecks, improve data quality, and enable faster development cycles. That’s why dbt Cloud supports multiple development experiences. The people closest to the business context may have a broad range of skillsets; they should all be able to work however they’re most comfortable, while contributing to the same central knowledge base. For more technical contributors — particularly those with a background in data engineering — the ability to use familiar, carefully configured tools to develop is a huge productivity unlock. The recently introduced dbt Cloud CLI (now generally available) is a command-line experience for such developers. It allows you to develop using any terminal or IDE of your choosing, backed by the power of dbt Cloud. You can use the dbt Cloud CLI to contribute to the same dbt projects as coworkers who may be using the in-browser [Cloud IDE](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/develop-in-the-cloud), or the new low-code visual editor (currently in beta). With dbt Cloud you get _multiple pathways_ to contribute, all powered by a single, central platform. ## Building a better developer experience When we announced the Cloud CLI back in October, developers flocked to try it and found that, with the power of dbt Cloud, [it provides a **superior local development experience** to dbt Core](https://www.getdbt.com/blog/a-closer-look-at-the-newly-launched-dbt-cloud-cli). With dbt Cloud CLI, you can: - **Easily save time and money.** With [deferral](https://docs.getdbt.com/docs/cloud/about-cloud-develop-defer) in the dbt Cloud CLI, you can build, run and test only models you’ve edited without having to first build the models upstream of them, significantly reducing resource consumption. - **Build for scale.** Build and seamlessly reference models across multiple inter-connected dbt projects using a [dbt Mesh architecture](https://www.getdbt.com/product/dbt-mesh). - **Enjoy improved performance**, particularly for CPU-intensive commands like dbt parse, and for larger dbt projects. Along with improved artifact download speeds, local development with the dbt Cloud CLI now offers >30% faster parse performance than dbt Core running on an M1 Mac. - (New) **Never worry about dbt version upgrades.** Once you [enable versionless upgrades in dbt Cloud](https://docs.getdbt.com/docs/dbt-versions/versionless-cloud), vetted new features and fixes will be delivered to your projects continuously. - (New) **Connect to your [Azure environments](https://docs.getdbt.com/docs/cloud/about-cloud/tenancy),** ensuring your team can operate with velocity regardless of your deployment architecture. - (New) **Enjoy** [**parallel command execution**](https://docs.getdbt.com/reference/dbt-commands#parallel-execution), allowing for multiple concurrent invocations at once. From the outset, we designed the Cloud CLI to be highly extensible, downloading the same artifacts and producing the same outputs as dbt Core does locally to ensure compatibility, making it easier to port workloads to dbt Cloud without disruption. Support for one familiar tool was missing, though: a rich VS Code extension, which was a request we heard frequently from customers. Extensions let third-party developers enhance VS Code’s native capabilities, and there are great ones out there for dbt. Above all, the [Power User for dbt](https://marketplace.visualstudio.com/items?itemName=innoverio.vscode-dbt-power-user) extension is much beloved by the dbt Community. So we called up [Altimate AI](https://www.altimate.ai/), the folks behind it, and got to work together. ## Introducing Power User for dbt support for the dbt Cloud CLI We’re pleased to say that the dbt Cloud CLI’s integration with Power User for dbt is now available to use. It’s one of the most feature-packed ways to build locally with dbt. Today, the extension includes: - All of the features of the local Power User for dbt extension, including query previews, smart-complete of code, and and accelerated development through AI agents - One-click, managed installation and configuration of Cloud CLI - All the capabilities of the dbt Cloud CLI mentioned above — deferral, dbt Mesh support, faster parse performance, versionless upgrades and more. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/5a247767c6d39dcc109ac844a5197bc2f130543b-2992x1646.png) Users can also access Altimate AI’s AI development agents via the DataPilot platform, which further accelerates and streamlines data development work. We’re excited to partner with Altimate to combine the powers of dbt Cloud and Power user for dbt. In just the couple of months since beta availability, hundreds of developers have used the Power User extension with dbt Cloud. If you’re interested in giving it a try, check out Altimate AI’s [documentation](https://docs.myaltimate.com/setup/reqdConfigCloud/) and get started now. ## Support for linting out of the box Alongside VS Code extension support, many users have been looking for a way to ensure linting parity locally for their dbt Mesh projects. To that end, the Cloud CLI now supports linting with SQLFluff via the `dbt sqlfluff` command — allowing you to ensure code consistency across your projects, prevent errors, and increase developer productivity. This is the same SQLFluff you know and love, now available across the Cloud IDE and Cloud CLI. ## **What’s next for dbt Cloud and the Power User for dbt integration?** We’re continuing to iterate and improve the local dbt Cloud development experience every week. In addition to improving speed and performance, we’re particularly focused on making the local development experience work more seamlessly with the rest of dbt Cloud, including: - Letting users define and invoke production jobs from the Cloud CLI directly. - A deeper integration with dbt Explorer. - More material on the popular Power User for dbt and dbt Cloud CLI integration, [including an upcoming talk at Coalesce](https://coalesce.getdbt.com/agenda)! Register now and catch us there. Have questions for us? Give our team a shout in the [#dbt-cloud-cli-and-ide](https://getdbt.slack.com/archives/C03SAHKKG2Z) channel in the dbt Community Slack. Happy CLI-ing! --- --- title: "The path to a mature Analytics Development Lifecycle" description: "A look at recent trends in data and analytics, the evolving AI landscape, and the importance of the ADLC." url: "https://www.getdbt.com/blog/the-analytics-development-lifecycle" date: "2024-07-21" authors: ["Tristan Handy"] categories: ["Insights"] --- # The path to a mature Analytics Development Lifecycle _This post first appeared in the [Analytics Engineering Roundup](https://roundup.getdbt.com/p/the-analytics-development-lifecycle)._ It’s been a fun, albeit _busy_, month since I got home from several weeks spent in the bowels of Moscone for Summit Season. Honestly, I’m having more fun than I’ve had in years. I wonder if the same is true of you? After four years in data that were defined by massive shifts in the macroeconomic landscape (first up, then down), it feels like the macro has settled. Private and public markets have rationalized, which means that we can mostly stop thinking (and talking!) about them. I ran a business from 2016-2019 and thought about macroeconomics basically never; I’m hopeful that we’re moving back to that world. It’s much more fun to think about data than about interest rates. Of course, it helps that good things are going on in the data ecosystem. I’m particularly excited by the embrace of open table formats and OSS catalogs. I [wrote about this](https://roundup.getdbt.com/p/summit-season) a few weeks ago. I think this trend is going to reshape quite a lot about our world in the coming 2-3-4 years. It will take some time but I think we’re at a point where a multi-data-cloud world is a foregone conclusion for companies of sufficient scale / complexity. (Expect me to talk about this at my [Coalesce keynote](https://coalesce.getdbt.com/)!) Finally: we’re starting to come out of the most extreme part of the hype cycle with AI. I’m a believer in AI, but I’m also a realist. New technology primitives take time to figure out how to deploy against ROI-positive use cases, and we’re still early in that process. For the last 12-18 months, the level of attention AI was getting in conversations with senior data leaders was out of sync with its level of real business value. That divergence is closing, which is good for my personal sanity. I sit writing this while my kids are at the pool, in the depths of summer on the east coast of the US, with some lofi jazz in my ears. So maybe I’m just in a good headspace personally! I think, though, it might be a great moment to be in data. Good things on the horizon, a lot of important work to do, and the space to do it. Can’t ask for more. ## The Analytics Development Lifecycle (ADLC) I’m working on one of my most important pieces of writing in several years. I’m still in the early phases, but it’s starting to take some shape. Over the coming months I’ll likely share bits and pieces of it here as a feedback mechanism. I would LOVE to hear your thoughts. For now, I want to share two specific sections. The first one explains what the paper is about and is from the intro: > In 2016, I authored a blog post entitled “[Building a Mature Analytics Workflow](https://www.getdbt.com/blog/building-a-mature-analytics-workflow).” That post helped launch a community and a product, and many of the assertions from that original post have been realized in the industry. However, eight years in, the original post is in need of an update. > > In this white paper, I propose a single, end-to-end model that I call the **Analytics Development Lifecycle** (ADLC). The ADLC is, I propose, the best path to building a mature analytics capability within an organization of any scale. **** Most of the paper will be about defining the ADLC. It’s not about technology, the ADLC is a _workflow_, and is one that leading data practitioners have continued to refine over the past decade since the advent of cloud-based data technology and devops-style tooling. One of the sections that I’ve had the most fun writing so far is the section about what expectations _users_ should have of a mature analytical system. User meaning: someone exploring data, consuming dashboards, etc., not someone building or maintaining the system itself. Here’s my current draft of this section: The ADLC does not express an opinion on exactly how analytical systems are used to generate business value. For example, it does not believe that there is a ‘right’ or a ‘wrong’ way to conduct exploratory data analysis. Rather, it specifies a set of assumptions that all users of analytical systems should have: - Users should be able to **discover and directly interact with** the artifacts from a mature analytical system without having to go through any intermediary humans. - Users should be able to **trust the correctness and timeliness** of data from a mature analytical system. - Users should be able to **delegate their own access** to a mature analytical system to their chosen tools and agents. - Users should be able to straightforwardly **investigate the provenance** of any data element in a mature analytical system. - Users should be able to **view a history** of all state changes to a mature analytical system. - Users should be able to **leave feedback** on any element of a mature analytical system. - Users should be able to **ignore the implementation details** of a mature analytical system. - Users should be **responsible for the costs** associated with their usage of a mature analytical system. - Users should be able to **use as many resources** as they are willing to pay for from a mature analytical system. - Users should be able to **choose the environment** of a mature analytical system they interact with. There might be a couple of items on this list that feel controversial, but I don’t honestly think these are groundbreaking statements. What is shocking to me, though, is just how infrequently the analytical systems deployed inside each of our companies live up to these requirements. Sometimes I think folks who have been in the data game for a few years now feel like the ecosystem is now ‘mature’ and that most of the important things have been built. I think that is wrong. Yes, we’ve come a long way. But **we have a long way yet to go**. When I stare at the above list it energizes me. There’s a lot of work to do. Hope you’re well. I’d love to hear from you, and would especially love to see you in person at [Coalesce](https://coalesce.getdbt.com/) in October! --- --- title: "Why you need a data control plane" description: "Data quality, data velocity, and cost overruns plague most data-driven organizations. A data control plane can fix that." url: "https://www.getdbt.com/blog/data-control-plane-why" date: "2024-07-20" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Why you need a data control plane Despite recent advancements in analytics—AI, cloud, self-service—we’ve still yet to solve core data problems plaguing our teams: data quality, data velocity, and keeping costs in check in the process. One reason is that today’s modern data ecosystem has created data and compute silos that need to be centralized in a data control plane so organizations can build the holistic context needed to confidently embrace data at scale. dbt Cloud is the data control plane that centralizes your analytics workflow metadata and makes it actionable, so your teams can ship and use trusted data, faster. Here’s how. ## The problem: Data doesn’t scale The cloud and AI have made data more accessible than ever. Organizations are prioritizing ways to translate raw data into trusted insights. However, as your data usage scales, the stakes get higher to ensure that data is accurate, timely, and well-governed. Without a standardized approach (control plane) to managing data at scale, your organization will face a few self-perpetuating problems: ### Data quality issues abound #### No visibility/ lineage You need a way to visualize data dependencies and trace lineage as data moves from source to model to metric. Without this, you’re flying blind. Your teams end up shipping analytics code into production that accrues data analytics debt exponentially. Stakeholders also lack the signals they need to trust that data is fresh and accurate. #### No testing or version control Without a built-in ability to test analytics code, teams push code into production and simply hope for the best. Or perhaps they wait for an angry stakeholder to surface an issue. Without version control, there’s no way to roll production code back to its prior state while a data engineer investigates the issue. #### Slow debugging Finding the root cause of an issue is toilsome without the ability to trace lineage (including column-level). #### No documentation Without a standardized way to document analytics code metadata (freshness, owner, column names, etc.), there’s no continuity as new or other team members try to understand lineage, what code had been built, etc. #### No support for [mesh architectures](https://www.getdbt.com/blog/what-is-data-mesh) Bringing governance logic and rules to disparate business domains helps improve and enforce data quality at scale. But it’s not enough for your organization to adopt a mesh architecture. The services that sit on top of it need to as well. ### Data pipelines get bogged down #### Starting from scratch Undocumented data analytics code offers zero visibility into what’s been built before and how various models interconnect. As a result, many teams end up building new solutions from scratch. They create stored procedures in a database, locally run analytics code, or other unmanaged and ungoverned solutions. #### Lack of automation AI co-pilots have proven effective at accelerating the productivity of knowledge workers. Without native AI copilots, data practitioners must conduct manual, tedious work by hand. #### No CI/CD In software engineering, [Continuous Integration and Continuous Deployment (CI/CD) pipelines](https://docs.getdbt.com/guides/custom-cicd-pipelines?step=1) deploy changes to production through stages that subject them to rigorous automated and manual verification procedures. Sadly, most analytics code still doesn’t follow this approach. Without a safe or automated way to build and merge analytics code into production, deploying changes to models becomes very manual and doesn’t scale. #### Limited number of developers Even if they had SQL skills, not enough data practitioners (analysts, for example) had the deep expertise to work with legacy transformation tools. This led to over-burdened data engineering teams who should have been spending more time on data architecture and platforming. #### No way to promote self-serve in a governed, scalable way Allowing end users to tap into trusted data models and metrics is the holy grail. However, without a way to centralize business logic and without first-class integrations into BI platforms and LLMs, you risk end-users accessing incorrect data. That further sabotages trust in data and data teams. ### Costs spiral out of control #### Vendor lock-in It’s more important than ever to avoid lock-in with any one vendor or approach.** **With tense dynamics across cloud and data platform vendors, most organizations don’t want to be price-gouged. #### Inefficient compute Getting data from its source to its consumer involves many hops between various models and tables. It’s easy to inadvertently and unnecessarily drive compute costs up without the proper context or controls into how these jobs are run. #### Optimization requires context Data estates are large, complex, and dynamic. You need intuitive ways of understanding how efficiently your models are performing and which models are most (and least) utilized so that you can take this information and assign resources to fine-tune and optimize your pipelines. ## The solution: A data control plane The solution to these problems: a data control plane powered by dbt Cloud. Business runs on data. And data runs on dbt. For years, customers have trusted dbt as their one-stop shop for transforming data into high-quality and accurate data sets. Now, you can leverage dbt Cloud to deliver high-quality data, faster, and at the lowest cost possible. dbt Cloud is your control plane for data: - It’s natively interoperable across various cloud and data platforms so you’re never locked-in - Its platform features support data developers and their stakeholders across various stages of the analytics development lifecycle to make data analytics a team sport - It provides the trust signals and observability features required to ensure all data outputs are accurate, governed, and trustworthy. With dbt Cloud as your data control plane, your data teams have a standardized and cost-efficient way to build, test, deploy, and discover analytics code. Meanwhile, data consumers have purpose-built interfaces and integrations to self-serve data that is governed and actionable. ## Do more with data With a data control plane powered by dbt Cloud, your organization can transform how it does data. With dbt Cloud, you can: ### Reliably deliver high-quality data to the business #### Testing and version control Improve the integrity of the SQL in each model by making assertions about the model logic ([unit tests](https://docs.getdbt.com/docs/build/data-tests)) or the expected results generated by a model. Embrace [version control](https://docs.getdbt.com/docs/collaborate/git/version-control-basics) with deep git integrations to track all code changes across dev, staging, and prod. Work collaboratively, safely, and simultaneously on a single project. Safely roll back to prior states when issues arise. #### Spot and fix issues quickly Monitor your [dbt jobs](https://docs.getdbt.com/docs/deploy/jobs) and set up proactive alerts to keep pipelines smooth. Use [column-level lineage](https://www.getdbt.com/blog/guide-to-data-lineage) to trace dependencies and debug issues quickly. Build trust with downstream teams by embedding health status tiles in analytics tools so everyone is aligned on data freshness and quality checks. Use audit logs to understand and troubleshoot user and system events quickly. #### Automated lineage and documentation Auto-generate the [documentation](https://docs.getdbt.com/docs/build/documentation) when your dbt project runs. dbt provides a mechanism to write, version-control, and share documentation for your dbt models. You can write descriptions (in plain text or markdown) for each model and field and navigate your entire detailed lineage in [dbt Explorer](https://docs.getdbt.com/docs/collaborate/explore-projects). #### Governance with guardrails dbt Cloud supports data mesh architectures with [dbt Mesh](https://docs.getdbt.com/best-practices/how-we-mesh/mesh-1-intro). Divide projects into data domains to reduce complexity and maintain data governance. [Role-based access controls](https://docs.getdbt.com/docs/cloud/manage-access/about-user-access) make it easy to configure who has access to what data. ### Build, test, and deploy data faster #### Reuse, don’t rebuild Create reusable (modular) [data models](https://docs.getdbt.com/docs/build/models) that can be referenced in future work instead of starting at the raw data with every analysis. This “DRY” (don’t repeat yourself) approach to code makes it maintainable by people other than yourself and scalable as the system load increases. Reuse models across projects to accelerate data delivery without unnecessary complexity. #### Automated end-to-end lineage Get an automated and holistic view of how data moves through your organization—where it comes from, how it’s transformed, and who consumes it. With this [visual graph](https://www.getdbt.com/blog/guide-to-dag) and data catalog, developers can build, troubleshoot, and analyze data workflows more efficiently and accelerate cross-org data literacy. #### Standardize and accelerate data development workflows Use [embedded AI-copilot experiences](https://www.getdbt.com/blog/introducing-dbt-copilot) to generate SQL on demand, and auto-generate tests, documentation, and metrics. Lean on custom rules for SQL formatting to standardize and optimize code development and trust those guidelines are natively enforced. #### Built-in CI/CD Run [CI jobs](https://docs.getdbt.com/docs/deploy/ci-jobs) to test your code before it’s merged to production to ensure it’s behaving as expected and won’t break anything downstream. Trigger jobs to run when a pull request (PR) is merged and defer to production to optimize compute cycles and velocity. Do all of this from the flow of your development and git environment. #### Democratize data development Get more data-literate stakeholders to participate in data development. With dbt, they don’t need to write boilerplate DML and DDL by managing transactions, dropping tables, and managing schema changes. They can write business logic with just a SQL select statement, or a Python DataFrame, that returns the dataset they need. dbt takes care of materialization. The new [Visual Editor](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/ide-user-interface) democratizes data development to _even less technical_ users with a visual drag-and-drop interface that’s all powered by version-controlled, governed SQL under the hood. #### Semantic layer integrations foster self-serve With [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), data teams can define business logic centrally, alongside their dbt models, and ship them to any endpoint…whether a BI tool, an embedded application, or LLM. This means that the people who rely on data to make decisions have it readily available at their fingertips in self-serve, accessible interfaces and everyone is confident that that data is accurate, governed, and consistent. ### Optimize data platform costs free from lock-in #### Flexible interoperability between cloud vendors With rapidly changing market dynamics, it’s more important than ever that teams don’t get locked-in to any one vendor or approach. There’s no need to hardcode logic at the platform level. dbt Cloud is an abstraction layer that’s interoperable across a variety of cloud data platforms. Use dbt Cloud to: - Enable cross-department mesh - Support centralized governance for complex projects running on multiple data platforms - Dynamically distribute workloads across platforms dbt Cloud provides unparalleled flexibility fueled by SQL and backed by a passionate open core community. #### Inject intelligence into your data builds Reduce data platform spend by only building models that have changed with [defer to production](https://docs.getdbt.com/docs/cloud/about-cloud-develop-defer). dbt Cloud supports auto-cancellation for stale CI builds. You can also use unit tests to validate model logic before the model is materialized. #### Quickly identify and resolve performance bottlenecks Pinpoint long-running or often-failing models and quickly identify opportunities to reduce infrastructure costs and save data team time. Discover and improve popular models, and divest from unused models. ## Conclusion dbt Cloud is your control plane for data. But it’s not just for data developers. With its purpose-built user interfaces and native integrations, stakeholders of _all_ technical stripes can participate in the data workflow and have the trusted insights needed to translate data into strategic decisions. Learn more about how dbt Cloud can deliver data at speed and scale—[contact us for a demo today](https://www.getdbt.com/contact). --- --- title: "Data governance for AI: What’s different?" description: "AI needs high-quality data—and that requires modern governance. Learn how AI reshapes data governance and how dbt can help." url: "https://www.getdbt.com/blog/understanding-data-governance-ai" date: "2024-07-19" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data governance for AI: What’s different? AI runs on data. Good AI runs on _good_ data. Artificial intelligence systems ingest massive amounts of training data to uncover patterns and correlations. These patterns become the basis for AI-generated predictions and decisions. That means the ultimate performance—and risk—of any AI model is directly tied to the quality of the data it’s trained on. If the data is inaccurate, incomplete, biased, or inconsistent, the model’s output can be unreliable—or even harmful. That’s why data governance is foundational to effective AI. Governance ensures that the data feeding your models is trustworthy, well-managed, and aligned with compliance and ethical standards. While data governance is essential for any data-driven business, it’s even more critical when AI is in play. In this article, we’ll explore the role of data governance in AI, how AI can strengthen governance practices, and why mastering both is a strategic advantage. ## What is data governance? Data governance is the framework an organization uses to manage the availability, integrity, security, and usability of its data. It defines and enforces policies, standards, roles, and responsibilities that govern how data is collected, stored, accessed, and used across the business. In traditional settings, data governance ensures trustworthy analytics, regulatory compliance, and operational efficiency. But when AI enters the picture, the stakes get higher. The quality of an AI system’s performance is directly tied to the quality of the data it’s trained on. Without well-governed data, AI can produce biased, inaccurate, or even harmful outcomes. **Data governance in AI is focused on managing the data used by AI systems: the quality, integrity, security, and usability of an organization’s data throughout its lifecycle. **AI data governance concentrates on the operational and technical aspects of data handling to produce clean and reliable data by creating standards and practices for data collection, storage, and use that prevent data mishandling and misuse. ## Overview: data governance vs. AI data governance AI data governance builds on traditional governance principles—but adds new dimensions of complexity. It’s not just about data management; it’s about managing the integrity of AI systems themselves. Here’s how the two approaches compare: AI data governance is a subset of broader AI governance. While traditional governance ensures clean, compliant data, AI data governance ensures responsible model behavior—tackling issues like: - Bias mitigation - Transparency and auditability - Explainability of model decisions - Continuous monitoring of outputs The goals are the same—trustworthy, usable data—but the challenges and responsibilities are new. ## Why data governance is essential for AI AI models are only as good as the data they’re trained on. If your input data is biased, incomplete, or inconsistent, the outputs will be too. That’s why strong data governance is foundational to building trustworthy, high-performing AI. Effective governance secures the full data lifecycle—ensuring that inputs don’t compromise the fairness, accuracy, or reliability of AI outputs. It also sets standards that help your organization meet ethical and regulatory obligations. Key responsibilities of AI data governance include: - **Bias mitigation: **Governance ensures that training data is diverse, representative, and actively monitored for bias—reducing the risk of skewed or discriminatory outcomes. - **Model transparency and accountability: **Clear governance policies support explainability by tracking data lineage and metadata. This enables users to understand how AI decisions are made and who's responsible for them. It's especially critical in regulated industries like healthcare, finance, and law. - **Accuracy and reliability: **Governance enforces data quality standards—ensuring that training data is accurate, complete, and consistent. This prevents misleading or unstable AI behavior. - **Ethical decision-making: **Sound governance surfaces potential ethical issues early, helping teams mitigate unintended consequences in areas like recruitment, lending, or insurance. - **Regulatory compliance: **With regulations like the EU AI Act emerging, governance frameworks help ensure that AI systems meet data privacy and compliance requirements—reducing [legal and financial risk](https://www.investopedia.com/terms/c/compliance-cost.asp). - **Data security and AI-specific threats: **Governance practices include safeguards to prevent unauthorized access, data breaches, and attacks specific to AI systems (like [prompt injection](https://www.nightfall.ai/ai-security-101/prompt-injection) or model inversion). - **Enabling ethical AI:** Ultimately, governance provides the structure to build and deploy AI systems that are not only high-performing—but also responsible, transparent, and just. ## How AI strengthens data governance AI isn’t just the reason we need better data governance—it can also help us achieve it. Machine learning and natural language processing offer distinct advantages: they work continuously, scale effortlessly, and surface issues faster than manual review ever could. Here’s how AI enhances key aspects of data governance: - **Improving data quality:** AI can detect anomalies, flag outdated records, and resolve inconsistencies across datasets—ensuring your data stays clean and reliable at scale. - **Automating quality control: **[By leveraging metadata](https://www.getdbt.com/blog/data-leadership-ai-2), AI systems can continuously monitor data for errors, duplications, or structural issues. ML models can even be trained to automatically correct common problems before they reach production. - **Accelerating data discovery and cataloging:** AI-powered tools can scan and classify massive volumes of data, tagging assets for easier search and access. This reduces bottlenecks and democratizes data across your organization. - **Enhacing data security and risk detection: **AI can identify unusual access patterns or flag suspicious behavior in real-time—helping prevent breaches and protect sensitive data. - **Supporting smart discovery with governance in mind:** AI tools can surface valuable insights across datasets while enforcing usage policies and access controls. This enables compliant, self-service exploration of your data. - **Enabling proactive regulatory compliance:** AI can track how data is handled across systems, ensuring your governance program aligns with evolving privacy laws and regulatory requirements. ## Unique challenges of governing data for AI AI governance introduces novel challenges that traditional data governance wasn’t designed to handle: - **Continuous change: **Unlike static datasets, AI systems learn and evolve. This demands continuous oversight—not just of the data, but also of how models behave over time. - **Ethical and societal impact:** AI raises complex questions around fairness, bias, explainability, and unintended consequences. These go far beyond the scope of most legacy governance programs. - **Annotated data quality:** AI models require large volumes of labeled training data. Governance must ensure that annotation processes are accurate, representative, and ethically sourced. - **Regulatory uncertainty: **AI-specific regulations (like the EU AI Act) are new, complex, and still evolving. Organizations need flexible, proactive governance frameworks to stay compliant as standards shift. ## The competitive edge of AI data governance Strong data governance doesn’t just reduce risk—it creates real competitive advantages for AI initiatives: - **Build trust and differentiate:** Transparent, explainable AI builds customer confidence. When users understand how your models make decisions—and trust those decisions—you gain an edge over less-governed competitors. - **Stay compliant and future-proof:** Governance frameworks help you meet emerging AI regulations like the [EU AI Act](https://artificialintelligenceact.eu/), protecting your organization from fines, legal risks, and reputational damage. - **Boost AI performance:** High-quality, well-governed data leads to better-trained models and more accurate outputs. That means faster insights, fewer errors, and smarter automation. - **Accelerate innovation—safely**: Clear governance policies reduce friction and fear around experimenting with AI. Teams can move quickly, knowing guardrails are in place to catch missteps early. - **Shorten development cycles:** Clean, well-documented, and version-controlled data pipelines remove roadblocks and rework, allowing AI teams to build and iterate faster. ## Conclusion: Good AI starts with great data governance Data governance might not be flashy—but it’s foundational. For AI to deliver value, it needs to be built on high-quality, well-managed data. Without it, models can produce flawed outputs, introduce bias, and create serious security and compliance risks. Strong governance doesn’t just protect your organization. It also powers better AI performance, accelerates development, and builds trust with stakeholders. In a fast-moving regulatory landscape, it’s a must-have—not a nice-to-have. AI is only as good as the data it learns from. And great data requires governance from the start. [**dbt**](https://www.getdbt.com/product/dbt) gives you a unified platform to govern all your analytics code and data products—from transformation to testing to documentation—so you can build responsibly, move faster, and trust your data every step of the way. 👉 Learn more about [how dbt supports data governance for AI](https://www.getdbt.com/product/governance) ## FAQs about data governance for AI **What is data governance in AI? ** Data governance in AI is a framework for managing the data used by AI systems across their lifecycle. It sets standards for how data is collected, stored, and used—ensuring that it’s accurate, secure, and ethically sourced. For example, tracking data lineage and metadata helps maintain transparency in how AI decisions are made. **How is AI data governance different from traditional data governance?** Traditional governance is often static and compliance-driven. AI data governance is more dynamic and continuous. Key differences include: AI models evolve over time—governance needs to evolve with them. **Why is data governance important for AI systems?** Poor data = poor predictions. Governance ensures that AI systems are trained on high-quality, representative data. This minimizes bias, increases reliability, supports ethical decision-making, and helps meet regulatory requirements—especially in high-stakes industries like finance, healthcare, and government. **How can AI improve data governance?** AI tools can: - Detect anomalies and inconsistencies at scale - Automate tagging and classification for faster data cataloging - Identify access risks or unusual usage patterns - Help organizations stay compliant by tracking how data is handled AI doesn’t just need good governance—it can power it. **What role does data governance play in mitigating AI bias?** Governance ensures that training data is representative and vetted. By monitoring inputs and auditing outputs, organizations can catch biased results early—and take corrective action. This is critical for building AI systems that treat all users fairly. **How can organizations balance innovation and compliance in AI?** A well-structured governance framework lets teams move fast without breaking things. Define roles, automate checks, and enforce policies that protect sensitive data. That way, teams can safely scale AI use cases without creating risk. **What are the key components of an effective AI data governance strategy?** - Clear standards for data quality and access - Transparent data lineage and metadata tracking - Assigned roles for stewardship and compliance - Processes for auditing and monitoring sensitive data use - Feedback loops for reviewing and improving governance policies **How does AI data governance create competitive advantage?** Organizations with strong data governance move faster and build more trustworthy AI. They spend less time fixing data issues and more time deploying high-impact use cases. Transparent, ethical AI also strengthens brand trust—critical in regulated markets. **What tools support AI data governance?** Platforms like dbt serve as control planes for managing data transformations, quality, documentation, and lineage—all essential components of AI governance. Additional tools for data cataloging, access controls, and monitoring can further enhance your governance stack. --- --- title: "Semantic Layer: What it is and when to adopt it" description: "The explosion of data sources can make it harder for employees to find, trust, and use data. Here’s how a semantic layer can help." url: "https://www.getdbt.com/blog/semantic-layer-introduction" date: "2024-07-16" authors: ["Alexis Jones"] categories: ["Learn"] --- # Semantic Layer: What it is and when to adopt it As data sources and use cases continue to explode within an organization, it’s critical to ensure data quality and consistency. With more stakeholders relying on accurate metrics to power their strategic decisions, companies need to ensure these metrics are actually accurate and consistent across teams, and that requires centralizing business logic in a single source of truth. A semantic layer is one such tool companies can use to ensure metric consistency and quality across the organization. In this article, we’ll look into what a semantic layer is, how it works, some common use cases, and how to get started adopting one. ## What is a semantic layer? A semantic layer is a framework that allows organizations to create a unified, business-friendly representation of their data. In other words, a semantic layer translates data into common language. It serves as a unified interface and repository that enables data teams to build and store metric logic centrally so data consumers can access consistent, high-quality, and governed data across a variety of endpoints. The semantic layer solves a key problem created by the explosion of data. Thanks to technologies such as cloud data warehouses and processes such as [ELT](https://www.getdbt.com/blog/etl-vs-elt), organizations can access, create, and manage more data than ever before. However, this explosion of data and the increase in data stakeholders introduces new challenges. [A Forrester survey in 2021](https://www.forrester.com/blogs/the-bi-fabric-baby-is-slowly-but-surely-growing-up/) found that over 61% of organizations use four or more BI tools, and a staggering 25% use 10 or more. Defining metrics and business logic within each of these disparate interfaces leads to workflow inefficiencies and introduces data quality risk. For example, you may start out with one approach to calculating and representing revenue. However, as time evolves and your approach to calculating this changes, older definitions become outdated. Updating this across multiple endpoints and reports is often unrealistic. The result is that different teams end up with different definitions and understandings of this key concept. This is where the semantic layer comes in. It acts as an API for data, supplying data stakeholders with a single source of truth from which to pull key data sets and calculations. A semantic model provides a “hub-and-spoke” architecture to your metric definitions. Metrics are written and defined centrally and can be queried from a number of “spokes”—analytics tools, APIs, LLMs, and more. Those endpoints are always accessing the same centralized data, so organizations can be sure that they’re working off the same metric everywhere, every time. ## Benefits of a semantic layer Implementing a semantic layer provides multiple benefits across your data stack, including: **Eliminates inconsistencies**. Differing data definitions can lead different teams to make divergent assessments of the same question. And when teams spot the discrepancy, it can lead to long debates about whose view of reality is right—this derails data-driven initiatives and sabotages trust in data teams. A semantic layer eliminates these debates and trust issues, driving consistency of terminology and calculation for key data structures across different projects. **Improves data democratization**. Stakeholders want to be empowered to get their own answers, not to mention, it isn’t scalable for data teams to be hands-on with every possible data request. However, data consumers may not know where to find the data they need or have the skills required to analyze it. Worse yet, they may end up using source data that’s out-of-date or incorrect. A semantic layer enables teams to pull insights from data even if they aren’t comfortable with writing complex SQL data transformations. Additionally, it makes this data available in a self-service format, which reduces support requests placed on data engineers. **Promotes data reusability**. Even when teams do have data engineers and analytics engineers on staff, that doesn’t mean they should spend half their time reinventing the wheel. A semantic model enables a single team to maintain one gold standard data set that other teams can leverage and build upon. This not only improves data consistency, it also streamlines data operations and optimizes costs. **Improves compliance**. A semantic model also helps with increasing compliance. By acting as a central point of access, the model can enforce role-based access controls to data, ensuring that sensitive data is protected and made available only to authorized stakeholders. ## Use cases for a semantic layer The notion of a semantic layer has been around for a while now. How are organizations taking advantage of consistent metrics across the business? We’ve seen five key patterns emerge. ### Reporting and BI Populating BI tools with relevant and accurate data is an obvious and urgent use case for the semantic layer. The use of multiple BI tools and interfaces across an org leads to a host of problems: - **Data maintenance**: Metrics must be manually built and updated across each BI tool as definitions evolve. It’s cumbersome to test and validate changes to logic or diagnose issues in a scalable, proactive manner. - **Data bottlenecks**: Adding or updating business logic across various tools results in slow development cycles. Troubleshooting discrepancies is time-consuming and impairs trust. - **Data trust**: Metrics drive business-critical decisions, and there is no room for error. Mistakes are noticed by C-level stakeholders and impair trust in data and data teams. By utilizing the “hub-and-spoke” architecture, a semantic layer provides, data teams can store semantic models and definitions centrally. They can then deliver_ _those metrics on demand into any downstream tool that queries it—whether from a first-class integrated tool [or via a custom integration](https://www.getdbt.com/product/integrations) powered by an export to your data platform. ### Embedded analytics Data is powerful. It shouldn’t be relegated to an internal KPI scorecard to align internal employees. It can - and should - be embedded into customer- and partner-facing applications to create delightful, personalized user experiences that build brand equity. With a semantic layer as your foundation, you can calculate complex metrics centrally right alongside the rest of your data. You can version control them and deliver them as up-to-date embedded visualizations in your app, website, or wherever your stakeholders consume data. Data teams can build custom web apps using developer-friendly APIs or SDKs and serve up relevant, personalized data to end-users - without driving up costs with legacy BI tools. ### AI and LLMs Everyone wants to embrace AI (or at least have a plan for adopting it). And no one wants bad data undermining those plans. In our most recent [State of Analytics Engineering Report](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024), the majority of respondents noted that they manage data for AI model training (either currently or within the next 12 months). Unsurprisingly, data quality is an undeniable prerequisite for the successful adoption of AI. [A recent Tableau / Salesforce study](https://www.salesforce.com/content/dam/web/en_us/www/documents/research/state-of-data-analytics.pdf) found that 86% of analytics and IT leaders agree that AI's outputs are only as good as its data inputs. Using a semantic layer, you can ensure high-quality outputs across your AI stack that reduce hallucinations, eliminate redundant data transformation work, and lower the barrier to analytics. Using this data, your AI apps can empower less technical users to self-service answers to their questions. ### Self-serve analytics To scalably adopt data-driven practices across an organization, data teams can’t be involved in every data request. Manual human intervention is neither scalable nor efficient. Self-serve analytics allows anyone—even those who aren’t technical—to get the data they need to make strategic decisions. By centralizing metrics with a semantic layer, data teams can minimize ad hoc requests while ensuring high-quality and governed data becomes easily accessible across the organization, whether through a spreadsheet, an AI chatbot, or any other accessible interface. As a result, data velocity, collaboration, and trust improve. That, in turn, further improves your data ROI. ### Exploratory analytics Exploratory analytics is a critical step in the data science workflow. It involves using Python libraries to inspect data, discover patterns, and verify hypotheses that ultimately inform winning business strategies and deliver data ROI. While highly strategic, exploratory analytics can be challenging: - Data sources are constantly expanding - You need the ability to combine data across various sources to get a complete picture of reality - You need the flexibility to continuously iterate your questions and quickly slice metrics across various dimensions to get to the bottom of something …all while ensuring data integrity and quality along the way. Using a semantic layer, data science teams can take advantage of centralized and governed metrics that live alongside other data models, querying and joining them to support iterative exploratory analytics workflows. **** ## Getting started with the semantic layer A semantic layer consists of four elements: - Varied data sources that flow into a central data repository - Data models - Metrics definitions - Endpoints (a BI tool, LMM, embedded site widget, etc.) dbt has long supported defining the data model layer, bringing a new level of velocity, data quality, and democratization to data. Using the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), you can define a common set of metrics that live alongside your data models, creating a unified data interface for all of your data stakeholders. What’s more, with the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), you can codify aggregation types and their underlying calculations, capturing business logic in a central location that can be consumed by downstream tools. This ensures that critical business concepts - like revenue - are defined and documented in one location, where everyone can both consume and maintain them. ![Semantic layer architecture diagram.](https://cdn.sanity.io/images/wl0ndo6t/main/4896fc98a8ec18fb1bd142f45cd2403b0baf136f-1686x1236.png) The dbt Semantic Layer also reduces redundancy by following [the DRY (Don’t Repeat Yourself) principle](https://docs.getdbt.com/terms/dry). Instead of calculating metrics in multiple places, you can define them in one location, alongside your dbt models and tap into them from any endpoint. This reduces duplicative work and code. It also prevents other teams from having to redo this work when they onboard a new downstream tool. Early adopters of the dbt Semantic Layer are taking advantage of consistent metrics across their businesses—and five key use cases are emerging. Check out this blog to learn more about the [five use cases for the dbt Semantic Layer.](https://www.getdbt.com/blog/five-use-cases-for-the-dbt-semantic-layer) --- --- title: "Five use cases for the dbt Semantic Layer" description: "How to use your organizational metrics to build lasting competitive advantage." url: "https://www.getdbt.com/blog/five-use-cases-for-the-dbt-semantic-layer" date: "2024-07-15" authors: ["Nick Handel", "Alexis Jones"] categories: ["Product"] --- # Five use cases for the dbt Semantic Layer We’ve written extensively about the [benefits of the dbt Semantic Layer](https://www.getdbt.com/blog/build-centralize-and-deliver-consistent-metrics-with-the-dbt-semantic-layer), [its technical architecture](https://www.getdbt.com/blog/how-the-dbt-semantic-layer-works), and our [integration partners](https://www.getdbt.com/blog/managing-data-source-changes-in-tableau)…but this is our first time formally documenting the various patterns for _the use cases for semantic models throughout a company._ But first, a quick level set on the dbt Semantic layer: Data teams use the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) to build metric definitions and store them centrally in code, right alongside their dbt models, and make them available across the enterprise through a consistent, tool-agnostic interface. It provides a “hub-and-spoke” architecture to your metric definitions: metrics are written and defined centrally (the "hub"), and can be queried from a number of “spokes”—analytics tools, APIs, LLMs, and more. Those end-points are always accessing the same centralized data, so organizations can be sure that they’re working off the same metric everywhere, every time. ![Architecture diagram for the dbt Semantic Layer, including common end points.](https://cdn.sanity.io/images/wl0ndo6t/main/4896fc98a8ec18fb1bd142f45cd2403b0baf136f-1686x1236.png) What’s more, metrics can also easily be [exported as tables back to your warehouse](https://docs.getdbt.com/docs/use-dbt-semantic-layer/exports) so that you have centralized, consistent, DRY business logic that’s traceable across your DAG and readily available as a critical building block to deliver on numerous use cases. - A note on exports ## Five use cases for the dbt Semantic Layer The data you capture and transform is a means to an end. So how are early adopters of the dbt Semantic Layer taking advantage of consistent metrics across the business? We’ve seen 5 key use cases emerge: reporting & BI, embedded analytics, AI and LLM integrations, self-serve analytics, and exploratory analytics. ![Visual representation of the five use cases for the dbt Semantic layer: reporting & BI, embedded analytics, AI/LLM integrations, self-sere analytics, exploratory analytics](https://cdn.sanity.io/images/wl0ndo6t/main/be951f3c7de681ab64247ea90b5e167ae3038684-1876x560.png) Let's dive in! ## Reporting & BI: Deliver accurate data to your boss Populating BI tools with relevant and accurate data is an obvious, and urgent, use case for the semantic layer. According to a [Forrester survey](https://www.forrester.com/blogs/the-bi-fabric-baby-is-slowly-but-surely-growing-up/), on average, a single organization utilizes four or more BI tools (and 25% of organizations use 10 or more!). When metric logic is housed within those individual tools, a few problems crop up: - Data delivery bottlenecks: - Metrics must be manually built within each BI tool, and as definitions evolve, they must be updated across each tool. This is undifferentiated heavy lifting that slows development time and introduces unnecessary data quality risk. - With dispersed metric logic, it’s cumbersome to test and validate changes to logic or diagnose issues in a scalable, proactive manner. - Impaired data trust: - Metrics drive business-critical decisions, and there is no room for error. Mistakes are noticed by C-level stakeholders and impair trust in data and data teams. - Discrepancies in metric outputs can derail cross-functional meetings, where energy is spent debating data accuracy and not on decision making. Together, these issues ultimately impair data quality, can damage trust in data and data teams, and sabotage organizational efforts to embrace data-driven practices. By utilizing the “hub-and-spoke” architecture a semantic layer provides, data teams can store semantic models and definitions _centrally,_ and those metrics can be delivered on demand into any downstream tool that queries it—whether from a [first-class integration with your analytics tool](https://docs.getdbt.com/docs/cloud-integrations/avail-sl-integrations), or via a custom integration powered by an [export](https://docs.getdbt.com/docs/use-dbt-semantic-layer/exports) to your data platform. That means that when the finance team uses Tableau to surface last month’s ARR for a board presentation, they’re seeing the same exact number as the Marketing team that uses PowerBI to analyze last month’s ARR by lead source. This consistency fosters trust in data, helps teams make decisions faster, and frees up data development cycles as teams no longer need to manually spelunk the root cause of a discrepancy. What’s more, teams can explore the metadata associated with a particular metric, like data lineage, data freshness, definitions, and joins, so they are empowered with the information they need to always be 100% confident in the metrics they rely on. > _“When you put everything on dbt, you ensure everyone is seeing the same number. You don’t get that message saying, ‘oh, my director got this GMV number and I’m getting this different one.’” ** > - Gabriel Marinho, Lead Analytics Engineer at Inventa**_ ## Embedded analytics: Power delightful in-app experiences Data is powerful. And so, it shouldn’t be relegated to an internal KPI scorecard to align internal employees; it can and should be embedded into customer- and partner-facing applications as real-time, personalized data is the bedrock of building delightful and differentiated customer experiences. A [recent report](https://media.thoughtspot.com/pdf/ThoughtSpot-PLA-Embedded-Analytics-Guide-2024.pdf) by Thoughtspot and Product Led Alliances found that companies that embed analytics with a differentiated user experience see increased user engagement and revenue. But there are few considerations to take into account before embracing an embedded strategy: 1. **Data integrity:** Once metrics are going outside the walls of your company, the stakes get higher. Unlike with an internal stakeholder, you can’t sort out discrepancies with a quick Slack message or Zoom alignment call: it’s CRITICAL that the data is correct. Otherwise, you risk impairing brand equity, losing customer trust, compromising revenue growth opportunities, and in some cases, breaking regulatory commitments. 2. **Costs:** Embedding dashboards from traditional BI tools can be prohibitively costly, especially when leaning on tools with seat-based pricing. Meanwhile, relying on internal engineering teams to stand up a solution requires tradeoffs on other product deliverables, not to mention a maintenance burden. 3. **Flexibility:** Leaning on a BI tool for embedded analytics limits a developer’s ability to build a customized and seamlessly branded product UI—and delivering a delightful customer experience is the driving force behind an embedded strategy in the first place. Teams need the consistency and control afforded by custom APIs. With the dbt Semantic Layer as your foundation, you can calculate complex metrics centrally—right alongside the rest of your data—version control them, and deliver them as up-to-date embedded visualizations in your app, website, or wherever your stakeholders consume data. Data teams can build custom web apps using developer-friendly APIs or SDKs and serve up relevant, personalized data to end-users…without driving up costs with legacy BI tools. And with built-in caching and performance-optimized SQL, you can always be sure those metrics are delivered lightning fast. What’s more, your teams can build confidence in the data that’s being served to your customers with the ability to visualize dependencies from source all the way to the metric. > _"The dbt Semantic Layer gives our data teams a scalable way to provide accurate, governed data that can be accessed in a variety of ways—an API call, a low-code query builder in a spreadsheet, or automatically embedded in a personalized in-app experience. Centralizing our metrics in dbt gives our data teams a ton of control and flexibility to define and disseminate data, and our business users and customers are happy to have the data they need, when and where they need it.” ** > - Hans Nelsen, Chief Data Officer, Brightside Health**_ ## AI & LLMs: Turn your questions into governed, high quality answers Everyone wants to embrace AI (or at least have a _plan_ for adopting it!)…and no one wants to have bad data. In our most recent [State of Analytics Engineering Report](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024), the majority of respondents noted that they manage data for AI model training (either currently or within the next 12 months). Unsurprisingly, data quality is an undeniable pre-requisite for the successful adoption of AI: [a recent Tableau / Salesforce study](https://www.salesforce.com/content/dam/web/en_us/www/documents/research/state-of-data-analytics.pdf) found that 86% of analytics and IT leaders agree that AI's outputs are only as good as its data inputs. Using the dbt Semantic Layer, you can set up your AI investments for success. By enriching your LLMs with high-context, well-governed data inputs, you can ensure high quality outputs across your AI stack: - **Reduce hallucinations:** Get high-quality AI responses, backed by real data which keeps your models on track, and enriched by business context - **Build once, use everywhere:** Define your business logic via metrics in the semantic layer, and access it through any connected LLM - **Lower the barrier to analytics:** Democratize data-driven decision making by empowering less technical users to self-service the answers to their questions using an agent that utilizes natural language. As an example of what can be accomplished by powering your LLM with governed semantic definitions, dbt Labs built an agent to interact with the dbt Semantic Layer using plain language text: Ask dbt. Unlike traditional AI chatbots, Ask dbt uses the dbt Semantic Layer to provide critical context about your dbt project, improving accuracy by 3x as observed in our [benchmark](https://roundup.getdbt.com/p/semantic-layer-as-the-data-interface). With Ask dbt, users can ask questions in natural language and receive insights in an understandable format, which can significantly speed up business processes and decision-making. Make decisions quickly with the confidence that the metrics you’re using are always consistent—regardless if you got them through Ask dbt, your AI chatbot, Tableau, Google Sheets, or anywhere else. ![Ask dbt chatbot in Snowflake native app](https://cdn.sanity.io/images/wl0ndo6t/main/262eaa2f581e862c463d4a2c6d242661e1f0410f-1600x1054.jpg) Ask dbt is currently in public beta for customers using the [dbt Snowflake Native App](https://app.snowflake.com/marketplace/listing/GZTYZSRT2R3/dbt-labs-dbt?search=dbt). ## Self-serve analytics: Democratize data where people already are (hint: that’s probably a spreadsheet) To scalably adopt data-driven practices across an organization, data teams can’t be involved in every data request. Manual human intervention isn’t scalable, nor is it efficient. Self-serve analytics makes it possible for anyone —~~even~~ especially those that aren’t technical—to get the data they need to make strategic decisions. And that means they need to access data from the tools they already use. Most likely, that’s a spreadsheet, but could be any interface they already gravitate towards. From there, users can tap into a no-code interface to view dashboards and reports or slide and dice data themselves to get reliable answers quickly. Predicated on the success of any self-serve strategy is that the data needs to be both **accessible** and **accurate**. dbt offers built-in integrations across a variety of data consumer end-points: - **BI & analytics tools:** Business users can jump into their BI tool of choice—including Tableau, Hex, and Google Sheets— build their own queries via intuitive drag-and-drop interfaces, or leverage the saved queries that their colleagues have used to expedite data discovery. - **LLMs:** Less technical users can ask natural language questions such as, “what was our revenue in EMEA last quarter?” and get a high-context, governed, consistent answer back. The dbt Semantic Layer currently offers a native integration with Snowflake Cortex via our Ask dbt feature. The data, metrics, and reports flowing into these end-points are sourced centrally and are fully governed, tested, and version controlled. This gives data teams the peace-of-mind to feel comfortable removing themselves from what otherwise would have been a ticket for a new data pull and an immediately outdated CSV file. dbt also streamlines authentication workflows with granular access controls that can be configured for individual users or groups. ![Query governed metrics from Google Sheets](https://cdn.sanity.io/images/wl0ndo6t/main/a4b212d6c3b260e5d881edcc16b0df12edf04741-3800x1662.png) By centralizing metrics on the dbt Semantic Layer, data teams can minimize ad hoc requests while ensuring high quality and governed data becomes easily accessible across the organization. As a result, data velocity, collaboration, and trust improve which further supports data ROI. ## Exploratory analytics: Surface new insights that drive competitive edge Exploratory analytics is a critical step in the data science workflow that involves using Python libraries to inspect data, discover patterns, and verify hypotheses that ultimately inform winning business strategies and deliver data ROI. For example, a data scientist may want to analyze trends for a particular metric over time and calculate summary statistics to understand historical data and better predict the future. While highly strategic, exploratory analytics can be challenging: data sources are constantly expanding, you need the ability to combine data across various sources to get a complete picture of reality, and you need the flexibility to continuously iterate your questions and quickly slice metrics across various dimensions to get to the bottom of something…all while ensuring data integrity and quality along the way. The inherent expansiveness and fluidity of this practice can introduce productivity challenges, and so it’s essential that data teams have the tooling and resources to make the iterative process of exploratory analytics as efficient and outcome-oriented as possible. We’ve all heard the phrase “garbage, in garbage out,” and the truism also applies when training ML models. If data scientists are building models starting from governed metrics definitions, they can be confident that the datasets they’re using accurately represents the business metric they are trying to improve. Without this assurance, all of the data modeling they conduct is moot. [Watch video](https://getdbt.wistia.com/medias/1mlx564bvf) Using the dbt Semantic Layer, data science teams can take advantage of centralized and governed metrics—that live alongside other data models—and query and join them to support exploratory analytics workflows. All data metrics and models are version controlled, lineage can be explored from source all the way through to metric, and teams can easily join semantic models with other dbt models. We offer native integrations to a number of workbooks and exploratory analytics tools (Hex, Mode), data can be exported back to your warehouse as centralized, governed metrics tables to then deliver downstream, or data teams can query data via our[ JDBC](https://docs.getdbt.com/docs/dbt-cloud-apis/sl-jdbc) or [GraphQL](https://docs.getdbt.com/docs/dbt-cloud-apis/sl-graphql) APIs in their notebook of choice. **** ## Get started To get started building your own semantic models and metrics, check out our [documentation](https://docs.getdbt.com/docs/use-dbt-semantic-layer/quickstart-sl), [schedule a call](https://www.getdbt.com/contact) with on of our product experts, or you can reach out in Community Slack (#dbt-cloud-semantic-layer). --- --- title: "What is a data control plane?" description: "Data quality, data ownership, and stakeholder literacy are pressing data problems. A data control plane solves them." url: "https://www.getdbt.com/blog/data-control-plane-introduction" date: "2024-07-12" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # What is a data control plane? Analytics engineering has come a long way in the past decade. Gone are the days when we fretted over how to build pipelines or whether we had access to or enough compute to handle our workloads. Cloud data platforms and the ecosystem that's emerged around them have helped make data analytics accessible to organizations large and small. ‌Despite this progress, today, we [face a whole new set of challenges](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024)—specifically, data quality, data literacy, and ambiguous ownership. At dbt, we’ve spent a lot of time thinking about how to tackle these modern problems. For us, the answer lies in the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) coupled with a [data control plane](https://www.getdbt.com/blog/data-control-plane-introduction). The ADLC is a vendor-agnostic process that helps teams develop mature analytics workflows using techniques similar to those in modern software engineering processes. The data control plane is an architectural layer that sits over an end-to-end set of data activities (e.g., integration, access, governance, and protection) to manage and control the holistic behavior of people and processes involved in distributed, diverse, and dynamic data. With [dbt Cloud as your data control plane](https://www.getdbt.com/resources/whitepaper-the-control-plane-for-data-collaboration-at-scale), your data teams have a standardized and cost-efficient way to build, test, deploy, and discover analytics code using the ADLC. dbt Cloud also gives data consumers purpose-built interfaces and integrations to self-serve data that is governed and actionable. We’ll dig into how the ADLC and the data control plane work together—and how to use dbt Cloud to drive both and improve data quality, promote data literacy, and clarify data ownership. ## Today's major data problems: quality, ownership, literacy In our [2024 State of Analytics Engineering report](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024), we asked over 450 data practitioners many questions about the state of analytics from their point of view. This was one of the most interesting findings from the survey: The question was: What do you find most challenging when preparing data for analysis? If this report were conducted in 2014, it wouldn’t be surprising to see building data transformations and constraints on compute resources cited as the biggest challenges. Remember: that was before cloud data platforms or industry standards for data transformation like dbt. But today, the community’s telling us that data transformation is mostly a solved problem. This is good news for all of us. That means it’s time to tackle the next order of problems. Our community is telling us that as data analytics takes off within their companies, new problems emerge downstream: data quality, data ownership, and stakeholder literacy are now the biggest challenges in the industry. ### Is your engineering team able to solve these problems? Here's a quick litmus test to prove the point: - Do your stakeholders always know where to go to find the right dashboard or the right metric? Do their numbers always agree? - Do upstream changes and data sources ever break your pipelines? - Do you have defined SLAs for the business—and are you hitting them? Many readers will feel a little squirmy when answering questions like this about their data workflows. That’s expected and okay, and just proves that we have more work to do as an industry. We’ve solved some big problems. But we’re not done. We should be proud of our progress, but we're still on a journey. Without clear, confident answers to questions like these, we can't help our business grow revenue, nor can we properly manage costs. And the big initiative that we’re all in service of—how to be a trusted strategic partner to the business—remains perpetually out of reach. ## A paradigm for maturing analytics That raises the question of what the next decade should look like for [maturing data analytics practices](https://www.getdbt.com/blog/analytics-next-era). That question has two answers: the first is to adopt the Analytics Development Lifecycle (ADLC) as a cultural and workflow paradigm, and the second is to choose technology solutions that help you do that successfully. The ADLC is a vendor-agnostic framework for mature analytics workflows. It encourages collaboration among various stakeholders and is designed to help data producers, data consumers, and—ultimately—the business ship and use trusted data products at speed and at scale. ### The ADLC workflow The ADLC promotes eight distinct workflow stages in the analytics development lifecycle: from planning analytics products to building, testing, and deploying them, to operating them in production and ensuring that they're reliable and discoverable. ![DataOps lifecycle infinity loop diagram showing stages: plan, develop, test, deploy, analyze, operate, observe, and discover.](https://cdn.sanity.io/images/wl0ndo6t/main/1fc981ff485ca62fb80c5c9d4bde3789904514f8-2400x1260.jpg) The ADLC borrows heavily from the [Software Development Lifecycle](https://aws.amazon.com/what-is/sdlc/) (“SDLC”) popularized in software engineering in the early 2000s. The SDLC helped cross-functional teams work together with more agility, velocity, accuracy, and, ultimately, business impact. The SDLC sought to erode the bifurcation between the software engineers who built software systems and the IT engineers who operationalized them. Similarly, the ADLC seeks to [erode the bifurcation of roles and responsibilities](https://www.getdbt.com/blog/what-is-dataops) between ‌data builders and ‌data consumers. The goal is to give all roles a standardized, repeatable framework they can use to work better together. It's high time that analytics professionals adopt a similar framework and rely on vendors who will accelerate and harden data workflows across these stages. Our CEO Tristan Handy has said, “We believe that implementing the ADLC is the best path to building a mature analytics practice within an organization of any size.” [We’ve written a paper going into more detail](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) about why—and how—the ADLC achieves this. We encourage you to read it and learn how to bring some of these best practices into your organization. ## The journey to the data control plane It’s important to remember that the ADLC isn’t owned by any one company, nor is it an explicit vendor solution. Rather, it’s a vendor-agnostic process and workflow methodology that promotes more mature analytics practices at any scale. We purport that adopting a data control plane like dbt Cloud is the best way to embrace the ADLC. (At the risk of an unnecessary history lesson) here’s why: A new data stack has emerged since the introduction of cloud-based data platforms over the past decade. The core workflow surrounding those cloud data platforms is the process of bringing data in from varied sources. These can be databases, applications, streaming data, or otherwise. That raw data is then transformed into clean data models. Finally, those data models are pushed to endpoints for AI, BI, or analytics so teams can understand and visualize trends and make informed business decisions with data. ### Workflow maturity had led to siloes dbt has sat in the middle of this workflow since our inception over eight years ago. However, in that time, with the increased adoption of modern data systems and the strategic insights they have delivered to the business, we've seen the emergence of a peripheral data ecosystem to solve next-order problems or optimize this core workflow. And that's because people started asking questions like: - How can we automate these workflows from data sources all the way through to the data consumer? - How can we get visibility into data pipeline performance and health, troubleshoot issues, and optimize velocity and costs in the process? - How can we increase data literacy across the company and improve data visibility and trust? - And how do we build and centralize business logic to ensure that consistent, reliable metrics are powering all reaches of our business? In response, we saw entire categories spring up for orchestration, data observability, data catalogs, and semantic stores to address these market needs. This has all been great progress for the maturity of our industry, helping automate workflows and ensure the freshest, highest-quality data is powering the business. But all of those add-on components have a unique and siloed way of surfacing subject-specific metadata with no centralized way to connect it or take holistic action. At the end of the day, your underlying data platform and pipeline become optimized, reliable, and cost-effective when it has context and awareness of all these varied work streams with metadata fragmented across tools, teams, and platforms. ## What is a data control plane? [A data control plane is an abstraction layer that sits across your data stack](https://www.getdbt.com/blog/data-control-plane-why), unifying capabilities for orchestration, observability, cataloging, semantics, and more. Perhaps more importantly, a data control plane centralizes metadata across your business, giving you a universal view of what's happening in your data estate. The data control plane provides signals to help you understand if your data is fresh and your platform is cost-optimized. You can also use it to verify that everyone is running from a common understanding of how business metrics are defined. ### Data plane vs. control plane vs. data control plane We’re dealing with a lot of similar terminology here. So let’s disentagle some of them. - A data plane handles data movement and processing. It executes queries and transforms datasets. - A control plane manages configurations and policies. It defines access controls and orchestrates workflows. - A data control plane unifies these functions. It centralizes metadata and standardizes workflows. A data plane focuses on execution. It processes, stores, and queries data. A control plane governs resources without handling data. It manages infrastructure and policies. A data control plane creates a centralized hub. It manages data workflows and governance. With this kind of integration, teams can build, test, and monitor analytics together. ## The three main criteria for a data control plane We believe that a data control plane should support and promote three things: 1. It should be **flexible and cross-platform to empower distributed teams**, helping them avoid vendor lock-in and manage data platform costs. 2. It should be **collaborative**. That means it should make data development more accessible, streamlined, and governed to more types of users. 3. Finally, it must **produce trustworthy outputs**. It should allow users to build and automate high-quality data pipelines so that the business can access, understand, and trust the data they receive. Weirdly (or not), those three characteristics map right back to what the market is telling us are the biggest challenges to solve: - The challenge of ambiguous data ownership begets the need to help disparate teams align on a common control plane for accelerating analytics, regardless of the underlying platforms that particular team relies on; - The challenge of poor stakeholder data literacy demands more governed inroads for data collaboration; and - Data quality challenges the imperative that data teams and stakeholders need a streamlined way to build trustworthy data products. Fortunately, you don't need to create this from scratch or with vendor-specific tools. You can build it today with [dbt Cloud](https://www.getdbt.com/product/dbt-cloud). [dbt Cloud is the data control plane](https://www.getdbt.com/blog/advancing-the-data-control-plane-vision-with-sdf-and-dbt) that centralizes your metadata and makes it actionable, so your teams can ship and use trusted data, faster. ‌It’s natively interoperable across various cloud and data platforms, so you’re never locked in. Its platform features support data developers and their stakeholders across various stages of the analytics development lifecycle, turning data analytics into a team sport. And it provides the trust signals and observability features required to ensure all data outputs are accurate, governed, and trustworthy. ## Conclusion A data control plane is a technological solution to the people and process approach laid out with the ADLC. dbt Cloud is a market-leading data control plane designed to help organizations successfully adopt the ADLC. dbt Cloud works across various cloud and data platform environments. It’s accessible to personas of varying technical backgrounds, providing a standardized, unified way to accelerate the Analytics Development Lifecycle. With dbt Cloud as your data control plane, you can: - **Abstract business logic into a flexible platform:** By standardizing on a platform-agnostic control plane, you can stay focused on shipping reliable data products while optimizing spend. - **Standardize on SQL:** All transformations are written in SQL (universal language, get more data people involved in transformation workflows) and dependencies and documentation are automatically built. - **Make data quality a habit: **Proactively prevent data issues with built-in testing and CI. If an issue occurs, find and fix it quickly with alerts and audit logs, roll changes back easily with version control, and use column-level lineage to quickly identify and resolve the root cause. - **Help your teams ship data faster: **Reduce bottlenecks and improve productivity with AI-assisted workflows and automated scheduling and orchestration of your end-to-end data pipelines. To learn more and see dbt Cloud and the data control plane in action, [view our recent webinar](https://www.getdbt.com/resources/webinars/one-dbt-the-control-plane-for-data-collaboration-at-scale-virtual-event). ## FAQs about data control planes **What exactly is a data control plane and how does it differ from traditional data solutions?** A data control plane is an architectural layer that centralizes governance across all data activities. It unifies capabilities for orchestration, observability, and semantics while providing a holistic view of your data landscape. Unlike traditional siloed solutions, a control plane connects fragmented metadata across tools and platforms. This comprehensive approach enables organizations to maintain data quality and literacy while clarifying ownership across distributed teams. **How does the Analytics Development Lifecycle (ADLC) work with a data control plane?** The ADLC provides a vendor-agnostic framework for developing mature analytics workflows across eight distinct stages. A data control plane, like dbt Cloud, serves as the implementation technology that makes ADLC principles actionable. The ADLC promotes collaboration between data producers and consumers through standardized processes. When powered by a data control plane, organizations can ship trusted data products efficiently while breaking down traditional role separations. **What are the most pressing data challenges that a control plane addresses?** Data quality concerns, poor stakeholder literacy, and ambiguous ownership represent today's most significant analytics challenges. A data control plane addresses quality issues through testing and lineage tracking capabilities. It improves literacy by standardizing access to trusted data sources for all business users. The control plane clarifies ownership by providing governance frameworks and metadata visibility, ensuring teams understand data responsibilities and relationships. **What benefits can organizations expect when implementing a data control plane?** Organizations gain platform flexibility, avoiding vendor lock-in while optimizing spend across data environments. Cross-functional collaboration improves as both technical and business users access standardized interfaces. Data quality becomes measurable and actionable through centralized testing and observability features. Ultimately, businesses achieve faster time-to-insight as teams standardize workflows and leverage automation to reduce bottlenecks across the entire analytics process. **How specifically does dbt Cloud function as a data control plane?** dbt Cloud centralizes metadata across platforms while providing actionable insights for data teams. It abstracts business logic into SQL transformations that work seamlessly across multiple data environments. The platform enables proactive data quality through built-in testing, version control, and column-level lineage tracking. dbt Cloud accelerates data delivery through AI-assisted workflows and automated orchestration, making it accessible to technical and non-technical users alike. --- --- title: "Getting started with data quality management" description: "Learn data quality management principles and how to apply them with dbt for reliable, high-quality data." url: "https://www.getdbt.com/blog/getting-started-data-quality-management" date: "2024-07-11" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Getting started with data quality management Poor quality data [has a domino effect](https://www.marshall.edu/irp/2023/10/04/importanceofdataquality/) that negatively impacts decision-making, compliance, operational efficiency, and trust. By contrast, working actively to improve data quality gives stakeholders the confidence to rely on it to make critical decisions and power new business ideas, such as [AI apps](https://www.getdbt.com/blog/data-leadership-ai). Achieving - and maintaining - a high level of data quality requires an active commitment from both data producers and consumers. In this article, we’ll introduce the principles of data quality management, discuss how to implement data quality management in practice, and look at some tools your teams can leverage in their day-to-day data workflows. ## What is data quality management? Data quality management is an organization's set of policies, practices, and methods to ensure its data is accurate, complete, and trustworthy. Ideally, these data quality management standards are a part of a larger [data governance framework](https://www.techrepublic.com/article/data-governance-framework/) that sets criteria for data quality and consistency across the entire company. The outputs of a data quality management process may include artifacts such as: ### Data retention policies How long to hold data and what happens to it after that time period. These policies will take into consideration, not just the business value of the data but compliance with applicable laws and regulations. ### Data formatting standards How to encode and represent data across the organization. This can include general rules (e.g., date formats) as well as guidelines for company-specific data (e.g., uniform customer IDs, acceptable ranges of values for a specific field). Data formatting standards help prevent downstream pipeline breakages and make data synchronization easier. ### [Data tests](https://www.getdbt.com/blog/data-testing) Code that checks a data set for various attributes of data quality. A data test might ensure, for example, that all of the required fields in a record are specified, that individual fields are formatted correctly, and that no anomalies exist between records. ### [Data quality metrics](https://www.getdbt.com/blog/data-quality-metrics) Data quality metrics provide a quantitative measure of data quality, enabling your organization to identify gaps and improve quality over time. Examples of data quality metrics include total number of data incidents, time to data incident detection, time since last data refresh, and number of passed/failed data tests for a table, among others. ## Data quality management in practice ### Data profiling Assesses the structure of your organization’s data and how each table and field relates to the others. This often takes the form of: - A repository of data models describing data sources, data destinations, and data transformation rules. - [Data lineage](https://www.getdbt.com/blog/getting-started-with-data-lineage) is typically represented as a [Directed Acyclic Graph (DAG) ](https://www.getdbt.com/blog/dag-use-cases-and-best-practices)that shows how both tables and columns relate to one another. A data profile gives you a complete picture of the data you own and how it flows throughout your organization. Data engineers can leverage it to identify the root causes of data quality issues and fix them at their source. ### [Data transformation](https://www.getdbt.com/blog/data-transformation) Corrals your data into tables and views that meet the business requirements of data consumers. Transformation also involves cleaning and formatting data to fit your organization and team’s data quality policies as outlined in your data governance framework. [Watch video](https://www.youtube.com/watch?v=FSc8jzdcmiw) ### Data validation Uses automated and manual validation to check data for its overall accuracy and sensibility. (For example, ensuring that an Age field never has a negative value.) Data validation ensures that the work done during the transformation phase is correct and that the data is free of obvious defects. ### [Metadata](https://guides.lib.unc.edu/metadata/definition) development Identifies characteristics such as table owner, date data was last modified, a description of the data and how it was calculated, related data tests, etc. By capturing its current status and business purpose, metadata enables better data discovery and usage. This helps you reduce “[dark data](https://www.getdbt.com/blog/key-components-of-data-mesh-creating-and-managing-data-products),” or unused data because no one can find it or validate its meaning. ### Monitoring and reporting Tools and alert systems that track effectiveness and proactively monitor for errors. This will include data quality metrics dashboards and data usage statistics. Data engineering team members can set up ongoing testing on production data and issue alerts immediately upon detecting a data anomaly. ## Implementing data quality management Implementing data quality management requires a combination of **processes** and **tools**. Processes ensure that all appropriate stakeholders—both technical and business line leaders—are involved in data quality management. In particular, processes need to involve [data domain owners](https://www.getdbt.com/blog/key-components-of-data-mesh-data-domains) to validate that data conforms to business requirements and outcomes. Tools support key elements of the data quality management process, including creating data transformation pipelines and automating data quality standards, testing, and metrics collection and reporting. Both processes and tools are necessary to implement data quality management at scale. ## Analytics Development Lifecycle (ADLC) Approaches such as the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) unite processes and tools into a united framework that enable data producers and consumers to improve data quality via short, rapid development cycles. Using an ADLC approach, companies can identify important data quality use cases and address them through iterative improvements. ![ALDC loop](https://cdn.sanity.io/images/wl0ndo6t/main/948eb5cda47eacf1bf78a2268c0666be43706ce5-4581x2126.png) ## Data quality management life cycle A typical data quality management life cycle will include: ### Plan Technical and business stakeholders work to identify existing data quality issues and prioritize them. For example, the team might identify that duplicate records in sales data are preventing an accurate analysis of purchasing trends. After identifying use cases, the team will establish procedures for reconciling records, preventing duplicates, and creating a data set that accurately reflects the current business reality. The team should also establish metrics to monitor and verify correctness (e.g., less than x% duplicates tested in the final data set, % data test pass rate). ### Develop The data engineering team will then create data transformation pipelines that produce a clean data set in line with data consumer’s requirements. They’ll also create tests to run against both pre-production and production data to verify the quality of the output. ### Test and deploy The data engineering team checks all of its data transformation and testing code into source control, using Pull Requests (PRs) to review data code changes internally before deployment. It also implements a Continuous Integration and Continuous Deployment (CI/CD) process to test data quality management code in a pre-production environment before releasing to production. ### Operate, observe, discover, and analyze Data consumers use the new, clean data set to create reports and data-driven applications. Along the way, they report any identifiable issues back to the data engineering team for fixing. Simultaneously, the data team tracks metrics and alerts to identify potential issues before they result in report or application downtime. All teams involved in the ADLC continue iterating over this cycle, delivering new data quality use cases with every release. ## Automating data quality management with dbt Implementing the processes and tools required for an effective data quality management program takes time. It takes even longer if you have to build all of your tooling and pipelines from scratch. dbt Cloud offers a host of features that significantly reduce the time and effort required to ship high-quality data. These include: - **Transformation**: [Create models](https://docs.getdbt.com/docs/build/models) that import data from multiple sources, cleaning and transforming them into new data sets that are ready to use. - **Documentation**: [Add descriptions ](https://docs.getdbt.com/docs/build/documentation)directly to data models. Publish new documentation automatically with every push to production, providing other data users with detailed information on the origin and meaning of your data. - **Testing**: [Leverage built-in tests (e.g., not-null checks) and create custom tests](https://docs.getdbt.com/docs/build/data-tests) that implement quality tests specific to each data domain. - **Version control integration**. [Check data models, transformation, and tests into source control](https://docs.getdbt.com/docs/collaborate/git-version-control) to ensure all changes are tracked and reviewed. Isolate in-development changes in branches so that data engineers can work freely on new features or fixes without affecting the current state of production. - **Job scheduling and orchestration**: [Regularly run your dbt models and tests](https://docs.getdbt.com/docs/deploy/job-scheduler) to bring data changes into production and continuously perform data quality checks. Unlike other tools, dbt Cloud easily enables automating data imports and testing in a single data pipeline. - [**CI/CD support**](https://docs.getdbt.com/docs/deploy/continuous-integration). Automatically run jobs fro your Git provider based on check-in or completed PRs. Test changes in a pre-production environment before releasing to users. - **Data cataloging**: Data producers and consumers alike can use [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) to find existing data sets and related documentation, as well as trace data lineage to verify the origin of data and troubleshoot upstream data issues. - **Monitoring and dashboards**: [Monitor metrics and fire alerts](https://www.getdbt.com/coalesce-2021/observability-within-dbt) in response to dbt Cloud test failures. Leveraging these tools in dbt Cloud, your team can build a robust data quality management process in a fraction of the time it’d take to build from scratch. Learn more about how dbt Cloud can kickstart your data quality management journey—[contact us today](https://www.getdbt.com/contact) for a demo. **** --- --- title: "DataOps vs DevOps: How do they differ?" description: "DataOps and DevOps are similar—but different. Learn what distinguishes them and how to implement DataOps easily." url: "https://www.getdbt.com/blog/dataops-devops-difference" date: "2024-07-05" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # DataOps vs DevOps: How do they differ? DataOps aims to do for the [Analytics Data Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle#stakeholders-of-the-adlc) what DevOps has done for the Software Development Lifecycle (SDLC). However, while similar, both processes have key differences in their personas and operations. Here’s what they share, how they diverge, and how to implement DatOps effectively. ## What is DevOps? DevOps is short for _Development _and _Operations_. The process was created to tear down a wall that historically existed between software engineers who create software (development) and the IT personnel who deploy, monitor, and maintain new releases (operations). In the past, development would spend months creating new software releases in isolation from operations. At the end of the cycle, they’d take what they built and “throw it over the wall” to the operations team. The result was chaos. Developers would use library versions, operating system features, memory and processing power, etc., with little to zero knowledge of what was running in production. That forced operations to spend weeks in a painful back-and-forth with development. Releases were often delayed and bug-ridden. By contrast, [in a typical DevOps lifecycle](https://aws.amazon.com/devops/what-is-devops/), deployment and release are viewed as two halves of the same release cycle, with all personas in both Dev and Ops working together as one team. Release cycles are shorter—typically one to two weeks—to enable rapid development, testing, and issue resolution. DevOps teams rely heavily on version control and automation to facilitate smooth and high-quality releases. All changes are committed to source control so they can be tracked, reviewed, and rolled back if necessary. The primary vehicle for deployment is a [Continuous Integration and Continuous Delivery (CI/CD) pipeline](https://www.redhat.com/en/topics/devops/what-is-ci-cd), which deploys and tests software changes from a source control check-in before releasing it to production. Everyone on a DevOps team shares responsibility for keeping the DevOps pipeline healthy and happy. ### The DevOps lifecycle For each deployable change, a DevOps lifecycle iterates rapidly over the following stages: ![DevOps lifecycle](https://cdn.sanity.io/images/wl0ndo6t/main/6e46d72cc8e129217deac3078767c9e6961c72f6-792x440.jpg) **Plan**: Decide which features to implement and how to measure success. **Code**: Author the features and all related tests. Obtain a code review from another team member before kicking off the deployment process. **Build**: Assemble the components of the system into a deployable package. **Test**: Run both automated and manual tests in multiple release environments (dev, test, staging, prod) to assess the quality of the change before release. **Release and deploy**: Make the feature available to customers. **Operate and monitor**: Observe metrics and logs to ensure smooth operations, firing a notification or alert if critical system values drop below an acceptable threshold. ### DevOps personas A DevOps lifecycle involves, at a minimum, the following personas: **Developer**. Software engineers responsible for the technical design, development, and testing of software. **IT administrator/system administrator**. Technical personnel responsible for the installation, configuration, monitoring, upkeep, and backup/recovery of technical assets. These include servers, storage (databases, object storage, data lakes, etc.), networks, load balancers, etc. **Systems Reliability Engineer (SRE)**: A role pioneered by Google in the early 2000s, SREs focus on creating features and automated solutions that enhance the reliability of a product. **Business stakeholder**. Business representatives who represent the voice of the customer, providing guidance on which features to develop next. ### Benefits of DevOps Done well, DevOps has numerous benefits. Most of these can be quantified as metrics that organizations can use to track improvements in the SDLC. **Increased deployment frequency**. By focusing on smaller work units, teams can ensure that a new feature works—and works well—before moving on to the next one. The tight cooperation between Dev and Ops enables better coordination on each release, leading to smoother releases with fewer obstacles. **Increased deployment quality**. By leveraging techniques such as automated testing in CI/CD pipelines and deploying to multiple environments (dev, test, stage, prod, etc.), a DevOps team can find and resolve critical issues before release. That results in higher-quality releases and less system downtime. **Increased scale of deployments**. Using automation to improve quality and velocity means that teams can build larger, more complex software systems without drowning themselves in a sea of defects. ## What is DataOps? Like DevOps, which inspired it, DataOps follows a similar approach of integrating Data —data acquisition, transformation, and deployment of new changes—with Operation— observation, metrics and logging, and discovery and analysis. Releasing new data products means merging and transforming data, usually from multiple sources. So, as in DevOps, DataOps uses version control to track changes to data transformation code and CI/CD automation to ship these changes to production. ### The DataOps lifecycle In DataOps, all personas who work with data work on the same team at each stage of the Analytics Development Lifecycle. The ADLC resembles the SDLC with a few minor changes: ![ADLC loop](https://cdn.sanity.io/images/wl0ndo6t/main/1fc981ff485ca62fb80c5c9d4bde3789904514f8-2400x1260.jpg) **Plan**: Decide which new data products to create or how to change an existing data product (e.g., adding a new field to a table) and define your KPIs and success factors. **Develop**: Create data pipelines and [data transformation models](https://docs.getdbt.com/docs/build/models) to create the new data set from one or more trusted sources. Submit model changes for code review. **Test**: Write and run unit, data, and integration tests that verify your transformations are running correctly. **Deploy**: Move the changes from development through production using [an automated CI/CD process](https://docs.getdbt.com/guides/custom-cicd-pipelines?step=1). **Operate** **and Observe**: Ensure data changes remain in a steady state by testing data in production and recovering quickly from failure to maintain always-on access to data. Ensure access to data is controlled by role-based access control (RBAC) and that sensitive data—e.g., customer’s Personally Identifiable Information (PII)—is restricted and audited. **Discover and Analyze**: Enable data stakeholders to find data products and standardized metrics so they can use them to answer questions and drive business decisions. ### DataOps personas Just as the process of DataOps differs from DevOps, so do the personas. The following personas aren’t fixed roles but, rather, hats that multiple people can wear at different times. **The engineer**: Creates reusable data assets—pipelines, models, metrics, etc. **The analyst**: Performs analysis on data sets that drive business decisions. **The decision-maker**: Takes the output from the engineer and the analyst and translates them into actions for the business. ### Benefits of DataOps DataOps has many of the same benefits as DevOps. It also adds a few additional benefits: **Focuses data teams on business outcomes**. Historically, data engineers and other technical roles have driven the definition of new data products. That’s led to new features being driven more by technical capabilities than by the needs of the business. By involving analysts and decision-makers throughout the ADLC, DataOps aligns all data changes with business KPIs and OKRs. **Democratizes access to data**. Up to 75% of data in a company may be “[dark data](https://www.ibm.com/topics/dark-data)” - i.e., data rotting away unused in undiscoverable [silos](https://www.getdbt.com/blog/what-are-the-four-principles-of-data-mesh). Thanks to its flexible personas and emphasis on data discovery, DataOps reduces dark data, driving additional business revenue with value-added data products. ## DataOps vs. DevOps breakdown Multiple similarities and differences likely jumped out at you while reading the descriptions above. Here’s a summary of some key points where both processes meet—and where they diverge. ### How DataOps and DevOps are similar **Agile**. Short development lifecycles with all team members participating at each step. Inclusion of business stakeholders to ensure alignment with business objectives and customer needs. **Emphasis on quality**. Use of version control, automated testing, code reviews, and multiple deployment environments to increase the quality of each release. **Automation**. Reliance on CI/CD pipelines to eliminate human error from the release process and increase release cadences. ### How DataOps and DevOps are different **Flexible personas**. In the data world, one person can be an engineer, an analyst, and a decision-maker on different projects. DataOps recognizes this reality and doesn’t associate personas with job titles. **Discoverability**. Good data doesn’t have value if people can’t find it. As such, DataOps devotes part of its process to ensuring new data products are published in a centralized repository for easy discovery. **Security and access control**. In DevOps, security is more about issues such as supply chain control and user application access. In DataOps, ADLC participants have to consider compliance with industry standards as well as national laws governing data privacy, such as [GDPR](https://gdpr-info.eu/). ## Driving DataOps with dbt Cloud Because they’re so heavily dependent on automation, DevOps and DataOps require great tooling to implement. Traditionally, this means teams must spend weeks or months building their own CI/CD pipelines from scratch. But not anymore. There’s a better way to do DataOps. With dbt Cloud as your data control plane, your data teams have a standardized and cost-efficient way to build, test, deploy, and discover analytics code. Meanwhile, data consumers have purpose-built interfaces and integrations to self-serve data that is governed and actionable. dbt Cloud makes implementing DataOps easy: - Represent all of your data transformation pipelines as [dbt models](https://docs.getdbt.com/docs/build/models) in SQL or Python, enabling anyone to develop data pipelines - Develop [data tests](https://docs.getdbt.com/docs/build/data-tests) - Store all changes in [version control](https://docs.getdbt.com/docs/collaborate/git/version-control-basics) to facilitate code reviews, versioned releases, and rollback - [Kick off a CI/CD pipeline](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud) to test and push changes from dev to stage to prod - Create and publish standardized metrics with [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) - Find data products and metrics using [dbt Explorer](https://docs.getdbt.com/docs/collaborate/explore-projects) Learn more about how dbt Cloud can bring DataOps to your organization—[schedule a demo today](https://www.getdbt.com/contact). --- --- title: "How to choose a data quality framework" description: "Do you need a data quality framework? An overview of popular frameworks - and how to adopt one tailor-fit to your business." url: "https://www.getdbt.com/blog/data-quality-framework-choosing" date: "2024-07-04" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # How to choose a data quality framework The phrase “garbage in, garbage out” is as true now as it was when it was coined 60-odd years ago. In fact it may be more true than ever in the age of Generative AI. Your analytics and AI applications require high-quality data to yield accurate, relevant, and timely outputs to drive business decision-making. High data quality requires a data quality framework. We’ll discuss what a data quality framework is and how to choose one—or, even better, create one customized to your business. **** ## What is a data quality framework? A data quality framework is a set of principles, standards, rules, and tools your organization uses to implement, test, and monitor the overall health of data. It provides a common foundation that individual teams can leverage to ensure their data is accurate, timely, and relevant. Establishing a data quality framework builds trust between data teams and business stakeholders. Furthermore, it creates an analytics foundation built on governance, scale, and accountability. ## How it fits within the Analytics Data Lifecycle A data quality framework isn’t a one-time band-aid or a silver bullet solution to your data issues. Rather, it’s a process that you implement throughout every stage of the [Analytics Data Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). Your data quality framework will grow and mature—adding new data cleansing and quality rules, providing better support for early detection of data errors, detecting anomalies through AI, etc.—as you gather on what works and what doesn’t. ## Components of a data quality framework Any set of rules, processes, and tools you set up to manage data quality constitutes a data quality framework. There are also several pre-built frameworks designed to support specific industries or data sets. Data quality frameworks differ in their focus and implementation (e.g., a focus on different dimensions of data, on a specific type of data, etc.). However, all of these frameworks will have the following components in common. ### Data pipelines Data pipelines define the workflow for ingesting data into your system and transforming, testing, and publishing it for consumption by data stakeholders. Data pipelines are the beating heart of your data quality framework, keeping high-quality data flowing to production. Data pipelines may consist of a mix of manual and automated processes. For example, a data engineer or analytics engineer may check-in data pipeline transformation code that another engineer reviews. This approval would [trigger a job](https://docs.getdbt.com/docs/deploy/jobs) that runs data tests in a pre-production environment before pushing the changes live. ### Data quality dimensions “Data quality” isn’t monolithic. It consists of a [number of dimensions](https://www.getdbt.com/blog/data-quality-dimensions) that include: - Usefulness - Accuracy - Validity - Uniqueness - Completeness - Consistency - Freshness ![Illustration showing seven dimensions of data quality: usefulness, accuracy, validity, uniqueness, completeness, consistency, and freshness, each with an icon and brief description.](https://cdn.sanity.io/images/wl0ndo6t/main/8111ea89a2d4bf939c59042d140a3834ced0623d-2178x548.png) A data quality framework may focus on different dimensions, emphasizing some over others. That focus may also shift over time as your program grows and evolves. ### Data quality and data cleansing rules [Data quality rules](https://www.getdbt.com/blog/data-quality-checks) check for conformity to one or more of the above data quality dimensions. Common rules include: - Unique record check - Non-null fields - Accepted values (e.g., status of an order, format of a customer ID) - Relationships and referential integrity - Freshness and recency ### Data governance rules [Data governance](https://www.getdbt.com/blog/what-is-enterprise-data-governance) is a set of rules and processes for collecting, storing, cleaning, and securing data for use. While it includes data quality, it also establishes rules for who owns data, how data is classified, and how your organization complies with the various data compliance and privacy regulations that govern your business (e.g., [GDPR](https://gdpr-info.eu/) or HIPAA). ### Data quality monitoring Detecting data quality issues proactively eliminates errors from data before they negatively impact data stakeholders and business decision-makers. Data quality monitoring constantly inspects incoming data and generates alerts at every part of the data lifecycle, testing data in its raw, transformed, and production form. ### Data quality tools Last, you need high-quality tools to power your data quality framework. Tools like dbt Cloud act as a data control plane, accelerating data delivery, tuning data quality, and optimizing compute costs. They help data teams scalably build, deploy, monitor, and discover data assets so organizations can move faster with trusted data. ## Popular data quality frameworks There are several data quality frameworks. Below is a short list of some of the more well-adopted ones, along with where and how data teams use them. ### Data Quality Assessment Framework (DGAF) [Developed by the International Monetary Fund (IMF)](https://www.imf.org/external/np/sta/dsbb/2003/eng/dqaf.htm), the DGAF defines a structure for evaluating your organization’s current practices against standard best practices for data quality. The DQAF tracks data quality across six dimensions—prerequisites, assurances, soundness, accuracy and reliability, serviceability, and accessibility—along with a set of elements and indicators for each dimension. It’s used primarily by governmental bodies, international organizations like the IMF and the UN, and organizations evaluating data for policy analysis or forecasts. ### Total Data Quality Management [Total Data Quality Management](https://www.sciencedirect.com/topics/computer-science/total-data-quality-management) is a holistic framework developed at MIT. Rather than insist on a defined set of metrics and data dimensions, it breaks down data quality into four stages of defining, measuring, analyzing, and improving the data quality dimensions that matter most to your business. ### ISO 8000 [ISO 8000](https://www.iso.org/obp/ui/#iso:std:iso:8000:-1:ed-1:v1:en) is an international standard that provides guidelines and best practices for improving data quality and creating enterprise master data (an authoritative version of the data most critical to your business). Governmental bodies and many worldwide companies—including Fortune 500 companies in the US—have used ISO 8000 to improve data quality and reduce costs. ### Data Quality Maturity Model (DQMM) The Data Quality Maturity Model (DQMM) is an umbrella name for a number of data quality frameworks that define different levels of data maturity, along with guidance on how to assess and improve your organization’s current level. One example is [ISACA’s CMMI](https://cmmiinstitute.com/cmmi/data), which most US software development contracts. It defines five levels of maturity—Initial, Managed, Defined, Quantitatively Managed, and Optimizing. ## Choosing a data quality framework Truth be told, most organizations these days shouldn’t bother with choosing an existing data quality framework unless it’s a business requirement—e.g., you do work for a government that makes adherence to the framework a condition of signing contracts. Most data quality frameworks were created a decade or more ago, when the data world dealt with fewer data sources and less overall data. In today’s world, where data growth is exploding exponentially, most companies need a more flexible and scalable approach. ## Defining your own data quality framework That doesn’t mean you don’t need a data quality framework! Rather, we recommend defining your own based on the needs of your business. This includes: - Defining which quality checks matter to your organization and individual teams. - Determining where to test. (We recommend continuous data quality testing across all environments, including raw sources.) - Establish a peer review process built on mechanisms such as [pull requests](https://docs.getdbt.com/blog/analytics-pull-request-template). - Repeating and iterating over this framework, adding new checks, procedures, and metrics. ## Using the ADLC to create a workflow [The Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle) is a process to create a mature analytics workflow that builds data quality into each step of the data lifecycle and iterates over your data quality framework. The ADLC consists of six stages that you apply to every data change: - **Plan**: Involving all stakeholders to define the business case for a data change, along with a data quality test plan and data quality metrics - **Develop**: Enable multiple stakeholders to build data transformation and business logic using the languages they already know - **Test**: Subject all code to a thorough code review and run unit, data, and integration tests in each environment - **Deploy**: Use an automated CI/CD process to push well-scoped changes to production; enable rollback mechanisms to revert to the status quo if you detect issues with the change in production - **Operate and Observe**: Catch errors before your stakeholders or customers do with continuous data quality monitoring - **Discover and Analyze**: Enable data stakeholders to find and use high-quality data sets easily to build analytics solutions, such as BI dashboards, reports, and data-driven applications ![Infinity loop diagram of the Analytics Development Life Cycle (ADLC) with phases: develop, test, deploy, plan, analyze, discover, observe, and operate.](https://cdn.sanity.io/images/wl0ndo6t/main/7fb61879a024adcb59b5f73665f7d804dbe8d223-1554x770.png) ## dbt: Your data quality framework tool Whether you adopt a data quality framework or create your own iteratively, your framework will only flourish if you have the right tools to implement it. With dbt Cloud, you can bring your data quality framework from theory to reality and reliably deliver high-quality data to your business: - Use [testing](https://docs.getdbt.com/docs/build/data-tests) and version control to validate assertions about your data and track code changes through all stages of deployment - Use [CI/CD pipelines](https://docs.getdbt.com/docs/deploy/continuous-integration) to automate data transformation code deployments - Spot and fix issues quickly using [column-level lineage](https://www.getdbt.com/blog/guide-to-dag) and embedded health status tiles in analytics tools - Auto-generate [documentation](https://docs.getdbt.com/docs/build/documentation) for all of your data transformation models - Implement data governance in a scalable manner with [dbt Mesh](https://docs.getdbt.com/best-practices/how-we-mesh/mesh-1-intro) - Deliver high-quality, tested metrics to stakeholders using the dbt [Semantic Layer](https://www.getdbt.com/product/semantic-layer) For more details, read our deep dive on [building a data quality framework with dbt Cloud](https://www.getdbt.com/blog/building-a-data-quality-framework-with-dbt-and-dbt-cloud). Or [book a demo today](https://www.getdbt.com/contact) to discuss how to bring a dbt Cloud-driven data quality framework to your business. [Watch video](https://www.youtube.com/watch?v=rItxGK0cYj8) --- --- title: "Data governance frameworks for AI-driven organizations" description: "Why robust data governance is critical for AI-driven organizations to ensure the integrity, security, and ethical use of data." url: "https://www.getdbt.com/blog/data-governance-frameworks-ai" date: "2024-07-03" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Data governance frameworks for AI-driven organizations Companies of every size and sector are racing to transform their operations and drive innovation with AI. Unfortunately, many dive into the AI deep end without considering the data that drives their AI initiatives. AI runs on data. The large language models that underpin AI systems need high-quality, consistent, and accurate data to return useful and accurate results. For AI to be a reliable and useful tool for your business, you need to have high-quality training data — and having high-quality data to feed the model depends on data governance. Good data governance is important for every business focused on making data-driven decisions. It becomes even more critical [when you integrate AI into your organization](https://www.getdbt.com/blog/understanding-data-governance-ai). Robust data governance is critical for AI-driven organizations because this is how you ensure the integrity, security, and ethical use of data that powers your AI systems. Adopting the right data governance framework is the essential first step to ensuring data quality, consistency, accessibility, and compliance. ## Why AI demands a strong data governance framework Adopting a strong data governance framework allows your organization to establish guidelines for data privacy, ethical use, and transparency of AI technologies no matter what business you’re in. (This is even more critical, though, for highly regulated sectors like healthcare and finance). Here’s why. ### Data complexity and scale AI makes governance more complex because it requires vast amounts of training data, coming in different forms from different domains within your organization—far too much data for casual or even manual management efforts. A data governance framework must handle large volumes and diverse data types (structured, unstructured, real-time) in order to deliver the benefits of AI. ### Ethics and bias If managed poorly, AI can generate results containing unintended biases or ethical issues. For instance, if a company implements an AI-driven chat agent trained on ungoverned data, the AI might pick up on existing biases in training data. This could lead to biased responses, like preferential treatment for certain demographic groups, or giving answers that unintentionally alienate or offend specific customers. Or if a company adopts AI to personalize customer recommendations but fails to put in place strong data governance, the AI might inadvertently misuse personal data. In an AI-powered enterprise, data governance frameworks provide (and help enforce) policies for the development and use of AI technologies that embed fairness, accountability, and model transparency. This helps make sure that your AI outcomes are unbiased and non-discriminatory, which in turn shields your company from risk. ### Increasing compliance and regulatory requirements Data-related laws like GDPR and AI-specific regulations like the EU’s new [AI Act](https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai), to name but two, are rapidly proliferating around the globe. The challenge for your company is to maintain compliance without stifling innovation. Harnessing a data governance framework that includes built-in compliance monitoring to align with both current and future regulations so you can stay focused on improving and delivering your product. ## Core components of data governance for AI Effective AI data governance involves data quality management, compliance with regulatory requirements, and continuous monitoring—all of which help organizations navigate the complexities of AI deployment, while adhering to ethical and legal standards. Look for a data governance framework that offers: ### Data ownership and stewardship Assign clear responsibilities for data handling. Data stewards, data owners, and custodians should understand their roles in managing and safeguarding data throughout its lifecycle, especially for high-risk AI use cases. ### Data quality management AI models rely heavily on high-quality data for accurate outcomes. This means investing in a framework with robust data quality monitoring and automated anomaly detection capabilities to ensure data completeness, accuracy, consistency, and relevance. ### Metadata management Maintaining rich metadata, including data lineage, provenance, and usage, is critical. Metadata allows you to trace data back to its source, which further supports data transparency, quality checks, and ethical AI practices. ### Bias detection and mitigation AI systems are prone to perpetuating biases. Our governance framework should include methodologies for detecting and reducing bias in data, algorithms, and model predictions. This could involve bias auditing tools and guidelines for creating diverse datasets. ### Privacy and security Privacy-preserving techniques (such as differential privacy, anonymization, and data masking) and robust data security measures (like encryption and access control) are essential for protecting sensitive information. ### Data lifecycle management A governance framework must manage the entire data lifecycle, from data acquisition to archiving and deletion, to ensure the ongoing relevance and compliance of data in your organization’s AI endeavors. ### AI model maintenance Governance doesn’t end at the data level—it also extends to the AI models themselves. To ensure that the models that drive your AI remain current and effective over time, AI-driven companies need a framework that incorporates model governance by tracking model versions, assessing performance, and monitoring for drift. ## How data governance can help AI initiatives By establishing a comprehensive governance framework, you ensure your company’s data and AI assets are managed effectively—and set the stage for positive outcomes. ### Scalability and flexibility With a data governance framework in place, your company can scale AI initiatives across the organization and, at the same time, adapt to changing data and regulatory landscapes. Data volumes are growing nonstop. New AI technologies can emerge at any time. A governance framework that allows for modular updates makes it easier to integrate new data sources, manage big data, and adopt emerging AI compliance standards. ### Risk mitigation By proactively addressing bias, privacy, and security concerns, our framework will help mitigate legal, ethical, and reputational risks. ### Enhanced decision-making Reliable data and AI outputs lead to better business decisions. Governance ensures data and AI models are accurate, relevant, and unbiased. ### Regulatory compliance With compliance baked into the framework, we reduce the risk of regulatory fines and foster trust among stakeholders. ## Frameworks and standards for AI-driven data governance There are many data governance frameworks out there. However, not all of them are suitable for defining a structure of guidelines, protocols, processes, and rules for enterprise data in a way that serves and supports AI. Here are three established data governance frameworks and standards specifically tailored to meet the demands of AI. - **[CDMC (Cloud Data Management Capabilities)](https://edmcouncil.org/frameworks/cdmc/):** This framework by the EDM Council helps organizations manage data in cloud environments, a crucial need for AI use cases that depend on vast datasets. It includes principles governing data in both cloud and hybrid environments, emphasizing data lineage and quality — key for AI implementations. - **[NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework):** The National Institute of Standards and Technology offers a data governance framework organized around risk mitigation for AI. It covers data integrity and explainability, along with bias management. - **[ISO/IEC 38505](https://www.iso.org/obp/ui#iso:std:iso-iec:ts:38505:-3:ed-1:v1:en):** The International Order of Standards issued this governance-oriented standard that specifically addresses governance of data for analytics and AI, with a focus on strategic alignment, accountability, and transparency. ## Data governance tools for AI A data governance framework helps your org define a set of standards and policies around your data assets—but you still need to implement them. This is where data governance tools come in, particularly for implementing auditability and control over AI applications. Because a data governance framework is only as strong as the tools (and people) supporting it, here are core features to prioritize when you evaluate a data governance solution: ### Data cataloging and discovery Look for one that supports data cataloging, which helps teams to easily discover and understand available data. Comprehensive metadata capabilities are also important for tracking data lineage, quality, and usage — all are essential for transparent AI operations. ### Automated data quality monitoring Real-time data quality monitoring and alerting are critical to ensuring your organization's AI models are trained solely on high-quality data. Look for a solution that can automatically detect and address issues like missing, duplicate, or inconsistent data. ### Data lineage and provenance tracking AI-driven decisions require full data traceability, especially for regulatory compliance. To provide transparency and accountability, any data governance tool must have data lineage capabilities so you can see where data originates, how it flows, and how it’s transformed. ### Privacy and security controls Because AI will likely handle sensitive data, whether personal privacy or company IP, tools must include privacy-preserving capabilities like data masking, encryption, and role-based access control, as well as compliance features for regulations (e.g., GDPR, CCPA). ### AI model governance and lifecycle management Data governance tools increasingly overlap with model management tools to track models from training to deployment. You need a solution that includes monitoring for model drift, performance, and adherence to fairness and accuracy standards. ## Conclusion Whether it’s applied to critical decision-making or just simplifying everyday tasks, the governance of data flow inside your organization is critical for safe and responsible AI utilization. As AI continues to redefine the boundaries of what's possible in business, the need to adopt an effective governance solution into your data stack has never been greater: - To scale governance across whatever AI initiatives your organization pursues, you need a data governance platform with automation tools for data lineage, data quality, and model tracking. - To ensure you’re in compliance with international and regional data regulations, you also need compliance tools with features such as automated audit trails. - For the most seamless governance possible, you need a tool that integrates with your existing data stack—whether that’s Snowflake, Databricks, or another data management tool. You can find all of these features, and more, in [dbt Cloud](https://www.getdbt.com/product/what-is-dbt), the standard for data transformation in modern environments. dbt serves as your data control plane so your organization can create the technical guardrails around data governance that let you adopt and scale AI technologies. --- --- title: "Guide to surrogate keys" description: "A surrogate key is a unique identifier derived from the data itself." url: "https://www.getdbt.com/blog/guide-to-surrogate-key" date: "2024-07-01" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Guide to surrogate keys A surrogate key is a unique identifier derived from the data itself. It often takes the form of a hashed value of multiple columns that will create a uniqueness constraint for each row. You will need to create a surrogate key for every table that doesn't have a natural primary key. Why would you ever need to make a surrogate key? Shouldn’t all tables innately just have a field that uniquely identifies each row? Now that would be too easy… Let’s say you have a table with all license plate numbers and the state of the plate. While license plate numbers are unique to their state, there could be duplicate license plate numbers across different states. So by default, there’s no natural key that can uniquely identify each row here. In order to uniquely identify each record in this table, you could create a surrogate key based on the unique combination of license plate number and its state. ## Surrogate keys, natural keys, and primary keys oh my! Primary keys can be established two ways: naturally or derived through the data in a surrogate key. - A **natural key** is a primary key that is innate to the data. Perhaps in some tables there’s a unique `id` field in each table that would act as the natural key. You can use documentation like entity relationship diagrams (ERDs) to help understand natural keys in APIs or backend application database tables. - A **surrogate key** is a hashed value of multiple fields in a dataset that create a uniqueness constraint on that dataset. You’ll essentially need to make a surrogate key in every table that lacks a natural key. Note - You may also hear about primary keys being a form of a _constraint_ on a database object. Column constraints are specified in the [DDL](https://www.getdbt.com/blog/guide-to-ddl) to create or alter a database object. For data warehouses that support the enforcement of primary key constraints, this means that an error would be raised if a field's uniqueness or non-nullness was broken upon an `INSERT` or `UPDATE` statement. Most modern data warehouses don’t support _and_ enforce primary key constraints, so it’s important to have [automated testing](https://docs.getdbt.com/blog/primary-key-testing#how-to-test-primary-keys-with-dbt) in-place to ensure your primary keys are unique and not null. ## How surrogate keys are created In analytics engineering, you can generate surrogate keys using a hashing method of your choice. Remember, in order to truly create a uniqueness constraint on a database object, you’ll need to hash the fields together that _make each row unique_; when you generate a correct surrogate key for a dataset, you’re really establishing the true [grain](https://www.getdbt.com/blog/guide-to-data-grain) of that dataset. Let’s take this to an example. Below, there is a table you pull from an ad platform that collects `calendar_date`, `ad_id`, and some performance columns. In this state, this table has no natural key that can act as a primary key. You know the grain of this table: this is showing performance for each `ad_id` per `calendar_date`. Therefore, hashing those two fields will create a uniqueness constraint on this table. To create a surrogate key for this table using the MD5 function, run the following: ```sql select md5(calendar_date || ad_id) as unique_id, * from {{ source('ad_platform', 'custom_daily_report')}} ``` After executing this, the table would now have the `unique_id` field now uniquely identifying each row. ## Testing surrogate keys Amazing, you just made a surrogate key! You can just move on to the next data model, right? No!! It’s critically important to test your surrogate keys for uniqueness and non-null values to ensure that the correct fields were chosen to create the surrogate key. In order to test for null and unique values you can utilize code-based data tests like [dbt tests](https://docs.getdbt.com/docs/build/data-tests), that can check fields for nullness and uniqueness. You can additionally utilize simple SQL queries or unit tests to check if surrogate key count and non-nullness is correct. ## A note on hashing algorithms Depending on your data warehouse, there’s several cryptographic hashing options to create surrogate keys. The primary hashing methods include MD5 or other algorithms, like HASH or SHA. Choosing the appropriate hashing function is dependent on your dataset and what your warehouse supports. Note - A collision occurs when two pieces of data that are different end up hashing to the same value. If a collision occurs, a different hashing method should be used. ## Why we like surrogate keys Let’s keep it brief: surrogate keys allow data folks to quickly understand the grain of the database object and are compatible across many different data warehouses. ### Readability Because surrogate keys are comprised of the fields that make a uniqueness constraint on the data, you can quickly identify the grain of the data. For example, if you see in your data model that the surrogate key field is created by hashing the `ad_id` and `calendar_date` fields, you can immediately know the true grain of the data. When you clearly understand the grain of a database object, this can make for an easier understanding of how entities join together and fan out. ### Compatibility Making a surrogate key involves a relatively straightforward usage of SQL: maybe some coalescing, concatenation, and a hashing method. Most, if not all, modern data warehouses support both the ability to concat, coalesce, and hash fields. They may not have the exact same syntax or hashing functions available, but their core functionality is the same. **** ## Performance concerns for surrogate keys In the past, you may have seen surrogate keys take the form of [monotonically increasing](https://docs.getdbt.com/terms/monotonically-increasing) integers (ex. 1, 2, 3, 4). These surrogate keys were often limited to 4-bit integers that could be indexed quickly. However, in the practice of analytics engineering, surrogate keys derived from the data often take the form of a hashed string value. Given this form, these surrogate keys are not necessarily optimized for performance for large table scans and complex joins. For large data models (millions, billions, trillions of rows) that have surrogate keys, you should materialize them as tables or [incremental models](https://docs.getdbt.com/docs/build/incremental-models) to help make joining entities more efficient. ## Surrogate Key FAQs **What is a surrogate key and when do I need one?** A surrogate key is a unique identifier derived from the data itself, typically created by hashing multiple columns together to ensure each row can be uniquely identified. You need to create a surrogate key for every table that doesn't have a natural primary key, such as when combining license plate numbers and states where the plate number alone isn't unique across states. **How do I create a surrogate key in my data models?** A surrogate key is created using a hashing method (like MD5, SHA256, or HASH) on the combination of fields that together make each row unique. For example, if you have a table tracking ad performance by date, you might create a surrogate key using `md5(calendar_date || ad_id)` to establish a unique identifier for each date-ad combination, which represents the true grain of your dataset. **What's the difference between surrogate keys and natural keys?** A natural key is a primary key that is innate to the data, such as a unique ID field that already exists in the table. A surrogate key, by contrast, is derived by hashing together multiple fields to create uniqueness when no natural key exists. While natural keys are preferable when available, surrogate keys provide a solution when tables lack a single field that can uniquely identify each row. **How should I test my surrogate keys?** After creating a surrogate key, it's critical to test it for both uniqueness and non-null values to ensure you've chosen the correct fields for the key. You can use code-based data tests like dbt tests to check fields for nullness and uniqueness, or utilize simple SQL queries and unit tests to verify that the surrogate key count is correct and contains no null values. **Are there performance concerns with surrogate keys?** Surrogate keys created through hashing methods (producing string values) may not be as performance-optimized as traditional integer-based surrogate keys for large table scans and complex joins. For large data models containing millions or billions of rows, it's recommended to materialize models with surrogate keys as tables or incremental models to make joining entities more efficient. ## Conclusion Surrogate keys are unique row identifiers that are created by using columns in a database object to create a uniqueness constraint on the data. To create a surrogate key, you will use a cryptographic algorithm usually in the form of the MD5 function to hash together fields that create a uniqueness constraint on the dataset. Ultimately, surrogate keys are a great way to create unique row identifiers for database objects that lack them naturally and allow folks to easily identify the grain of the data. ## Further reading Want to learn more about keys, dbt, and everything in-between? Check out the following: - [Generating surrogate keys across warehouses](https://docs.getdbt.com/blog/sql-surrogate-keys) - [Generating an auto-incrementing ID in dbt](https://discourse.getdbt.com/t/generating-an-auto-incrementing-id-in-dbt/579/2) - [The most underutilized function in SQL](https://www.getdbt.com/blog/the-most-underutilized-function-in-sql/) --- --- title: "June dbt Community update" description: "Stay updated with the dbt Community. In June we had an AMA on BI reporting, 13 in-person Meetups, and more." url: "https://www.getdbt.com/blog/june-dbt-community-update" date: "2024-07-01" authors: ["Kathryn Chubb"] categories: ["Community"] --- # June dbt Community update Welcome to the dbt Community Update, a monthly blog about everything happening in the [dbt Community](https://www.getdbt.com/community)! This month we hosted an AMA on BI reporting and the semantic layer, had 13 in-person [Meetups,](https://www.meetup.com/pro/dbt/) and facilitated a ton of great discussions on our Slack channel. Are you ready for the recap? Let’s get started. ## dbt Community Slack AMA Each month we host a live Ask Me Anything event. This month, Alex Welch and Roxi Dahlke hosted the AMA and discussed BI reporting and the semantic layer. Here’s a recap if you missed it. You can also check out the [full video recording](https://www.getdbt.com/resources/community-slack-ama/confirmation). ### Understanding the concept The conversation started with Alex and Roxi sharing their optimism about pushing the envelope in the BI space. Their excitement was rooted in the increasing inclusion of non-technical users into the analytical fold, thanks to the emergence of the semantic layer. This development fuels various use cases, leading far beyond traditional BI, and making it an interesting paradigm shift to witness. They also made an interesting comparison - the metrics store is to the semantic layer what a chapter is to a book. Imagine the semantic layer as a broad abstraction of all data objects, including their relationships, while a metrics store zooms into the metrics in particular. Alex added a dash of machine learning parlance – metrics akin to features, with semantics bringing a higher level of business context. ### Building and maintaining semantic layers Building a semantic layer isn't just a technical endeavor – it's a dance between business and data teams, discussing and translating objectives into data. By weaving in YAML configurations for metrics definition, dbt Labs encourages these sometimes tricky but vital conversations. But an essential challenge Alex and Roxi raised was the difficulty of maintaining consistency. Tooling concerns surfaced here, as they warned us against outdated and fragmented documentation in specific tools. ### Navigating BI Tools Speaking of tools, there's a lesson or two to learn from some personal anecdotes. Alex reminded attendees to remain skeptical of BI tools and their promises, pointing out the risks of data sprawl and contradictory figures. Roxi shared a harrowing experience of misdelivering metrics and the lesson it carries. ### A journey of data literacy Both speakers raved about the journey of enhancing data literacy in dbt Labs, celebrating the contributions from even non-data teams. This highlighted their focus on community, encouraging a culture of testing, using, and providing feedback on their products. ### Doing semantic right Making the semantic layer work depends heavily on understanding the subtleties between the metrics layer and the semantic layer. Also, tools like dbt Cloud are crucial for maintaining visibility into your metrics. They also touched on the additional possibilities of the data modeling layer in providing a comprehensive perspective on data. ### Keeping up with industry developments Both Alex and Roxi shared their tactics for maintaining relevance with industry trends, ranging from webinars and blogs to interactions with industry professionals. Both emphasized the importance of continuous learning in an ever-changing field like data management. ### Forward-looking insights Among the final interesting points from the session was Alex’s belief that BI platforms would become more personalized with AI interfaces, providing contextualized insights. Roxi foresaw AI helping users discover new data within BI tools. Last but not least, they discussed the implementation of dbt Cloud on teams or Enterprises. Their advice is to start small, with vital metrics, and gradually build it. ### Get ready for another AMA in July [Join us for next month’s AMA](https://www.getdbt.com/resources/webinars/community-ama) on July 30th at 3pm EST with special guest Tristan Handy, Founder & CEO of dbt Labs. Register now to get the link to watch live and join the [#dbt-community-merge channel in Slack](https://getdbt.slack.com/archives/C025ZN1L679/p1716308109072499) to participate in the conversation! ## June dbt Meetups In June we had 13 Meetups in nine countries: Washington DC, Oslo, Belgium, Medellín, Chicago, Atlanta, Taipei, Barcelona, Tokyo, Sāo Paulo, Halifax, Bogotá, and Madrid. We also ran a new Community Event Series, dbt Labs on dbt, in Philadelphia and Dublin at our offices. We were especially excited to bring back the São Paulo dbt Meetup after a little over a year of not having one. On June 24th, over 70 attendees gathered and presented on the following topics: - How to optimize dbt-core setup in VS Code - Migrating a pipeline with more than 1,000 tables to the modern data stack with dbt - From code to insight: The power of dbt Artifacts in the data-driven culture You can see photos from the meetups below, including São Paulo (organized by Bruno Souza de Lima and Thales Donizet), Barcelona (organized by Infinite Lambda,), Belgium (organized by Sam Debruyn at the dataroots offices), and Medellín (organized by Factored). ![Sao Paulo 2](https://cdn.sanity.io/images/wl0ndo6t/main/b5a527ead08d5f527b5ba6d731c1d5995e5bb473-2048x1365.jpg) ![Barcelona](https://cdn.sanity.io/images/wl0ndo6t/main/9f4c032aee92ea171b64c2754a98f70ab52e553d-1600x1200.jpg) ![Belguim](https://cdn.sanity.io/images/wl0ndo6t/main/de6c0444536c0f4855a0848287bab797b9b21256-4032x3024.jpg) ![Medellin](https://cdn.sanity.io/images/wl0ndo6t/main/4f84a69e3d2fe5b98d07baa0735ab7d49b8294da-1024x768.jpg) ## dbt Community announcements We’ll wrap up this month’s update with some of the exciting announcements that are regularly posted in our [#announcements](https://getdbt.slack.com/archives/C0VLZM3U2/p1715777392876319) channel on Slack. ### dbt swag shop To celebrate Pride Month, we released a pride drop on the dbt swag shop. Check out the exclusive designs here: [shop.getdbt.com](http://shop.getdbt.com) ### Upcoming events - July 30th - [Register for our next Community AMA](https://www.getdbt.com/resources/webinars/community-ama) featuring Tristan Handy, Founder & CEO of dbt Labs - Ongoing [Cloud Demo with Experts](https://www.getdbt.com/resources/dbt-cloud-demos-with-experts/) in North America, EMEA, and APAC-friendly times - October 7-10th, 2024, [Coalesce](https://coalesce.getdbt.com/register-2024), by dbt Labs in Las Vegas and Online - Add [dbt Events](https://www.addevent.com/calendar/Tb314369) to your calendar! ### July dbt Meetups We’ve got a busy month coming up with 6 [in-person dbt Meetups](https://www.meetup.com/pro/dbt) scheduled. If you’re looking for opportunities to learn with fellow members of the dbt Community, and have fun while doing so, join us at one of the sessions listed below: - 🇯🇵 Tokyo | Thursday, July 4th, organized by dbt Labs, Shinya Takimoto, and DATUM STUDIO Co. Ltd. This event will be available to watch online. It will be in English with live Japanese translation. - 🇻🇳 Ho Chi Minh City | Saturday, July 6th, organized by Joon Solutions and Infinite Lambda - 🇹🇼 Taipei | Thursday, July 11th, organized by community members Karen Hsieh, Laurence Chen, and Allen Wang - 🇺🇸 San Francisco | Monday, July 15th, organized by dbt Labs - 🇺🇸 Minneapolis | Tuesday, July 16th, organized by Improving - 🇺🇸 Boston | Wednesday, July 17th, organized by Cleartelligence There are so many exciting things going on in the dbt Community, and we can’t wait to see you all there! If you haven’t yet, [join the community](https://www.getdbt.com/community) today. --- --- title: "Guide to subquery" description: "A subquery is what the name suggests: a query within another query." url: "https://www.getdbt.com/blog/guide-to-subquery" date: "2024-07-01" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Guide to subquery A subquery is what the name suggests: a query within another query. _The true inception of SQL_. Subqueries are often used when you need to process data in several steps. For the majority of subqueries you’ll see in actual practice, the inner query will execute first and pass its result to the outer query it's nested in. Subqueries are usually contrasted with [Common Table Expressions (CTEs)](https://www.getdbt.com/blog/guide-to-cte) as they have similar use cases. Unlike CTEs, which are usually separate `SELECT` statements within a query, subqueries are usually `SELECT` statements nested within a `JOIN`, `FROM`, or `WHERE` statement in a query. To be honest, we rarely write subqueries here at dbt Labs since we prefer to use CTEs. We find that CTEs, in general, support better query readability, organization, and debugging. However, subqueries are a foundational concept in SQL and still widely used. We hope you can use this glossary to better understand how to use subqueries and how they differ from CTEs. ## Subquery syntax While there are technically several types of subqueries, the general syntax to build them is the same. A subquery usually consists of the following: - Enclosing parentheses - A name - An actual SELECT statement - A main query it is nested in via a FROM, WHERE, or JOIN clause Let’s take this to an example, using the [sample jaffle_shop dataset](https://github.com/dbt-labs/jaffle_shop). ```sql select customer_id, count(order_id) as cnt_orders from ( select * from {{ ref('orders') }} ) all_orders group by 1 ``` Given the elements of subqueries laid out in the beginning, let’s break down this example into its respective parts. When this query is actually executed, it will start by running the innermost query first. In this case, it would run `select * from {{ ref('orders') }}` first. Then, it would pass those results to the outer query, which is where you grab the count of orders by `customer_id`. If you want to learn more about what a `ref` is, [check out our documentation on it.](https://docs.getdbt.com/reference/dbt-jinja-functions/ref) This is a relatively straightforward example, but should hopefully show you that subqueries start off like most other queries. As you nest more subqueries together, that’s when you unearth the power of subqueries, but also when you start to notice some readability tradeoffs. If you are using subqueries regularly, you'll want to leverage indenting and [strong naming conventions](https://docs.getdbt.com/blog/on-the-importance-of-naming) for your subqueries to clearly distinguish code functionality. ## Types of subqueries In your day-to-day, you won’t normally formalize the names of the different types of subqueries you can write, but when someone uses the term “correlated subquery” at a data conference, you'll want to know what that means! ### Nested subqueries Nested subqueries are subqueries like the one you saw in the first example: a subquery where the inner query is executed first (and once) and passes its result to the main query. The majority of subqueries you will see in the real world are likely to be a nested subquery. These are most useful when you need to process data in multiple steps. ### Debugging subqueries tip It’s important to note that since the inner query is executed first in a nested subquery, the inner query must be able to execute by itself. If it’s unable to successfully run independently, it cannot pass results to the outer query. ### Correlated subqueries A correlated subquery is a nested subquery’s counterpart. If nested subqueries execute the inner query first and pass their result to the outer query, correlated subqueries execute the outer query first and pass their result to their inner query. For correlated subqueries, it’s useful to think about how the code is actually executed. In a correlated subquery, the outer query will execute row-by-row. For each row, that result from the outer query will be passed to the inner query. Compare this to nested queries: in a nested query, the inner query is executed first and only once before being passed to the outer query. These types of subqueries are most useful when you need to conduct analysis on a row-level. ### Scalar and non-scalar subqueries Scalar subqueries are queries that only return a single value. More specifically, this means if you execute a scalar subquery, it would return one column value of one specific row. Non-scalar subqueries, however, can return single or multiple rows and may contain multiple columns. You may want to use a scalar subquery if you’re interested in passing only a single-row value into an outer query. This type of subquery can be useful when you’re trying to remove or update a specific row’s value using a [Data Manipulation Language (DML)](https://www.getdbt.com/blog/guide-to-dml) statement. ## Subquery examples You may often see subqueries in joins and DML statements. The following sections contain examples for each scenario. ### Subquery in a join In this example, you want to get the lifetime value per customer using your `raw_orders` and `raw_payments` table. Let’s take a look at how you can do that with a subquery in a join: ```sql select orders.user_id, sum(payments.amount) as lifetime_value from {{ ref('raw_orders') }} as orders left join ( select order_id, amount from {{ ref('raw_payments') }} ) all_payments on orders.id = payments.order_id group by 1 ``` Similar to what you saw in the first example, let’s break down the elements of this query. In this example, the `all_payments` subquery will execute first. you use the data from this query to join on the `raw_orders` table to calculate lifetime value per user. Unlike the first example, the subquery happens in the join statement. Subqueries can happen in `JOIN`, `FROM`, and `WHERE` clauses. ### Subquery in a DML command You may also see subqueries used in DML commands. As a jogger, DML commands are a series of SQL statements that you can write to access and manipulate row-level data in database objects. Oftentimes, you’ll want to use a query result in a qualifying `WHERE` clause to only delete, update, or manipulate certain rows of data. In the following example, you'll attempt to update the status of certain orders based on the payment method used in the `raw_payments` table. ```sql UPDATE raw_orders set status = 'returned' where order_id in ( select order_id from raw_payments where payment_method = 'bank_transfer') ``` ## Subquery vs CTE A subquery is a nested query that can oftentimes be used in place of a CTE. Subqueries have different syntax than CTEs, but often have similar use cases. The content won’t go too deep into CTEs here, but it’ll highlight some of the main differences between CTEs and subqueries below. ### Subquery vs CTE example The following example demonstrates the similarities and differences between subqueries and CTEs. Using the first subquery example, you can compare how you would perform that query using subquery or a CTE: #### Subquery example ```sql select customer_id, count(order_id) as cnt_orders from ( select * from {{ ref('orders') }} ) all_orders group by 1 ``` #### CTE example ```sql with all_orders as ( select * from {{ ref('orders') }} ), aggregate_orders as ( select customer_id, count(order_id) as cnt_orders from all_orders group by 1 ) select * from aggregate_orders ``` While the code for the query involving CTEs may be longer in lines, it also allows us to explicitly define code functionality using the CTE name. Unlike the subquery example that executes its inner query and then the outer query, the query using CTEs executes moving down the code. Again, choosing to use CTEs over subqueries is a personal choice. It may help to write out the same code functionality in a subquery and with CTEs and see what is more understandable to you. ## Data warehouse support for subqueries Subqueries are likely to be supported across most, if not all, modern data warehouses. Please use this table to see more information about using subqueries in your specific data warehouse. ## Conclusion I’m going to be honest, I was hesitant to start writing the glossary page for SQL subqueries. As someone who has been using CTEs almost exclusively in their data career, I was intimidated by this concept. However, I am excited to say: Subqueries are not as scary as I expected them to be! At their core, subqueries are nested queries within a main query. They are often implemented in `FROM`, `WHERE`, and `JOIN` clauses and are used to write code that builds on itself. Despite the fact that subqueries are SQL like any other query, it is important to note that subqueries can struggle in their readability, structure, and debugging process due to their nested nature. Because of these downsides, we recommend leveraging CTEs over subqueries whenever possible. I have not been made a subquery convert, but I’m walking away from this a little less intimidated by subqueries and I hope you are too. ## Further reading Please check out some of our favorite readings related to subqueries! - [Glossary: CTE](https://www.getdbt.com/blog/guide-to-cte) - [On the importance of naming: model naming conventions (Part 1)](https://docs.getdbt.com/blog/on-the-importance-of-naming) --- --- title: "Guide to data grain" description: "Data grain is the combination of columns at which records in a table are unique." url: "https://www.getdbt.com/blog/guide-to-data-grain" date: "2024-07-01" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Guide to data grain Data grain is the combination of columns at which records in a table are unique. Ideally, this is captured in a single column, a unique primary key, but even then, there is descriptive grain behind that unique id. Let’s look at some examples to better understand this concept. In the above table, each `user_id` is unique. This table is at the _user grain_. In the above table, `user_id` is no longer unique. The combination of `user_id` and `address` creates a unique combination, thus this table is at the user address grain. We generally describe the grain conceptually based on the names of the columns that make it unique. A more realistic combination you might see in the wild would be a table that capture the state of all users at midnight every day. The combination of the captured `updated_date` and `user_id` would be unique, meaning our table is at user per day grain. ### Creating surrogate keys In both examples listed in the previous paragraph, we’d want to create a [surrogate key](https://www.getdbt.com/blog/guide-to-surrogate-key) of some sort from the combination of columns that comprise the grain. This gives our table a primary key, which is crucial for testing and optimization, and always a best practice. Typically, we’ll name this primary key based on the verbal description of the grain. For the latter example, we’d have `user_per_day_id`. This will be more solid and efficient than testing than repeatedly relying on the combination of those two columns. ## The importance of data grain in data modeling Thinking deeply about grain is a crucial part of data modeling. As we design models, we need to consider the entities we’re describing, and what dimensions (time, attributes, etc.) might fan those entities out so they’re no longer unique, as well as how we want to deal with those. Do we need to apply transformations to deduplicate and collapse the grain? Or do we intentionally want to expand the grain out, like in our user per day example? **** There’s no right answer here, we have the power to do either as it meets our needs. The key is just to make sure we have a clear sense of our grain for every model we create, that we’ve captured it in a primary key, and that we’re applying tests to ensure that our primary key column is unique and not null. ## Choosing the right type of data grain for your use case Selecting appropriate grain is critical for effective data modeling success. Here are some considerations to keep in mind: - **Analysis requirements**: The questions you need to answer determine your grain selection. Deeper questions often require finer grain structure. - **Data volume**: Larger data volumes challenge storage and processing capabilities. Higher grain levels reduce storage needs but sacrifice analytical detail. - **Query performance**: Finer grain means more records to process during queries. Performance tuning might require pre-aggregation at higher grain levels. - **Update frequency**: More frequent data updates favor transactional grain. Less frequent updates work well with snapshot approaches. - **Process visibility**: Tracking multi-stage processes requires accumulating snapshots. through workflows. - **Reporting cadence**: Regular reporting schedules influence grain selection. Daily reporting needs different grain than quarterly analysis. - **Historical analysis**: Long-term trend analysis has unique grain requirements. Consider how historical comparisons will work with your chosen grain. #### Ready to bring clarity and consistency to your data models? Start building trusted, well-documented data pipelines with dbt. [Try dbt for free and join the data teams building with confidence at every grain. ](https://www.getdbt.com/product/dbt) **Related reading:** - [A complete guide to dimensional modeling with dbt](https://www.getdbt.com/blog/guide-to-dimensional-modeling) - [Successful data transformation: Six steps | dbt Labs](https://www.getdbt.com/blog/successful-data-transformation) - [Five real data transformation examples | dbt Labs](https://www.getdbt.com/blog/data-transformation-examples) ## FAQs for data grain **What is data grain? ** Data grain is the combination of columns at which records in a table are unique. Ideally this uniqueness is captured in a single column called a primary key, but even then, there is descriptive grain behind that unique ID. For example, in a table where each `user_id` is unique, the table is at the "use grain"; but if `user_id` appears multiple times with different addresses, the table is at the "user address grain". **How does grain influence data modeling? ** Grain is critically important because it: - Determines what questions can be answered with your data - Impacts storage requirements and query performance - Affects the complexity of your data model - Influences data integration capabilities. **What are examples of different grains? ** Here's a really basic example: a table with unique `user_id`s (user grain) compared to a table where users can have multiple addresses (use address grain). A more realistic example would be a table that captures the state of all users at midnight every day, where the combination of `updated_date` and `user_id` would be unique (user per day grain). In both cases, creating a surrogate key that combines these fields establishes a proper primary key. **What best practices should I follow with data grain? ** Make sure you have a clear sense of the grain for every model you create and capture it in a primary key column. Apply tests to ensure that your primary key column is unique and not null. When dealing with multiple columns that define uniqueness, create a surrogate key named based on the verbal description of the grain (like `user_per_day_id`). There's no single right approach — you can choose to collapse or expand grain as needed — but clarity and consistency are essential. **Why is creating a surrogate key important? ** Creating a surrogate key from the combination of columns that comprise the grain gives your table a primary key, which is crucial for testing and optimization. This practice is more solid and efficient than repeatedly relying on testing combinations of columns. For example, in a table at the "user per day" grain, you might create a surrogate key called "user_per_day_id" that combines `user_id` and `updated_date` to ensure uniqueness. --- --- title: "DRY principles: How to write efficient SQL" description: "DRY is a software development principle that stands for “Don’t Repeat Yourself.”" url: "https://www.getdbt.com/blog/dry-principles" date: "2024-07-01" authors: ["Daniel Poppy"] categories: ["Learn"] --- # DRY principles: How to write efficient SQL DRY is a software development principle that stands for “Don’t Repeat Yourself.” Living by this principle means that your aim is to reduce repetitive patterns and duplicate code and logic in favor of modular and referenceable code. The DRY code principle was originally made with software engineering in mind and coined by Andy Hunt and Dave Thomas in their book, _The Pragmatic Programmer_. They believed that “every piece of knowledge must have a single, unambiguous, authoritative representation within a system.” As the field of analytics engineering and [data transformation](https://www.getdbt.com/blog/data-transformation) develops, there’s a growing need to adopt [software engineering best practices](https://www.getdbt.com/product/what-is-dbt/), including writing DRY code. ## Why write DRY code? DRY code is one of the practices that makes a good developer, a great developer. Solving a problem by any means is great to a point, but eventually, you need to be able to write code that's maintainable by people other than yourself and scalable as system load increases. That's the essence of DRY code. But what's so great about being DRY as a bone anyway, when you can be WET? ### Don’t be WET WET, which stands for “Write Everything Twice,” is the opposite of DRY. It's a tongue-in-cheek reference to code that doesn’t exactly meet the DRY standard. In a practical sense, WET code typically involves the repeated _writing_ of the same code throughout a project, whereas DRY code would represent the repeated _reference_ of that code. Well, how would you know if your code isn't DRY enough? That’s kind of subjective and will vary by the norms set within your organization. That said, a good rule of thumb is [the Rule of Three](https://en.wikipedia.org/wiki/Rule_of_three_(writing)#:~:text=The%20rule%20of%20three%20is,or%20effective%20than%20other%20numbers.). This rule states that the _third_ time you encounter a certain pattern, you should probably abstract it into some reusable unit. There is, of course, a tradeoff between simplicity and conciseness in code. The more abstractions you create, the harder it can be for others to understand and maintain your code without proper documentation. So, the moral of the story is: DRY code is great as long as you [write great documentation.](https://docs.getdbt.com/docs/build/documentation) ### Save time & energy DRY code means you get to write duplicate code less often. You're saving lots of time writing the same thing over and over. Not only that, but you're saving your cognitive energy for bigger problems you'll end up needing to solve, instead of wasting that time and energy on tedious syntax. Sure, you might have to frontload some of your cognitive energy to create a good abstraction. But in the long run, it'll save you a lot of headaches. Especially if you're building something complex and one typo can be your undoing. ### Create more consistent definitions Let's go back to what Andy and Dave said in _The Pragmatic Programmer_: “Every piece of knowledge must have a single, unambiguous, authoritative representation within a system.” As a data person, the words “single” and “unambiguous” might have stood out to you. Most teams have essential business logic that defines the successes and failures of a business. For a subscription-based DTC company, this could be [monthly recurring revenue (MRR)](https://www.getdbt.com/blog/modeling-subscription-revenue/) and for a SaaS product, this could look like customer lifetime value (CLV). Standardizing the SQL that generates those metrics is essential to creating consistent definitions and values. By writing DRY definitions for key business logic and metrics that are referenced throughout a dbt project and/or BI (business intelligence) tool, data teams can create those single, unambiguous, and authoritative representations for their essential transformations. Gone are the days of 15 different definitions and values for churn, and in are the days of standardization and DRYness. ### dbt Semantic Layer, powered by MetricFlow The [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl), powered by [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow), simplifies the process of defining and using critical business metrics, like revenue in the modeling layer (your dbt project). By centralizing metric definitions, data teams can ensure consistent self-service access to these metrics in downstream data tools and applications. The dbt Semantic Layer eliminates duplicate coding by allowing data teams to define metrics on top of existing models and automatically handles data joins. ## Tools to help you write DRY code Let’s just say it: Writing DRY code is easier said than done. For classical software engineers, there’s a ton of resources out there to help them write DRY code. In the world of data transformation, there are also some tools and methodologies that can help folks in [the field of analytics engineering](https://www.getdbt.com/blog/what-is-analytics-engineering) write more DRY and [modular code](https://www.getdbt.com/blog/modular-data-modeling-techniques). ### Common Table Expressions (CTEs) [CTEs](https://www.getdbt.com/blog/guide-to-cte) are a great way to help you write more DRY code in your data analysis and dbt models. In a formal sense, a CTE is a temporary results set that can be used in a query. In a much more human and practical sense, we like to think of CTEs as separate, smaller queries within the larger query you’re building up. Essentially, you can use CTEs to break up complex queries into simpler blocks of code that are easier to debug and can connect and build off of each other. If you’re referencing a specific query, perhaps for aggregations that join back to an unaggregated view, CTEs can simply be referenced throughout a query with its `CTE_EXPRESSION_NAME.` ### View materializations View [materializations](https://docs.getdbt.com/docs/build/materializations) are also extremely useful for abstracting code that might otherwise be repeated often. A view is a defined passthrough SQL query that can be run against a database. Unlike a table, it doesn’t store data, but it defines the logic that you need to use to fetch the underlying data. If you’re referencing the same query, CTE, or block of code, throughout multiple data models, that’s probably a good sign that code should be its own view. For example, you might define a SQL view to count new users created in a day: ```sql select created_date, count(distinct(user_id)) as new_users from {{ ref('users') }} group by created_date ``` While this is a simple query, writing this logic every time you need it would be super tedious. And what if the `user_id` field changed to a new name? If you’d written this in a WET way, you’d have to find every instance of this code and make the change to the new field versus just updating it once in the code for the view. To make any subsequent references to this view DRY-er, you simply reference the view in your data model or query. ### dbt macros and packages dbt also supports the use of [macros](https://docs.getdbt.com/docs/build/jinja-macros) and [packages](https://docs.getdbt.com/docs/build/packages) to help data folks write DRY code in their dbt projects. Macros are Jinja-supported functions that can be reused and applied throughout a dbt project. Packages are libraries of dbt code, typically models, macros, and/or tests, that can be referenced and used in a dbt project. They are a great way to use transformations for common data sources (like [ad platforms](https://hub.getdbt.com/dbt-labs/facebook_ads/latest/)) or use more [custom tests for your data models](https://hub.getdbt.com/calogica/dbt_expectations/0.1.2/) _without having to write out the code yourself_. At the end of the day, is there really anything more DRY than that? ## Conclusion DRY code is a principle that you should always be striving for. It saves you time and energy. It makes your code more maintainable and extensible. And potentially most importantly, it’s the fine line that can help transform you from a good analytics engineer to a great one. ## Further reading - [Data modeling technique for more modularity](https://www.getdbt.com/analytics-engineering/modular-data-modeling-technique/) - [Why we use so many CTEs](https://docs.getdbt.com/docs/best-practices) - [Getting started with CTEs](https://www.getdbt.com/blog/guide-to-cte) ## DRY principle FAQs **What is the "Don't Repeat Yourself" (DRY) principle and why is it important in software development?** DRY means each piece of knowledge or logic is defined once and reused everywhere it’s needed. It matters because it cuts duplication, reduces bugs from copy‑paste drift, speeds up changes (edit in one place), and keeps systems easier to understand, maintain, and scale. **What does the DRY principle assert about how knowledge should be represented within a software system?** It asserts that every piece of knowledge must have a single, unambiguous, authoritative representation. In practice, that means centralizing definitions so updates propagate consistently and you avoid conflicting implementations of the same idea. **How does applying DRY improve maintainability, readability, consistency, and reduce errors in a codebase?** Maintainability improves because you change one source instead of hunting many copies; readability improves through clear, well‑named abstractions; consistency improves when everyone references the same logic; and errors drop because copy‑paste mistakes and divergent fixes are eliminated. **What practical techniques can developers use to enforce DRY, such as creating functions, using classes/inheritance, extracting constants, or modularizing code?** Use reusable functions/methods and modules; encapsulate domain logic in classes or components; extract constants/configuration to a single location; factor repeated patterns into templates/macros; and package shared code into libraries. In SQL and analytics, apply DRY via common table expressions (CTEs) to factor subqueries, views/materializations to centralize reusable logic, macros/packages to parameterize repeated patterns, and a semantic layer to define metrics once for consistent reuse. **How would you know if your code isn't DRY enough?** Look for repetition you’ve copy‑pasted, “change amplification” where a small update requires edits in many places, and inconsistencies in outputs that should match. A practical heuristic is the Rule of Three: by the third time you write the same pattern, abstract it. Balance DRY with clarity—don’t over‑abstract—and document the abstractions you create. --- --- title: "Guide to DML" description: "Data Manipulation Language (DML) is a class of SQL statements that are used to query, edit, add and delete row-level data." url: "https://www.getdbt.com/blog/guide-to-dml" date: "2024-07-01" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Guide to DML Data Manipulation Language (DML) is a class of SQL statements that are used to query, edit, add and delete row-level data from database tables or views. The main DML statements are `SELECT`, `INSERT`, `DELETE`, and `UPDATE`. DML is contrasted with [Data Definition Language (DDL)](https://www.getdbt.com/blog/guide-to-ddl) which is a series of SQL statements that you can use to edit and manipulate the _structure_ of databases and the objects in them. Similar to DDL, DML can be a _tad_ bit boring. However, DML statements are what allows analysts and analytics engineers to do their work. We hope you can use this glossary to understand when and why DML statements are used and how they may contrast with similar DDL commands. ## Types of DML Statements The primary DML statements are `SELECT`, `INSERT`, `DELETE`, and `UPDATE`. With the exception of `SELECT` statements, all of the others are only applicable to data within tables in a database. The primary difference between `SELECT` and all the other DML statements is its impact to row-level data: - To _change_ the actual data that lives in tables, use `INSERT`, `DELETE`, and `UPDATE` statements - To _access_ the data in database object, use `SELECT` statements **** ### SELECT Ah, our favorite of DML statements! This is the SQL we all know and love (most of the time). Because the `SELECT` statement allows you to access and manipulate data that exists in database objects, it makes it the true powerhouse in data analysis and analytics engineering. You write `SELECT` statements to create queries that build data models and perform robust analysis. With `SELECT` statements, you can join different views and tables, qualify data by setting filters, apply functions and operators on the data, and more. `SELECT` statements, unlike `INSERT`, `DELETE`, and `UPDATE`, don’t actually change the row-level value stored in the tables/views. Instead, you write `SELECT` statements to express the business logic needed to perform analysis. All `SELECT` statements need three elements: a `SELECT` clause in the beginning, the actual field selection and manipulation, and a `FROM` statement which is specifying which database object you’re trying to access. Here’s an example `SELECT` statement: ```sql select payment_method, sum(amount) AS amount from {{ ref('raw_payments') }} group by 1 ``` In this example, your selection of the `payment_method` column and summation of the `amount` column is the meat of your query. The `from {{ ref('raw_payments') }}` specifies the actual table you want to do the selecting from. ### INSERT Using the `INSERT` DML command, you can add rows to a table that exists in your database. To be honest, data folks are rarely inserting data into tables manually with the `INSERT` command. Instead, data team members will most often use data that’s already been inserted by an ELT tool or other data ingestion process. You can insert a record [in jaffle_shop’s](https://github.com/dbt-labs/jaffle_shop) `raw_customers` table like this: ```sql INSERT INTO raw_customers VALUES (101, 'Kira', 'F.'); ``` As you can see from this example, you clearly set all the column values that exist in your `raw_customers` table. For `INSERT` statements, you can explicitly specify the values you want to insert or use a query result to set the column values. ### DELETE The `DELETE` command will remove rows in an existing table in your database. In practice, you will usually specify a `WHERE` clause with your `DELETE` statement to only remove specific rows from a table. But, you shouldn't really ever delete rows from tables. Instead, you should apply filters on queries themselves to remove rows from your modeling or analysis. For the most part, if you wanted to remove all existing rows in a table, but keep the underlying table structure, you would use the `TRUNCATE` DDL command. If you wanted to remove all rows and drop the entire table, you could use the `DROP` DDL command. You can delete the record for any Henry W. in jaffle_shop’s `customers` table by executing this statement: ```sql DELETE FROM customers WHERE first_name = 'Henry' AND last_name = 'W.'; ``` ### UPDATE With the `UPDATE` statement, you can change the actual data in existing rows in a table. Unlike the `ALTER` DDL command that changes the underlying structure or naming of database objects, the `UPDATE` statement will alter the actual row-level data. You can qualify an `UPDATE` command with a `WHERE` statement to change the values of columns of only specific rows. You can manually update the status column of an order in your orders table like this: ```sql UPDATE orders SET status = 'returned' WHERE order_id = 7; ``` Tip - The `UPDATE` statement is often compared to the `MERGE` statement. With `MERGE` statements, you can insert, update, _and_ delete records in a single command. Merges are often utilized when there is data between two tables that needs to be reconciled or updated. You'll see merges most commonly executed when a source table is updated and a downstream table needs to be updated as a result of this change. Learn more about [how dbt uses merges in incremental models here](https://docs.getdbt.com/docs/build/incremental-models-overview#how-incremental-models-work-in-dbt). ## Conclusion DML statements allow you to query, edit, add, and remove data stored in database objects. The primary DML commands are `SELECT`, `INSERT`, `DELETE`, and `UPDATE`. Using DML statements, you can perform powerful actions on the actual data stored in your system. You'll typically see DML `SELECT` statements written in data models to conduct data analysis or create new tables and views. In many ways, DML is the air that us data folks breathe! ## DML in dbt FAQ **What is DML and how does it differ from DDL?** Data Manipulation Language (DML) is a class of SQL statements used to query, edit, add, and delete row-level data from database tables or views. While DML (SELECT, INSERT, DELETE, and UPDATE) handles the data within database objects, Data Definition Language (DDL) manipulates the structure of databases and their objects. DML is essentially the foundation for analytics engineers and analysts to perform their daily work with data. **How is the SELECT statement used in dbt?** The SELECT statement is the primary DML statement used in dbt for creating queries that build data models and performing analysis. It allows you to access and manipulate data from database objects without changing the underlying row-level values, joining different views and tables, applying filters, and implementing business logic. **How do MERGE statements relate to DML in dbt?** MERGE statements combine the functionality of INSERT, UPDATE, and DELETE operations in a single command, making them particularly useful for incremental models in dbt. When implementing incremental models, dbt uses merge operations to reconcile data between source and target tables, efficiently updating only what has changed. This approach is more efficient than rebuilding entire tables and represents a powerful application of DML operations in modern data workflows. **Why don't dbt users typically write DML statements directly?** dbt users generally don't write direct DML statements because dbt abstracts these operations through its modeling framework and materialization strategies. Instead of manually writing INSERT or UPDATE statements, dbt users define the desired state of data through SELECT statements in models, and dbt handles the underlying DML operations necessary to create or update tables. This approach allows analysts to focus on business logic rather than database manipulation details. ## Further reading For more resources on why people who use dbt don’t write DML, check out the following: - [Why not write DML](https://docs.getdbt.com/faqs/Project/why-not-write-dml) - [SQL dialect](https://docs.getdbt.com/faqs/Models/sql-dialect) For database-specific DML documents, please check out the resources below: - [DML in Snowflake](https://docs.snowflake.com/en/sql-reference/sql-dml.html) - [Updating tables with DML commands in Redshift](https://docs.aws.amazon.com/redshift/latest/dg/t_Updating_tables_with_DML_commands.html) - [DML in Google BigQuery](https://cloud.google.com/bigquery/docs/reference/standard-sql/data-manipulation-language) - [Delta Lake DML for Databricks](https://databricks.com/blog/2020/09/29/diving-into-delta-lake-dml-internals-update-delete-merge.html) --- --- title: "Guide to dimensional modeling" description: "Dimensional modeling is a data modeling technique where you break data up into “facts” and “dimensions”." url: "https://www.getdbt.com/blog/guide-to-dimensional-modeling" date: "2024-07-01" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Guide to dimensional modeling Dimensional modeling is a data modeling technique where you break data up into “facts” and “dimensions” to organize and describe entities within your data warehouse. The result is a staging layer in the data warehouse that cleans and organizes the data into the business end of the warehouse that is more accessible to data consumers. By breaking your data down into clearly defined and organized entities, your consumers can make sense of what that data is, what it’s used for, and how to join it with new or additional data. Ultimately, using dimensional modeling for your data can help create the appropriate layer of models to expose in an end business intelligence (BI) tool. There are a few different methodologies for dimensional modeling that have evolved over the years. The big hitters are the Kimball methodology and the Inmon methodology. Ralph Kimball’s work formed much of the foundation for how data teams approached data management and data modeling. Here, we’ll focus on dimensional modeling from Kimball’s perspective—why it exists, where it drives value for teams, and how it’s evolved in recent years. ## What are we trying to do here? Let’s take a step back for a second and ask ourselves: why should you read this glossary page? What are you trying to accomplish with dimensional modeling and data modeling in general? Why have you taken up this rewarding, but challenging career? Why are _you_ here? This may come as a surprise to you, but we’re not trying to build a top-notch foundation for analytics—we’re actually trying to build a bakery. Not the answer you expected? Well, let’s open up our minds a bit and explore this analogy. If you run a bakery (and we’d be interested in seeing the data person + baker venn diagram), you may not realize you’re doing a form of dimensional modeling. What’s the final output from a bakery? It’s that glittering, glass display of delicious-looking cupcakes, cakes, cookies, and everything in between. But a cupcake just didn’t magically appear in the display case! Raw ingredients went through a rigorous process of preparation, mixing, melting, and baking before they got there. Just as eating raw flour isn’t that appetizing, neither is deriving insights from raw data since it rarely has a nice structure that makes it poised for analytics. There’s some considerable work that’s needed to organize data and make it usable for business users. This is where dimensional modeling comes into play; it’s a method that can help data folks create meaningful entities (cupcakes and cookies) to live inside their [data mart](https://docs.getdbt.com/best-practices/how-we-structure/4-marts) (your glass display) and eventually use for business intelligence purposes (eating said cookies). So I guess we take it back—you’re not just trying to build a bakery, you’re also trying to build a top-notch foundation for meaningful analytics. Dimensional modeling can be a method to get you part of the way there. ## Facts vs. dimensions The ultimate goal of dimensional modeling is to be able to categorize your data into their fact or dimension models, making them the key components to understand. So what are these components? ### Facts A fact is a collection of information that typically refers to an action, event, or result of a business process. As such, people typically liken facts to verbs. In terms of a real business, some facts may look like account creations, payments, or emails sent. It’s important to note that fact tables act as a historical record of those actions. You should almost never overwrite that data when it needs updating. Instead, you add new data as additional rows onto that table. For many businesses, marketing and finance teams need to understand all the touchpoints leading up to a sale or conversion. A fact table for a scenario like this might look like a `fct_account_touchpoints` table: Accounts may have many touch points and this table acts as a true log of events leading up to an account conversion. This table is great and all for helping understanding what might have led to a conversion or account creation, but what if business users need additional context on these accounts or touchpoints? That’s where dimensions come into play. ### Dimensions A dimension is a collection of data that describe who or what took action or was affected by the action. Dimensions are typically likened to nouns. They add context to the stored events in fact tables. In terms of a business, some dimensions may look like users, accounts, customers, and invoices. A noun can take multiple actions or be affected by multiple actions. It’s important to call out: a noun doesn’t become a new thing whenever it does something. As such, when updating dimension tables, you should overwrite that data instead of duplicating them, like you would in a fact table. Following the example from above, a dimension table for this business would look like an `dim_accounts` table with some descriptors: In this table, each account only has one row. If an account’s name or status were to be updated, new values would overwrite existing records versus appending new rows. **** ### Facts and dimensions at play with each other Cool, you think you’ve got some facts and dimensions that can be used to qualify your business. There’s one big consideration left to think about: how do these facts and dimensions interact with each other? ![Fact star](https://cdn.sanity.io/images/wl0ndo6t/main/9af3743dc43f5db3beb29d177745b2594c9158a5-1540x1234.png) Pre-cloud data warehouses, there were two dominant design options, star schemas and snowflake schemas, that were used to concretely separate out the lines between fact and dimension tables. - In a star schema, there’s one central fact table that can join to relevant dimension tables. - A snowflake schema is simply an extension of a star schema; dimension tables link to other dimension tables making it form a snowflake-esque shape. It sounds really nice to have this clean setup with star or snowflake schemas. Almost as if it’s too good to be true (and it very well could be). The development of cheap cloud storage, BI tools great at handling joins, the evolution of SQL capabilities, and data analysts with growing skill sets have changed the way data folks use to look at dimensional modeling and star schemas. Wide tables consisting of fact and dimension tables joined together are now a competitive option for data teams. Below, we’ll dig more into the design process of dimensional modeling, wide tables, and the beautiful ambiguity of it all. ## The dimensional modeling design process According to the Kimball Group, the official(™) four-step design process is (1) selecting a business process to analyze, (2) declaring the [grain](https://www.getdbt.com/blog/guide-to-data-grain), (3) Identifying the dimensions, and (4) Identifying the facts. That makes dimensional modeling sound really easy, but in reality, it’s packed full of nuance. Coming back down to planet Earth, your design process is how you make decisions about: - Whether something should be a fact or a dimension - Whether you should keep fact and dimension tables separate or create wide, joined tables This is something that data philosophers and thinkers could debate long after we’re all gone, but let’s explore some of the major questions to hold you over in the meantime. ### Should this entity be a fact or dimension? Time to put on your consultant hat because that dreaded answer is coming: it depends. This is what makes dimensional modeling a challenge! Kimball would say that a fact must be numeric. The inconvenient truth is: an entity can be viewed as a fact or a dimension depending on the analysis you are trying to run. ### Birds of a feather If you ran a clinic, you would probably have a log of appointments by patient. At first, you could think of appointments as facts—they are, after all, events that happen and patients can have multiple appointments—and patients as dimensions. But what if your business team really cared about the appointment data itself—how well it went, when it happened, the duration of the visit. You could, in this scenario, make the case for treating this appointments table as a dimension table. If you cared more about looking at your data at a patient-level, it probably makes sense to keep appointments as facts and patients as dimensions. All this to say is that there’s inherent complexity in dimensional modeling, and it’s up to you to draw those lines and build those models. So then, how do you know which is which if there aren’t any hard rules!? Life is a gray area, my friend. Get used to it. A general rule of thumb: go with your gut! If something feels like it should be a fact to meet your stakeholders' needs, then it’s a fact. If it feels like a dimension, it’s a dimension. The world is your oyster. If you find that you made the wrong decision down the road, (it’s usually) no big deal. You can remodel that data. Just remember: you’re not a surgeon. No one will die if you mess up (hopefully). So, just go with what feels right because you’re the expert on your data 👉😎👉 Also, this is why we have data teams. Dimensional modeling and data modeling is usually a collaborative effort; working with folks on your team to understand the data and stakeholder wants will ultimately lead to some rad data marts. ### Should I make a wide table or keep them separate? Yet again, it depends. Don’t roll your eyes. Strap in for a quick history lesson because the answer to this harkens back to the very inception of dimensional modeling. Back in the day before cloud technology adoption was accessible and prolific, storing data was expensive and joining data was relatively cheap. Dimensional modeling came about as a solution to these issues. Separating collections of data into smaller, individual tables (star schema-esque) made the data cheaper to store and easier to understand. So, individual tables were the thing to do back then. Things are different today. Cloud storage costs have gotten really inexpensive. Instead, computing is the primary cost driver. Now, keeping all of your tables separate can be expensive because every time you join those tables, you’re spending usage credits. Should you just add everything to one, wide table? No. One table will never rule them all. Knowing whether something should be its own fact table or get added on to an existing table generally comes down to understanding who will be your primary end consumers. For end business users who are writing their own SQL, feel comfortable performing joins, or use a tool that joins tables for them, keeping your data as separate fact and dimension tables is pretty on-par. In this setup, these users have the freedom and flexibility to join and explore as they please. If your end data consumers are less comfortable with SQL and your BI tool doesn’t handle joins well, you should consider joining several fact and dimension tables into wide tables. Another consideration: these wide, heavily joined tables can tend to wind up pretty specialized and specific to business departments. Would these types of wide tables be helpful for you, your data team, and your business users? Well, that’s for you to unpack. ## Advantages and disadvantages of dimensional modeling The benefits and drawbacks of dimensional modeling are pretty straightforward. Generally, the main advantages can be boiled down to: - **More accessibility**: Since the output of good dimensional modeling is a [data mart](https://docs.getdbt.com/best-practices/how-we-structure/4-marts), the tables created are easier to understand and more accessible to end consumers. - **More flexibility**: Easy to slice, dice, filter, and view your data in whatever way suits your purpose. - **Performance**: Fact and dimension models are typically materialized as tables or [incremental models](https://docs.getdbt.com/docs/build/incremental-models). Since these often form the core understanding of a business, they are queried often. Materializing them as tables allows them to be more performant in downstream BI platforms. The disadvantages include: - **Navigating ambiguity**: You need to rely on your understanding of your data and stakeholder wants to model your data in a comprehensible and useful way. What you know about your data and what people really need out of the data are two of the most fundamental and difficult things to understand and balance as a data person. - **Utility limited by your BI tool**: Some BI tools don’t handle joins well, which can make queries from separated fact and dimensional tables painful. Other tools have long query times, which can make querying from ultra-wide tables not fun. ## Conclusion Dimensional data modeling is a data modeling technique that allows you to organize your data into distinct entities that can be mixed and matched in many ways. That can give your stakeholders a lot of flexibility. [While the exact methodologies have changed](https://www.getdbt.com/analytics-engineering/modular-data-modeling-technique/)—and will continue to, the philosophical principle of having tables that are sources of truth and tables that describe them will continue to be important in the work of analytics engineering practitioners. ## Dimensional Modeling FAQ **What is dimensional modeling?** Dimensional modeling is a data modeling technique that organizes data into "facts" and "dimensions" within your data warehouse. Facts typically represent business events or actions (like account creations or payments), while dimensions provide context about those events (such as user accounts or customers). This approach creates a more accessible staging layer in the data warehouse that makes data easier for business users to understand and analyze. **How do facts and dimensions differ from each other?** Facts are collections of information that typically refer to actions, events, or business processes (think of them as verbs) and are stored as historical records that shouldn't be overwritten. Dimensions, on the other hand, describe who or what took the action (think of them as nouns) and provide context to the events in fact tables, with updates typically overwriting existing records rather than adding new rows. **What are star and snowflake schemas in dimensional modeling?** A star schema features one central fact table that can join to relevant dimension tables, creating a structure that resembles a star. A snowflake schema is an extension of this approach where dimension tables link to other dimension tables, forming a snowflake-like shape. Both are traditional approaches to organizing facts and dimensions, though modern data practices sometimes favor wide tables combining elements of both. **What are the main advantages of dimensional modeling?** Dimensional modeling increases data accessibility by creating data marts with tables that are easier for end users to understand. It offers greater flexibility for slicing, filtering, and analyzing data in various ways, and typically improves performance since fact and dimension models are materialized as tables or incremental models, making them more efficient for downstream BI platforms. **Should I use separate fact and dimension tables or create wide tables?** The decision depends largely on your end users and technical environment. If your users are comfortable with SQL and your BI tool handles joins well, separate fact and dimension tables provide more flexibility. For users less comfortable with SQL or when using BI tools that struggle with joins, wide tables that combine facts and dimensions may be more appropriate, despite potentially becoming specialized to specific business departments. ## Additional Reading Dimensional modeling is a tough, complex, and opinionated topic in the data world. Below you’ll find some additional resources that may help you identify the data modeling approach that works best for you, your data team, and your end business users: - [Modular data modeling techniques](https://www.getdbt.com/analytics-engineering/modular-data-modeling-technique/) - [Stakeholder-friendly model naming conventions](https://docs.getdbt.com/blog/stakeholder-friendly-model-names/) - [How we structure our dbt projects guide](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview) --- --- title: "Structuring databases with DDL" description: "Data Definition Language (DDL) is a group of SQL statements that you can execute to manage database objects." url: "https://www.getdbt.com/blog/structuring-databases-with-ddl" date: "2024-07-01" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Structuring databases with DDL Data Definition Language (DDL) is a group of SQL statements that you can execute to manage database objects, including tables, views, and more. Using DDL statements, you can perform powerful commands in your database such as creating, modifying, and dropping objects. DDL commands are usually executed in a SQL browser or stored procedure. DDL is contrasted with [Data Manipulation Language (DML)](https://www.getdbt.com/blog/guide-to-dml) which is the SQL that is used to actually access and manipulate data in database objects. The majority of data analysts will rarely execute DDL commands and will do the majority of their work creating DML statements to model and analyze data. Note - Data folks don’t typically write DDL [since dbt will do it for them](https://docs.getdbt.com/docs/about/overview#:~:text=dbt%20allows%20analysts%20avoid%20writing,dbt%20takes%20care%20of%20materialization.). To be honest, DDL is definitely some of the drier content that exists out there in the greater data world. However, because DDL commands are often uncompromising and should be used with caution, it’s incredibly important to understand how they work and when they should be used. We hope you can use this page to learn about the basics, strengths, and limitations of DDL statements. ## Types of DDL Statements DDL statements are used to create, drop, and manipulate objects in your database. They are often, but not always, unforgiving and irreversible. “With great power comes great responsibility,” is usually the first thing I think of before I execute a DDL command. We’ll highlight some of the primary DDL commands that are used by analytics engineers below. Important - The syntax for DDL commands can be pretty database-specific. We are trying to make this glossary page as generic as possible, but please use the “Further Reading” section to see the specifics on how the following DDL commands would be implemented in your database of interest! ### ALTER Using the `ALTER` DDL command, you can change an object in your database that already exists. By "change", we specifically mean you can: - Add new, remove, and rename columns to views and tables - Rename a view or table - Modify the structure of a view or table - And more! The generic syntax to use the ALTER command is as follows: ```sql ALTER ; ``` To alter a table’s column, you may do something like this: ```sql ALTER TABLE customers rename column last_name as last_initial; ``` In this example, you have to rename the `last_name` column [in jaffle_shop’s](https://github.com/dbt-labs/jaffle_shop) `customers` table to be called `last_initial`. ### DROP The `DROP` command. Probably the most high-stakes DDL statement one can execute. One that should be used with the _utmost_ of care. At its core, an executed `DROP` statement will remove that object from the data warehouse. You can drop tables, views, schemas, databases, users, functions, and more. Some data warehouses such as Snowflake allow you to add restrictions to `DROP` statements to caution you about the impact of dropping a table, view, or schema before it’s actually dropped. In practice, we recommend you never drop raw source tables as they are often your baseline of truth. Your database user also usually needs the correct permissions to drop database objects. The syntax to use the `DROP` command is as follows: ```sql DROP ; ``` You can drop your `customer` table like this: ```sql DROP TABLE customers; ``` ### CREATE With the `CREATE` statement, you can create new objects in your data warehouse. The most common objects created with this statement are tables, schemas, views, and functions. Unlike `DROP`, `ALTER`, and `TRUNCATE` commands, there’s little risk with running `CREATE` statements since you can always drop what you create. Creating tables and views with the `CREATE` command requires a strong understanding of how you want the data structured, including column name and data type. Using the `CREATE` command to establish tables and views can be laborious and repetitive, especially if the schema objects contain many columns, but is an effective way to create new objects in a database. After you create a table, you can use DML `INSERT` statements and/or a transformation tool such as dbt to actually get data in it. The generic syntax to use the `CREATE` command is as follows: ```sql CREATE ; ``` Creating a table using the `CREATE` statement may look a something like this: ```sql CREATE TABLE prod.jaffle_shop.jaffles ( id varchar(255), jaffle_name varchar(255) created_at timestamp, ingredients_list varchar(255), is_active boolean ); ``` Note that you had to explicitly define column names and column data type here. _You must have a strong understanding of your data’s structure when using the CREATE command for tables and views._ ### TRUNCATE The `TRUNCATE` command will remove all rows from a table while maintaining the underlying table structure. The `TRUNCATE` command is only applicable for table objects in a database. Unlike `DROP` statements, `TRUNCATE` statements don’t remove the actual table from the database, just the data stored in them. The syntax to use the `TRUNCATE` command is as follows: ```sql TRUNCATE TABLE ; ``` You can truncate your jaffle_shop’s `payments` table by executing this statement: ```sql TRUNCATE TABLE payments; ``` Previously, this table was 113 rows. After executing this statement, the table is still in your database, but now has zero rows. ## Conclusion DDL statements allow you to remove, edit, and add database objects. Some of the most common DDL statements you’ll execute include `CREATE`, `DROP`, `COMMENT`, `ALTER`, and more. DDL commands are typically executed in a SQL browser or stored procedure. Ultimately, DDL commands are all-powerful and potentially high-risk and should be used with the greatest of care. In the case of DDL, **do not** throw caution to the wind… ## DDL on dbt FAQs **What is Data Definition Language (DDL) and how does it differ from DML?** Data Definition Language (DDL) is a group of SQL statements used to manage database objects like tables and views. Unlike Data Manipulation Language (DML) which accesses and manipulates data within objects, DDL focuses on creating, modifying, and removing the objects themselves. Most data analysts primarily work with DML statements, often written by analytics engineers or automated through tools like dbt. **What are the main types of DDL statements used by analytics engineers?** The primary DDL commands include CREATE (to make new database objects), DROP (to remove objects entirely), ALTER (to modify existing objects), and TRUNCATE (to remove all rows while keeping the table structure). Each command serves a specific purpose in database management, with varying levels of risk and impact on your data infrastructure. **Why should DDL commands be used with caution?** DDL commands are often unforgiving and irreversible, particularly the DROP statement which permanently removes objects from your data warehouse. As noted in the article, "With great power comes great responsibility" applies to executing DDL commands. It's recommended never to drop raw source tables as they often represent your baseline of truth, and you should ensure you have proper permissions before executing these high-stakes commands. **Do data professionals typically need to write DDL when using dbt?** Data professionals using dbt typically don't need to write DDL statements directly since dbt handles the generation and execution of DDL for them. This automation eliminates the laborious and repetitive process of manually creating complex schema objects with many columns. By abstracting away the DDL implementation details, dbt allows data teams to focus on modeling and analysis rather than database structure management. ## Further reading For database-specific DDL resources, check out the following: - [DDL commands in Snowflake](https://docs.snowflake.com/en/sql-reference/sql-ddl-summary.html) - [SQL commands in Amazon Redshift](https://docs.aws.amazon.com/redshift/latest/dg/c_SQL_commands.html) (contains DDL) - [DDL statements in Google BigQuery](https://cloud.google.com/bigquery/docs/reference/standard-sql/data-definition-language) - [DDL statements in Databricks](https://docs.databricks.com/sql/language-manual/index.html#ddl-statements) - [DDL in Amazon Athena](https://docs.aws.amazon.com/athena/latest/ug/language-reference.html) --- --- title: "Getting started with data lineage" description: "Data lineage provides a holistic view of how data moves through an organization, where it’s transformed and consumed." url: "https://www.getdbt.com/blog/getting-started-with-data-lineage" date: "2024-07-01" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Getting started with data lineage Data lineage provides a holistic view of how data moves through an organization, where it’s transformed and consumed. Overall, data lineage is a fundamental concept to understand in the practice of analytics engineering and modern data work. At a high level, a data lineage system typically provides data teams and consumers with one or both of the following resources: - A visual graph (DAG) of sequential workflows at the data set or column level - A data catalog of data asset origins, owners, definitions, and policies This holistic view of the data pipeline allows data teams to build, troubleshoot, and analyze workflows more efficiently. It also enables business users to understand the origins of reporting data and provides a means for data discovery. We’ll unpack why data lineage is important, how it works in the context of analytics engineering, and where some existing challenges still exist for data lineage. ## Why is data lineage important? As a data landscape grows in size and complexity, the benefits of data lineage become more apparent. For data teams, the three main advantages of data lineage include reducing root-cause analysis headaches, minimizing unexpected downstream headaches when making upstream changes, and empowering business users. ### Root cause analysis It happens: dashboards and reporting fall victim to data pipeline breaks. Data teams quickly need to diagnose what’s wrong, fix where things may be broken, and provide up-to-date numbers to their end business users. But when these breaks happen (and they surely do) how can teams quickly identify the root cause of the problem? If data teams have some form of data lineage in place, they can more easily identify the root cause of the broken pipeline or data quality issue. By backing out into the data models, sources, and pipelines powering a dashboard a report, data teams can understand all the upstream elements impacting that work and see where the issues lie. Will a data lineage or a DAG solve your breaking pipelines? Definitely not. Will it potentially make your life easier to find problems in your data work? Heck yes. ### Downstream impacts on upstream changes You may have been here—your backend engineering team drops the `customers` table to create a newer, more accurate `users` table. The only bad thing is…[they forgot to tell the data team about the change](https://docs.getdbt.com/blog/when-backend-devs-spark-joy). When you have a data lineage system, you can visually see which downstream models, nodes, and exposures are impacted by big upstream changes such as source or model renaming or removals. Referring to your DAG or data lineage system before any significant change to your analytics work is a great way to help prevent accidental downstream issues. ### Value to business users While data lineage makes it easier for data teams to manage pipelines, stakeholders and leaders also benefit from data lineage, primarily around promoting data transparency into the data pipelines. **Shared data literacy** New hires, existing team members, and internal data practitioners can independently explore a holistic view of the data pipeline with a data lineage system. For data teams using a DAG to encapsulate their data work, business users have a clear visual representation of how data flows from different sources to the dashboards they consume in their BI tool, providing an increased level of transparency in data work. At the end of the day, the added visibility makes it easier for everyone to be on the same page. **Pipeline cleanliness** A visual graph (DAG) of how data flows through various workflows makes it easy to identify redundant loads of source system data or workflows that produce identical reporting insights. Spotlighting redundant data models can help trim down on WET (write every time/write everything twice) code, non-performant joins, and ultimately help promote reusability, modularity, and standardization within a data pipeline. Overall, data lineage and data-driven business go hand-in-hand. A data lineage system allows data teams to be more organized and efficient, business users to be more confident, and data pipelines to be more modular. ## How does data lineage work? In the greater data world, you may often hear of data lineage systems based on tagging, patterns or parsing-based systems. In analytics engineering however, you’ll often see data lineage implemented in a DAG or through third-party tooling that integrates into your data pipeline. ### DAGs (directed acyclic graphs) If you use a transformation tool such as dbt that automatically infers relationships between data sources and models, a DAG automatically populates to show you the lineage that exists for your [data transformations](https://www.getdbt.com/analytics-engineering/transformation/). ![dbt Cloud Project with generated DAG](https://cdn.sanity.io/images/wl0ndo6t/main/320f8f822dc8a595dffdf5e635a8d8266111dce2-1537x982.png) Your [DAG](https://www.getdbt.com/blog/guide-to-dag) is used to visually show upstream dependencies, the nodes that must come before a current model, and downstream relationships, the work that is impacted by the current model. DAGs are also directional—they show a defined flow of movement and form non-cyclical loops. Ultimately, DAGs are an effective way to see relationships between data sources, models, and dashboards. DAGs are also a great way to see visual bottlenecks, or inefficiencies in your data work (see image below for a DAG with...many bottlenecks). Data teams can additionally add [meta fields](https://docs.getdbt.com/reference/resource-configs/meta) and documentation to nodes in the DAG to add an additional layer of governance to their dbt project. ![A bad DAG](https://cdn.sanity.io/images/wl0ndo6t/main/03b33d53cc9ffa35ba36407b9af9ea2a8229fac6-1999x1124.png) ### Automatic > Manual DAGs shouldn’t be dependent on manual updates. Instead, your DAG should be automatically inferred and created with your data transformation and pipelines. Leverage tools such as dbt to build your own version-controlled DAG as you develop your data models. ### Third-party tooling Data teams may also choose to use third-party tools with lineage capabilities such as [Atlan](https://ask.atlan.com/hc/en-us/articles/4433673207313-How-to-set-up-dbt-Cloud), Alation, Collibra, [Datafold](https://www.datafold.com/column-level-lineage), Metaphor, [Monte Carlo](https://docs.getmontecarlo.com/docs/dbt-cloud), [Select Star](https://docs.selectstar.com/integrations/dbt/dbt-cloud), or Stemma. These tools often integrate directly with your data pipelines and dbt workflows and offer zoomed-in data lineage capabilities such as column-level or business logic-level lineage. ## Data lineage challenges The biggest challenges around data lineage become more apparent as your data, systems, and business questions grow. ### Data lineage challenge #1: Scaling data pipelines As dbt projects scale with data and organization growth, the number of sources, models, macros, seeds, and [exposures](https://docs.getdbt.com/docs/build/exposures) invariably grow. And with an increasing number of nodes in your DAG, it can become harder to audit your DAG for WET code or inefficiencies. Working with dbt projects with thousands of models and nodes can feel overwhelming, but remember: your DAG and data lineage are meant to help you, not be your enemy. Tackle DAG audits in chunks, document all models, and [leverage strong structure conventions](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview). **** ### Data lineage challenge #2: Column-level lineage Complex workflows also add to the difficulties a data lineage system will encounter. For example, consider the challenges in describing a data source's movement through a pipeline as it's filtered, pivoted, and joined with other tables. These challenges increase when the granularity of the data lineage shifts from the table to the column level. As data lineage graphs mature and grow, it becomes clear that column- or field-level lineage is often a needed layer of specificity that is not typically built in to data lineage systems. Learn more about the [column-level lineage](https://docs.getdbt.com/docs/collaborate/column-level-lineage) feature in [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) and how it can help you gain insights. ## Conclusion Data lineage is the holistic overview of how data moves through an organization or system, and is typically represented by a DAG. Analytics engineering practitioners use their DAG and data lineage to unpack root causes in broken pipelines, audit their models for inefficiencies, and promote greater transparency in their data work to business users. Overall, using your data lineage and DAG to know when your data is transformed and where it’s consumed is the foundation for good analytics work. ## Data Lineage FAQ **What is data lineage and why is it important?** Data lineage provides a holistic view of how data moves through an organization, where it's transformed and consumed. It typically includes a visual graph (DAG) of workflows and a data catalog of data assets. Data lineage is crucial because it helps with root cause analysis when pipelines break, allows teams to anticipate downstream impacts of upstream changes, and increases transparency for business users. **How does data lineage work in analytics engineering?** In analytics engineering, data lineage is often implemented through a Directed Acyclic Graph (DAG) or through third-party tooling that integrates with your data pipeline. Tools like dbt automatically infer relationships between data sources and models, creating a DAG that visually shows upstream dependencies and downstream relationships without requiring manual updates. **What are the benefits of data lineage for business users?** Data lineage enhances business users' experience by promoting shared data literacy and pipeline cleanliness. It provides a clear visual representation of how data flows from different sources to the dashboards they consume, increasing transparency and helping everyone understand the data journey. Additionally, it helps identify redundant data models, promoting reusability and standardization. **What challenges exist with data lineage systems?** Two major challenges with data lineage are scaling data pipelines and implementing column-level lineage. As dbt projects grow with thousands of models and nodes, auditing DAGs for inefficiencies becomes more difficult. ## Further reading DAGs, data lineage, and root cause analysis…tell me more! Check out some of our favorite resources of writing modular models, DRY code, and data modeling best practices: - [Glossary: DRY](https://www.getdbt.com/blog/guide-to-dry) - [Data techniques for modularity](https://www.getdbt.com/analytics-engineering/modular-data-modeling-technique/) - [How we structure our dbt projects](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview) --- --- title: "DAG use cases and best practices" description: "Use DAGs to visualize dependencies, debug models, and improve performance in your analytics pipelines." url: "https://www.getdbt.com/blog/dag-use-cases-and-best-practices" date: "2024-07-01" authors: ["Daniel Poppy"] categories: ["Learn"] --- # DAG use cases and best practices A DAG is a **D**irected **A**cyclic **G**raph, a type of graph whose nodes are directionally related to each other and don’t form a directional closed loop. In the practice of analytics engineering, DAGs are often used to visually represent the relationships between your data models. While the concept of a DAG originated in mathematics and gained popularity in computational work, DAGs have found a home in the modern data world. They offer a great way to visualize data pipelines and [lineage](https://www.getdbt.com/blog/guide-to-data-lineage), and they offer an easy way to understand dependencies between [data models](https://www.getdbt.com/product/develop). ## DAG use cases and best practices DAGs are an effective tool to help you understand relationships between your data models and areas of improvement for your overall [data transformations](https://www.getdbt.com/analytics-engineering/transformation/), including opportunities for [cost optimization](https://www.getdbt.com/product/cost-optimization). ### Unpacking relationships and data lineage Can you look at one of your data models today and quickly identify all the upstream and downstream models? If you can’t, that’s probably a good sign to start building or looking at your existing DAG. **Upstream or downstream?** How do you know if a model is upstream or downstream from the model you’re currently looking at? Upstream models are models that must be performed prior to the current model. In simple terms, the current model depends on upstream models in order to exist. Downstream relationships are the outputs from your current model. In a visual DAG, such as the dbt Lineage Graph, upstream models are to the left of your selected model and downstream models are to the right of your selected model. Ever confused? Use the arrows that create the directedness of a DAG to understand the direction of movement. One of the great things about DAGs is that they are _visual_. You can clearly identify the nodes that connect to each other and follow the lines of directions. When looking at a DAG, you should be able to identify where your data sources are going and where that data is potentially being referenced. Take this mini-DAG for an example: ![A miniature DAG](https://cdn.sanity.io/images/wl0ndo6t/main/bc7fac1a96e70f12da5e125a08f3bff54adae543-1878x838.png) What can you learn from this DAG? Immediately, you may notice a handful of things: - `stg_users` and `stg_user_groups` models are the parent models for `int_users` - A join is happening between `stg_users` and `stg_user_groups` to form the `int_users` model - `stg_orgs` and `int_users` are the parent models for `dim_users` - `dim_users` is at the end of the DAG and is therefore downstream from a total of four different models Within 10 seconds of looking at this DAG, you can quickly unpack some of the most important elements about a project: dependencies and data lineage. Obviously, this is a simplified version of DAGs you may see in real life, but the practice of identifying relationships and data flows remains very much the same, regardless of the size of the DAG. What happens if `stg_user_groups` just up and disappears one day? How would you know which models are potentially impacted by this change? Look at your DAG and understand model dependencies to mitigate downstream impacts with [testing and observability](https://www.getdbt.com/product/test-and-observe). ### Auditing projects A potentially bold statement, but there is no such thing as a perfect DAG. DAGs are special in-part because they are unique to your business, data, and data models. There’s usually always room for improvement, whether that means making a [CTE](https://www.getdbt.com/blog/guide-to-cte) into its own view or performing a join earlier upstream, and your DAG can be an effective way to diagnose inefficient data models and relationships. You can additionally use your DAG to help identify bottlenecks, long-running data models that severely impact the performance of your data pipeline. Bottlenecks can happen for multiple reasons: - Expensive joins - Extensive filtering or [use of window functions](https://docs.getdbt.com/blog/how-we-shaved-90-minutes-off-model) - Complex logic stored in views - Good old large volumes of data ...to name just a few. Understanding the factors impacting model performance can help you decide on [refactoring approaches](https://learn.getdbt.com/courses/refactoring-sql-for-modularity), [changing model materialization](https://docs.getdbt.com/blog/how-we-shaved-90-minutes-off-model#attempt-2-moving-to-an-incremental-model)s, replacing multiple joins with [surrogate keys](https://www.getdbt.com/blog/guide-to-surrogate-key), or other methods. ![A bad DAG, one that follows non-modular data modeling techniques](https://cdn.sanity.io/images/wl0ndo6t/main/03b33d53cc9ffa35ba36407b9af9ea2a8229fac6-1999x1124.png) ### Modular data modeling best practices See the DAG above? It follows a more traditional approach to data modeling where new data models are often built from raw sources instead of relying on intermediary and reusable data models. This type of project does not scale with team or data growth, making [data modernization](https://www.getdbt.com/product/data-modernization) essential. As a result, analytics engineers tend to aim to have their DAGs not look like this. Instead, there are some key elements that can help you create a more streamlined DAG and [modular data models](https://www.getdbt.com/analytics-engineering/modular-data-modeling-technique/) that [build trust in data and data teams](https://www.getdbt.com/product/build-trust-in-data-and-data-teams): - Leveraging [staging, intermediate, and mart layers](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview) to create layers of distinction between sources and transformed data - Abstracting code that’s used across multiple models to its own model - Joining on surrogate keys versus on multiple values These are only a few examples of some best practices to help you organize your data models, business logic, and DAG. **** ## dbt and DAGs The marketing team at dbt Labs would be upset with us if we told you we think dbt actually stood for “dag build tool,” but one of the key elements of dbt is its ability to generate documentation and infer relationships between models. And one of the hallmark features of [dbt Docs](https://docs.getdbt.com/docs/build/documentation) is the Lineage Graph (DAG) of your dbt project. Whether you’re using dbt or Core, dbt docs and the Lineage Graph are available to all [dbt developers](https://www.getdbt.com/product/develop). The Lineage Graph in dbt Docs can show a model or source’s entire lineage, all within a visual frame, providing comprehensive visibility through [dbt Catalog](https://www.getdbt.com/product/dbt-catalog). Clicking within a model, you can view the Lineage Graph and adjust selectors to only show certain models within the DAG. Analyzing the DAG here is a great way to diagnose potential inefficiencies or lack of modularity in your dbt project, helping [analysts](https://www.getdbt.com/product/analyst) work more effectively. ![The Lineage Graph in dbt Docs](https://cdn.sanity.io/images/wl0ndo6t/main/9a24352e6c6fd1bc8f27c8138a174c6868839173-1999x1066.png) The DAG is also [available in the dbt IDE](https://www.getdbt.com/blog/on-dags-hierarchies-and-ides/), so you and your team can refer to your lineage while you [build your models](https://www.getdbt.com/product/develop). **Leverage exposures** One of the newer features of dbt is [exposures](https://docs.getdbt.com/docs/build/exposures), which allow you to define downstream use of your data models outside of your dbt project _within your dbt project_. What does this mean? This means you can add key dashboards, machine learning or data science pipelines, reverse ETL syncs, or other downstream use cases to your dbt project’s DAG. This level of interconnectivity and transparency can help boost data governance (who has access to and who [owns](https://docs.getdbt.com/reference/resource-configs/meta#designate-a-model-owner) this data) and transparency (what are the data sources and models affecting your key reports). ## Conclusion A Directed acyclic graph (DAG) is a visual representation of your data models and their connection to each other. The key components of a DAG are that nodes (sources/models/exposures) are directionally linked and don’t form acyclic loops. Overall, DAGs are an effective tool for understanding data lineage, dependencies, and areas of improvement in your data models. > _Get started with [dbt today](https://www.getdbt.com/signup/) to start building your own DAG!_ ## Further reading Ready to restructure (or create your first) DAG? Check out some of the resources below to better understand data modularity, data lineage, and how dbt helps bring it all together: - [Data modeling techniques for more modularity](https://www.getdbt.com/analytics-engineering/modular-data-modeling-technique/) - [How we structure our dbt projects](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview) - [How to audit your DAG](https://www.youtube.com/watch?v=5W6VrnHVkCA) - [Refactoring legacy SQL to dbt](https://docs.getdbt.com/guides/refactoring-legacy-sql) --- --- title: "Getting started with CTEs" description: "How Common Table Expressions enhance SQL readability and efficiency in dbt for cleaner, maintainable queries." url: "https://www.getdbt.com/blog/getting-started-with-cte" date: "2024-07-01" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Getting started with CTEs In a formal sense, a Common Table Expression (CTE), is a temporary result set that can be used in a SQL query. You can use CTEs to break up complex queries into simpler blocks of code that can connect and build on each other. In a less formal, more human-sense, you can think of a CTE as a separate, smaller query within the larger query you’re building up. Creating a CTE is essentially like making a temporary view that you can access throughout the rest of the query you are writing. There are two-types of CTEs: recursive and non-recursive. This glossary focuses on non-recursive CTEs. ## Why you should care about CTEs Have you ever read through a query and thought: - “What does this part of the query do?” - “What are all the sources referenced in this query? Why did I reference this dependency?” - “My query is not producing the results I expect and I’m not sure which part of the query is causing that.” These thoughts often arise when we’ve written SQL queries and models that utilize complex business logic, references and joins multiple upstream dependencies, and are not outputting expected results. In a nutshell, these thoughts can occur often when you’re trying to write data models! How can you make these complexities in your code more digestible and usable? CTEs to the rescue! ## CTE Syntax: How it works To use CTEs, you begin by defining your first CTE using the `WITH` statement followed by a `SELECT` statement. Let’s break down this example involving a `rename_columns` CTE below: ```sql with rename_columns as ( select id as customer_id, lower(first_name) as customer_first_name, lower(last_name) as customer_last_initial from {{ ref('raw_customers') }} ) select * from rename_columns ``` In this query above, you first create a CTE called `rename_columns` where you conduct a simple `SELECT` statement that renames and lower cases some columns from a `raw_customers` table/model. The final `select * from rename_columns` selects all results from the `rename_columns` CTE. While you shouldn't always think of CTEs as having classical arguments like SQL functions, you’ve got to call the necessary inputs for CTEs something. - CTE_EXPRESSION_NAME: This is the name of the CTE you can reference in other CTEs or SELECT statements. In our example, `rename_columns` is the CTE_EXPRESSION_NAME. **If you are using multiple CTEs in one query, it’s important to note that each CTE_EXPRESSION_NAME must be unique.** - CTE_QUERY: This is the `SELECT` statement whose result set is produced by the CTE. In our example above, the `select … from {{ ref('raw_customers') }}` is the CTE_QUERY. The CTE_QUERY is framed by parenthesis. ## When to use CTEs The primary motivation to implement CTEs in your code is to simplify the complexity of your queries and increase your code’s readability. There are other great benefits to using CTEs in your queries which we’ll outline below. ### Simplification When people talk about how CTEs can simplify your queries, they specifically mean how CTEs can help simplify the structure, readability, and debugging process of your code. #### Establish Structure In leveraging CTEs, you can break complex code into smaller segments, ultimately helping provide structure to your code. At dbt Labs, we often like to use the [import, logical, and final structure](https://docs.getdbt.com/guides/refactoring-legacy-sql?step=5#implement-cte-groupings) for CTEs which creates a predictable and organized structure to your dbt models. #### Easily identify dependencies When you import all of your dependencies as CTEs in the beginning of your query/model, you can automatically see which models, tables, or views your model relies on. #### Clearly label code blocks Utilizing the CTE_EXPRESSION_NAME, you can title what your CTE is accomplishing. This provides greater insight into what each block of code is performing and can help contextualize why that code is needed. This is incredibly helpful for both the developer who writes the query and the future developer who may inherit it. #### Test and debug more easily When queries are long, involve multiple joins, and/or complex business logic, it can be hard to understand why your query is not outputting the result you expect. By breaking your query into CTEs, you can separately test that each CTE is working properly. Using the process of elimination of your CTEs, you can more easily identify the root cause. ### Substitution for a view Oftentimes you want to reference data in a query that could, or may have existed at one point, as a view. Instead of worrying about the view actually existing, you can leverage CTEs to create the temporary result you would want from the view. ### Support reusability Using CTEs, you can reference the same resulting set multiple times in one query without having to duplicate your work by referencing the CTE_EXPRESSION_NAME in your from statement. ## CTE example Time to dive into an example using CTEs! For this example, you'll be using the data from our [jaffle_shop demo dbt](https://github.com/dbt-labs/jaffle_shop) project. In the `jaffle_shop`, you have three tables: one for customers, orders, and payments. In this query, you're creating three CTEs to ultimately allow you to segment buyers by how many times they’ve purchased. ```sql with import_orders as ( select * from {{ ref('orders') }} ), aggregate_orders as ( select customer_id, count(order_id) as count_orders from import_orders where status not in ('returned', 'return pending') group by 1 ), segment_users as ( select *, case when count_orders >= 3 then 'super_buyer' when count_orders <3 and count_orders >= 2 then 'regular_buyer' else 'single_buyer' end as buyer_type from aggregate_orders ) select * from segment_users ``` Let’s break this query down a bit: 1. In the first `import_orders` CTE, you are simply importing the `orders` table which holds the data I’m interested in creating the customer segment off of. Note that this first CTE starts with a `WITH` statement and no following CTEs begin with a `WITH` statement. 2. The second `aggregate_orders` CTE utilizes the `import_orders` CTE to get a count of orders per user with a filter applied. 3. The last `segment_users` CTE builds off of the `aggregate_orders` by selecting the `customer_id`, `count_orders`, and creating your `buyer_type` segment. Note that the final `segment_users` CTE does not have a comma after its closing parenthesis. 4. The final `select * from segment_users` statement simply selects all results from the `segment_users` CTE. Your results from running this query look a little like this: **** ## CTE vs Subquery A [subquery](https://www.getdbt.com/blog/guide-to-subquery) is a nested query that can oftentimes be used in place of a CTE. Subqueries have different syntax than CTEs, but often have similar use cases. This content won’t go too deep into subqueries here, but it'll highlight some of the main differences between CTEs and subqueries below. ## Data warehouse support for CTEs CTEs are likely to be supported across most, if not all, [modern data warehouses](https://www.getdbt.com/blog/future-of-the-modern-data-stack). Please use this table to see more information about using CTEs in your specific data warehouse. ## Conclusion CTEs are essentially temporary views that can be used throughout a query. They are a great way to give your SQL more structure and readability, and offer simplified ways to debug your code. You can leverage appropriately named CTEs to easily identify upstream dependencies and code functionality. CTEs also support recursiveness and reusability in the same query. Overall, CTEs can be an effective way to level-up your SQL to be more organized and understandable. ## Further Reading If you’re interested in reading more about CTE best practices, check out some of our favorite content around model refactoring and style: - [Refactoring Legacy SQL to dbt](https://docs.getdbt.com/guides/refactoring-legacy-sql?step=5#implement-cte-groupings) - [dbt Labs Style Guide](https://docs.getdbt.com/best-practices/how-we-style/0-how-we-style-our-dbt-projects) - [Modular Data Modeling Technique](https://www.getdbt.com/analytics-engineering/modular-data-modeling-technique/) Want to know why dbt Labs loves CTEs? Check out the following pieces: - [Why we use so many CTEs](https://discourse.getdbt.com/t/why-the-fishtown-sql-style-guide-uses-so-many-ctes/1091) - [CTEs are Passthroughs](https://discourse.getdbt.com/t/ctes-are-passthroughs-some-research/155) --- --- title: "Data integration vs. data transformation: What’s the difference?" description: "Learn about the relationship between data integration and data transformation and how to manage them across your enterprise." url: "https://www.getdbt.com/blog/data-integration-vs-data-transformation" date: "2024-06-30" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Data integration vs. data transformation: What’s the difference? Businesses are becoming increasingly data-driven. That makes producing high-quality data a make-or-break proposition, no matter your field. Unfortunately, high-quality data doesn’t come for free. Raw data has to be collected from numerous sources and turned into a format suitable for driving day-to-day decision-making. Data integration and data transformation are two interrelated but different concepts that are integral to turning raw data into business insights. In this article, we’ll discuss what they are, how they’re related, and how to manage them at scale across your enterprise. ## What is data transformation? Data transformation converts raw data from its original format into one or more readily usable by business decision-makers. This includes normalizing, cleaning, and validating the data to ensure it's ready for analysis. Data transformation happens as part of a data pipeline, an automated process that syncs new data from its source, transforms it, and stores it in a new destination. Data pipelines are often run on a set schedule, importing new data from their sources on a regular basis. Analysts and decision-makers can then use languages such as SQL and Python to query the data and generate reports. Typically, data transformation occurs inside of a process using one of two methodologies: - **ETL** (Extract, Transform, Load): Data is taken from its source(s), changed into its new format, and then loaded into a new destination. - **ELT** (Extract, Load, Transform): A similar process except that data is transformed after it’s been loaded into its destination. Over the years, [ETL has been replaced with ELT](https://www.getdbt.com/blog/etl-vs-elt), as cloud computing has made it easier and more cost-efficient to load data prior to transformation. Data is loaded (typically into a data warehouse) and then transformed and stored in a separate location. This approach makes the raw data available immediately to everyone with data warehouse access. It also allows teams with different needs to transform the raw data however they see fit. ### Benefits of data transformation No business that relies on data can afford to lose out on the benefits of data transformation, a process that: **Increases data quality**. Raw data can almost never be used as is. It’s full of malformatted or missing values, redundancies, inconsistencies, and sometimes outright incorrect information. This poor-quality data can cost an organization [up to 30% of its yearly revenue](https://pragmaticworks.com/blog/the-cost-of-bad-data-infographic). Besides poor decision-making based on incorrect values, unexpected values in new data (e.g., a different date format, a malformed customer ID number) can cause data pipelines to break, making reports and data-driven applications unavailable until they’re fixed. Bad data can also result in filing incorrect regulatory reports, which could result in fines. **Produces organized, easy-to-use data**. Without data transformation, analysts and decision-makers would have to reinvent the wheel whenever they wanted to create a new report, fixing data issues anew each time. Data transformation provides a clean data set in an accessible location, making it easier to generate new reports and data-driven apps. **Paves the way for machine learning and AI workloads**. Analytics is a deterministic approach to data. The volume of data doesn’t matter as much as its accuracy and performance for querying. By contrast, machine learning and AI take a probabilistic approach to data, using neural networks and statistical inference to generate new outputs. This requires a large volume of high-quality data - whether you’re training your own Large [Language Models (LLMs)](https://aws.amazon.com/what-is/large-language-model/) or using data for context with techniques such as [retrieval-augmented generation (RAG)](https://blogs.nvidia.com/blog/what-is-retrieval-augmented-generation/). ## What is data integration? Data integration is a _type_ of data transformation. As part of the data transformation process, a data pipeline may bring data in from multiple sources. It then combines this data to provide a single, unified view of a data set across the enterprise. For example, data about customers in a retail business may be split across multiple systems - e-commerce sales, marketing emails, website analytics, advertising campaigns, web search, etc. Bringing this data together and presenting it as a unified customer record can help businesses answer questions such as where they acquire their highest-spend customers, which types of users are most likely to refer their friends, and more. ### Benefits of data integration **Promotes better decision-making**. Having a 360-degree view of the business makes it easier for decision-makers to connect the dots and see trends that otherwise might have been obscured by keeping data in separate systems. **Can improve performance**. Performing joins across data residing in multiple external systems can result in slow query performance. Unifying the data into a single location eliminates the need for cross-system joins, meaning ad-hoc queries return results orders of magnitude faster. **Increases data discoverability**. [According to a 2022 Matillion and IDG survey](https://www.matillion.com/blog/matillion-and-idg-survey-data-growth-is-real-and-3-other-key-findings), most large companies draw data from over 400 sources. 20 percent are drawing data from over 1,000 sources. It can be hard for employees to find precisely what they need in these vast data oceans. Data integration brings critical data together into a single location, which helps eliminate [data silos](https://www.techtarget.com/searchdatamanagement/definition/data-silo)—islands of data that don’t produce business value because data consumers can’t find them. **Encourages data democratization**. All teams need data. However, not every team can afford to hire a small crew of data engineers or [analytics engineers](https://www.getdbt.com/blog/what-is-analytics-engineering) to wrangle it from across the enterprise into a usable format. Consolidating such data into a single location makes it easier for data consumers who are conversant in SQL to find and make use of data. ### Types of data integration There are several ways to perform data integration: **Data warehousing**. The traditional approach in which engineers physically move data into a single location and a minimal number of tables. Data warehousing uses [different data modeling techniques](https://www.getdbt.com/blog/data-modeling-techniques) than relational database systems that enable faster querying for BI use cases. **Virtualized integration**. Virtualized integration provides access to numerous data sources from within a single location. This approach misses out on some of the performance benefits provided by data warehousing (though it may utilize techniques such as caching to speed up subsequent access). The upside is that business users can access data where it currently lives ‌without waiting for engineers to import it into the warehouse. **Data mesh**. Bringing data into a data warehouse typically depends on relying on a centralized data engineering team to create new data pipelines. This can create a bottleneck as the team’s queue fills up with new requests. [A data mesh architecture](https://www.getdbt.com/blog/what-are-the-four-principles-of-data-mesh) solves this by modeling data as a set of interconnected domains. A mesh architecture provides a self-service data platform that makes it easier for teams to create their data pipelines and for business users to find and use the data sets they produce. ## The challenges with data transformation and data integration Data integration, in short, is one data transformation technique you'll use among many others in a data pipeline to ensure that data is reliable, accurate, and performant. Other [data transformation techniques](https://www.getdbt.com/blog/implementing-common-data-transformation-techniques-dbt) include cleaning, aggregating, generalization, validation, normalization, and enrichment. The nature of enterprise data makes managing both data transformations and data integration at scale a challenge. Your company, like many others, is probably managing data across hundreds of sources, including data warehouses, analytics tools, marketing platforms, relational databases, NoSQL data stores, data lakes and lakehouses, etc. Lacking a single approach to manage transformation and integration across such disparate systems, engineering teams often end up implementing data pipelines in an ad hoc manner, using technologies and languages that lock them into vendor-specific solutions. This creates numerous inefficiencies: - Engineers reuse little code, instead solving the same problem redundantly over and over in different languages. - As an organization, you have little visibility into what data assets you have and what data pipelines might currently be running. This makes it impossible to ensure data quality and consistency across data stores. It almost makes it impossible to manage data pipeline spend and control costs. - There’s no consistency in how data transformation and data integration code is tested before it’s put in front of users. This can result in injecting bad data into production data sets. - There’s also a lack of consistency and how changes are rolled out to production. While software engineering has focused on techniques such as [Continuous Integration and Continuous Deployment (CI/CD)](https://www.redhat.com/en/topics/devops/what-is-ci-cd) to verify changes before release, most analytics code is still deployed in an ad hoc, one-off fashion. ## dbt Cloud: A data control plane for your data pipelines Managing data transformation and data integrations across your enterprise requires a [data control ](https://www.getdbt.com/blog/data-control-plane-why)plane—a single toolset that enables you to manage all data transformations across the enterprise. dbt Cloud is a data control plane for managing your data pipelines that enables everyone in your enterprise to work with data no matter where it lives. Using [dbt models](https://docs.getdbt.com/docs/build/models), data engineers, analytics engineers, and even business users can model data transformations uniformly using SQL or Python code, avoiding vendor lock-in. This enables data engineers to monitor, manage, and fine-tune data transformation workloads across the enterprise. dbt Cloud provides out-of-the-box features that support shipping and using high-quality data sets, including: - [Version control](https://docs.getdbt.com/docs/collaborate/git/version-control-basics) and peer reviews - Support for creating [DRY code](https://www.getdbt.com/blog/guide-to-dry) that can be reused across projects - [Testing](https://docs.getdbt.com/docs/build/data-tests) and [documentation](https://docs.getdbt.com/docs/build/documentation), including automatically generated [data lineage](https://www.getdbt.com/blog/guide-to-data-lineage) - [CI/CD deployment](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud) and automated testing of changes in pre-production environments before release - Data discovery and management via [dbt Explorer](https://www.getdbt.com/product/dbt-explorer), which enables data consumers to find and learn about data sets while giving data producers insight into usage and performance - AI-assisted support for generating new data models, tests, and docs using [dbt Copilot](https://www.getdbt.com/blog/dbt-copilot-is-ga), which embeds context-aware AI into every phase of your analytics workflow Learn more about how dbt Cloud can manage data transformations at scale across your enterprise—ask us for a demo today. --- --- title: "Data quality best practices: Six essential principles for analytics and AI" description: "Use these six essential principles to produce quality data for both analytics and AI." url: "https://www.getdbt.com/blog/data-quality-best-practices" date: "2024-06-30" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Data quality best practices: Six essential principles for analytics and AI Poor-quality data can torpedo your analytics and AI initiatives. [One survey by FiveTran](https://www.fivetran.com/blog/new-ai-survey-poor-data-quality-leads-to-406-million-in-losses) found that bad data can cost companies an average of USD $406M yearly. Basic data hygiene techniques, such as testing, play a huge role in improving data quality. However, writing a few tests isn't enough. You need a comprehensive approach that implements checks and balances and multiple points in the process. In this article, we'll look at six principles you can use to enhance data quality at every step of the data life cycle. ## Data quality best practices In working with hundreds of clients in the data space, we’ve distilled a few common best practices any company can benefit from. These include: - Build a quality-oriented analytics workflow - Work with data through a single data control plane - Shift data quality testing left - Monitor and report continuously in production - Control access to data - Don’t over-index on data quality Let’s take a look at each one in turn, what it means, and how you can implement it. ### Build a quality-oriented analytics workflow Historically, a lot of work in data analytics has been ad hoc. Data engineers write data transformation scripts in a variety of languages and often for their own personal use (i.e., not checked into source control). ‌Changes may get run against production systems with minimal verification or testing. Over time, this leads to poor data quality and data corruption. The first step towards improving data quality is creating a mature analytics workflow. This is a workflow that: - Represents all analytics changes as code and puts all assets under version control - Enables collaboration at scale - Is accessible to the widest variety of data users possible - Enables a high deployment velocity - Supports security, auditing, and governance We've laid out our approach to a mature analytic workflow, which we call the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). ‌The ADLC is a vendor agnostic approach that builds quality into every part of an analytics workflow: - At the **plan** stage, a unified team - including data engineers, analysts, and business users - define data quality explicitly, as well as determine who should have access to what data - During **development**, engineers track all of their changes using version control. They save their code to non-production branches of a source repo and can only merge their changes to the production branch by creating a [Pull Request (PR)](https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/proposing-changes-to-your-work-with-pull-requests/about-pull-requests) and undergoing a peer review - **Test** requires all engineers to validate their changes to ensure their code does what they intend it to do under a variety of circumstances, and fails gracefully when it encounters an unexpected situation - **Deploy** uses [automated Continuous Integration/Continuous Development (CI/CD) pipelines](https://docs.getdbt.com/docs/deploy/continuous-integration) to run tests in pre-production environments before exposing changes to users, ensuring changes can be deployed both quickly _and_ safely - **Observe and analyze** provides always-on access to analytics and continuously tests in production to find errors before users do - **Discover and analyze** makes all data models and their accompanying documentation easily discoverable and available to data stakeholders, enabling them to validate both the correctness and timeliness of the data they’re using A mature analytics process does more than test code. It provides transparency and visibility to all data stakeholders throughout the entire development process, so that everyone understands and trusts the data they’re using. ### Work with data through a single data control plane In the past, data pipelines were complicated. They are written in a plethora of languages, were often buggy and temperamental, and required detailed knowledge of arcane processes to run. ‌This meant there was no single, common approach to implementing and verifying data quality. The explosion of data and growth of data-driven applications — including analytics and, now AI—means companies need easy-to-use tools for transforming data and verifying data quality that they can use across the business. ‌Ideally, they need a single place that data stakeholders can use to manage all data transformations and verify the accuracy, timeliness, and origins of data. We call this the **data control plane**—an open, cross-platform approach that provides a single point of access for transforming, storing, discovering, using, and managing data. With a data control plane, you should be able to: - Model, transform, and test data - Democratize data so that data engineers, analysts, and even savvy business users can build and deploy analytics code changes with high velocity - Optimize data platform costs across all workloads - Discover data and its accompanying documentation and see a map of your data estate using automatically generated [data lineage](https://www.getdbt.com/blog/guide-to-data-lineage) dbt Cloud functions as a data control plane by providing a common framework for managing data. Engineers and analysts can create models and tests using SQL or Python with the same toolset and process each time, reusing models and tests across projects. Meanwhile, data stakeholders can view data, documentation, and important data quality information— such as lineage and data freshness—via a single portal. ### Shift data quality testing left [Shift-left testing](https://www.bmc.com/blogs/what-is-shift-left-shift-left-testing-explained/) has been an important concept for years in the software development world. It refers to moving testing earlier in the development process rather than treating testing as a separate phase that happens at the end. Shifting testing left recognizes two realities. First, bugs detected earlier in the development process are easier and cheaper to fix than bugs discovered in production. Second, doing testing “at the end” usually means testing never happens at all. Starting testing earlier —for example, during the development phase—ensures that testing is an integral part of the process rather than an afterthought. With dbt Cloud, developers can create tests alongside their code. Teams new to shift-left testing can start by implementing [the essential data quality checks](https://www.getdbt.com/blog/data-quality-checks). From there, they can implement more advanced testing, including domain-specific custom tests. One aspect of shift-left testing that's remained challenging for data has been testing locally during development. Developers have either had to maintain their own data warehouse instances or run their changes through a full deployment pipeline iteration to test them fully. To accelerate this, [dbt is integrating SDF](https://www.getdbt.com/blog/advancing-the-data-control-plane-vision-with-sdf-and-dbt), which emulates data stores like Snowflake locally to enable full end-to-end testing before developers even check in their code. ‌This will enable accelerated testing and reduce time spent in the dev cycle fixing broken check-ins. The result is higher data quality delivered in less time. ### Monitor and report continuously in production CI/CD pipelines contribute to data quality by running your data tests against mock data in pre-production environments. This identifies the majority of issues in your code before you make your changes live for stakeholders. However, you can never fully predict how your changes will act in production against real world data. Especially in long-running data systems, it's almost impossible to foresee every exception and edge case you may encounter. You can account for this by continuing to test even after deployment. [Testing in production](https://www.getdbt.com/blog/adlc-operate-observe) runs your tests against incoming data changes, raising alerts whenever an issue is discovered. ‌This keeps your data clean and helps find errors before they impact customers. You can further keep your data clean by using [data profiling](https://www.ibm.com/think/topics/data-profiling) to assess and track overall data quality along different [data quality dimensions](https://www.getdbt.com/blog/data-quality-dimensions). Assessing your overall data quality means you can set benchmarks, measure progress across teams, and mentor individual teams on how to improve their overall data quality. ### Control access to data An often overlooked aspect of data quality is preventing unauthorized changes. Ad hoc changes that don't go through your analytics workflow development process can inject errors that might take weeks or even months to discover. Rather than provide direct access to data systems, use your data control plane to ensure all analytics code changes go through the entire ADLC process, which includes peer review and automated testing. This ensures accountability and conformance to your company’s data quality standards. Your data control plane should provide a way to control access to data based on roles versus hard-coding access to individuals. Use role-based access control (RBAC) wherever possible to both grant and revoke access based on a user’s job role. ([dbt Cloud supports RBAC](https://docs.getdbt.com/docs/collaborate/govern/model-access) for controlling access to dbt assets on a per-project basis.) Use Single Sign-On (SSO) so that a user’s permissions are revoked as soon as their corporate ID is decommissioned. This prevents situations where, for example, individuals maintain access to critical systems even after their employment is terminated. ### Don’t over-index on data quality [The 10x9 rule](https://blog.alexewerlof.com/p/10x9) in systems engineering says that, for every [nine of reliability](https://www.splunk.com/en_us/blog/learn/five-nines-availability.html) you add, you increase reliability by 10x— but at 10x of the total cost of your solution. What does this mean for data testing? It means that, at some point, more testing is too much. ‌The exact definition of “too much” will differ between teams and depend on your use cases. Use metrics to measure your data reliability across the various data quality dimensions and establish a KPI for each dimension that's easy to achieve. Once you hit that, focus more on making it easier and faster to ship high-quality analytics code rather than blowing money on chasing perfection. ## Conclusion A serious commitment to data quality requires a cultural shift. It means establishing processes and inculcating habits that may feel new or strange to all parties involved. Done well, however, they enable data stakeholders to create, deploy, and use new data products with higher quality and in less time than ever. While data quality is more than just tools, a good tools framework can simplify and democratize access to data across your entire organization. Using dbt Cloud as your data control plane, you can provide a consistent approach to data quality to all data stakeholders with minimal overhead. To learn more, [ask for a demo today](https://www.getdbt.com/contact). --- --- title: "Implementing common data transformation techniques with dbt" description: "Explore common data transformations—and how dbt brings consistency, testing, and self-serve access to your analytics pipeline." url: "https://www.getdbt.com/blog/implementing-common-data-transformation-techniques-dbt" date: "2024-06-30" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Implementing common data transformation techniques with dbt High-quality data is the foundation for every data-driven decision. But clean, trusted, and analysis-ready data doesn’t happen on its own—it requires deliberate transformation. Data transformation is the process of cleaning, verifying, and reshaping raw data into a format stakeholders can actually use. While writing individual transformations in SQL or Python might be simple, managing them at scale is not. Designing, testing, documenting, and deploying code across your entire data estate requires structure. That’s where a data control plane comes in. In this article, we’ll explore common data transformation techniques—and how [dbt](https://www.getdbt.com/product/what-is-dbt) helps teams implement them in a scalable, governed, and cost-efficient way. ## Data transformation: core concepts Data transformation is the process of cleaning, restructuring, and optimizing raw data so it’s usable for analysis. It takes data from multiple sources and turns it into a consistent format that business users and systems can trust. Think of it like translating between languages. Raw data arrives in different formats, schemas, and structures—each its own “dialect.” Transformation turns that into a shared, standardized language that analytics tools and teams can understand. ### Two common pipeline architectures You can build transformation pipelines in different ways depending on your team’s needs. Here are two typical approaches: #### Approach 1: Sequential processing pipeline 1. Extract raw customer data from a CRM 2. Clean and validate email addresses and phone numbers 3. Aggregate purchase history by customer ID 4. Enrich with demographic data from external sources 5. Load the final dataset into your analytics warehouse #### Approach 2: Parallel processing pipeline 1. Extract raw customer data from a CRM 2. Split into three parallel streams: - A: Clean contact information - B: Aggregate purchase metrics - C: Enrich with external demographics 3. Merge all streams back together 4. Apply final validation checks 5. Load into your analytics warehouse Both methods deliver the same end result. Sequential pipelines are easier to debug and maintain. Parallel pipelines offer faster performance, but require more complex orchestration—something that tools like dbt can help simplify and standardize. ## Types of data transformation Transforming data usually means applying one of a fixed set of operations to change data to a more usable format. Sometimes, this means addressing errors or inconsistencies in your data. ‌It also involves changing data into a format that's more readily usable for your use cases. Data transformation pipelines [save time and money](https://www.getdbt.com/blog/successful-data-transformation) by bringing consistency to data. Without data transformation pipelines, everyone—analysts, data engineers, and business users—would be slicing and dicing data their own way, wasting time and introducing data inconsistencies. Most of the time, you'll be applying one of the following transformations to your data: - **Cleaning**. ‌Removing errors in inconsistencies from your data—missing fields, inaccurate entries, duplicated data, etc. - **Aggregation**. ‌Rolling up critical values for faster access—for example, sales data for a given customer or time period. - **Generalization**. ‌Breaking up a single data unit into a hierarchy, such as an address. - **Discretization**. Transforming continuous data, such as ages, into a set of ranges (e.g., ages 18-29) to make it easier to drive initiatives such as targeted marketing. - **Normalization**. ‌Enforcing standards for the format of certain fields and rationalizing data types and identifiers. Example: converting currency data into a single standard currency, such as USD. - **Validation**. Ensuring that data is in the correct format. One example is verifying that phone numbers have the correct number of digits, that they have a valid country code, etc. - **Enrichment**. Also called attribute construction, enrichment adds additional data to enable enhanced decision-making — e.g., adding weather data to scheduled shipment information to warn customers about potential delays. - [**Integration**](https://www.getdbt.com/blog/data-integration-vs-data-transformation). Bringing in data from multiple sources to create a single, consistent data set that doesn’t require complex joins or high-latency connections across different databases. When done well, transformation turns chaotic inputs into structured, trusted assets—and dbt helps you manage and scale these workflows like code. ## dbt: A control plane for data transformations Data teams can write transformations in many ways—from ad hoc SQL to custom Python scripts. But managing these workflows at scale requires more than just code—it needs structure, visibility, and repeatability. That’s where dbt comes in. As a transformation control plane, dbt helps teams manage the entire lifecycle of analytics code: development, testing, documentation, and deployment. Here’s how: - **Treats analytics like software**. ‌dbt lets you write transformations in SQL or Python, version them in Git, and manage them like code—so changes can be tracked, reviewed, reused, and rolled back with confidence. - **Works across warehouses**. dbt provides a consistent, vendor-agnostic framework that supports all major data platforms. With support for both SQL and Python, contributors from across the team can build and maintain transformations without learning new tools. - **Enables built-in testing**. ‌Testing transformations isn't an afterthought. With dbt, [you can write tests alongside your data transformations](https://docs.getdbt.com/docs/build/data-tests) that are run automatically at various points to ensure the code is correct before it touches production data. - **Generates documentation and lineage**. ‌[dbt builds documentation](https://docs.getdbt.com/docs/build/documentation) as part of your workflow and [visualizes lineage](https://www.getdbt.com/blog/guide-to-data-lineage) across your models. This makes it easier for stakeholders to trust and understand the data—without needing to ask engineers how it works. With dbt, your transformation workflows are structured, tested, documented, and discoverable—by default. ## Avoiding common pitfalls in data transformation with dbt Writing transformation code is just the beginning. Teams also need safe, repeatable ways to deploy changes—and ensure data consumers can easily find and trust what’s been built. dbt helps organizations avoid common transformation challenges like: - Corrupted values from bugs - Unreliable deployment processes - No rollback strategy - Low data discoverability Let’s look at each of these issues in detail and how dbt addresses them. ### Corrupted values from bugs Unchecked bugs can break reports, mislead stakeholders, and create costly cleanup. dbt addresses this with multiple layers of quality control: - All transformations live in [version-controlled code ](https://docs.getdbt.com/docs/collaborate/git/version-control-basics)(SQL or Python) - Developers work in isolated branches, then open pull requests (PRs) - PRs trigger automated tests and peer reviews before merging to production - [dbt supports reusable code modules](https://docs.getdbt.com/docs/build/enhance-your-code), reducing the chance of duplicating flawed logic With [dbt’s acquisition of SDF](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs), developers can now emulate popular data warehouses locally—enabling early, fast feedback before pushing any code to Git. This shifts testing left and cuts down PR churn. ### Unreliable deployment processes Historically, deploying data changes often meant running manual scripts in production—a risky and opaque process. dbt brings the [rigor of DevOps](https://www.getdbt.com/resources/the-analytics-development-lifecycle) to data workflows. You can implement CI/CD pipelines to validate, test, and promote code changes through dev, staging, and prod environments. This approach improves security, increases confidence, and reduces bottlenecks. ### No rollback strategy Even with robust testing, some issues only show up in production. With dbt, every transformation is versioned and tracked—making it simple to revert to a known good state while your team troubleshoots and fixes the issue. ### Lack of data discoverability The best models are useless if no one knows they exist. dbt solves this with built-in documentation and [dbt Catalog](https://www.getdbt.com/product/dbt-catalog). Data stakeholders can search, explore, and adopt trusted models—complete with lineage and descriptions—without needing engineering help. This drives adoption, increases data trust, and supports true self-service analytics. ## Get started with better data transformation today You can write transformation logic in SQL or Python—but managing it at scale requires more than code. A control plane powered by dbt brings consistency, quality, and reusability to your analytics workflows. With dbt, data teams gain a standardized, cost-effective way to build, test, and deploy models—while business users get governed, self-serve access to the data they need. [Start for free and see how dbt can power your modern data transformation workflows.](https://www.getdbt.com/signup) ## FAQs about data transformation **What are the benefits of using dbt for data transformations over ad hoc approaches?** dbt brings consistency, quality, and reusability to your analytics code. It turns all data transformations into code that can be tracked, reviewed, and rolled back when needed. The platform provides built-in testing capabilities to ensure code correctness before touching production data. Additionally, dbt generates automatic documentation and lineage maps, making datasets easier to discover and increasing stakeholder confidence. **How does dbt help prevent data transformation errors?** dbt implements multiple safeguards to prevent errors from reaching production. Changes are managed through Git-based version control, keeping work-in-progress isolated from tested production code. Pull requests trigger code reviews and automated testing against pre-production data. With SDF integration, developers can also test transformations locally before committing code, significantly reducing defect rates in the deployment pipeline. **What common data transformation techniques can be implemented with dbt?** With dbt, you can implement cleaning operations to remove inconsistencies and errors from datasets. Aggregation and normalization techniques help standardize data formats and roll up critical values. You can also perform validation to ensure data correctness, enrichment to add valuable context, and integration to combine multiple data sources. These transformations are implemented as SQL or Python code within dbt's consistent framework. **How does dbt support DataOps practices?** dbt enables true DataOps by bringing software engineering rigor to analytics workflows. It supports Continuous Integration/Continuous Deployment pipelines that automatically test changes before production deployment. The platform reduces manual steps through automation, allowing more frequent and higher-quality deployments. When issues occur, version control makes it easy to roll back to previous states while engineers address root causes. **How does dbt improve data discoverability for stakeholders?** dbt Catalog allows stakeholders to find and use data models without engineering assistance. Users can access comprehensive documentation that's automatically generated with each release. Data lineage maps show how data flows from source to destination across the entire data estate. This self-service capability ensures valuable datasets don't go unused and increases trust in data quality. --- --- title: "Data quality dimensions: What they are and how to incorporate them" description: "Learn how to use data quality dimensions to create a general profile of your data’s quality by understanding how the data is used." url: "https://www.getdbt.com/blog/data-quality-dimensions" date: "2024-06-30" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Data quality dimensions: What they are and how to incorporate them Poor-quality data can kill a data-driven project before it even starts. ‌Unfortunately, data quality remains a struggle across all industries. Eight years ago, the situation was dire. A Harvard Business Review study conducted in 2017 found that [only 3% of corporate data met basic quality standards](https://hbr.org/2017/09/only-3-of-companies-data-meets-basic-quality-standards). While things have improved, [Gartner still estimates that companies lose USD $12 million annually](https://www.gartner.com/smarterwithgartner/how-to-improve-your-data-quality) to poor-quality data. Improving data quality requires a proactive approach. ‌This, in turn, requires knowing how to measure data quality and what the trade-offs are. We'll look at data quality dimensions, why they're vital to improving data quality, and discuss the key quality dimensions along with some examples. ## What are data quality dimensions? A dimension is some aspect of your data that has meaning or value to your business. Breaking down data quality into dimensions enables you to determine which aspects are most important to your teams and the company as a whole. This enables you to do several things: - Create a comprehensive framework for data quality - Measure and establish a baseline for quality across teams so you can focus on the areas of greatest impact - Compare data quality across teams using common terminology and definitions Data quality dimensions comprise one part of an overall [data quality framework](https://www.getdbt.com/blog/data-quality-framework-choosing) that should also include data pipelines, data cleansing rules, data governance rules, data quality monitoring, and data quality tools. ## What are the key data quality dimensions? While there are many different ways to categorize data quality dimensions, one useful taxonomy is **usefulness**, **accuracy**, **completeness**, **consistency**, **uniqueness**, **validity**, and** freshness**. Most taxonomies leave off usefulness. ‌We view defining data's use and business value early on as a key part of the [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/blog/adlc-plan). ‌Other breakdowns may look at data quality dimensions from slightly different angles—there's no single way to think about this. What's important is to have a consistent and useful taxonomy. Let's take a look at each category of data quality dimensions, along with some examples. ### Usefulness We feel people often leave usefulness out because it's the hardest category to define. You can define usefulness by answering the question: **is the data generating value for the business?** One way to measure this is through the overall presence of dark data in your company. Dark data is all the data in your company that's lying dormant, unused. Sadly, for most companies, this is the majority of their assets. [One survey by Splunk](https://www.splunk.com/en_us/form/the-state-of-dark-data.html) estimates as much as 55% of any company's data might be dark. Dark data generates no revenue or business value. Worse, it **loses** money, as you have to pay for the compute and storage you use to transform and preserve the data. You can identify existing dark data by tracking data usage statistics—for example, by using a data catalog. ‌You can then take two approaches with this data: - Address any issues and data quality, discoverability, or documentation that led it to become dark in the first place; or - Obsolete the data, shutting down its transformation pipelines and storing existing data in a cheaper form of cold storage (e.g., [Amazon S3 Glacier storage classes](https://aws.amazon.com/s3/storage-classes/glacier/)). ### Accuracy Accuracy defines how closely a data result matches existing reality. As data practitioners, it’s your job to ensure that the data you expose to end users is accurate, as in it contains the values that reflect reality. If your business sold 198 new subscriptions today, 198 should be represented in your raw data. Accuracy issues can arise due to conflicting upstream data sources, out-of-date data, buggy analytics code, or other technical issues. You can identify inaccurate data currently in your system using techniques such as statistical analysis, data profiling, consistency checks, spot checking, and sampling. The best approach, of course, is to disallow any accuracies in the first place. You can get closer to this ideal state by using a data transformation framework like dbt to implement [data tests](https://docs.getdbt.com/docs/build/data-tests) that verify all values for a sample data set are correct after transformation. ### Completeness Completeness means you have all the required records—and fields within those records—needed to answer a given set of business questions. In other words, completeness is another dimension that should be defined during the planning stage of the [Analytics Development Lifecycle](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle). As with accuracy, you can measure completeness using techniques such as data profiling, sampling, and attribute profiling. You can also use tests to check for completeness in data transformation — e.g., checking that key fields don’t have null or empty values. If you implement tests in dbt, you can speed up this process by implementing generic data tests that data engineers can reuse across projects. The following test, for example, verifies that a given value is not null. (This is just an example; dbt supports four generic tests out of the box, including not_null, unique, accepted_values, and relationships.) {% test not_null(model, column_name) %} select * from {{ model }} where {{ column_name }} is null {% endtest %} ### Consistency Consistency measures whether data is self-consistent and whether it remains consistent across upstream and downstream data sources as it passes through the data lifecycle. Data consistency may also be called data integrity. Consistency can have severe negative consequences depending on the data. Consider, for example, patient medical systems. An inconsistency across systems in something like a patient ID can result in overlapping or missing records, such as a patient’s list of allergies or currently prescribed medications. A decision made with such data could put someone’s life at risk. You can use data transformation tools like dbt to track and improve consistency. dbt is built upon the concept of a [DAG](https://docs.getdbt.com/terms/dag), or a directed acyclic graph. When someone makes changes to upstream models—say a condition on a case when a statement changed, so there’s a new value for a user_bucket column—you can rebuild your table with the new column value and impact all downstream cases of user_bucket thanks to [dbt’s reference capabilities](https://docs.getdbt.com/reference/node-selection/graph-operators). dbt also supports the use of [metrics via its Semantic Layer](https://www.getdbt.com/product/semantic-layer/), allowing you to define a core metric once (in a version-controlled, code-based environment). You can then expose that same core metric across your database and BI tools. At the end of the month, the CFO and COO shouldn’t have two different values for ARR. With a semantic layer approach to metrics, you can govern and maintain key KPIs in a way that creates unparalleled consistency. ### Uniqueness Uniqueness ensures that data is non-duplicative. It prevents the havoc that can be caused by having multiple records, each with slightly different information. You can test for uniqueness primarily by defining good primary keys and enforcing unique values. However, the challenge is _maintaining_ uniqueness across different data systems. A data transformation system like dbt can help with this by encouraging the reuse of critical data models across projects. ### Validity Validity ensures that a data value is correct for its column type. For some values, this can mean the value is in an accepted range—e.g., an integer representing a month should only have values between 1 and 12. For string values, it might require checking the format of structured data—such as verifying a field contains a properly-formatted ZIP code, well-formatted JSON, or a valid GUID. Validity issues can also cross data fields. For example, a person marked as living in a row probably shouldn’t have a birth date more than 120 years in the past. You can use [accepted value tests](https://docs.getdbt.com/reference/resource-properties/tests#accepted_values) to ensure invalid values don’t enter your data stores. For example, the following dbt test checks that a refund dollar amount isn’t less than 0: -- Refunds have a negative amount, so the total amount should always be >= 0. -- Therefore return records where total_amount < 0 to make the test fail. select order_id, sum(amount) as total_amount from {{ ref('fct_payments') }} group by 1 having total_amount < 0 To detect validity issues in existing data, you can run data audits that apply these same checks to historical data and calculate a validity score. As with all data issues, engineers can then use data lineage to find and fix these issues at their upstream source. ### Freshness Also known as timeliness, freshness measures data that has been updated within a target timeframe. The “acceptable” Service Level Agreement (SLA) for data KPIs will differ depending on the use case. For example, you may only need new sales data sent to a data warehouse table to fuel a sales report every week, meaning a 1 day SLA will suffice. By contrast, an open product order should update whenever the order status changes; since this can happen several times daily, you’ll likely need a 1 hour SLA. You can measure freshness via data freshness (how recently data was updated) and data latency (how long it takes for an upstream change to propagate downstream). You can [implement data freshness reports easily in dbt Cloud](https://docs.getdbt.com/docs/deploy/source-freshness) by checking a checkbox on your data transformation jobs. ## Creating high-quality data across your enterprise Defining data quality dimensions and associated metrics can give you a sense of where your data quality weak points are. With these metrics in hand, you can identify the improvements that will have the most immediate impact on your overall data quality. The challenge is implementing this new approach to data quality _consistently_. In most companies, every team takes its own ad hoc approach to data transformation. That results in inconsistencies in quality from one project to the next. [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) is your data control plane for data. It provides a single, uniform, and vendor-agnostic approach to data transformation you can use to model, test, and verify data across your data ecosystem, no matter where it lives. Using dbt Cloud, you can give teams an easy-to-use toolset for guaranteeing high-quality data without locking yourself into a specific vendor or data architecture. Learn more about how dbt Cloud can improve data quality by [contacting us for a demo today](https://www.getdbt.com/contact). --- --- title: "Data transformation: Six critical best practices" description: "Learn six data transformation best practices that act as high-octane fuel powering your entire business engine." url: "https://www.getdbt.com/blog/data-transformation-best-practices" date: "2024-06-29" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Data transformation: Six critical best practices Today, every business aspires to be data-driven. However, the data that drives business decisions and insights comes from different sources and arrives in many different formats. Identifying meaningful patterns within a typical organization’s jumbled masses of raw data is a herculean task. Data transformation is how we make sense of this data chaos. Let’s explore the fundamental concepts of data transformation, and learn best practices for effectively transforming unprocessed data into actionable intelligence. ## What is data transformation? [Data transformation ](https://www.getdbt.com/analytics-engineering/transformation)is the process of converting raw data from its original format or structure into a standardized, ready to use format—cleaning, normalizing, validating, and enriching data into a state where it is consistent and ready for analysis. This “cleansed” data can be used to derive meaningful insights from any and all of the various different data asset types within an organization. Functionally, data transformation means taking raw source data and using SQL or Python to clean, join, aggregate, and implement business logic on the data to create relevant and usable datasets. These end datasets are then typically fed into a business intelligence (BI) tool, where they form the backbone of an organization’s data-driven business decisions. Data transformation is the “T” in the[ ETL/ELT process](https://www.getdbt.com/blog/etl-vs-elt), where the “T” represents the data transformation stage. The key difference between ETL (Extract, Transform, Load) and ELT (Extract, Load, Transfer) is when and where the data transformation occurs. With ETL, data is transformed before loading. In ELT, data is transformed after loading into the data warehouse. Regardless of when data transformation happens, though, transformation is typically handled by an analytics engineer, data analyst, or data engineer. Other roles may also contribute, depending on the usability and maturity of a company’s data pipeline tools. ## Why is data transformation important? Without data transformation: - Analysts would be writing ad hoc custom queries against raw data sources - Data engineers would get bogged down in maintaining deeply technical pipelines - Business users wouldn't be able to make responsive data-informed decisions in an accessible and scalable way - Everyone would be slicing and dicing data their own way, wasting time and introducing data inconsistencies This is why data transformation is so important to successful organizations: Good transformation creates high-quality, useful datasets that your end users can trust. It’s how raw data gets turned into actionable insights that can propel businesses forward. ## Business benefits of data transformation Data transformation isn’t boring but necessary technical mechanics—it’s the critical bridge that connects raw information and strategic intelligence. The process of data transformation is comparable to refining crude oil into high-octane fuel that powers your entire business engine. The real-world benefits include: ### Strategic value Raw data is essentially unprocessed potential. Data transformation is far more than simply organizing information: By standardizing, cleaning, and enriching data, you can create a strategic asset that can drive competitive advantage. Imagine being able to see precise customer behavior patterns, predict market trends, or identify operational inefficiencies with crystal clarity. ### Operational efficiency Low-quality data [costs organizations an estimated 20-30% of their revenue](https://pragmaticworks.com/blog/the-cost-of-bad-data-infographic). By transforming data, we're not just improving analysis—we're directly impacting the bottom line. Clean, normalized data means faster decision-making, reduced redundancy, and more efficient cross-departmental collaboration. ### Advanced analytics and AI Machine learning and AI are only as good as the data they're trained on. Data transformation prepares your data for advanced analytics, predictive modeling, and AI-driven insights. It's the foundation that allows sophisticated algorithms to generate meaningful, reliable predictions. ### Deeper customer understanding By integrating and cleaning data from multiple sources—for example, sales, support, and marketing—data transformation makes it possible to create comprehensive user/customer profiles that inform laser-targeted strategies and enable hyper-personalized user experiences. ## Data transformation fundamentals Data transformation is the foundation of data analytics. But what are the foundations of data transformation itself? ### Understanding your data Before you dive straight into data transformation, you need to understand two things: the data you are working with, and the needs of the end users who will ultimately consume this data for business purposes.**Data attributes: **The first step is to catalog your organization’s entire data estate. Assess the existing data structures, identify their key attributes, and determine the quality of each data asset. **User attributes: **Next, your data engineers should interview the different data stakeholders within your organization. Develop an understanding of their specific requirements and assess how to align your org’s data assets with business needs and opportunities. ### Core concepts of data transformation Once you have a clear concept of both your current data holdings and what you need to do with that data, it’s time to commence the data transformation process. There are five core data transformation concepts to understand: data cleaning, transformation, validation, enrichment, and integration. **1. Data cleaning: **This is the process of finding and fixing errors and inconsistencies in your organization’s data, such as finding and filling in missing values, correcting inaccuracies, deleting duplicated data, etc. Data cleaning ensures that your data is accurate and reliable from the start, because “dirty” data can lead to poor or just plain wrong analysis results and conclusions. **2. Normalization:** Data normalization is the process of transforming data into a standard range or format to ensure consistency and comparability (and without introducing distortion). ‌Normalization helps adjust applicable data attributes to a common scale, making it easier to compare data and derive insights. - **Example: **A global retail company may normalize transaction data by converting all currency values to USD, enabling accurate financial reporting and analysis across regions. - **Overlap:** Since the normalization process helps ensure consistency in data formats (particularly when dealing with inconsistent units or divergent data sources), data normalization often occurs in tandem with data cleaning. **3. Validation:** Validation verifies that your data adheres to your specified criteria, rules, or standards before it’s eligible for analytics use. This is crucial for maintaining data integrity and quality. - **Example: **Before launching a loyalty program, this global retail company could validate customer data to ensure accurate contact details for efficient communication and engagement.Common validation checks include data format validation (e.g., phone numbers should always be in the format (xxx) xxx-xxxx); validating unique constraints (e.g., ensuring unique identifiers, like customer IDs aren't duplicated), completeness, to ensure that no critical fields are empty or null ; data type verification, and data range validation. - **Overlap:** Similar to normalization, validation can happen concurrently with data cleaning since it ensures that data is in the correct state before any further transformation happens. **4. Enrichment:** During data transformation, doing data enrichment allows you to enhance your internal data with external sources for deeper insights. - **Example:** Our retailer could enrich shipment data with real-time weather information to predict delivery delays and improve customer communication. **5. Integration: **Data integration merges data from different sources into a unified and “apples to apples” data set for comparison and analysis. - **Example: **Our global retailer might want to integrate data from CRM systems, online store accounts, and loyalty program databases to create a 360-degree customer view. This information lets the company offer personalized marketing campaigns, such as recommending products based on purchase history, regardless of the shopping channel, for a seamless shopping experience across online and physical stores. ## Six data transformation best practices So how does a company turn these five core concepts of data transformation into an actual initiative? ‌Implementing the following best practices will help you optimize your transformation processes—and provide the necessary safeguards for reliable and efficient data operations. ### 1. Know your use cases To get the most out of your data transformation initiative, you need a firm view of your business objectives. How will the transformed data be used? The goal is to align the data you have with the real-world outcomes that you want to achieve. - Identify the use cases that will help you reach your business goals, such as improving customer insights, enabling predictive analytics, or ensuring regulatory compliance. ### 2. Use a DataOps approach to data products Applying a [DataOps framework](https://www.getdbt.com/blog/dataops-dbt-guide) ensures data quality and consistency across your entire organization by removing silos between data producers (the creators of data products, like data engineers and data stewards), and data consumers (data end users, like analysts). In [DataOps,](https://www.getdbt.com/blog/dataops-devops-difference) data producers work closely with data consumers in short, rapid deployment cycles to design, develop, deploy, observe, and maintain new data products that fulfill your data consumers’ evolving needs and serve your organization's business goals. - DataOps requires you to define clear standards for data management and transparency, and build accountability for the data processes in your organization. - A well-implemented DataOps program helps keep high-quality data flowing throughout your company, while making sure you are in [compliance](https://www.getdbt.com/security) with any applicable data regulations. ### 3. Automate through CI/CD Implementing [Continuous Integration and Continuous Deployment (CI/CD) pipelines](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud) as part of your DataOps helps to automate and streamline the data transformation process. - Within the CI/CD pipeline, teams can automate deployment and testing to reduce manual intervention and improve overall efficiency. - Automated pipelines facilitate rapid iteration while minimizing risks, making it easy to scale and evolve your global data processes. - This makes sure that any changes in data workflows are tested, validated, and deployed consistently, improving quality and reducing time to production. ### 4. Design for scalability Data volumes are increasing steadily every year in nearly every org, so it’s essential that your data transformation processes are designed for flexible and built-in scalability. Build a modular data architecture that leverages [cloud infrastructure](https://www.getdbt.com/blog/how-to-build-the-business-case-for-dbt-cloud) and codifies best-practice data management policies. This ensures that your data transformation workflows can handle increasing data volumes and complexity. ### 5. Set up continuous monitoring and optimization [Monitoring data transformation ](https://docs.getdbt.com/docs/deploy/monitor-jobs)jobs, whether in real-time or through scheduled checks, is key to seamless orchestration of your data processes as well as detecting issues early. Tools like logging frameworks and performance dashboards can help track the health of transformation pipelines. Continuous optimization, such as tweaking workflows for improved performance or updating data validation rules. That ensures that processes remain efficient as data volumes and requirements evolve ### 6. Choosing the right solutions Teams need to look for tools and solutions that build in these best practices as part of the platform. [A platform like dbt](https://www.getdbt.com/product/what-is-dbt) that offers built-in capabilities for governance, scalability, and automation lets you transform raw data into analysis-ready insights, and make data-driven decisions with confidence: - Represent all of your data transformation pipelines as [dbt models](https://docs.getdbt.com/docs/build/models) in SQL or Python, enabling anyone to develop data pipelines - Develop [data tests](https://docs.getdbt.com/docs/build/data-tests) - [Store all changes in version control ](https://docs.getdbt.com/docs/collaborate/git/version-control-basics)to facilitate code reviews, versioned releases, and rollback - [Kick off a CI/CD pipeline](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud) to test and push changes from dev to stage to prod - Create and publish standardized metrics with [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) - Find data products and metrics using [dbt Explorer](https://docs.getdbt.com/docs/collaborate/explore-projects) ### Conclusion Data transformation isn’t just a technical step in data management—it’s a cornerstone for achieving your business goals. As a fundamental part of the ETL/ELT process within the modern data stack, data transformation allows you to take your raw source data and find meaning in it for your end business users. With dbt Cloud as your data control plane, your data teams have a standardized and cost-efficient way to build, test, deploy, and discover analytics code. Meanwhile, data consumers have purpose-built interfaces and integrations to self-serve data that's governed and actionable. And you, knowing that you can trust dbt to maintain consistency, efficiency, and security throughout your data transformation lifecycle, can stop worrying and turn your focus to the things that really matter. Learn more about how dbt Cloud can bring DataOps to your organization—[schedule a demo ](https://www.getdbt.com/contact)today. --- --- title: "Pride Month at dbt Labs: LGBTQ+ Identity at work" description: "Learn how dbt Labs honors Pride Month and supports LGBTQIA identities year-round through policies, events, and our Queeries ERG." url: "https://www.getdbt.com/blog/pride-month-at-dbt-labs" date: "2024-06-28" authors: ["Finn Russell"] categories: ["Company"] --- # Pride Month at dbt Labs: LGBTQ+ Identity at work Data transformation is all about taking some raw stuff you were given and building your vision for something more useful, beautiful, and joyful, using all the tools and language available to _Become_. A queer metaphor if there ever was one. At dbt Labs, about 1/8 of our employees identify as LGBTQIA. For the most part, you wouldn’t be able to identify the exact impact of that on our product— there’s no specific feature where our queerness is particularly clockable– except that it’s _all about transformation_. Our queer identities impact our culture, our policies, our values, and the experience we create for our community and users. So here’s a blog post about how we’re honoring and celebrating Pride this month, and how we support a vibrant and diverse team year-round, including: - Queeries, our LGBTQIA employee resource group - dbt Labs’ Trans-Inclusion Policy and Health-Related Travel Policy - How our LGBTQIA identities are reflected in our values ## Queeries: Building community and inclusion Queeries, our employee-led ERG, is an incredible source of community at dbt Labs. Its mission is to foster an inclusive and empowering workplace, where all people can thrive as their authentic selves. For Pride this month, Queeries leaders planned a combo of sync and async ways to celebrate and learn something new. There’s an Understanding Pride Month event that’s a primer on LGBTQIA+ history, including the [Stonewall Uprising](https://guides.loc.gov/lgbtq-studies/stonewall-era) and its impact, along with themed Zoom backgrounds, resources (scroll to the bottom for an abridged list!), and a [swag drop](https://shop.getdbt.com/products?s%5Bf%5D%5Bc%5D%5B%5D=%2FPride). Queeries is run by a rotating group of around three employees. They have run book clubs, led fundraisers for the [National Center for Transgender Equality](https://transequality.org/) and the [Transgender Law Center](https://transgenderlawcenter.org/donate/), and hosted educational and community-building events. Queeries members mirror internally what dbt community members do externally: connect, make our work better and more fun, and put time and effort into making dbt Labs everything we want it to be. Many Queeries members come from work or personal backgrounds where being out wasn’t easy and wasn’t safe. At dbt Labs, we get to focus on making things great instead of making things bearable. That’s a big deal to a lot of us. ## Trans-Inclusion Policy Employee-driven projects and events are important for fostering connections and reinforcing a sense of social safety, and backing that up with formal company processes is essential. In addition to a robust budget for ERGs and other Diversity, Equity, and Inclusion work, dbt Labs prioritizes a supportive environment for all our employees through our policies and benefits. This month, dbt Labs added a Trans-Inclusion Policy to our Employee Handbook that lays out our commitment to creating a workplace where all employees feel welcome, valued, and supported. All employees have the right to express their gender identity, characteristics, or expression without fear of consequences, have the right to use the restroom that corresponds to their gender identity when they’re in-person at dbt Labs offices, and have the right to be addressed by their preferred name and pronouns and to have those preferences updated in company systems and records. Employees also have access to transition-related medical care through our healthcare benefits. This matters, because while we're celebrating Pride Month right now, anti-trans laws and rhetoric are on the rise globally. In the US alone, [44 anti-trans laws have been passed so far this year](https://translegislation.com/), with another 328 under review currently. Trans people, and especially trans women of color, are disproportionally likely to be targets of violence [in the US](https://www.hrw.org/report/2021/11/18/i-just-try-make-it-home-safe/violence-and-human-rights-transgender-people-united) and [globally](https://transrespect.org/en/), and disproportionally likely to die of suicide. The major factors contributing to the suicide risk are discrimination, bullying, violence and harassment, ill-treatment by the healthcare system, and being rejected by family, friends, and community. One policy at one company cannot solve global inequity, but it can help us take care of our people, and show people looking to join dbt Labs that the company is committed to being a safe place for trans people. The Trans-Inclusion Policy in our handbook complements the rest of our policies. It lives alongside our Health-Related Travel benefit that will reimburse up to USD$4,000 in expenses for travel, lodging, meals, and a caregiver for any team member or dependent who has to travel to obtain medical care that isn’t available in their country or state of residence. The policy was originally conceived in response to Roe v. Wade being overturned in 2022, but it supports trans people who need to travel for transition-related care, along with anyone else who can’t access the care they need in the place where they live. These policies, alongside our Equal Employment Opportunity Policy, our Anti-Harassment and Discrimination Policy, and all the rest of our Code of Conduct, are designed to keep our employees safe and set a standard of care, respect, and inclusion in our workplace. ## Our Values: We are human > [Our Values](https://www.getdbt.com/about-us/values) are core to our work at dbt Labs. The “We are human” value includes this in the description: > > We bring our whole selves to work, and we recognize that our identities extend beyond the work that we do. Bringing our whole selves to work means showing up authentically and honoring the identities of our peers, but it also means bringing all the things that make us _us_ into the work that we do. For LGBTQIA employees, that shows up in a lot of ways. We bring experience from queer community organizing to inclusive event planning for Coalesce and other dbt Labs events. We bring our ethos (and our memes) to the dbt slack. We strive to bring curiosity, empathy, and a sense of communal responsibility to the work we do, which makes us better at working together, solving problems, and supporting the people who use our products. I’m very proud of dbt Labs this month, and very glad to share that pride with you. [Reach out on Slack](https://www.getdbt.com/community/join-the-community) to give us feedback on how we can make our product, our community, and our events more inclusive and supportive! ## Resources An abridged list of resources collected by our Queeries ERG: - **History:** - [NASCSP: The history and significance of Pride Month](https://nascsp.org/in-honor-of-pride-month-a-little-history/) - [Smithsonian: Marsha Johnson, Sylvia Rivera, and the history of Pride Month](https://www.si.edu/stories/marsha-johnson-sylvia-rivera-and-history-pride-month) - **Advice on coming out:** - [Advice on coming out, and what to do if someone comes out to you](https://www.npr.org/2020/06/01/867059156/navigating-the-coming-out-conversation-from-both-sides) - **Pronouns:** - [Personal pronouns in the workplace](https://www.stonewall.org.uk/workplace-trans-inclusion-hub/beginner%E2%80%99s-guide-pronouns-and-using-pronouns-workplace) - [Pronouns 101: introduction to your loved one’s new pronouns](https://kconrod.medium.com/pronouns-101-introduction-to-your-loved-ones-new-pronouns-3fef080266d0) - [Pronouns 102: how to stop messing up pronouns](https://kconrod.medium.com/pronouns-102-how-to-stop-messing-up-pronouns-9bd66911118) - **For managers:** - [LGBT Inclusion at Work: The 7 Habits of Highly Effective Managers](https://dojpride.org/wp-content/uploads/2014/04/for-viewing-online-lgbt-tips-for-managers-brochure-accessible-fina.pdf) --- --- title: "16 must-have data engineer skills" description: "Discover 16 essential data engineer soft and technical skills, from SQL to cloud platforms, to thrive in modern data engineering." url: "https://www.getdbt.com/blog/data-engineer-skills" date: "2024-06-24" authors: ["Daniel Poppy"] categories: ["Learn"] --- # 16 must-have data engineer skills Data engineering is a critical function in ensuring that data pipelines are efficient, reliable, and scalable. Data engineers play a vital role in bridging the gap between raw data and actionable insights. In this article, we'll explore what a data engineer does, why data engineering is crucial for businesses, and the essential skills—both technical and soft—that every data engineer should know. We'll also examine how tools like dbt Cloud can help data engineers excel in their roles. **** ## What is a data engineer? Simply put, a data engineer designs, builds, and maintains the infrastructure that supports the collection, storage, and transformation of data. This infrastructure includes data pipelines that move data from various sources into a central data warehouse or data lake where it can be analyzed. Data engineers ensure efficient, error-free, and scalable data flow, enabling data scientists and analysts to focus on generating insights. They often work with large datasets and are expected to build systems that can handle both structured and unstructured data. ## Why is data engineering important? [Data engineering](https://www.getdbt.com/blog/what-is-data-engineering) is crucial because it ensures that data is accessible, accurate, and reliable. Without proper engineering, data can become a bottleneck, slowing down analytics and impeding decision-making. In an age where data-driven decisions are key to staying competitive, organizations rely on data engineers to maintain the health and efficiency of their data ecosystems. With the rise of machine learning and artificial intelligence, clean, well-organized, and structured data is more important than ever. Data engineers make this possible by building robust data pipelines, often using both [ETL (extract, transform, load)](https://www.getdbt.com/blog/extract-transform-load) and [ELT (extract, load, transform)](https://www.getdbt.com/blog/extract-load-transform) methods to move and structure data. ## Technical data engineer skills ### SQL proficiency Structured Query Language (SQL) is the backbone of most data engineering work. Data engineers need to be adept at writing efficient SQL queries to manipulate and retrieve data from databases. SQL is essential for working with relational databases like PostgreSQL, MySQL, and others. ### Data warehousing Understanding data warehouse design and architecture is fundamental for data engineers. A good grasp of how data is stored, organized, and accessed in data warehouses (like Snowflake, Redshift, or BigQuery) allows engineers to optimize data retrieval for analytics. ### ETL and ELT frameworks Extract, Transform, Load (ETL) and Extract, Load, Transform (ELT) frameworks are crucial for building data pipelines. ETL involves extracting data from source systems, transforming it to fit the data model, and loading it into a target system. ![ETL data pipeline diagram showing extraction from sources like Email CRM, Netsuite, Facebook Ads, and Backend DB, transformation steps including rename, cast, join, and enrich, followed by loading into a data warehouse.](https://cdn.sanity.io/images/wl0ndo6t/main/8ebc1795380b891973bec53f9410998e4e21e006-1216x698.webp) ELT, on the other hand, focuses on loading raw data into the target system first and then transforming it. Both approaches have their use cases, and understanding when to apply each is critical for data engineers. ![ELT data pipeline diagram showing data extracted from sources like Email CRM, Netsuite, Facebook Ads, and Backend DB, loaded into a data warehouse as raw data, then transformed through steps such as rename, cast, join, and enrich.](https://cdn.sanity.io/images/wl0ndo6t/main/37560d5362949a8d4de4090389003ecab617c6ef-1706x748.webp) **** ### Programming languages Data engineers should be proficient in at least one programming language, with Python and Java being the most common. Python, in particular, is favored due to its versatility and the rich ecosystem of libraries like Pandas and NumPy, which are useful for data manipulation and transformation. ### Cloud platforms Cloud infrastructure plays a significant role in modern data engineering. Familiarity with platforms like AWS, Google Cloud, or Microsoft Azure is essential, as more companies are moving their data workloads to the cloud for scalability and cost-effectiveness. Knowing how to set up and manage services like S3, Redshift, or Google BigQuery is invaluable. ### Data modeling [Data modeling](https://www.getdbt.com/blog/data-modeling-techniques) skills allow engineers to define how data is organized within databases. Engineers need to know how to design efficient schemas, choose appropriate data types, and understand normalization vs. denormalization. This ensures that data is stored in a way that is easy to query and analyze. ### Data governance and security As organizations deal with growing volumes of sensitive data, security and [governance](https://www.getdbt.com/blog/data-governance-best-practices) are becoming increasingly important. Data engineers must implement best practices for data encryption, masking, and role-based access control. They also need to be familiar with compliance regulations, such as GDPR, to ensure that data is handled appropriately. ### Automation and orchestration tools Managing data pipelines manually is time-consuming and error-prone. Tools like [Apache Airflow](https://www.getdbt.com/blog/dbt-airflow) or [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) can automate and orchestrate these processes, ensuring data flows seamlessly between systems. Data engineers should understand how to set up and manage these tools for pipeline automation. ### Big data tools Handling massive datasets requires specialized tools. Data engineers should be familiar with technologies like Hadoop, Spark, and Kafka, which are designed to process and manage large volumes of data in real-time or batch processing environments. ### API integration Many data pipelines require extracting data from APIs. Understanding how to work with RESTful APIs and tools like Postman is essential for building robust pipelines that can pull data from third-party sources. ## Soft data engineer skills It's easy to overemphasize the technical side of data engineering. But data engineers are _always_ part of a team. Being able to manage the human side of an engineering team is also essential. ### Communication Data engineers often work in cross-functional teams with data scientists, analysts, and business stakeholders. Strong communication skills are essential for understanding requirements and explaining technical concepts to non-technical stakeholders. ### Problem-solving The ability to troubleshoot and solve complex problems is crucial in data engineering. Whether it's debugging a failing pipeline or optimizing a slow-running query, data engineers need to approach problems with creativity and persistence. ### Collaboration Data engineers need to work closely with data analysts, data scientists, and IT teams. Strong collaboration skills ensure that everyone is aligned and that data infrastructure supports the broader business goals. ### Adaptability The data landscape constantly evolves with new technologies and methodologies. Data engineers must be adaptable and open to learning new tools, frameworks, and techniques to stay current in the field. ### Attention to detail Data engineers must be detail-oriented, as even small errors in a data pipeline can lead to incorrect analyses and flawed business decisions. Ensuring data integrity and accuracy is paramount. ### Project management Data engineers often manage multiple projects simultaneously, from building new pipelines to maintaining existing infrastructure. Having strong project management skills allows engineers to prioritize tasks, meet deadlines, and ensure smooth delivery of projects. ## How dbt Cloud helps data engineers do their best work Data engineers can use [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) to transform raw data into actionable insights more efficiently. While dbt Core is an open-source tool that focuses on transformation within the ELT process, dbt Cloud enhances this by providing additional features like automated workflows, version control, and collaboration tools. By using dbt Cloud, data engineers can automate much of their pipeline, reducing the need for manual intervention. It integrates seamlessly with existing data warehouses, making it easier to manage large-scale transformations without overcomplicating workflows. dbt Cloud also provides robust monitoring and alerting features, ensuring that pipelines are always running smoothly. Additionally, dbt Cloud’s ability to enforce data testing and documentation helps ensure the quality and reliability of the data, which is critical for maintaining trust in the outputs of analytics and machine learning models. ## Next action Tools like [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) play a crucial role in helping data engineers manage the complexities of modern data pipelines, making their work more efficient and scalable. As data continues to play an increasingly central role in business decision-making, the demand for skilled data engineers will only grow. To learn more about how engineers, and data engineering, fit into broader workflows, try these related resources: - [Analytics engineering and data engineering: Do you need both? | dbt Labs](https://www.getdbt.com/blog/analytics-vs-data-engineering) - [Understanding AI data engineering | dbt Labs](https://www.getdbt.com/blog/ai-data-engineering) - [How AI will disrupt data engineering as we know it | dbt Labs](https://www.getdbt.com/blog/how-ai-will-disrupt-data-engineering) ## FAQs **What technical skills are most important to prioritize as a beginner data engineer?** SQL proficiency should be your first priority as it's fundamental for data manipulation and querying. Python comes next, as it's versatile for building data pipelines and automation scripts. Understanding ETL/ELT processes is crucial for moving data effectively between systems. Cloud platform familiarity rounds out your foundation, with AWS, Google Cloud, or Azure being most common in the industry. **How do soft skills impact a data engineer's career progression?** Communication skills help data engineers collaborate with cross-functional teams and explain complex concepts to stakeholders. Problem-solving abilities are essential for troubleshooting pipeline issues and optimizing data workflows. Adaptability ensures you can keep pace with rapidly evolving technologies and methodologies. These soft skills often differentiate senior engineers who can drive projects and establish best practices within organizations. **How does dbt Cloud enhance a data engineer's workflow?** dbt Cloud automates transformation workflows, reducing manual interventions and potential errors in the process. It integrates with existing data warehouses while providing robust testing and documentation capabilities. Its monitoring features ensure pipelines operate reliably by alerting engineers to potential issues. These efficiencies allow data engineers to focus on higher-value tasks rather than managing repetitive processes. **What's the difference between ETL and ELT approaches in data engineering?** ETL transforms data before loading it into target systems, making it ideal for legacy systems with processing limitations. ELT loads raw data first and transforms it within the target system, leveraging modern data warehouse computing power. The ELT approach often provides more flexibility for different analytical needs. Your choice between these approaches depends on your specific use case and technological infrastructure. **How can I gain practical experience as an aspiring data engineer?** Create personal projects that extract data from public APIs and transform it using Python. Practice building data pipelines with open-source tools like Apache Airflow or dbt Core. Set up a cloud environment using free tiers from major providers to experiment with data warehousing. These hands-on experiences will demonstrate your capabilities to potential employers better than theoretical knowledge alone. --- --- title: "Getting started with Data Vault on AutomateDV and dbt Cloud" description: "Learn how to boost your data warehouse development with AutomateDV and dbt Cloud in this Data Vault guide." url: "https://www.getdbt.com/blog/getting-started-with-data-vault-on-automatedv-and-dbt-cloud" date: "2024-06-21" authors: ["Alex Higgs"] categories: ["Product"] --- # Getting started with Data Vault on AutomateDV and dbt Cloud _This is a guest post authored by Alex Higgs, AutomateDV product manager at Datavault._ *** [**Learn how to build a Data Vault in an on-demand walkthrough for scalable data architecture.**](https://www.getdbt.com/resources/webinars/build-for-scale-agility-and-reliability-with-dbt-cloud-automatedv-and-data-vault) *** If you are looking for a trusted and reliable tool to streamline your data warehouse development, leveraging AutomateDV, dbt Cloud, and Data Vault might be your solution. The Data Vault method is a proven data warehousing methodology for building reliable and integrated data warehouses on your data platform. It is an open-source method available to anyone. One of the key advantages of Data Vault is that it is pattern-based and therefore lends itself to automation. AutomateDV is a dbt package providing Data Vault automation to dbt. In this AutomateDV quickstart guide, you will understand exactly what is required for a successful project with the Data Vault automation tool built on dbt. ## **What is a Data Vault?** Data Vault provides an approach for deploying enterprise data warehouses which includes three core pillars: architecture, modelling and method. Importantly it is well-suited for Agile implementation, which is vital for the rapid scalability required by modern organisations. Successful Data Vault implementations focus on integration at the business concept level rather than being source-system driven, so developing some understanding of Data Vault modeling and standards helps prevent wasted time and false starts. ### **Why use Data Vault?** **Long-term historical storage**: Data Vault supports storage of data in an integrated way from multiple operational systems for the history of your data and business. All the data, all the time, within scope. **Auditing and Traceability**: Data Vault provides built-in auditing and traceability of data over time which is vital for certain industry sectors such as financial, healthcare, education and more. Get access to historical insights and understand what you knew when as a result of Data Vault’s carefully considered design. **Optimized for load**: The Data Vault model is designed to optimize loading time by supporting highly-parallel loading, which is crucial for handling large datasets, multiple source systems and real-time feeds. This also reduces the time to insight for the business. **Scalability**: The modular structure of Data Vault allows it to scale alongside your business as it grows, allowing rapid adaptation to changing business requirements. Build a data warehouse with agility. **Single version of truth**: By separating business rules from the raw data, Data Vault ensures a single version of truth, enhancing data trustworthiness and auditability. Reduce the risk of business transformations being calculated inconsistently, increasing understanding and trust in how metrics are calculated. ## **How does dbt Cloud and Data Vault work together?** **Enhanced Visibility and Governance**: Combining dbt and AutomateDV is smart—they work seamlessly together. dbt features included in the [dbt Mesh ](https://www.getdbt.com/product/dbt-mesh)offering include data security, lineage and contracts as well as model versioning support many of the paradigms also built-in to Data Vault. All of this is defined using files, so it is all tracked in version control as well. **Scalability and Efficiency**: AutomateDV on dbt Cloud supports scalable Data Vault components and has been tried and tested in larger companies and production environments. Combined with enterprise features in dbt Cloud, it provides effective management of extensive Data Vault projects. **Higher Standards and Reliability**: The dbt ecosystem of packages and features provide a solid foundation for doing Data Vault projects consistently and to a high standard with less maintenance overhead. [dbt Cloud streamlines development, ensuring speed without sacrificing quality](https://www.getdbt.com/product/dbt-cloud). Paired with Data Vault's pattern-based approach, it lets developers focus on business needs, not just technical tasks. ## **What is AutomateDV?** [AutomateDV](https://automate-dv.com/) is the open-source tool that, when combined with dbt Cloud, provides Data Vault templates to your dbt Project. The dbt package allows developers to automatemany of the tasks involved in setting up and maintaining a Data Vault, making it easier for data engineers to manage their projects. Don’t just take our word for it, though. Organizations like [McDonald’s Nordic](https://www.getdbt.com/case-studies/mcdonalds-nordics), Betway, and NHS Digital have used AutomateDV to develop andmaintainlarge-scale Data warehouse platforms, allowing them to integrate many disparate systems and meet business needs faster than a traditional approach. AutomateDV is a package for dbt available on the dbt hub [here](https://hub.getdbt.com/Datavault-UK/automate_dv/latest/). ## **Recommended content** Here are several books that can help you get started with Data Vault and AutomateDV, including: - [‘Building a Scalable Data Warehouse with Data Vault 2.0’ – Dan Linstedt](https://www.amazon.co.uk/Building-Scalable-Data-Warehouse-Vault/dp/0128025107/ref=asc_df_0128025107/?tag=googshopuk-21&linkCode=df0&hvadid=310819191513&hvpos=&hvnetw=g&hvrand=2682553738356535476&hvpone=&hvptwo=&hvqmt=&hvdev=c&hvdvcmdl=&hvlocint=&hvlocphy=1007203&hvtargid=pla-450204322531&psc=1&mcid=db5b05e95fe83a8a98ba18a6743d77be&th=1&psc=1) - [‘The Data Vault Guru’ - Patrick Cuba](https://www.amazon.co.uk/Data-Vault-Guru-pragmatic-building/dp/B08KJLJW9Q/ref=asc_df_B08KJLJW9Q/?tag=googshopuk-21&linkCode=df0&hvadid=463092568931&hvpos=&hvnetw=g&hvrand=7191281729089180479&hvpone=&hvptwo=&hvqmt=&hvdev=c&hvdvcmdl=&hvlocint=&hvlocphy=1007203&hvtargid=pla-1004699066823&psc=1&mcid=1c5845a8fcb238fd981c1d0713b3c26b&th=1&psc=1) - [‘The Elephant in the Fridge’ – John Giles](https://www.amazon.co.uk/Elephant-Fridge-Success-Building-Business-Centered/dp/1634624890/ref=asc_df_1634624890/?tag=googshopuk-21&linkCode=df0&hvadid=310819191513&hvpos=&hvnetw=g&hvrand=3188481295510491746&hvpone=&hvptwo=&hvqmt=&hvdev=c&hvdvcmdl=&hvlocint=&hvlocphy=1007203&hvtargid=pla-750266742918&psc=1&mcid=5c93af354333351b817f9e0d9a55a2d8) We also recommend the following online resources: - [Demystifying Data Vault with dbt - Coalesce 2023 (youtube.com)](https://www.youtube.com/watch?v=XIG3m-2O_Lg) - [What is Data Vault? (data-vault.com)](https://data-vault.com/what-is-data-vault-core-concept/) - [How to learn Data Vault in 2024 (data-vault.com)](https://data-vault.com/how-to-learn-data-vault-in-2024/) - Patrick Cuba articles ## **Recommended training** Before diving into AutomateDV, we recommend that you familiarize yourself with the core concepts of Data Vault and undergo AutomateDV training. This will ensure you have a solid foundation to work from. - [Data Vault: Core Concepts](https://data-vault.com/free-data-vault-training/) – a free, 2-hour live introduction to Data Vault training. - [Data Vault for Developers Course](https://data-vault.com/data-vault-training-for-developers/) – a 20-hour live online course to learn the fundamentals of successful Data Vault projects. - [AutomateDV training](https://automate-dv.com/training-courses/) – a specialist AutomateDV online course providing on-demand material and hands-on exercises to get you ready to use AutomateDV on a project. ## **AutomateDV documentation** Once you’ve gone through the recommended content and training, it’s time to put your knowledge into practice. A worked example can help you understand how to apply what you’ve learned. The [AutomateDV ReadTheDocs site](https://automate-dv.readthedocs.io/en/latest/worked_example/) is a comprehensive resource that provides detailed information about the tool. **What’s included in the AutomateDV documentation?** - **Worked example** – Demonstrates AutomateDV. Guiding you through developing a Data Vault based on the Snowflake TPC-H dataset, step-by-step using pre-written dbt models using AutomateDV templates. - **Best practices** – Outlines the standards for using AutomateDV, including hashing, loading, and NULL handling. - **Tutorials** – Provides a detailed understanding of the Data Vault concepts which AutomateDV supports, and how to use AutomateDV to create a Data Vault. - **Much more**: Explore additional resources to deepen your understanding and proficiency with AutomateDV and Data Vault. [**Get started with Data Vault in this on-demand walkthrough by AutomateDV and dbt Labs.**](https://www.getdbt.com/resources/webinars/build-for-scale-agility-and-reliability-with-dbt-cloud-automatedv-and-data-vault) --- --- title: "Data orchestration vs. ETL: What’s the difference?" description: "What separates data orchestration from ETL? A look at how both processes overlap - and where they differ." url: "https://www.getdbt.com/blog/data-orchestration-vs-etl" date: "2024-06-20" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data orchestration vs. ETL: What’s the difference? Both data orchestration and ETL enable collecting, uniting, and transforming data to extract business value. However, both of these processes have very different scopes and take different approaches. Here’s how they overlap, how they differ, and which one provides a more modern and reliable approach to managing data at scale. ## What is ETL? [ETL, or Extract-Transform-Load](https://docs.getdbt.com/terms/etl), is a process used to extract data from one location, convert it into another format, and place it into another. It’s used most often to transfer data from multiple sources —e.g., relational databases, flat files, CSVs, etc.—into a target system. As denoted by the acronym, the process has three central parts: ### Extract Pull the data from the host system or systems and put it into a temporary staging area. This can be done as a batch process (e.g., a job that runs and queries data every x hour) or via a real-time replication mechanism such as Change Data Capture (CDC). ### Transform Coerce the data into a different format suitable for the new workload. At this stage, the process also cleans the data, standardizing the use of types and formats and correcting any errors (e.g., null values in required fields). ### Load Data is brought into the destination, where it’s available for query by data consumers. ETL was primarily a process used before the advent of cloud computing. Originally, it loaded data from online transaction processing (OLTP) systems into a relational format optimized for reporting. ## What is data orchestration? Data orchestration is the process of automating the flow of data systematically throughout an enterprise. It gathers data that was previously siloed and unites it into a single location, where it can be discovered and used by multiple teams. Data orchestration consists of three phases: ### Organization A data pipeline gathers data from various places —flat files, relational databases, APIs, etc.—and collects it in a data warehouse or a data lake. ### Transformation As in ETL, data is unified, cleansed, and corrected across its multiple formats. Newly arriving data goes through a series of manual and automated quality checks before it’s considered valid and promoted to a production database. ### Activation Data is delivered and put to use by Business Intelligence tools or other data-driven applications. In contrast to ETL, data orchestration typically relies on a related method called ELT (Extract-Load-Transform). In ELT, data is loaded as is into a data warehouse and then transformed into multiple formats depending on the use case. This enables different teams to use the same core data in different ways. **** ## Data orchestration vs. ETL: The key differences Data orchestration and ETL differ in two key facets: scope and management. ### Scope ETL jobs are often one-off processes that solve a specific problem—e.g., producing a quarterly sales report, a monthly report on patient outcomes at a hospital, etc. By contrast, data orchestration is about managing the flow of data across the enterprise, not just for one single purpose. It aims to eliminate [data silos](https://www.getdbt.com/blog/how-dbt-can-help-solve-4-common-data-engineering-pain-points), islands of independent data that are hard to discover and, more often than not, not properly monitored or governed. ### Management ETL jobs have a reputation for being brittle and requiring constant manual care to keep running. Traditional ETL tools also have difficulty managing the volume of data that most enterprises manage. Data orchestration focuses on creating automated data pipelines that run periodically—either on a schedule or in real-time as new data arrives in its source systems. A well-designed data orchestration system provides tools for tracking the status of data pipeline jobs, issuing alerts when errors occur, and providing centralized observability via metrics and logs. ### Development ETL jobs are often developed solely by data engineers working alone based on loose requirements provided by a business user. This often results in a mismatch between what engineering delivers and the needs of the business. As a result, it can take weeks or even months to deliver a new pipeline to production. Data orchestration platforms provide tools that enable engineers, analysts, and decision-makers to work together closely in short, rapid cycles. We call this the [Analytics Development Lifecycle](https://www.getdbt.com/resources/guides/the-analytics-development-lifecycle), or ADLC. In the ADLC, all three of these personas take an active role in designing, implementing, approving, and unblocking the creation of new data transformation pipelines. Because data pipelines are highly automated, teams can deploy, test, and re-deploy changes rapidly, shaving weeks off of traditional data pipeline development. ## The benefits of migrating from ETL to data orchestration Some companies still use ETL for processing data in highly regulated industries, such as health care. Most companies, however, can benefit from moving to a comprehensive data orchestration platform. For years, dbt has provided industry-standard technology for modeling and transforming data across the enterprise. dbt Cloud expands upon these base capabilities, creating a data control plane you can use to manage the flow of data across your enterprise. Using [dbt Cloud](https://www.getdbt.com/product/dbt-cloud), you can forge a data orchestration platform that provides numerous benefits over using ETL: - Accelerated data delivery - Improved data quality - Democratization of data - Improved compliance and data governance ### Accelerated data delivery Because data orchestration pipelines are automated, they can sync more data, more reliably, and in a shorter time frame. If a problem arises, the data orchestration platform provides tools to issue a notification, perform root cause analysis, and get the pipeline back up and running quickly. Using dbt Cloud, data engineers, analytics engineers, and even tech-savvy business users can create [dbt models](https://docs.getdbt.com/docs/build/models) using SQL code. They can check their code to a source control system like GitHub and automatically deploy any data model changes through a [Continuous Integration (CI) job](https://docs.getdbt.com/docs/deploy/continuous-integration). ### Improved data quality Because they’re often one-off jobs, data quality in ETL pipelines can be hit or miss. By contrast, data orchestration emphasizes providing data consumers with high-quality data with every job run. Using dbt Cloud, teams can ensure that all data model changes are reviewed by other team members before they go live. Engineers can also create [data tests](https://docs.getdbt.com/docs/build/data-tests) that are run and verified in multiple environments before new transformations are run on production data. ### Democratization of data ETL pipelines don’t make it easy to find and work with data. A data orchestration platform provides easy-to-use tools to find, transform, and activate data. With dbt Cloud, analysts and decision-makers don’t have to file a trouble ticket and wait days—or weeks—for a dedicated engineer to resolve an issue in a data pipeline. Since dbt uses familiar languages like SQL, any authorized user with the requisite knowledge can apply changes to a data pipeline. Once data models are published to production, engineers, analysts, and decision-makers can discover them easily using [dbt Explorer](https://docs.getdbt.com/docs/collaborate/explore-projects). Users can find data, read any associated documentation, view the data’s [lineage graph](https://www.getdbt.com/blog/guide-to-data-lineage) to discover its origins, and see how the source data was transformed. This boosts ‌users’ trust and confidence in the quality of their data. ### Improved compliance and data governance Accidental exposure of sensitive information can have a devastating impact on customer trust. It can also result in expensive regulatory fines. By providing a comprehensive data transformation solution, a data orchestration platform can provide better data security and governance than a more atomized, siloed approach. [dbt Cloud provides a number of tools](https://www.getdbt.com/product/governance) to enforce and monitor data governance: - **Workflow governance**. Engineers, analysts, and decision-makers can standardize on the same platform. Using [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), they can define a single source of truth for key organizational metrics. - **Role and access governance**. Model and project owners can control who can access what via role-based access control. - **Data governance**. Auto-generated documentation, version control, and integrated testing ensure that all changes are thoroughly vetted, recorded, and traceable. ## Conclusion ETL is an older, more traditional approach to managing data transformations on an ad hoc basis. Data orchestration provides a comprehensive approach to automating and monitoring the flow of data across your company. dbt Cloud acts as a data orchestration platform that any company can use to accelerate data delivery, increase data quality, and increase data democratization while simultaneously improving compliance and security. Learn more about how dbt Cloud can serve as your data orchestration platform—[ask us for a demo today](https://www.getdbt.com/contact). --- --- title: "dbt Labs on dbt: How dbt Labs solves marketing challenges with the Campaign 360 Dashboard" description: "See how dbt Labs uses the Campaign 360 Dashboard to solve marketing challenges, driving efficient decision-making and cost savings" url: "https://www.getdbt.com/blog/how-dbt-labs-solves-marketing-challenges-with-the-campaign-360-dashboard" date: "2024-06-19" authors: ["Kathryn Chubb"] categories: ["Product"] --- # dbt Labs on dbt: How dbt Labs solves marketing challenges with the Campaign 360 Dashboard We’re big fans of dbt Cloud for our own projects here at dbt Labs. We’re hosting an entire “dbt Labs on dbt” series to [showcase how various teams at our company use dbt Cloud](https://www.getdbt.com/resources/webinars/dbt-labs-on-dbt-streamlining-kpi-dashboards-with-the-dbt-semantic-layer) for trusted data. This blog is focused on how dbt Labs’ marketing team uses dbt Cloud. The current macroeconomic conditions have been unfavorable for everyone, including marketing teams. Marketing teams are under pressure to **reduce customer acquisition costs**, i**dentify and eliminate inefficient channels**, and **increase average contract values**—all while working with reduced headcount and budget. These pressures have caused a shift toward more data-informed decision-making processes. To address these needs, dbt Labs developed the Campaign 360 Dashboard. ## Introducing the Campaign 360 Dashboard The Campaign 360 Dashboard is a comprehensive tool designed to analyze paid media, organic social, email, web, marketing automation, pipeline, and product usage behavior—all in one place. This dashboard enables marketers to quickly access metrics like campaign performance, engaged contacts, trials, and pre-sourced pipelines, fostering a more data-driven approach to marketing. It's user-friendly and intuitive, enabling every member of the marketing team to derive value from it. ### Goals of the Campaign 360 Dashboard 1. **Improve accuracy and consistency**: Ensure measurement reporting is precise and uniform. 2. **Empower marketing teams**: Enable marketing team members to engage with and own their reporting. 3. **Reduce dependency on data teams**: Allow data and operations teams to focus on developing new data products. ## Using the Campaign 360 Dashboard Let's dive into how our marketing team uses this dashboard, following a typical analysis flow. ### High-level overview At the top of the dashboard, users can get a brief campaign overview. For example, a campaign manager can filter by campaign name and see how many campaigns were run last quarter. This simple question, which might stump many campaign managers, is easily answered here. ### KPI analysis Next, users can focus on key performance indicators (KPIs) for these campaigns. Questions like "How many contacts engaged with these campaigns?" or "How many trials and pipeline opportunities were generated?" are answered in specific sections of the dashboard. ### Campaign effectiveness To determine which campaigns were the most effective, users can compare campaigns against one another. By sorting by source pre-pipeline, users can see which campaigns drove the most pre-pipeline opportunities. ### Detailed follow-up For even deeper analysis, users can follow links to accessory dashboards for granular metrics. Users can identify where conversions came from, how many accounts were qualified for the sales team, and how many of these accounts have been followed up. ### Operationalizing insights Users can operationalize this information by identifying the contacts that weren't followed up and prompting the SDR team to reach out to them. This process has transformed the behavior of our campaign team, making them more confident in data conversations and incorporating daily dashboard checks into their routine. ## How to implement your own Campaigns 360 Dashboard ### Key steps 1. **Standardize and optimize**: First, standardize your measurement framework by collaborating with marketing, operations, and data teams to define campaign structures, KPIs, and dimensions. 2. **Embed business logic**: Use dbt to embed this business logic into your data model. 3. **Drive self-service enablement**: Provide knowledge assets and training to ensure team members can use the dashboard independently. ### Leveraging dbt features - **Macros**: We use macros to codify our marketing campaign nomenclature, parsing custom dimensions from campaign fields to create multiple columns of useful data. - **Upstream tools with guardrails**: We've implemented upstream tools in Notion for stakeholder-controlled mapping of sources and mediums, enabling seamless updates without coding; integrated with our data warehouse for consistent downstream data updates, and ensuring collaborative input with built-in guardrails. - **Packages**: Packages like AdReporting, OrganicSocialPosts, and HubSpotEmailSends streamline the process, allowing us to focus on adding business logic to the tail end of these models. - **DAG (Directed Acyclic Graph):** Visualizing data relationships to enhance understanding and self-serve capabilities. ## Measuring success **Over 94% of dashboard usage for the Campaign 360 Dashboard comes from non-data team members. **This shift has allowed our data and operations teams to focus on building new data products while enabling stakeholders to analyze and optimize their campaigns independently. ## Streamline data operations and elevate marketing insights with dbt The Campaign 360 Dashboard has significantly improved our marketing team's efficiency and confidence in data-driven decision-making. By standardizing measurement frameworks, embedding business logic with dbt, and driving self-service enablement, we've empowered our team and reduced dependency on data and operations teams. We hope you can leverage these insights to implement a similar solution in your organization. Interested in learning more about how dbt Labs uses dbt? Check out our [on-demand webinar](https://www.getdbt.com/resources/webinars/dbt-labs-on-dbt-streamlining-kpi-dashboards-with-the-dbt-semantic-layer) where we dive into how dbt Labs leverages dbt Cloud—and specifically the dbt Semantic Layer—to automate our business KPI dashboards with real-time, accurate metrics across various BI tools. Find out how we transformed a previously manual, hectic, and error-prone process into a streamlined, efficient workflow. --- --- title: "Critical tools for data governance" description: "Scale your data governance for Generative AI with these essential tools to meet growing demands." url: "https://www.getdbt.com/blog/critical-tools-data-governance" date: "2024-06-18" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Critical tools for data governance No organization that relies heavily on data can survive without a detailed data governance strategy. However, implementing that strategy - particularly at scale - requires having the right tools. In this article, we’ll take a look at the tools that are indispensable for maintaining the quality, security, and transparency of data. ## Why data governance requires data governance tools “Data governance” is a broad term that encompasses several aspects of data: - **Data quality**. Poor data quality [costs companies up to USD $12.9 million yearly](https://www.gartner.com/smarterwithgartner/how-to-improve-your-data-quality). Low-quality data undermines data trust, which prevents many new data projects from ever getting off the ground. Maintaining standards and tools to format, clean, verify, and track the movement of data increases trust. It also saves millions in expensive and time-consuming troubleshooting. - **Data security**. [The average cost of a data breach](https://www.ibm.com/reports/data-breach) is USD $4.88 million. Security breaches - both external and internal - can undermine customer trust and incur lasting reputational damage. A data governance security strategy uses role-based permission controls to manage and audit access to data. - **Data compliance**. Companies need to manage sensitive data - company financials, medical records, personally identifiable information (PII) - in accordance with both industry best practices and local regulations. This requires classifying all system data according to its sensitivity levels, tracking its movement across the organization, and maintaining thorough logging and audit trails. - **Data interoperability**. The same critical data often ends up replicated in multiple formats across an organization. Interoperability efforts standardize core data structures and calculations, reducing broken data pipelines and confusion over diverging data. The trick is implementing these practices at scale. [Data continues to grow exponentially year over year](https://www.statista.com/statistics/871513/worldwide-data-created/). The rise of [Generative AI](https://www.getdbt.com/blog/modern-data-strategy-ai) - which requires large data volumes for accuracy - is only increasing the demand for accurate, high-quality data sets. Data governance tools enable governance at scale by combining automation with human oversight. They enable data producers and consumers to manage large volumes of data effectively - e.g., by finding data easily wherever it lives in the organization, or automatically allowing or denying access to data based on user role and data classification. Automating governance using data governance tools enables teams to work independently while ensuring and demonstrating compliance with all applicable data standards and regulations. This [federated computational approach](https://www.getdbt.com/blog/key-components-of-data-mesh-federated-computational-governance) makes data governance a shared, community effort in which data producers, consumers, and governance experts collaborate to create high-quality data sets. ## Critical data governance tools There are a number of data governance tools, with more appearing on the market by the day. However, the following have proven indispensable to any organization looking to implement governance at scale. ### Data catalog The first step in data governance is getting a handle on everything you own. That means somehow keeping track of data across hundreds or thousands of different data stores. This is where a data catalog comes in. A data catalog is the single source of truth for data in an organization. It provides a repository that describes, not just the data itself, but all of its associated metadata - e.g., who owns it, when it was last updated, etc. Data catalogs work by connecting directly to data sources or to an intermediary description of your data (e.g., a dbt [data model](https://docs.getdbt.com/docs/build/models)). They provide an interface where other users across the company (depending on their permissions) can then discover and use this data in their own data projects. Capturing all data in a data catalog lays the foundation for all other data governance practices. It provides a central location where the organization can monitor data security, classification, and quality, no matter where the data itself lives. ### Data lineage Seeing how data flows throughout your company is critical to increasing trust in data. A [data lineage](https://www.getdbt.com/blog/guide-to-data-lineage) tool shows, in visual form, a sequential workflow of how data travels through your system at a data set or columnar level. You can track data back to its source and see who’s consuming it - via reports, applications, etc. - throughout your organization. You can use data lineage to improve data governance and quality in a number of ways: - **Root cause analysis**: When a data problem occurs (e.g., a data pipeline breaks due to a malformatted value), data engineers can use lineage to find and fix the problem at its source. - **Impact analysis**. Leverage data lineage to see when a change to a data set’s values might result in downstream breakages, so you can work with the impacted stakeholders before releasing it. - **Verify data provenance**. Data consumers and business stakeholders can use data lineage to assess ### Data security management A data catalog enables anyone to discover data across the company. However, that doesn’t mean everyone should have permission to access any and all data. Every teams needs controls they can leverage to control which data they expose, and to whom. With data security management, teams can establish fine-grained access controls around data. They can expose certain data sets to company users based on the user’s roles, while keeping other data sets private for their own internal use. ### Data classification Regulations such as the [General Data Protection Act (GDPR)](https://gdpr-info.eu/) require strict handling of sensitive customer information. Penalties for violating these regulations are often stiff and costly. Data classification provides tools to tag data according to its classification. Using data classification, you can identify sensitive data and enforce appropriate data governance policies, such as restricting access or expunging data after a certain time period. For example, say a data set stores a customer’s email address along with their credit card information. You can classify the email address as Moderate sensitivity and the credit card info as High sensitivity. In turn, data governance tools can automatically restrict access to this data, as well as audit any access to it. ### Data quality management Data quality often suffers because data engineers lack the tools required to control, track, approve, test, and track changes to data. Data quality management tools enable data producers and consumers to use a [DataOps](https://www.getdbt.com/blog/what-is-dataops) methodology to manage data changes. In a DataOps framework, data changes are modeled as code, so that data teams can commit, review, test, and approve every change prior to release. With DataOps, every change goes through a Continuous Integration and Continuous Deployment (CI/CD) process that tests it in multiple environments before pushing it to production. Once live, data teams can continuously run their tests against incoming data, raising alerts and sending notifications if their tests detect any anomalies. ## How dbt Cloud supports data governance For years, data teams have relied on dbt to model and transform their data at scale. With [dbt Cloud](https://www.getdbt.com/product/dbt-cloud), your company can build a data control plane that provides a firm foundation for data governance. dbt Cloud provides numerous features to enable data governance at scale, including: - **dbt Explorer**: See a bird’s-eye view of all of your end-to-end data pipelines, along with all of their dependencies. See the flow of your data down to the column level via an intuitive data lineage graph to diagnose data quality issues and perform impact analysis. - **dbt Mesh**: Enable each data domain team to create, maintain, and ship its own dbt projects. Find and import models from other data teams. Give teams the tools to define and manage access permissions to their dbt models. Track global standards for data governance across all teams to measure and improve the overall quality and compliance of your data estate. - **Continuous integration**. Automatically test your changes against a test schema before merging new code to production. dbt Cloud builds and tests only the code that changed. - **Tests and alerts**. Develop assertions against your data using a few lines of code and test them with every job run. Run tests continuously in production to validate incoming data, and raise alerts and notifications via Slack, email, or webhook if your tests catch an anomaly. - **dbt Semantic Layer**. Standardize on key metrics by defining them in a single, centralized location alongside your dbt models. Enable anyone across the organization to access metrics via Tableau, Google Sheets, Hex, and a host of other analytics and BI tools. To learn more about how dbt Cloud can support your data governance initiatives, [ask us for a demo today](https://www.getdbt.com/contact). --- --- title: "Summit season: Key takeaways from the Snowflake and Databricks Summits" description: "I just spent two weeks in a hotel room near Moscone. Here's what I learned." url: "https://www.getdbt.com/blog/key-takeaways-snowflake-databricks-summits" date: "2024-06-16" authors: ["Tristan Handy"] categories: ["Company"] --- # Summit season: Key takeaways from the Snowflake and Databricks Summits _This post first appeared in_ _[The Analytics Engineering Roundup](https://roundup.getdbt.com/p/summit-season)._ Hi! It’s been a minute since I’ve been able to sit down and write; while podcasting is fun (and my [most recent episode](https://roundup.getdbt.com/p/the-rapid-experimentation-of-ai-agents) with Yohei Nakajima was particularly fascinating!), there’s nothing quite like the opportunity to sit down and collect my thoughts via long-form writing. The fact that a big group of humans read this newsletter is just a side benefit; really, this is my excuse to block a big chunk of time on my calendar to do nothing but learn and think. It’s been the only hope I’ve had of staying current over the past decade as the ecosystem has seen multiple massive shifts. I do want to take a second to say, as I dive into a newsletter issue that is about two specific vendors, what my personal stance is around talking about vendors in this space. In this newsletter I will share news about vendors in the data space, I will extrapolate and anticipate and analyze, I will express excitement when it’s legitimately held. **I will never say negative things about vendors in public.** Including competitors. I used to! When this newsletter was only read by a small group and the powers-that-be in the industry couldn’t care less what I said here, I consistently dunked on vendors that I felt like weren’t innovating or were captive to an outdated mindset. Never unkindly, but very…directly. But our world is very small, my voice now gets real attention throughout the industry, and _dbt Labs partners with basically all of these vendors_. You can guarantee that every time I say anything remotely negative about _any vendor_ in the industry, I hear about it. Fortunately, there are plenty of voices who can fill this role. If you want a “who won the summit wars?” post, this isn’t it. But I really do think there’s enough very positive stuff going on that the horserace doesn’t have to be the story. Sorry for the long preamble, let’s get into it. ## Overall Both events were _very_ well-attended and well-produced. IMO by far the best experience that I’ve had at either event. I don’t know attendee numbers for both events, but I think each of them had to have over 10k folks in attendance based on my knowledge of other events and crowd sizes. As crowd size increases, attendee composition changes. Two years ago these were events where I would catch up with industry friends. Snowflake Summit two years ago was a bunch of MDS nerds thinking about their stacks and team structures. Databricks Summit two years ago was just starting to move beyond a group of Spark nerds. Both communities have evolved significantly. If you’re a long-timer, you may have experienced some nostalgia this year. But if you wanted to hear about new innovations and meet potential customers, these events have never delivered more. ## Iceberg, Delta, and the Metastore As much as everyone arrived at both events assuming that AI was going to be The Thing, the big f#$!ing deal this year was actually open table formats. The news items on this front: - [Databricks bought Tabular](https://www.databricks.com/company/newsroom/press-releases/databricks-agrees-acquire-tabular-company-founded-original-creators), the makers of a metastore built on top of Iceberg. - Snowflake [announced a new open source metastore called Polaris](https://investors.snowflake.com/news/news-details/2024/Snowflake-Unveils-Polaris-Catalog-and-Emphasizes-Commitment-to-Interoperability-with-AWS-Google-Cloud-Microsoft-Azure-Salesforce-and-More/default.aspx) based on the [Iceberg REST spec](https://iceberg.apache.org/concepts/catalog/). There appears to be wide cross-ecosystem support for this project, and they committed to releasing / open sourcing the code within 90 days. - Databricks [open sourced Unity Catalog](https://www.databricks.com/company/newsroom/press-releases/databricks-open-sources-unity-catalog-creating-industrys-only-open), a metastore that supports all of the leading table formats, including Delta, Iceberg, and Hudi. Widely used today within the Databricks ecosystem. Both events were replete with rumors about this stuff. It’s very clear that it is not a coincidence that it all happened within a two-week timespan, but I’m not going to speculate here on exactly what happened behind-the-scenes or the motivation of the various players. Honestly who cares. What matters to me is that shared, cross-platform open file/table formats and open metastores have the ability to dramatically shift the dynamics inside of the data ecosystem. (FYI: if the differences between file formats, table formats, and metastores are a bit of a mystery to you, [this article from Starburst](https://www.starburst.io/data-glossary/open-table-formats/) does a good job of explaining the differences.) Let me just make some broad statements that I believe to be true. 1. For most companies, the biggest data-related infrastructure spend is in their compute layer. It’s not in ingest or transform or BI or storage, it’s in compute. And it doesn’t matter what you’re using—a commercial product like Snowflake or Databricks or OSS Presto on raw EC2 nodes—this is nearly always true. Sometimes it can be 10x bigger than the next biggest cost. 2. Companies are therefore highly incentivized to exert downwards pressure on this spend. This can be a board-level priority for Fortune 500 companies. 3. But, this is hard to do because of data gravity. If you load all of your data into one system in one proprietary file format, you lose all of your negotiating leverage against that vendor. It is a TON of work to replatform, and potentially career-limiting for the relevant executive if done poorly. 4. The majority of all workloads that run in modern platforms are defined in languages / frameworks that are not specific to the platforms themselves: SQL, Python, Spark. Many of these workloads _could_ be ported between platforms with modest code changes. This is notably different than the prior era, where logic was locked up inside things like stored procedures. 5. If you eliminate data lock-in and allow workloads to “travel” between platforms based on cost / performance characteristics, you create a more efficient market for workloads. This allows competition to naturally push prices down over time. Notably, both Sridhar and Ali said—on stage in their keynotes!—some version of “may the best engine win.” So this is competition that they’re both ready for and (seemingly!) welcome. I think this is truly the only customer-centric stance to have and am very happy to see both sides embracing it. Now, this low-friction workload portability doesn’t happen automatically just because you have an open file format, table format, and metastore. From what I can tell, in order to make this a reality, you need: 1. An ability to transpile workloads between execution engines’ dialects / environments with accuracy guarantees. 2. An ability to route workloads automatically between multiple execution engines. 3. An ability to decision which engine is best suited to execute a given workload. 4. The platforms themselves have to have a minimum shared level of support for the various table formats and metastores, with appropriate performance characteristics. The big gate-keeper here in the past has been point #4, and the reason this week was so interesting was that it represented a major new commitment from both Snowflake and Databricks to support these open standards more completely, specifically around Iceberg. I expect that this will open up a flood of innovation over the coming months and will be watching this space closely. ## GenAI GenAI took up a lot of airtime at both events. For both, there was a lot of attendee talk about how much of this stuff is “real” (i.e. driving real consumption and customer value) vs. investing out ahead of where actual customers are today. My read is that there is certainly a bit of both going on, but use cases are starting to emerge. My expectation is that the coming year is the year where this balance flips and we start reading about a lot more real production use cases, because the platform features really do work at this point. Snowflake announced [a number of updates to their Cortex AI platform](https://www.snowflake.com/blog/cortex-ai-advances-enterprise-ai-no-code-development/) including: - **Cortex Analyst:** Allows business users to interact with data in Snowflake using natural language - **Cortex Search:** an enterprise search offering that uses Neeva retrieval and ranking technology with Arctic LLMs - **Cortex Fine-Tuning:** A suite of tools for customizing LLMs - **AI & ML Studio:** A no-code studio for non-technical users to build with AI and ML - **Cortex Guard:** Flags and filters out harmful content in your data to help ensure your LLM-powered experiences are safe and usable. The word of the day for Databricks was “compound AI system”, and most of their [updates](https://www.databricks.com/company/newsroom/press-releases/databricks-unveils-new-mosaic-ai-capabilities-help-customers-build) to [Moasic AI](https://www.databricks.com/product/machine-learning) (the rebranded MosaicML platform) had to do with how an agent-based approach to AI could improve model quality and reliability. I’m not sure I’ve ever seen a production software system, AI or not, that wasn’t “compound,” so I’m not totally sure why this is a useful distinction…? But if we’re just talking about agents, I’m all in. New features included: - **Mosaic AI Agent Framework:** Allows developers to build their own RAG-based applications on top of Mosaic AI Vector Search, [which went GA last month](https://www.databricks.com/blog/announcing-mosaic-ai-vector-search-general-availability-databricks#:~:text=Vector%20Search%20is%20fully%20integrated,source%20data%20with%20vector%20indexes) - **Mosaic AI Agent Evaluation:** A tool for testing how well AI does in production, which includes some components from the [Lilac acquisition](https://www.databricks.com/blog/lilac-joins-databricks-simplify-unstructured-data-evaluation-generative-ai) earlier this year - **Mosaic AI Model Training:** Fine tuning for open source foundation models - **Mosaic AI Gateway:** A unified access point for LLMs within an application, allowing customers to switch between models without writing new code I found it interesting that both companies made a point to highlight experiences targeted at less technical users, like Snowflake’s AI & ML Studio (where they brought a random audience member onstage to build a chatbot in real time…kinda fun!), and [Databricks’ AI/BI experience](https://www.databricks.com/company/newsroom/press-releases/introducing-databricks-aibi-intelligent-analytics-real-world-data). There are some beliefs that both companies seem to share about AI, and they may be true, but it’s worth at least spelling them out: - “Enterprise data” is a critical component of the AI story. Attending these events both companies want you to believe you’re at the center of the AI revolution. But given how big AI is, I think enterprise data is only a modest part of it. _Will AI impact graphic designers or data analysts more?_ - Data gravity and security / governance will be the biggest differentiator in enterprise AI, and model quality will matter somewhat less. While both DBRX and SNOW both have their own models, they are not fundamentally transforming themselves into LLM training companies. This feels like a solid bet to me, at least for the next couple of years. - People want to ask questions of their data in natural language. Again, this may very well be true, but it is such a widely-assumed belief that it’s worth at least putting it out there that we could all be wrong about natural language as a good way to ask questions of data. Are we not seeing this behavior in volume today simply because the performance isn’t yet good enough? Or do people not actually want this? ## NVIDIA Both companies talked about deepening their partnership with NVIDIA. Jensen showed up as a part of the keynote at both events, wearing his now-emblematic leather jacket. I feel like a major part of Jensen’s job description these days is showing up at others’ conferences and being an AI booster. - [Snowflake highlighted](https://investors.snowflake.com/news/news-details/2024/Snowflake-and-NVIDIA-Power-Customized-AI-Applications-for-Customers-and-Partners/default.aspx) the integration of the Nvidia NeMo Retriever and Inference Server into Cortex AI. I hadn’t heard of [NeMo](https://developer.nvidia.com/nemo-microservices) before this announcement. Looks like it’s a very new NVIDIA product that provides developer tooling, but that it’s pre-release right now. I do not have a personal opinion on whether this is needle-moving for practitioners and the sense I got from others at the event was that there were a lot of folks asking what NeMo was. Looks like we all have some learning to do. - [Databricks highlighted](https://www.databricks.com/company/newsroom/press-releases/databricks-and-nvidia-strengthen-partnership-accelerate-enterprise) the integration of NVIDIA’s Cuda computing platform into the Databricks stack and the availability of the DBRX open source LLM as a NIM microservice. The single line in this announcement that I was most curious about was this: _“Databricks plans to develop native support for NVIDIA-accelerated computing in Databricks’ next-generation vectorized query engine,_ _[Photon](https://www.databricks.com/product/photon), to deliver improved speed and efficiency for customers’ data warehousing and analytics workloads.”_ This is very interesting. I wrote many years ago about whether or not there was potential to accelerate SQL workloads with GPUs, and to-date, that has not been a meaningful thread of innovation in the industry. I’m not sure what the details are behind what inside of Photon is being accelerated with GPUs, but I’m very curious. ## Other announcements **Snowflake** - [Snowflake Native App Framework integration with Snowpark Container Services](https://www.snowflake.com/blog/expanding-range-of-apps-deploy-distribute-on-snowflake/). Over 160 applications were launched on the Snowflake Marketplace, including [dbt for Snowflake](https://app.snowflake.com/marketplace/listing/GZTYZSRT2UA/dbt-labs-dbt). - [More tools for developers](https://www.snowflake.com/blog/simplified-end-to-end-development/) including [Pandas API support for data scientists using Python](https://www.snowflake.com/blog/snowpark-pandas-api-run-at-scale/), [Notebooks](https://docs.snowflake.com/en/user-guide/ui-snowsight/notebooks), an improved [CLI](https://docs.snowflake.com/en/developer-guide/snowflake-cli-v2/index), and an observability suite called [Snowflake Trail](https://www.snowflake.com/en/data-cloud/snowflake-trail/). - Horizon, Snowflake’s data governance offering, [added a private preview of an internal marketplace for data products](https://www.snowflake.com/blog/horizon-leading-governance-data-discovery/), plus some additional privacy and security features. If you want to go deeper on any of these announcements, Snowflake’s [Cameron Wasilewsky](https://www.linkedin.com/in/cameronwasilewsky/) did a fantastic writeup [here](https://medium.com/snowflake/snowflake-summit-2024-highlights-game-changing-feature-announcements-9897002d1de2#4d6f). You can also dive into the Snowflake documentation on new features [here](https://docs.snowflake.com/en/release-notes/2024/june-summit). **Databricks** - [Databricks going 100% serverless](https://x.com/databricks/status/1800924292387107157) on July 1 got a loud ovation from the crowd, aka no more worrying about clusters or what version of Spark you’re running. - [General availability of Predictive Optimization](https://www.databricks.com/blog/announcing-general-availability-predictive-optimization), a capability that optimizes table data layouts for faster queries and improved performance. - [Previewed LakeFlow](https://techcrunch.com/2024/06/12/databricks-launches-lakeflow-for-building-data-pipelines/), a three-part solution for data ingestion (LakeFlow Connect), transformation (Flow Pipelines…from what I can tell this will be [DLT](https://www.databricks.com/product/delta-live-tables) under the hood), and orchestration (LakeFlow Jobs). The solution is rolling out in phases, starting with [LakeFlow Connect](https://www.databricks.com/product/data-ingestion), which will be in preview soon. - [Previewed Databricks AI/BI](https://www.databricks.com/company/newsroom/press-releases/introducing-databricks-aibi-intelligent-analytics-real-world-data), a BI experience that includes a natural language interface, called Genie, to interrogate your data. Under the hood, it uses agents to learn the semantics of your business and update its understanding of metrics on the fly, instead of relying on a static semantic layer. --- --- title: "Data quality metrics: what to pay attention to" description: "Data quality metrics help organizations assess their data. Here’s how to define the metrics that matter to you." url: "https://www.getdbt.com/blog/data-quality-metrics" date: "2024-06-14" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Data quality metrics: what to pay attention to Data quality is the cornerstone of an organization’s data systems. Low-quality data can undermine trust in data, leaving employees feeling lost and rudderless. By contrast, high-quality data fosters trust, encourages quick and confident decision-making, and helps drive new data projects that increase business value. Data quality metrics provide a quantitative measure of data quality, enabling your organization to identify gaps and improve quality over time. However, defining a framework for data quality takes time—and requires changing how your organization works with data on a daily basis. We’ll look at the dimensions of data quality, the best metrics for measuring it, and the tools and processes you can use to measure, improve, and maintain the quality of your data. ## Dimensions of data quality Researchers into data quality define it according to several different factors, or dimensions of data. These can include the following categories and metrics: ### Accuracy A combination of accuracy of the data (i.e., is it free of error? Does it correspond to reality?), believability, and completeness. ### Freshness A measure of data _timeliness_, or whether the data is recent enough for a given business purpose. ### Usability A combination of understandability, interpretability, and accessibility. ### Security and compliance These measures can include whether the data is available only to authorized users, how much data is tagged with sensitivity labels, and whether that tagging is accurate. ## What matters in data quality The above is a brief summary of the dimensions you might utilize when measuring data quality. In addition, your team may face trade-offs when defining key performance indicators for specific metrics. For example, defining a low metric (like minutes or seconds) for timeliness might mean you can’t achieve the results you want for the accuracy or completeness of your data. Or, your attempts to make data more accessible and usable across the organization may conflict with your need to keep certain sensitive data sets secured against unauthorized access. The nature of your data and its business use will drive not just what metrics you need but how much importance you attach to them. For example, [when Rocket Money needed a quote-to-cash system](https://www.getdbt.com/case-studies/rocket-money) to ensure accuracy of their financial data, they placed additional weight on accuracy as a metric (and spent more time in testing to guarantee it). The team had an obligation, not just to ensure the numbers were accurate, but to _prove_ they were accurate. ## Measuring data quality While dimensions help to identify data quality attributes, they’re not metrics. To effectively utilize data quality dimensions as a data quality framework, you first need to identify the associated metrics you want to capture. Here’s an incomplete list of examples. You may use these alongside other metrics to get a full picture of a given data set's quality and overall value to the business: - Metrics related to data incidents such as total number of data incidents, time to detection, time to resolution, table health (number of incidents/table) (Accuracy) - Number of empty/incomplete values, data transformation error rates (Accuracy, Completeness) - Hours since last data refresh, data ingestion delay, tables with the most recent/oldest data, min/max/average data delays (Timeliness) - Number of tests passed/failed over time (Accuracy, Completeness) - Data importance score, number of users of a given table/asset/query, percentage of “dark” or unused data (Accessibility, Believability, Interpretability, Understandability) - Dashboard uptime, table uptime, data storage costs, time to value (Accessibility) - Number of assets tagged for sensitivity levels, number of data-related security incidents (Security) Once you’ve identified the metrics you want to capture, you need some way to hold yourself and the organization accountable for meeting them. This requires defining key performance indicators (KPIs) for your data and then [verifying them through data quality testing](https://www.getdbt.com/blog/data-quality-testing). For example, for an empty/incomplete values test, you could set a threshold at which point a record is considered incomplete—e.g., if 10% of its values are missing. This value may fluctuate depending on your use case. A voluntary customer survey may have 50% or more values missing and still contain worthwhile data. On the other hand, critical financial data could be considered incomplete if even 2% of its total values are missing. ## Data quality tools Implementing a rigorous data quality metrics framework that works requires more than defining metrics and writing tests. It requires both creating a data quality culture and adapting the right technology and tools. A data quality culture at the organizational level ensures that all teams are aligned on the connection between value and quality. This is usually laid out as a set of definitions and principles—e.g., “we prioritize accuracy in order to maintain client trust.” A set of common KPIs, along with associated definitions of data quality, is also crucial to ensure everyone is assessing and measuring quality in the same way. Another component of a data quality culture is making data quality an integral part of everyone’s daily workflow. This can comprise measures such as requiring comprehensive tests for every new data set, setting standards and goals for data classification, providing publicly available data quality dashboards, and streamlining the processing for reporting and resolving data errors. Poor data quality at any point in a data set’s lifecycle can cause downstream failures in reports and data-driven applications. Such issues undermine everyone’s trust in the data. Over time, that decreases its use and its value. A data quality culture emphasizes that everyone who works with data is responsible for ensuring its accuracy and usefulness. Technology and tools are also critical in ensuring data quality. There are three ways in which tools help drive data quality: ### Data catalogs Tools such as a [data catalog](https://docs.getdbt.com/terms/data-catalog) enable everyone in a company to find what they need by serving as a single source of truth for an organization’s data. Data catalogs leverage metadata—such as ownership, date last modified, data types, description, etc. —to help users both find and understand data. ### Data lineage [Data lineage](https://docs.getdbt.com/terms/data-lineage) shows how data moves through an organization. It shows business users where a given piece of data comes from so they can have confidence in the data’s origins and accuracy. It also enables data engineers to find and resolve data quality errors at their source. ### Data testing Data quality tests ensure that extracted and transformed data conforms to business users’ expectations and understanding of how the data should look. A robust data testing framework should run tests automatically every time new data is imported into the system to ensure it’s complete and error-free. ## Data quality with dbt dbt is a data transformation tool that provides a framework for modeling, transforming, and storing data. With dbt, you can define data models that combine data from various sources into useful data products that business stakeholders can use to drive decision-making. At dbt Labs, we’ve long recognized how critical it is to maintain high-quality data sets. That’s why we’ve built multiple mechanisms for verifying quality with every data import. These include: ### Test framework Using dbt, you can define [data tests](https://docs.getdbt.com/docs/build/data-tests) to accompany your models using a combination of SQL and Jinja templating. You can define both single data tests, which test a specific table and set of records; or generic data tests, which implement general checks applicable to multiple data sets (e.g., a not-null test for a column in a table). ### Continuous Integration Preventing errors from creeping into your data requires continuous testing. Using dbt Cloud’s [Continuous Integration (CI)](https://docs.getdbt.com/docs/deploy/continuous-integration) feature, you can configure a migration pipeline that re-tests your data models with every code change. A check-in to source control or a Pull Request (PR) completion triggers a build and runs all of the relevant tests you’ve written automatically in a staging environment. If the staging build completes without error and your tests all pass, you can promote the change to production, running your tests again to verify you haven’t injected any new errors into production data. If you do detect errors, you can roll back your changes to the previous version of your model while you identify the root cause. ### Monitoring and alerting You can use dbt Cloud to [monitor all of your data transformation jobs](https://docs.getdbt.com/docs/deploy/monitor-jobs). You can raise alerts on failure, monitor source freshness, retry failed jobs, and send notifications and alerts on both successful and failed job runs. ## Conclusion The metrics you use to define and measure data quality will be specific to your organization—and will likely even differ between teams and projects. No matter the metrics you select, establishing a data quality culture and selecting the right tools to measure data quality are essential to building a robust, repeatable process that makes data quality an organizational priority. See how tools like dbt Cloud can help you create a data quality culture—[contact us for a demo today](https://www.getdbt.com/contact). --- --- title: "The best Data+AI Summit is in the books!" description: "dbt + Databricks is an amazing combo!" url: "https://www.getdbt.com/blog/the-best-data-ai-summit-yet-is-in-the-books" date: "2024-06-14" authors: ["Jeff Mills"] categories: ["Community"] --- # The best Data+AI Summit is in the books! ## What a week! This week we joined over 16,000 people in San Francisco at the annual Databricks Data+AI Summit. The energy was fantastic, talks were engaging and the announcements just kept on coming. We really enjoyed meeting the thousands of people (including Databricks CEO, Ali Ghodsi) who came by our booth and events to talk about how to democratize data and AI across their organizations - and how data transformations are critical to that effort. Let’s jump right in and relive the magic. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/12b646662bb6b0af1009aaa10f656a20409d20cf-910x777.jpg) ## Minnesota DHS and Slalom ([Eric](https://www.linkedin.com/in/ejmccool/) and [Gabriel](https://www.linkedin.com/in/gabriel-eckers/)) The Minnesota Department of Human Services is over ⅓ of the entire state budget, and they were looking to modernize their data stack and utilize AI to make the best use of state dollars and have the most impact on the people they serve. Eric and Gabriel shared how they are using Databricks and dbt to identify in-home childcare facilities that are non compliant and may result in child harm. It was a super impactful conversation and a model for how more public agencies can use data to improve outcomes. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/e0a4b2170f5d19d4f5e572bce6650b97ada6e6a2-2016x1512.jpg) ## Metadata and AI with dbt and Databricks ([Harsh Jetly](https://www.linkedin.com/in/harshjetly/)) Our very own Harsh Jetly shared with a packed Theater 5 our point of view on the role metadata, dbt and Databricks play in the future of AI. He shared our research and finding from the [State of Analytics Engineering](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024) report, and provided a live demo of dbt on Databricks. The audience interaction and questions were wonderful. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/8c08a045d7e8da010e685e7c67985cd4542786a1-2016x1512.jpg) ## The Data Practitioner for the AI Era At this Summit it became increasingly clear that the data practitioner is going to be critical to most, if not all, AI initiatives. We partnered with MIT Technology Review and Databricks to talk to some of the most cutting edge companies to learn what skills and strategies are needed to be the best Data+AI practitioner. It’s a must read for anyone on the journey of using AI in their enterprise. Get the report [[HERE](https://lnkd.in/dE-i5qyC)](https://lnkd.in/dE-i5qyC) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/b35a44d87d996d9e80aa5269af177fe419a7d148-720x720.jpg) ## Top 10 Data & AI Product on Databricks Over the past few years, Databricks has polled their customer base about what tools they're using to be wildly successful with Databricks, and this year we accelerated our presence inside Databricks accounts and are now the 3rd most used tool in the Databricks ecosystem. Be sure to read all findings and reach out to learn how you could use dbt inside your organization. Check out the news [HERE](https://www.databricks.com/discover/state-of-data-ai). ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/64e29daf25e31723e73e9e76787951dc12c5fe55-1394x1050.png) ## Launch partner to Databricks for their OSS Unity Catalog Data & metadata assets have risen as some of the most critical fuel needed for the AI fire and this week we were proud to be a technology launch partner with Databricks announcing the Open Sourcing of their Unity Catalog. This will make the metrics and metadata you define in dbt Semantic Layer easier than ever to share with Unity Catalog, and visa versa. See how our Semantic Layer works with Unity Catalog [HERE](https://www.getdbt.com/product/semantic-layer) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/7af4025a12224bf80e290bd4f72ad19b2e4cc16d-4080x3072.jpg) ## The momentum continues Databricks will be with us in Las Vegas this year where we’ll announce the latest innovations to power our customer’s success - Don’t miss our annual customer conference [Coalesce](https://coalesce.getdbt.com/). Oct 7-10. You’ll join thousands of data practitioners and leaders to sharpen your Data+AI skills to continue to drive amazing results in your organization. --- --- title: "Data transformation: The foundation of analytics work" description: "Data transformation cleans, standardizes, and automates raw data, ensuring high-quality, trustworthy datasets for analysts." url: "https://www.getdbt.com/blog/analytics-engineering-transformation" date: "2024-06-11" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Data transformation: The foundation of analytics work Data transformation is a key component of the [ETL](https://docs.getdbt.com/terms/etl)/[ELT](https://docs.getdbt.com/terms/elt) process where the “T” represents the data transformation stage, and is typically performed by the analytics engineer on the team or, depending on the organizational structure and needs, data analysts or data engineers. Without data transformation, analysts would be writing ad hoc queries against raw data sources, data engineers will be bogged down in maintaining deeply technical pipelines, and business users will not be able to make data-informed decisions in a scalable way. ![Extract, Load, Transform process diagram](https://cdn.sanity.io/images/wl0ndo6t/main/fbd7272a44ac5eb193cff61045b52d96695b68b3-1706x748.png) As a result, data transformation is at the heart of a business: good transformation creates clear, concise datasets that don’t have to be questioned when used, empowers data analysts to take part in the analytics workflow, and presents a series of challenges that keeps analytics work interesting 😉 ## What is data transformation? Data transformation is the process of taking raw source data and using SQL and Python to clean, join, aggregate, and implement business logic to create important datasets. These end datasets are often exposed in a business intelligence (BI) tool and form the backbone of data-driven business decisions. ## Benefits of data transformation Why is data transformation the foundation for modern data analytics? Because it’s the baseline for increasing the data quality of your business and creating meaningful data for your end users. ### Increases data quality Data transformation can increase data quality through the process of standardization, testing, and automation. During the transformation process, raw data is cleaned, casted, converted, joined, and aggregated using SQL and Python to create end datasets that are consumed by business users. In an ideal world, these transformations are version-controlled and peer-reviewed. This transformation process should also follow automated testing practices, ultimately creating tables data analysts and end business users can trust. By transforming your data with tooling that supports standardization, version control, integrated documentation, modularity, and testing, you leave little room for error. Data analysts don’t need to remember which dataset is in which time zone or currency; they know the data is high-quality because of the standardization that has taken place. In addition to standardizing raw data sources, metrics can be properly created during the transformation process. dbt supports the creation of [metrics](https://docs.getdbt.com/docs/build/metrics) and exposing them via the [Semantic Layer](https://www.getdbt.com/product/semantic-layer/), ultimately allowing you to create and apply the same metric calculation across different models, datasets, and BI tools, ensuring consistency across your stack. As you develop consistent metric definitions, your data quality increases, trust in your data work increases, and the ROI of a data team becomes much more apparent. ### Creates reusable, complex datasets Data transformation allows you to automate various data cleaning and metric calculations. This ensures consistent, accurate, and meaningful datasets are being generated in the data warehouse each day, or on whatever time cadence your business chooses. By automating certain data models, data analysts do not need to repeat the same calculations over and over again within the BI layer. These data sets can be referenced directly within a report or dashboard instead, speeding up compute time. Data transformation also activates the [reverse ETL process](https://docs.getdbt.com/terms/reverse-etl/). Transformation allows analytics engineers to join different datasets into one data model, providing all the needed data in one dataset. Because datasets are being automated using data transformation, this data can be ingested into different reverse ETL tools, giving stakeholders the data they need, where and when they need it. ## Challenges of data transformation Data transformation is fun, but tough work for analytics practitioners. The difficulty often varies given the complexity and volume of your data, the number of sources you’re pulling from, and the needs of your stakeholders. Some of the biggest challenges you’ll face during the [data transformation process](https://www.getdbt.com/blog/data-transformation-process/) are: creating consistency, standardizing core metrics, and defining your data modeling conventions. ### Consistency across multiple datasets During the transformation process, it can be challenging to ensure your datasets are being built with standardized naming conventions, following SQL best practices, and conforming to consistent testing standards. You may often find yourself checking if time zones are the same across tables, whether primary keys are named in a consistent format, and if there’s duplicative work across your data transformations. How you determine what consistency and standardization look like in your data transformation process is unique to your team and organization. However, we recommend using a tool, such as dbt, that encourages data transformation DRYness and modularity, code-based and automatic tests for key columns, and explorable documentation to help you keep consistent and governable data pipelines. Here are some other dimensions you should keep in mind when trying to create consistency in your data: - What timezone are your dates and timestamps in? - Are similar values the same data type? - Are all numeric values rounded to the same number of decimal points? - Are your column names named using the same format? - Are all primary keys being regularly tested for uniqueness and non-nullness? These are all different factors to consider when creating consistent datasets. Doing so in the transformation stages will ensure analysts are creating accurate dashboards and reports for stakeholders, and analytics practitioners can more easily understand the requirements to contribute to future transformations. ### Defining data modeling conventions Defining data modeling conventions is a must when utilizing data transformation within your business. One of the reasons data transformation is so powerful is because of its potential to create consistent, standardized data. However, if you have multiple analytics engineers or data analysts working on your data models, this can prove difficult. In order to create high-quality, valuable datasets your data team must decide on style conventions to follow before the transformation process begins. If proper style guidelines are not in place, you may end up with various datasets all following different standards. The goal is to ensure your data is consistent across all datasets, not just across one engineer’s code. We recommend creating a style guide before jumping into the code. This way you can write all of your standards for timezones, data types, column naming, and code comments ahead of time. This will allow your team to create more consistent, scalable, and readable data transformations, ultimately lowering the barrier for contribution to your analytics work. ### Standardization of core KPIs A lack of consistency in key metrics across your business is one of the largest pain points felt by data teams and organizations. Core organizational metrics should be version-controlled, defined in code, have identifiable lineage, and be accessible in the tools business users actually use. Metrics should sit within the transformation layer, abstracting out the possibility of business users writing inaccurate queries or conducting incorrect filtering in their BI tools. When you use modern data transformation techniques and tools, such as dbt, that help you standardize the upstream datasets for these key KPIs and create consistent metrics in a version-controlled setting, you create data that is truly governable and auditable. There is no longer a world where your CFO and head of accounting are pulling different numbers: [there is only one world where one singular metric definition is exposed to downstream users](https://www.getdbt.com/product/semantic-layer/). The time, energy, and cost benefit savings that comes from a standardized system like this are almost incalculable. ## Data transformation tools Just like any other part of the modern data stack, there are different data transformation tools depending on different factors like budget, resources, organization structure, and specific use cases. Below are some considerations to keep in mind when looking for a data transformation tool. ### Enable engineering best practices One of the greatest developments in recent years in the analytics space has been the emphasis on bringing software engineering best practices to analytics work. But what does that really mean? This means that data transformation tools should conform to the practices that allow software engineers to ship faster and more reliable code—practices such as version control, automatic testing, robust documentation, and collaborative working spaces. ### Version Control You should consider whether or not a data transformation tool offers version control, and supports the basic Git flow, so that you can keep track of transformation code changes over time. A data transformation tool should support the following seven steps of a basic Git flow: 1. Clone the original codebase 2. Create your development branch 3. Stage your updates to files 4. Commit your staged changes to the local repository 5. Push your changes to the remote repo + open a pull request 6. Merge your changes with the master codebase. 7. Pull down a new clone of the main repo ### Transformations-as-code Your data transformation tool should also support transformations-as-code, allowing anyone who knows SQL can partake in the data transformation process. ### Build vs buy Like all internal tooling, there will come a time and place when your team needs to determine whether to build or buy the software and tooling your team needs to succeed. When considering building your own tool, it’s vital to look at your budget and available resources: - Is it cheaper to pay for an externally managed tool or hire data engineers to do so in-house? - Do you have enough engineers to dedicate the time to building this tool? - What do the maintenance costs and times look like for a home-grown tool? - How easily can you hire for skills required to build and maintain your tool? - What is the lift required by non-technical users to contribute to your analytics pipelines and work? Factors such as company size, technical ability, and available resources and staffing will all impact this decision, but if you do come to the conclusion that an external tool will be appropriate for your team, it’s important to break down the difference in open source and SaaS offerings. ### Open source vs SaaS If you decide to use an externally created data transformation tool, you’ll need to decide whether you want to use an open source tool or SaaS offering. For highly technical teams, open source can be budget-friendly, yet will require more maintenance and skilled technical team members. SaaS tools have dedicated infrastructures, resources, and support members to help you set up the tool, integrate it into your already-existing stack, and scale your analytics efficiently. Whether you choose open source or SaaS will again depend on your specific budget and resources available to you: - Do you have the time to integrate and maintain an open source offering? - What does your budget look like to work with a SaaS provider? - Do you want to depend on someone else to debug errors in your system? - What is the technical savviness of your team and end business users? _dbt offers two primary options for data transformation: dbt Core, an open source Python library to help you develop your transformations-as-code using SQL and the command line. dbt Cloud, the SaaS offering of dbt Core, includes an integrated development environment (IDE), orchestrator, hosted documentation site, CI/CD capabilities, and more for your transformations defined in dbt. Learn more about the three flexible [dbt Cloud pricing options here](https://www.getdbt.com/pricing/)._ ### Data documentation in data catalogs Analytics engineers are data librarians, and documentation is our Dewey decimal system to catalog information. Someone may not know precisely where to begin on a project, so they’ll ask the analytics engineer, aka the librarian, to point them in the right direction. And at other times, they’ll know exactly what they’re looking for and delve directly into the Dewey decimal card catalog (aka the documentation) themselves. Either way, both the analytics engineer and the documentation are here to help. Relying 100% on an analytics engineer is inefficient, yet relying entirely on documentation lacks comprehensiveness. Having both empowers people to find the information they need in the most direct, efficient way possible. The argument for documentation revolves around four crucial points we hold dear: 1. **People only use code they trust:** Testing and documentation provide the coverage the code needs to gain others’ trust. Those who can understand your code and view the tests performed will use it. Code that has no documentation will never be used, resulting in wasted time and wasted code. 2. **Reliable documentation is key to scaling a data team: **Unrecorded experiential knowledge and outdated documentation are a recipe for an ineffective team, where individuals ask the same questions repeatedly and waste each other’s time. Documented data takes the knowledge out of people’s heads and arranges it so everyone has access. No matter if you’re working in a small team or as part of a team of 40, you’ll all be on the same page. 3. **Automation is key:** One of the reasons data documentation is so often unsuccessful is because it relies on someone manually entering explanations and notes — possibly one of the most boring things in the world to do. dbt automates most documentation by making inferences about your codebase, thus taking away the uncomfortable task of manually maintaining a Google sheet or Confluence article. There is still a need for some manual input, especially, for example, when you just can’t convey everything one needs to know about a particular column in the column title itself. If you’ve ever looked in Snowflake explorer and not understood what a column means, you know why manual documentation is still necessary on some level. 4. **The documentation process shouldn’t be a game of catch-up:** Imagine an inchworm making its way along a leaf — just as its back-end catches up, its front end moves forward a little more, so both parts of the body are moving, but the tail is always one step behind. Documentation should always step in tandem with the work; it must be a part of the same workflow as the transformations it seeks to describe. It can’t if the two are separate; the documentation will always lag behind. ### Technical ramp period Last but definitely not least, you must weigh the technical learning and adoption curves that come with choosing a data transformation tool. You need to ask yourself if those on your data team have the technical expertise to use whatever tool you choose to implement. Can they code in the language the tool uses? If not, how long will it take them to learn? If it’s an open source tool you decide on you may need to consider whether or not your data team is familiar with hosting that tool on their own infrastructure. It’s also important to note the lift required by your end business users: will they have difficulty understanding how your data pipelines work? Is transformation documentation accessible, understandable, and easily explorable by business users? For folks who likely only know baseline SQL, what is the barrier to contributing? After all, the countless hours of time and energy spent by data practitioners are to help empower their business users to make the most informed decisions they can using the data and infrastructure they maintain. ## Conclusion Data transformation is a fundamental part of the ETL/ELT process within the modern data stack. It allows you to take your raw source data and find meaning in it for your end business users; this transformation often takes the form of modular data modeling techniques that encourage standardization, governance, and testing. When you utilize modern data transformation tools and practices, you produce higher-quality data and reusable datasets that will help propel your team and business forward. While there are very real challenges, the benefits of following modern data transformation practices far outweigh the hurdles you will jump. --- --- title: "Five real data transformation examples" description: "Discover five real data transformation examples and learn how dbt Cloud enhances data workflows for better business insights." url: "https://www.getdbt.com/blog/data-transformation-examples" date: "2024-06-11" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Five real data transformation examples Businesses today are swimming in data—whether it’s from customer interactions, transactions, or digital behaviors. But raw data by itself isn’t useful until it’s organized and transformed into something actionable. That’s where [data transformation](https://www.getdbt.com/blog/what-is-data-transformation) comes in. By converting data into a structured, consistent format, companies can unlock insights that drive decision-making, reporting, and analysis. Data transformation is key to getting the most value out of your data. Whether it’s improving customer engagement, optimizing internal processes, or even detecting fraud, transforming data into usable information helps businesses move forward. In this article, we’ll break down what data transformation really means, explore the tools that make it happen, and dive into five examples that show its impact in the real world. Plus, we’ll see how [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) makes the whole process smoother and more scalable. **** ## What is data transformation? [Data transformation](https://www.getdbt.com/blog/what-is-data-transformation) is the process of converting raw data into a more meaningful format to meet the needs of business operations, analytics, and other processes. As the volume and variety of data sources grow, organizations need efficient methods to integrate and manage this data in a way that allows them to gain actionable insights. Data transformation involves several steps, including cleansing, filtering, and enriching data. It’s essential for converting data from different systems and formats into a unified model, ensuring consistency, accuracy, and usability. This process often involves moving data into a centralized location, like a data warehouse, where it can be used for business intelligence and decision-making. Transforming raw data into an actionable asset is a critical component of modern data strategies. ## The data transformation tech stack Before diving into examples, it’s important to understand the common components of a modern data transformation stack. Below are some of the key tools and technologies involved in a typical data transformation process: - **[ETL (Extract, Transform, Load)](https://www.getdbt.com/blog/extract-transform-load) tools**: ETL tools like [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) help extract raw data from multiple sources, transform it into the desired structure, and load it into a destination like a data warehouse or data lake. - **Data warehouses**: Platforms like Snowflake, Google BigQuery, and Amazon Redshift are often the final destination for transformed data, where it’s stored and made available for analysis. - **Data visualization tools**: After data has been transformed, tools like Looker, Tableau, or Power BI help translate it into visual dashboards that make data insights more accessible to decision-makers. - **Data governance and quality**: Tools like [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) enhance the governance process by ensuring that transformed data is consistent, well-documented, and maintained with version control. This stack can vary depending on the organization's infrastructure and goals, but the end result is the same: ensuring that data is in the right format to provide valuable insights. ## Five data transformation examples Understanding how businesses transform their data is essential for illustrating the importance of this process. Below, we look at five real-world examples of how companies have successfully used data transformation to drive efficiency, profitability, and customer satisfaction. ### 1. Transforming e-commerce data for better customer insights An e-commerce company managing multiple product lines across various regions has to deal with vast amounts of transactional and customer behavior data. Initially, the data might be scattered across different platforms—point-of-sale systems, customer relationship management (CRM) tools, and third-party logistics platforms. This raw data needs to be aggregated and standardized before it can offer value. Using an ETL tool, the company can extract data from these disparate sources and standardize it in a central data warehouse like Google BigQuery. During transformation, the company might clean up customer data, removing duplicates and fixing inconsistencies, before enriching the data with additional attributes like customer lifetime value or purchase frequency. The transformed data can then be used to create personalized marketing campaigns and offer more relevant recommendations, leading to improved customer engagement and higher sales. ### 2. Data transformation in financial services for fraud detection A global financial services firm processes millions of transactions daily. This vast amount of data contains valuable insights but also creates an environment ripe for fraudulent activity. To detect anomalies and potential fraud in real-time, the company must transform raw transactional data into a format that can be analyzed by machine learning models. By using an automated data transformation tool like dbt, the company extracts transaction data from multiple systems, normalizes it, and enriches it with contextual information such as the location of the transaction, time of day, and customer profile. This transformed data is fed into a machine learning model, which then flags unusual patterns for further investigation. As a result, the company reduces fraud-related losses while maintaining a seamless customer experience. ### 3. Marketing analytics at scale for a global retailer A global retail chain with both online and offline channels collects large quantities of data on customer behavior, sales, and inventory. For their marketing team to create effective campaigns, they need to merge and analyze this data in real time. However, data from different departments (such as sales, customer service, and marketing) exist in silos, complicating the process. Using dbt Cloud, the retailer extracts this data and transforms it by aligning customer IDs, merging transaction histories, and ensuring that inventory levels are accurately reflected in marketing offers. This unified view enables marketers to run highly targeted campaigns, measure their success, and quickly pivot based on data insights. In addition, having a centralized, transformed data source allows the marketing team to run advanced analyses, such as customer segmentation and churn prediction. ### 4. Healthcare analytics for patient outcomes In the healthcare industry, data is gathered from numerous sources like electronic health records (EHRs), laboratory results, and billing systems. For a healthcare provider, transforming this data into a unified format allows for better patient care and operational efficiency. For example, a large hospital chain uses dbt to transform patient data from these multiple systems into a comprehensive health profile. The transformed data includes patients’ treatment history, medication prescriptions, and diagnostic results, which are used to predict patient outcomes and make personalized treatment recommendations. This not only improves patient care but also helps optimize hospital resources by predicting admission rates and staffing needs. ### 5. Real-time analytics for a media streaming service A media streaming company tracks real-time user engagement metrics, such as the number of active users, content consumption patterns, and subscriber growth. To offer personalized content recommendations and ensure the best possible user experience, the company must constantly transform raw event data into a format that can be used for analytics and machine learning models. With the help of a data transformation tool like dbt Cloud, the company can rapidly transform raw data into a structured dataset that captures key metrics, such as user preferences, viewing times, and content popularity. This transformed data feeds directly into recommendation engines, helping the service retain users and increase overall engagement by providing personalized content suggestions. ### How dbt Cloud enhances the data transformation workflow The [data transformation](https://www.getdbt.com/blog/what-is-data-transformation) process is complex, but tools like [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) make it significantly more efficient and scalable. Here are some key ways dbt Cloud adds value to the transformation process: ### Version control and collaboration dbt Cloud provides built-in version control, allowing data teams to collaborate on transformation projects. This feature ensures that all transformations are well-documented, auditable, and easy to roll back if necessary. This enhances the governance process, reducing errors and ensuring that data is always trustworthy. ### Automated testing and deployment One of the most significant benefits of dbt Cloud is its automated testing and deployment features. Before deploying a transformation to production, dbt Cloud can automatically test the data model, ensuring that the transformations are accurate and adhere to business rules. This minimizes the risk of introducing errors into the data pipeline and speeds up the overall transformation process. ### Scalability As businesses grow, so do their data needs. dbt Cloud supports scalable data transformation workflows, meaning it can handle large volumes of data from different sources, whether you're running transformations in a small startup or a global enterprise. With its cloud-based infrastructure, dbt Cloud also eliminates the need for businesses to manage their own servers or infrastructure, reducing the complexity of scaling data operations. ### Integration with modern data warehouses dbt Cloud integrates seamlessly with modern data warehouses like Snowflake, BigQuery, and Redshift, making it easier for organizations to load and transform data in a central location. This integration ensures that businesses can run transformations in real time, driving timely insights and decision-making. ## Conclusion Data transformation is an essential process for organizations looking to unlock the full potential of their data. From improving customer insights to optimizing operational efficiency, transforming raw data into actionable insights provides measurable value across industries. As seen in the examples above, businesses are leveraging transformation tools like dbt Cloud to streamline and enhance their data workflows. By automating key steps like testing, version control, and deployment, dbt Cloud enables businesses to scale their data transformation processes, ensuring they remain competitive in an increasingly data-driven world. If you want to see how dbt Cloud can enhance your data transformation process, [book a demo with a dbt expert](https://www.getdbt.com/contact) or [create a free dbt Cloud account](https://www.getdbt.com/signup). --- --- title: "The data practitioner for the AI era" description: "The dbt perspective on AI from MIT Technology Review's report on data practitioners." url: "https://www.getdbt.com/blog/the-data-practitioner-for-the-ai-era" date: "2024-06-11" authors: ["Drew Banin"] categories: ["Insights"] --- # The data practitioner for the AI era _The following is an _excerpt_ from MIT Technology Review's "The Data Practitioner for the AI Era," a report co-sponsored by dbt Labs and Databricks. [**Read the full report here**](https://www.databricks.com/resources/whitepaper/databricks-and-dbt-labs-data-practitioner-ai-era?utm_source=mitpr&utm_medium=partner&utm_campaign=701vp000003p66xiaq)._ In today’s rapidly evolving business landscape, the integration of artificial intelligence (AI) into operations is no longer a luxury but an outright necessity. Companies worldwide are increasingly adopting AI initiatives to drive innovation, improve decision-making, enhance customer experiences, and deliver a competitive edge. But for all its promise and investment, AI adoption is fraught with challenges, particularly in managing and leveraging the vast amounts of data that fuel these initiatives. The enthusiasm for incorporating AI into business operations is often dampened by the complexity of managing the underlying data infrastructure. A common challenge is the integration of large language models (LLMs) with cloud data platforms. Without high-quality, continuously updated data connected to clear semantic definitions, businesses risk exposing inaccurate or outdated information to their LLM applications. Further, strong data governance practices are required to ensure that LLMs operate on data appropriately and that sensitive or regulated data is not misused or misconstrued by LLMs. Ultimately, LLMs are great at internalizing mountains of data and spitting out answers; if we prompt them with rich, high-quality inputs, we can expect to see high-quality outputs. Likewise, if we prompt with low-quality or non-compliant inputs, they may jeopardize business integrity and erode customer trust by outputting incorrect, invalid, or otherwise inappropriate outputs. [dbt Cloud is a powerful ally in this context, providing a suite of tools designed to maintain the integrity and trustworthiness of data powering AI initiatives.](https://www.getdbt.com/product/ai) ## Semantic layer implementation [The creation of a semantic layer](https://www.getdbt.com/product/semantic-layer), which maps data to business concepts, ensures that the data exposed to AI models is accurate, relevant, and consistent. A semantic layer can significantly reduce the risk of hallucinations and inaccuracies. ## Metadata framework [A built-in metadata framework](https://www.getdbt.com/product/dbt-cloud) enables data to be enriched with a wealth of context and meaning, and it magnifies AI’s ability to yield reliable answers to critical business questions. ## Data quality with contracts Data contracts enforce clear definitions of data quality, structure, and relationships across teams. This ensures that only compliant data feeds into AI projects, even if that data crosses team boundaries, data stores, or domains. ## Version control and testing With built-in version control and testing capabilities, teams can track changes, test data models rigorously, and ensure that the data infrastructure remains stable and reliable. ## Alerting and continuous integration Real-time alerting mechanisms to flag data quality issues, paired with continuous integration processes to catch issues before they hit production, ensure that data quality or model performance problems are promptly identified and addressed, preventing potential setbacks in AI initiatives. ## Accelerating AI projects AI initiatives can only move as fast as the development of the data that underlies them. [dbt Cloud fortifies the data foundation of AI projects and accelerates development and deployment by streamlining data transformation and modeling processes](https://www.getdbt.com/product/ai). At dbt Labs, our mission is to empower data practitioners to safely create and disseminate organizational knowledge. These data practitioners will shape how AI is deployed in the enterprise and drive the strategy that leads to higher-quality results and designing data workflows that reduce risks in AI implementations. For them to be successful, we believe practitioners must adopt a structured and reliable approach to data management. By ensuring data integrity, fostering trust, and facilitating rapid development, dbt Cloud empowers businesses and data practitioners to leverage AI with confidence on top of their cloud data platforms. --- --- title: "New dbt Labs and Databricks Report Highlights The Evolving Roles of Data Practitioners in the AI Era" description: "dbt Labs and Databricks release report on the strategies and processes needed to effectively deploy AI in the enterprise" url: "https://www.getdbt.com/blog/new-dbt-labs-and-databricks-report-highlights-the-evolving-roles-of-data-practitioners-in-the-ai" date: "2024-06-10" authors: ["Alyssa Smrekar"] categories: ["Press"] --- # New dbt Labs and Databricks Report Highlights The Evolving Roles of Data Practitioners in the AI Era **New dbt Labs and Databricks Report Highlights The Evolving Roles of Data Practitioners in the AI Era** _Organizations in need of better strategies and processes to improve data quality and AI outputs_ **PHILADELPHIA – **June 10, 2024 – [dbt Labs](https://www.getdbt.com/), the pioneer in analytics engineering, and [Databricks](https://www.databricks.com/), the Data and AI company, today unveiled “The Data Practitioner for the AI Era” report. Produced by [MIT Technology Review](https://www.technologyreview.com/) and revealed at the [Databricks Data + AI Summit](https://www.databricks.com/dataaisummit), the findings shed new light on the critical role data practitioners play in managing data pipelines and processes, supporting business strategy and operations, and improving the quality of AI outputs. “Data is the foundation of large language models and therefore data professionals have the power to make or break AI applications,” said Roger Murff, VP of Technology Partners at Databricks. "The role of data teams – and the groundwork they set – is foundational to building AI models. That’s why it’s important that we solve the challenges posed to data teams so they can focus on creating reliable, transparent, and high-quality data." Key takeaways from the “The Data Practitioner for the AI Era” report include: - **Data jobs are quickly evolving **as practitioners transition from traditional, siloed roles to more integrated, strategic positions within an organization. Data teams are integrating data intelligence platforms that enhance data visibility and transparency, bridging the gap between technical and business units. - **Data is the cornerstone of AI strategy **and organizations are increasingly investing in it to remain competitive. Data practitioners are central to this effort, responsible for ensuring data quality, governance, and the effective deployment of AI technologies. - **The concept of “Data as a Product"** is emerging as it requires practitioners to adopt a product management mindset, prioritizing tasks that have significant business impact. The result is better data practices and teams deeply embedded into business strategy and operations. - **Data challenges remain before organizations can realize their AI aspirations**, including data silos, processing speed, data sufficiency, and monitoring lineage. Companies need to have a unified way to query and understand all their data across these silos. - **Retrieval-augmented generation (RAG) is a key technique** to improve data quality** **for AI outputs, as organizations seek to mitigate issues like bias and hallucinations. By grounding AI models with verifiable external knowledge sources, organizations can ensure the production of higher-quality, more specialized, and verifiable AI outputs. “AI is a gamechanger, and data roles, responsibilities, and workflows are transforming as it takes shape and becomes embedded in how teams work,” said Drew Banin, Cofounder of dbt Labs. "Analytics engineers for example, a new breed of data practitioners, are bridging the gap between data and business needs in a way that we haven’t seen before. They play a vital role in translating business requirements into effective data transformations, making data more accessible, actionable, and essential to business strategy and operations." To learn more, download the [“The Data Practitioner for the AI Era” report](https://www.databricks.com/resources/whitepaper/databricks-and-dbt-labs-data-practitioner-ai-era?utm_source=mitpr&utm_medium=partner&utm_campaign=701vp000003p66xiaq). **About dbt Labs** Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 40,000 companies using dbt every week. To learn more about dbt Labs, visit [getdbt.com](https://www.getdbt.com/) and follow us on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/). **Press Contact** Method Communications for dbt Labs [dbtlabs@methodcommunications.com](mailto:dbtlabs@methodcommunications.com) --- --- title: "The rapid experimentation of AI agents" description: "Yohei Nakajima, general partner at Untapped Capital and creator of BabyAGI, on AI agents, and where they might take us." url: "https://www.getdbt.com/blog/rapid-experimentation-ai-agents" date: "2024-06-09" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # The rapid experimentation of AI agents _This post first appeared in_ _[The Analytics Engineering Roundup](https://roundup.getdbt.com/p/the-rapid-experimentation-of-ai-agents)._ Yohei Nakajima is an investor by day and coder by night. In his day job, he invests in early stage companies as a general partner at [Untapped Capital](https://www.untapped.vc/). As a hacker, he's been focused of late on the applications of AI. In particular, one of his projects, called [BabyAGI](https://babyagi.org/), got a ton of attention about a year ago. BabyAGI is an AI agent framework, creating a plan-execute loop. If you give it a goal, it will create a plan to achieve this goal, and then go execute on this plan. All of this plays out in a long chain of LLM API calls, and you can observe every step along the way. The truth is that this is an extremely experimental space, and depending on how strict you want to be with your definition, there aren't a lot of production use cases to point to today of AI agents. When you watch the demo videos on Yohei's Twitter feed, you can immediately see the promise. **Listen & subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Google Podcasts](https://podcasts.google.com/feed/aHR0cHM6Ly9hbmFseXRpY3NlbmdpbmVlcmluZ3JvdW5kdXAubGlic3luLmNvbS9yc3M) - [Stitcher](https://www.stitcher.com/show/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) ## Key takeaways from this episode. ### What do you spend your time on these days? **Yohei Nakajima:** I'm a VC by day, builder by night. It's always been a hobby I've always had on my side. I've never been an engineer by trade, but I've been coding as a hobby since high school. And I became more of a no-coder for a while, just because with the limited time to build, no-code was easier. But when I started using AI, I realized that I could pump out some pretty cool code in a matter of two to three hours, so I switched back to code about a year after I started using OpenAI. That's where all my GitHub projects come from. ### Do you want to take a minute to plug a company or two that you've invested in recently, for the audience to see what you're seeing on the VC side? One of the tools that I use every day, which is the easiest plug, is a company called [Wokelo.ai](https://www.wokelo.ai/). They do AI research and due diligence. When I'm interested in a company, I can just type in the company name and click Generate Report. Thirty minutes later, I get a 20-50 page report that includes market insights, recent news, competitors, management profiles, all that kind of stuff. It really is the kind of work you’d ask analysts to do and it’d take them two weeks to do, but it's in 30 minutes. ### Based on the deal flow that you're seeing, what are some underreported things that will be true about the world in two or three years that no one's paying attention to? I've done a couple of talks on autonomous agents, and one of the things that gets more of a reaction is the idea of an agent who is a CEO. When people think about AI right now, you imagine AI at the bottom of the rung. But an example I give is Amazon Mechanical Turk or Upwork. These are technology platforms that help you manage people. And now imagine taking this capability of LLMs and rebuilding something like an Amazon Mechanical Turk. You can imagine how you could build an agent that can manage a college ambassador program and that AI can do that more efficiently than a human, in theory. If you extend that, then why can't you have a CEO who's available 24-7, has access to every single piece of data in the company, can take feedback from every single employee and synthesize it, and has bias that's at least measurable and transparent, so you can adjust it, versus wondering if they're being unfair. ### One of our values as a company is we move up the stack. It's an invitation to always replace yourself. And if there is a very effective AI CEO, I'm ready. The initial idea around BabyAGI was to prototype an autonomous startup founder. I think AI is far from being able to build an innovative startup that becomes a unicorn. But if we pick something that's a little bit more straightforward and purely digital, like an e-commerce dropshipping business, I think it's within reach to build an AI system that can run and manage something like that. ### Tell us about BabyAGI. What is it? And what have people built with it? I think the reason BabyAGI went viral was at that point, ChatGPT was still relatively new. People were still building on top of that chat interface. BabyAGI was one of the first popular open-source projects to loop in LLM. I added some capabilities so when you're given an objective, I’d first have a task list creation agent to generate a task list based on that objective. And then I would use code to parse that out. And then send each task one by one to an execution agent to execute that. When I was generating new tasks, I had a task prioritization agent, which would review past results and update the task list. It would then check for the most similar tasks first, so that it would try to generate new kind of tasks. And what happens when you press run, build a business, it would just come up with things to do. At this point it was just LLM calls; it was just generating tasks. But when I asked it to start a business—I need a marketing plan, I need to build a website, I need to come up with product ideas—it just kept going one by one on, like we do around building a business. ### This has become commonly referred to as an agent, this idea that there's an AI that operates in a loop and it plans and it executes, and it can to a certain extent take actions in the real world. I think this is not a new idea, but it is an idea that maybe it's time has come. Is that a fair way to think about it? I agree. It's not a new idea to let a software program autonomously run. I think with LLM capabilities you could finally get it to reason, do similarity search, and do it in a way that was much more robust. ### Can you talk at all about agent architectures? Are there better and worse ways to construct agents? When new technology emerges, there's a period of rapid experimentation. If you look at old cars, there are three-wheeled cars, there are steam-engine cars. But if you look today, a lot of the cars or phones, they all look the same today. With autonomous agents, we're in the rapid experimentation phase, where there's a whole bunch of different frameworks, a whole bunch of different architectures, a whole bunch of different approaches. It's hard to say which one's going to stick, but ultimately there’ll be some consolidation. So I think it helps if you're thinking about architecture, just look at them each one by one. Task planning is probably one of the main ones. The two major ways is the react style, which is to do one thing at a time and reflect on it. And then there's the BabyAGI style—plan and execute—which is generate a task list first and then go through them. But of course there's a fuzzy line because you can generate a task list first and then reflect on each result to update the task list. It's not one or the other. ### How should we think about where we will see agents impacting our lives first? I think about it in three buckets. I think of the first bucket as what I call handcrafted agents, where you are writing each problem. And you're chaining it specifically with API calls and it's a very specific flow. Some people would call that not autonomous because it's a human generating the task list and you're not AI running it autonomously. That being said, with the idea of Wokelo, if I can give an AI a company and it's going to give me back a 30-50 page report, from a usage standpoint, it feels very autonomous to me. And then there's what I call the specialized agent, which is the I think where it becomes truly autonomous, it's dynamically generating its own task list or recursively figuring out what to do, but it's within a specific set of skills and tools. So you can imagine a coding-specialized agent for a VC specialization that just knows how to look up CrunchBase. And then there's the general-purpose, fully generalized autonomous agent. Handcrafted agents are useful today. I mean, Wokelo is one specific example, but I know many companies that are doing handcrafted agents. This is humans generating the task list and they're charging lots of money with low churn, and they're creating value. With specialized agents, what I'm seeing right now is pretty interesting demos with a lot of promise. These companies are raising capital; they're starting to have conversations with enterprises because they're interested, and they are starting to line up pilots because these companies are willing to test it. But I haven't seen anything that's blowing me away in terms of reliability and value creation. And then when it comes to general automation, I haven't seen anything remotely close to reliable. I also don't think of them as three fully separate buckets, but actually more of starting here and slowly moving to there. ### Can you talk at all about why agent designers switch between different models for different tasks? Different models have different strengths is the shortest answer. Some models are better than others at writing code. Some models are better than others at writing long pieces in an eloquent manner. And it's constantly shifting. Developers are looking at the cost, the speed, and the quality of each model, and when I say quality, specific to the use case. So with an agent system, you have parts that need to write code, parts that need to respond to the user, so depending on the specific need you're swapping out models to see which one is optimal. ### Now that language models can write code, can formulate their outputs and inputs as JSON or other structured data, is there any particular limit on what types of actions they could take in the world? The first thing that popped to mind was anything that requires a physical body will be hard for a digital agent to do unless you give them the robotic parts to do so, but again, I wouldn't say that's a limit, like we can build the robotic parts and we can give them division capabilities and we can do those things. ### What's so interesting to me is messaging. Are there agents in the world today that are connected to the Twilio API or to SMTP and sending emails? Yeah, Definitely. There are tons of emails that are automatically sent. But I think the question you're asking is, are there any that are fully dynamically managing it on its own? Probably. ### Do you think that most AI workflows are going to be created by software engineers? Or because the capabilities of AI are so powerful in the realm of writing code, are they actually going to be created by less technical folks who are closer to the business? This is a little bit more of a hypothesis. I don't know is the short answer, but I would guess, based on my experience building it, that you need a really good core framework, and that has to be built by the engineers. And the whole point of the framework should really be around a good agent should be one that the more you use it, the better it gets. In the future, when you say, who's building the task list, I don't think it's the engineers. It's the users who are going to ask an AI to do something, and then when an AI doesn't do it the right way, it'll give it guidance, and the AI will remember how to do it that way and keep you getting better. The engineers will be building essentially how the brain works itself, but just like how you and I have gotten better at things, the AI is going to get better by working for somebody and getting feedback from that person. ### What do you hope is true about the data or in this case, the AI space in five years? I hope that we’ll see more use cases of AI with the goal of helping people better understand each other. [I actually did a TED talk on it](https://www.ted.com/talks/yohei_nakajima_how_ai_will_help_us_connect_with_ourselves_and_each_other), the idea that AI can help us better understand ourselves and each other, both through the usage and building of it. --- --- title: "Snowflake Data Cloud Summit 24' is a wrap!" description: "Snowflake Data Cloud Summit 24' was one for the books. See what we learned!" url: "https://www.getdbt.com/blog/snowflake-summit-24-is-a-wrap" date: "2024-06-07" authors: ["Jeff Mills"] categories: ["Company"] --- # Snowflake Data Cloud Summit 24' is a wrap! Phew - what a week! School may be out for the summer, but the learnings continued in SF at Snowflake Summit. While AI took center stage we learned through all the conversations on the show floor that everyone is at a different stage of their data journey. And no matter where people are in their journey - they are looking for ways to transform their data for the step that they’re on. We held thousands of conversations this week and that theme was in every one of them. ![Team dbt ](https://cdn.sanity.io/images/wl0ndo6t/main/2512a2e1078a545a8b1417862ea6d53619ad1b2b-5712x4284.jpg) ## Our customers taught us all how to dbt **Techstyle ([Yigit](https://www.linkedin.com/in/yyoruk/) and [Rachana](https://www.linkedin.com/in/rachana-mukherjee-6190555/))** shared how they constructed their modern data infrastructure, how dbt Cloud was part of the process and critical to their success. They are the team behind some of today’s hottest brands and they needed to move quickly. They migrated from Airflow to dbt Cloud - helping with their remaining scaling challenges by implementing dbt Mesh and data contracts. This led to a faster data products development cycle and greater alignment with their business. ![Techstyle](https://cdn.sanity.io/images/wl0ndo6t/main/3eee8ecfc043f146139956c9ce8bd1c840fb1a19-2016x1512.jpg) **Ally and Podium ([Prasanth](https://www.linkedin.com/in/prasanthv/) and [Tanumoy](https://www.linkedin.com/in/tanumoysamanta/) with [Collin](https://www.linkedin.com/in/collin-austad-a360116a/) and [Wilson](https://www.linkedin.com/in/wilson-parrish-7aabbbb6/)) **shared the stage! Ally shared how they’re building data mesh architectures to support their business and drive agility moving from mono repos to multi repos. Then Podium shared how they’re using dbt Cloud, Cortex AI and Snowflake to keep tabs on product reviews over time so their customers can know the real-time sentiment of their product offerings. Both are having huge impacts on their companies and showed how you can get started and really modernize your data strategies very quickly with dbt Cloud ![Ally / Podium](https://cdn.sanity.io/images/wl0ndo6t/main/7620f8f657bdc9b5157750ba0b6554bc028bbb07-2016x1512.jpg) **Symend ([Raman](https://www.linkedin.com/in/ramanpreetsingh1/)) **wrapped up the week with a talk that drove amazing engagement. They showcased how they leverage Snowflake and dbt Cloud to drive customer-facing data products and deliver unparalleled client value. Raman was surrounded after the session by people wanting to learn more. ![Symend](https://cdn.sanity.io/images/wl0ndo6t/main/88de475b19c86646e473b765707af9fd5eca607e-1230x918.png) ## We announced our most meaningful Snowflake integration to date On the back of our repeat as Snowflake Data Integration Partner of the year, we announced our first Native Application on the Snowflake Marketplace. Our joint customers can now find, purchase and deploy dbt Cloud inside of Snowflake. And, it comes with dbt Explorer, dbt Semantic Layer and our new chatbot, Ask dbt, to interact with Snowflake Cortex AI to deploy smart, accurate and fast AI applications in the enterprise. Read about our launch [HERE](https://www.getdbt.com/blog/introducing-dbt-for-snowflake). And sign up for our global webinar [HERE](https://www.getdbt.com/resources/webinars/boost-ai-reliability-with-dbt-cloud-and-snowflake-cortex-ai). ![Native App Presentation](https://cdn.sanity.io/images/wl0ndo6t/main/3ceef5a75237fbd397aec380563f02e01ff85fde-2016x1512.jpg) ## Momentum through the summer leading to Coalesce And last but not least - a few weeks back we announced our huge semi annual release in the middle of May ([learn more](https://www.getdbt.com/blog/whats-new-in-dbt-cloud-june-2024)). There’s so much good stuff out there we can’t wait for you to try! If all this goodness is making you excited to get back to school in the fall, and continue your learning with dbt - don’t miss our annual customer conference [Coalesce](https://coalesce.getdbt.com/), happening October 7-10 in Las Vegas. You’ll join thousands of data practitioners and leaders to sharpen your skills to continue to drive amazing results in your organization. --- --- title: "What's new in dbt Cloud - June 2024" description: "Learn about all the newest features and functionality now live in dbt Cloud." url: "https://www.getdbt.com/blog/whats-new-in-dbt-cloud-june-2024" date: "2024-06-05" authors: ["Alexis Jones", "Azzam Aijazi"] categories: ["Product"] --- # What's new in dbt Cloud - June 2024 It's that time again...settle in for the latest regular installment of our product announcements blog post 🌞. We've had a very busy couple of months, including hosting our inaugural [dbt Cloud Launch Showcase](https://www.getdbt.com/resources/webinars/dbt-cloud-launch-showcase) where we announced [dozens of new features](https://www.getdbt.com/blog/dbt-cloud-launch-showcase-2024): an AI-assist co-pilot, a visual editor, unit testing, automatic exposures, and much more! Some features are still in private beta, and trust that you'll hear from us when they're ready for prime time. For now, we've consolidated all of the latest features you _can_ get your hands around today—including features that we didn't announce at the Showcase—in one easy-to-read blog post. Let's dive in 👇! ## 🔎 dbt Explorer Our vision is for [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) to be the best place for data teams to discover, understand, and improve their dbt Cloud projects so downstream teams can leverage trusted data assets with confidence. Here’s what’s new: ✅ **dbt Explorer is GA (generally available):** The [foundational dbt Explorer experience](https://www.getdbt.com/blog/proactively-improve-your-dbt-projects-with-new-dbt-explorer-features)—including column-level lineage, model performance analysis, and project recommendations—is now generally available! Get started by clicking the “Explore” tab in dbt Cloud. **📊 [Column-level lineage](https://docs.getdbt.com/docs/collaborate/column-level-lineage)** now includes new features like a lineage lens to view when columns are transformed, icons to identify primary keys, and propagating descriptions for reused columns.  ![Column-level lineage in dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/209fd3d2998407a5e4918514195a330027aa45eb-3058x1100.png) 🎦 **Support for staging environments:** In addition to production environments, dbt Explorer now supports staging environments! Staging environment support is also available for cross-project refs through dbt Mesh. This makes it easier to understand pre-prod/QA state and catch issues before they hit production, while allowing users to bolster their data isolation and governance posture. 💻 **Azure support:** dbt Explorer is now generally available for customers running dbt Cloud on Azure single tenant. **📝 Open in IDE:** Enjoy more cohesive and streamlined developer workflows by jumping directly from dbt Explorer into the dbt Cloud IDE to edit a resource. **📈 Performance and search improvements:** We continue to invest in the underlying experience to deliver improved performance for large lineage graphs, including faster load times and a new default loading state. We also now support more dbt selector methods and you can find auto-suggested selectors in the lineage search bar. ## 📈 dbt Semantic Layer Check out what’s new in the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), and peruse customer FAQs [here](https://docs.getdbt.com/docs/use-dbt-semantic-layer/sl-faqs): 🗄️ **Declarative caching:** Save relevant queries to “pre-warm” the cache and significantly improve the performance of key dashboards or common ad-hoc query requests while also reducing compute costs for frequently-queried metrics. [Declarative caching](https://docs.getdbt.com/docs/use-dbt-semantic-layer/sl-cache#declarative-caching) is now GA for all Semantic Layer customers. 🌎 **Integrations:** As part of the GA of our Tableau integration, you can now find and install the dbt Semantic Layer directly from the [Tableau Exchange](https://exchange.tableau.com/connectors). Also, our product experts recently wrote a [blog post](https://www.getdbt.com/blog/managing-data-source-changes-in-tableau) all about how you can improve data change management in Tableau using the semantic layer. ![Find and install dbt Cloud on the Tableau Exchange](https://cdn.sanity.io/images/wl0ndo6t/main/bab9ebd56cb735adb49496bab3512833aa2bd6ed-2796x1386.png) Additionally, our Google Sheets integration is now GA. You can find all of the dbt Semantic layer integrations [here](https://docs.getdbt.com/docs/use-dbt-semantic-layer/avail-sl-integrations). 🔢 **Metrics as dimensions in MetricFlow:** MetricFlow is the powerful SQL query generation tool behind our semantic layer, and we continue to make improvements to it to make your workflows more streamlined and flexible. Case in point: with metrics as dimensions, you can use the value of _another_ metric in your metric definition. For example, say you want to count “activated accounts,” which is defined as (1) an account (2) with more than five log-ins to your platform. To express this metric in SQL, you’d first write a query to calculate the number of log-ins per account, then count the number of accounts who have with more than 5 log-ins. Now, you can do this natively in MetricFlow by adding metrics as filters to other metrics! [Read the docs](https://docs.getdbt.com/docs/build/ref-metrics-in-filters) to learn more. ## 🌐 dbt Mesh [dbt Mesh](https://www.getdbt.com/product/dbt-mesh), a pattern for collaboration at scale in dbt Cloud, is now GA. It enables teams to make use of multiple, inter-connected dbt projects, each aligned to a domain — boosting collaboration without compromising governance. New capabilities include: 👉 **Trigger on job completion, _across projects_:** Get even more more flexibility in how you deploy your dbt models into production. [Read the docs](https://docs.getdbt.com/docs/deploy/deploy-jobs#trigger-on-job-completion) to learn more. 🎭 **Support for canonical staging environments:** Improve data isolation and build in dbt Cloud without access to production data. [Read the docs](https://docs.getdbt.com/docs/deploy/deploy-environments#staging-environment) to learn more. ☁️ **Azure support:** dbt Mesh is now generally available for customers running dbt Cloud on Azure single tenant. ## ✏️ Develop We shipped lots of exciting improvements to both the dbt Cloud CLI and IDE. Check 'em out! ✅ **The [dbt Cloud CLI](https://docs.getdbt.com/docs/cloud/cloud-cli-installation) is now GA:** Develop anywhere using your code editor of choice, bolstered by dbt Cloud, including capabilities like [dbt Mesh support](https://www.getdbt.com/product/dbt-mesh), [defer to production](https://docs.getdbt.com/docs/cloud/about-cloud-develop-defer), and improved performance. Other features at GA include: - **💪 Support for dbt Power User:** If you use VS Code, you can now also use the [Power User for dbt Core and dbt Cloud](https://docs.myaltimate.com/setup/reqdConfigCloud/) extension with the dbt Cloud CLI to bolster your productivity. - **☁️ Azure support:** Additionally, the dbt Cloud CLI is now available to organizations running dbt Cloud on Azure single tenant. ![dbt Cloud CLI mockup](https://cdn.sanity.io/images/wl0ndo6t/main/9a66bc76734e8a1f3160b0e00bd362b32603fca7-942x450.png) 🔢 **Unit testing is now GA:** Use [unit tests](https://docs.getdbt.com/docs/build/unit-tests) to validate the behavior of model logic _before_ the model is materialized with real data. If a test fails, the model won’t build—saving you from unnecessary data platform spend, while improving data product reliability. 🌱 **Prune branches in the IDE:** Using this [Git button](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/ide-user-interface#prune-branches-modal), you can delete local branches that have been deleted from the remote repository, keeping your branch management tidy. Available in all regions now and will be released to single tenant accounts during the next release cycle. 🎋 **Git branch as an environment variable:** You can now [reference your current Git branch](https://docs.getdbt.com/docs/build/environment-variables#special-environment-variables) as an environment variable, allowing you to do things like dynamically use the Git branch name as a prefix for a development schema. 🧹 **Support for SQLFluff v3 in the IDE**. In addition to other benefits, you’ll now get better feedback on .sqlfluff configuration errors directly in the dbt Cloud UI as logs and toasts. ⚠️ **Better notifications around invocation failures in the IDE.** Now, when an invocation fails, the IDE will surface a prominent notification banner above the system log, making it easier to immediately see when a job has failed 👍 **Other Cloud IDE improvements:** You can now make changes to multiple projects at the same time, which is really helpful for users operating in a mesh, and we’ve also made improvements to our backend Cloud IDE file system to improve overall performance. ## 🔄 Deploy Ship pipelines faster, and more reliably, with these workflow improvements. **🔀 Merge jobs.** Immediately trigger a job to run when a pull request is merged, and enjoy [native functionality for continuous deployment (CD)](https://docs.getdbt.com/docs/deploy/merge-jobs) in dbt Cloud. Coupled with deferral, you can be sure the latest data is always reflected in production…without driving up data platform spend. ![Screenshot: Run on merge](https://cdn.sanity.io/images/wl0ndo6t/main/03c9ec08036a8dfe92a9e59394ab70b1fe50deaa-1116x166.png) 🛑 **Job deactivation**: Runs with repeated failures are automatically deactivated so they don’t continue to run and fail indefinitely. They can be easily reactivated by editing a deactivated job. ## 💻 Platform improvements The team is always hard at work to make dbt Cloud more performant, reliable, and interoperable. 💫 **Keep on latest version:** dbt Cloud should feel and function like the other SaaS apps your team uses: you shouldn’t have to manually upgrade versions under the hood. Now generally available, just select “[Keep on latest version](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.8#keep-on-the-latest-version-in-dbt-cloud)” in your environments and jobs to get immediate access to the latest and greatest functionality in dbt. Going forward, this is how we will be delivering dbt to our customers: reliably and continuously. Read our [recent blog post](https://www.getdbt.com/blog/seamless-scalability-effortless-upgrades-the-enhanced-dbt-cloud-platform) for more! ![Screenshot: Keep on latest version](https://cdn.sanity.io/images/wl0ndo6t/main/465b96234ef4f661bf64c628156ae12702f75ad0-1620x1116.png) ⏰ **Parse time improvements.** We’ve also made optimizations under the hood to significantly improve parse performance in dbt Cloud. These are available today to everyone running on “Keep on latest version.” 🔒**Databricks OAuth:** Now generally available, dbt Cloud supports developer [OAuth with Databricks](https://docs.getdbt.com/docs/cloud/manage-access/set-up-databricks-oauth), providing an additional layer of security for dbt Enterprise users. ## 👭 Partnerships It's been an exciting few weeks on the partnership front! ❄️ **Snowflake native app:** dbt is now available on the [Snowflake Marketplace](https://www.snowflake.com/en/data-cloud/marketplace/) as a native app! The dbt for Snowflake Native App brings dbt Cloud’s discovery and semantic capabilities to Snowflake's robust, governed architecture. Now, the dbt Cloud experience extends directly to the Snowflake UI, allowing users to jump in and gain insights from their dbt projects with one Snowflake login. Moreover, you can use your Snowflake committed spend to pay for dbt Cloud and sign on Snowflake paper (no additional vendor approvals required! 🙌). Read the [blog post](https://www.getdbt.com/blog/introducing-dbt-for-snowflake) to learn more. 🗣️ **Ask dbt:** Now in open beta for dbt users on the Snowflake native app, we also launched an AI chatbot designed to help users get trusted answers to their questions, faster—without writing a single line of SQL. Ask dbt combines the power of a Snowflake Cortex LLM with the dbt Semantic Layer to translate natural language questions into a semantic query. Check out the [blog post](https://www.getdbt.com/blog/introducing-dbt-for-snowflake) to learn more. ![Ask dbt chatbot in the Snowflake native app](https://cdn.sanity.io/images/wl0ndo6t/main/262eaa2f581e862c463d4a2c6d242661e1f0410f-1600x1054.jpg) **👋 Microsoft adapters:** In addition to the GA of our [Microsoft Fabric adapter](https://docs.getdbt.com/docs/cloud/connect-data-platform/connect-microsoft-fabric), dbt Cloud now supports [Microsoft Azure Synapse Analytics](https://docs.getdbt.com/docs/cloud/connect-data-platform/connect-azure-synapse-analytics) (in Preview). To get started, create a new dbt project in dbt Cloud and choose Fabric or Synapse as your data platform. ## Wrapping up We're so excited to get these features in your hands and as always, look forward to hearing your feedback. Until next time! --- --- title: "Introducing dbt for Snowflake" description: "Learn about the dbt native app on Snowflake Marketplace." url: "https://www.getdbt.com/blog/introducing-dbt-for-snowflake" date: "2024-06-04" authors: ["Amy Chen"] categories: ["Product"] --- # Introducing dbt for Snowflake Today, we’re excited to announce that dbt is now available on the [Snowflake Marketplace](https://app.snowflake.com/marketplace/listing/GZTYZSRT2UA) as a Native App! The dbt Snowflake Native App brings dbt Cloud’s discovery and semantic capabilities to Snowflake's robust, governed architecture. Now, the dbt Cloud experience extends directly to the Snowflake UI, allowing users to jump in and gain insights from their dbt projects with one Snowflake login. Whether you're a data analyst or business decision-maker, the dbt Snowflake Native App empowers teams to access and interact with data more effectively, reducing the cost of producing insights. ## How it works Our dbt Snowflake Native App will launch with the initial experience to empower users closer to the business. In the initial release, you will gain access to Orchestration Observability, dbt Explorer, and Ask dbt, new dbt-assisted chatbot powered by the dbt Semantic Layer and Snowflake Cortex (Ask dbt is currently in beta). [Watch video](https://www.loom.com/share/cd2f75380e9f45cabd8abe9b768941fe?sid=dd7f835e-7cf2-4c7c-897e-72e5344d42bf) ## Start with discovery dbt Explorer brings data catalog capabilities directly to Snowflake and helps business analysts navigate and leverage data products to deliver trustworthy insights. Snowflake users can use dbt Explorer directly in the Snowflake UI after installing the dbt Snowflake Native App in their account. Users can quickly jump in and see the lineage and status of their dbt assets, from source to metric. Reuse what you have, not duplicate, to optimize compute costs and engineering time. Users will have access to: - Column level lineage - dbt Mesh cross-project lineage - Rich metadata on the latest state of your sources and models, including freshness and last executed at - Project Recommendations and Model Performance ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/d3c3b05eb0d90787cac6cbd79a0d8d83127e8661-1514x883.jpg) ## Ask questions of your data A new beta experience that we are introducing in this app is Ask dbt. Ask questions like “What is ARR growth over time?” or “What is the count of customers by plan type?” and get answers without writing a single line of SQL. It works behind the scenes by converting the natural language request into a dbt Semantic Layer query. This is done by combining Snowflake Cortex AI and dbt Semantic Layer. That query is then used to generate and execute SQL directly in Snowflake. Unlike traditional AI chatbots, Ask dbt uses the dbt Semantic Layer to provide a critical layer of context about your dbt project, improving accuracy by 3X as observed in our [benchmark](https://roundup.getdbt.com/p/semantic-layer-as-the-data-interface). With Ask dbt, users can ask questions in natural language and receive insights in an understandable format, which can significantly speed up business processes and decision-making. What’s more magical is that the dbt Semantic Layer also powers your insights for Tableau, Google Sheets, and other BI tools, so you know that numbers are the same throughout your stack. Users can also use the SQL query Ask dbt generates to get started in their own discovery and querying. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/57f7a754603e4fac67307acaa3547cbabc67ca26-1600x1054.jpg) ## Trust your source freshness When you look into dbt Explorer, you can tell the freshness of the data easily. If you work with stale data, you can immediately lose trust in the result. By providing visibility into each job run, Orchestration Observability means analysts can identify and address issues at their inception. See the last time the job ran successfully and empower yourself to fix it or find the right person. If you want to kick off a job when new data has been loaded into Snowflake, check out the sample code to set up a Snowflake Triggered Job to orchestrate dbt Cloud jobs using Snowflake Tasks. Additional orchestration functionality is coming soon. ## Get started Accessing the power of dbt Cloud has never been easier. You can purchase the dbt Snowflake Native App via the public listing on [Snowflake Marketplace](https://app.snowflake.com/marketplace/listing/GZTYZSRT2UA) or contact us for custom pricing. Once purchased, you can install the app and a dbt Cloud enterprise account, all within your Snowflake environment. For procurement, you can use your Snowflake committed spend to pay for the purchase and sign on Snowflake paper (No additional vendor approvals required!). ## Want more? We are so excited to continue building this experience, and we want to hear from you. Let’s chat! Can't get enough of dbt and Snowflake? Join our upcoming webinar,[ Boost AI Reliability with dbt Cloud and Snowflake Cortex](https://www.getdbt.com/resources/webinars/boost-ai-reliability-with-dbt-cloud-and-snowflake-cortex-ai), to learn how integrating Snowflake Cortex with dbt Cloud can enhance your AI initiatives with reliable and accurate data, reduce AI hallucinations, accelerate data-driven projects, and lower the barrier to analytics. --- --- title: "What is data engineering?" description: "Data engineering powers your analytics. Learn the key roles, challenges, and tools driving today’s modern data workflows." url: "https://www.getdbt.com/blog/what-is-data-engineering" date: "2024-06-03" authors: ["Daniel Poppy"] categories: ["Learn"] --- # What is data engineering? Data engineering is the foundation of modern data work. As data volumes grow and business decisions become more reliant on real-time insights, the role of the data engineer has become critical to organizational success. **But what exactly is data engineering, and why does it matter?** In this guide, we’ll define data engineering, outline the key responsibilities of data engineers, and explore its growing impact across industries. We’ll also share how [dbt](https://www.getdbt.com/product/dbt) helps modern data teams streamline engineering workflows and deliver trusted, analysis-ready data. ## What is data engineering? Data engineering is the practice of designing, building, and managing the infrastructure that enables efficient data collection, storage, transformation, and analysis. At its core, it’s about making raw data usable—so analytics engineers, data scientists, and business teams can turn it into insight. Data engineers create and optimize the pipelines that move data from source systems into centralized platforms like cloud data warehouses. Their work ensures data is timely, accurate, and accessible. It’s important to distinguish data engineering from data science. While data scientists analyze and interpret data to build models or answer questions, data engineers build the systems that make that work possible. Think of data engineering as the backbone of the modern data stack. Without it, none of the higher-value data work can happen. ### Data engineering vs. data science Data engineering and data science are often mentioned together, but they serve distinct, complementary roles in the modern data team. Data science is about deriving insights from data: building predictive models, uncovering patterns, and supporting business decisions through analysis. Data engineering focuses on making that analysis possible: designing the systems and pipelines that ensure data is clean, structured, and ready for use. Without solid infrastructure, even the best analysis can’t get where it needs to go. ## Why is data engineering important? Strong data engineering is the difference between having data and actually using it. When data is fragmented, messy, or slow to access, teams waste time and make decisions based on incomplete information. A modern data engineering function ensures data is well-organized, governed, and available—so teams can focus on delivering insights. By building scalable pipelines and standardizing data transformation, data engineers enable everything from operational dashboards to advanced machine learning. ## The demand for data engineers is only growing Data engineering is one of the fastest-growing roles in tech—and it’s no surprise why. Organizations are generating more data than ever. [By 2027, global data creation is expected to hit 394 zettabytes per day](https://www.digitalsilk.com/digital-trends/how-much-data-is-generated-per-day/). As data volumes surge, so does the need for skilled professionals to manage, structure, and transform that information. The [World Economic Forum lists data engineering as one of the top growth jobs through 2030](https://reports.weforum.org/docs/WEF_Future_of_Jobs_Report_2025.pdf). It’s a core function for any company that wants to be data-driven. That demand comes with strong financial incentives. According to [Glassdoor, average salaries for data engineers](https://www.glassdoor.com/Salaries/data-engineer-salary-SRCH_KO0,13.htm) consistently reach six figures—reflecting both the importance and the scarcity of this expertise. ## Key responsibilities of a data engineer The role of a data engineer spans a wide range of technical and strategic responsibilities. At a high level, they ensure data flows smoothly from source systems to destinations where it can be analyzed and acted on. Core responsibilities include: - **Building and maintaining data pipelines**: Data engineers design and manage the pipelines that move data from source systems (like databases, APIs, or streaming platforms) into centralized storage. These pipelines must be reliable, efficient, and scalable. - **Ensuring data quality**: Clean, accurate data is non-negotiable. Data engineers build validation and cleaning processes to catch errors, handle anomalies, and maintain trust in downstream reporting and analytics. - [**Data transformation**](https://www.getdbt.com/blog/what-is-data-transformation): Raw data often needs significant reshaping before it’s ready for use. Engineers write transformation logic to normalize formats, combine datasets, apply business rules, and prepare data for analysis. - **Optimizing data storage**: Choosing the right storage solution—whether a relational database, cloud warehouse, or NoSQL system—is critical. Data engineers design schemas and partitioning strategies that balance speed, cost, and scalability. - **Scaling infrastructure**: As data volumes grow, infrastructure must scale with them. Data engineers automate workflows, optimize performance, and ensure the stack can handle increasing complexity without slowing down. ## Data pipelines: the heart of data engineering Data pipelines are the foundation of modern data engineering. These automated workflows move raw data from its source, apply necessary transformations, and load it into systems where it can be stored, analyzed, and acted on. Pipelines ensure that data flows continuously and reliably across the organization. There are several types of data pipelines, each suited to different use cases: - **Batch processing pipelines**: These pipelines process large volumes of data at scheduled intervals—for example, aggregating daily sales from a retail system every night. - **Streaming pipelines**: Designed for real-time processing, these pipelines handle data as it arrives. Industries like finance rely on streaming pipelines for up-to-the-second decision-making. - [**ETL (Extract, Transform, Load)**](https://www.getdbt.com/blog/extract-transform-load): In ETL workflows, data is extracted from source systems, transformed outside the warehouse, and then loaded into a destination. [dbt](https://www.getdbt.com/product/what-is-dbt) helps simplify the transformation layer by enabling engineers to build modular, testable SQL logic. - [**ELT (Extract, Load, Transform)**](https://www.getdbt.com/blog/extract-load-transform): A modern alternative, ELT workflows load raw data directly into the warehouse first, then apply transformations inside the warehouse. This approach is more scalable—and it’s where dbt excels. With [dbt](https://www.getdbt.com/product/dbt), teams can transform data using version-controlled SQL models, automated testing, and lineage documentation. ## The data engineering lifecycle Data engineering follows a structured lifecycle to ensure data is handled consistently from collection to analysis. Each stage plays a critical role in transforming raw inputs into trusted insights: ### 1. Data collection Data engineers gather data from various sources—databases, APIs, web services, and streaming platforms. These inputs are often unstructured or semi-structured, requiring processing before they’re analytics-ready. ### 2. Data storage Once collected, data needs to be stored in a way that supports scalability and accessibility. Depending on volume and use case, this could involve data lakes, cloud storage, or modern data warehouses. ### 3. Data transformation After storage, raw data must be cleaned, standardized, and modeled for analysis. This includes filtering, formatting, and joining datasets. Tools like dbt enable engineers to automate this process with version-controlled, modular SQL logic—making [transformations](https://www.getdbt.com/blog/what-is-data-transformation) more transparent and reliable. ### 4. Data analysis Once transformed, data is ready for use by analysts and data scientists. Data engineers continue to support this phase by maintaining infrastructure that enables fast, dependable querying. ### 5. Data governance Data engineers implement [governance policies](https://www.getdbt.com/blog/data-governance-best-practices) to ensure data is secure, compliant, and discoverable. This includes managing access controls, tracking lineage, and enforcing standards around data quality and retention. ## Common challenges in data engineering Data engineers navigate a wide range of technical and organizational challenges. Each requires sharp problem-solving and a strong foundation in data infrastructure. - **Data sprawl and diverse sources. **Engineers must integrate data from many sources—often in incompatible formats. Bringing this data together is essential to create a unified, trusted foundation for analysis. - **Ensuring consistency and quality. **Reliable insights depend on clean, accurate data. Engineers build systems that validate and standardize data as it flows through pipelines, catching errors early and maintaining trust in downstream outputs. - **Scaling data infrastructure. **As organizations grow, so does the volume and complexity of their data. Engineers must continuously optimize pipelines and storage to ensure performance, reduce cost, and support increasing user demands. - **Governance and security. **Engineers are responsible for implementing access controls, tracking data lineage, and enforcing compliance. Strong governance safeguards sensitive data and supports auditability. - **Balancing batch and real-time needs. **Many teams rely on a mix of real-time and scheduled data workflows. Engineers must design flexible systems that can support both modes, enabling timely insights without sacrificing efficiency. - **Automating complex workflows. **Modern data pipelines often include multiple steps, dependencies, and stakeholders. Engineers use orchestration tools to automate these workflows, reducing manual effort and increasing reliability. - **Keeping pace with evolving tools. **The data ecosystem evolves quickly. Engineers need to stay current with emerging frameworks, tools, and best practices to build scalable systems and support innovation. **** ## The tools of the data engineer Data engineers use a wide range of tools to build reliable, scalable data systems. Here are some of the most common: - [**dbt**](https://www.getdbt.com/product/dbt): dbt is a SQL-based transformation tool that lets engineers build modular, tested, and version-controlled data pipelines directly in the data warehouse. It supports cloud platforms like Snowflake, BigQuery, Redshift, and Databricks, and brings software engineering best practices—like CI/CD and documentation—to analytics workflows. - [**Apache Kafka**](https://kafka.apache.org/): Kafka is a distributed event streaming platform used for ingesting and processing large volumes of real-time data. It’s commonly used for building high-throughput, fault-tolerant data pipelines. - [**Apache Spark**](https://spark.apache.org/): Spark is a powerful processing engine for large-scale data workloads. It supports both batch and streaming data and is often used for machine learning, ETL, and analytics workflows requiring distributed computing. - [**Apache Airflow**](https://airflow.apache.org/): Airflow is a workflow orchestration platform that allows data engineers to programmatically author, schedule, and monitor complex pipelines. It’s commonly paired with dbt to coordinate jobs across the modern data stack. ## Conclusion Data engineering powers the [modern data stack](https://www.getdbt.com/blog/ai-data-engineering). It ensures organizations can collect, transform, and deliver reliable data at scale—fueling everything from operational reporting to advanced analytics. By building robust pipelines, enforcing data quality, and streamlining transformation, data engineers lay the groundwork for trusted, actionable insights. And with modern tools like dbt, teams can bring software engineering practices—like version control, testing, and modular development—into the data workflow. Whether you’re scaling a mature data platform or just starting out, a strong foundation in data engineering principles—and the right tools—will set your organization up for long-term success in a data-driven world. ## The role of dbt in data engineering dbt helps data engineers transform data directly in the warehouse using modular, SQL-based code. It brings software engineering best practices—like testing, version control, and documentation—into the analytics workflow. With dbt, data teams can: - **Build modular transformations**: Reuse logic across models to reduce duplication and make pipelines easier to manage. - **Collaborate through version control**: Built-in Git workflows allow teams to review, test, and track changes before they go live. - **Scale confidently**: dbt works natively with platforms like Snowflake, BigQuery, and Redshift to support large, complex data environments. dbt makes it easier to manage ELT workflows and ensures your pipelines are reliable, governed, and built to scale. [**Try dbt for free**](https://www.getdbt.com/signup/) to see how it simplifies the way data engineers build and maintain trusted data pipelines. ## FAQs about data engineering **What is data engineering? ** Data engineering is the practice of designing, building, and maintaining the infrastructure that allows organizations to collect, store, and analyze data efficiently. It’s distinct from data science—while data scientists analyze and interpret data, data engineers build and optimize the pipelines that make analysis possible. Think of data engineering as the foundation of any data-driven operation: it ensures your data is available, accurate, and analysis-ready. **What’s the difference between ETL and ELT in data engineering?** ETL (Extract, Transform, Load) transforms data before loading it into a warehouse. ELT (Extract, Load, Transform) loads raw data into the warehouse first, then transforms it there. ELT is more common in modern stacks, thanks to cloud data warehouses that can handle transformations at scale. dbt is especially effective for ELT workflows, enabling teams to manage transformations directly in the warehouse. **What are the key responsibilities of a data engineer?** Data engineers build and maintain pipelines that move raw data from various sources into centralized storage. They validate and clean data to ensure quality, transform it into usable formats, and optimize storage for performance and cost. As data volumes grow, they scale infrastructure and automate processes to keep pipelines efficient and resilient. **What are data pipelines and why are they important?** Data pipelines are automated workflows that extract, transform, and load data for analysis. They can operate in batches or real time. Pipelines are essential for moving data reliably across systems and ensuring it’s clean, consistent, and ready for use in decision-making. **What are the main stages of the data engineering lifecycle?** The lifecycle includes: 1. **Data collection** from sources like APIs and databases 2. **Storage** in cloud data lakes or warehouses 3. **Transformation** into structured, analysis-ready formats 4. **Analysis** by data scientists and business users 5. **Governance** to ensure data is secure, accessible, and compliant **How does data engineering differ from data science?** Data engineers build the systems that collect, clean, and deliver data. Data scientists use that data to create models and generate insights. A common analogy: engineers build the roads; scientists drive on them. **How does dbt benefit data engineering teams?** dbt brings version control, testing, and modular development to the transformation layer. It supports collaboration with Git workflows, simplifies model maintenance, and integrates with cloud warehouses like Snowflake, BigQuery, and Redshift. It’s how modern teams manage data transformations like software. --- --- title: "May dbt Community update" description: "In this monthly Community update, we share insights from our Slack AMA, new community spotlights, meetups, and announcements." url: "https://www.getdbt.com/blog/may-community-update" date: "2024-06-01" authors: ["Kathryn Chubb"] categories: ["Community"] --- # May dbt Community update Welcome to the dbt Community update, a monthly blog about everything happening in the [dbt Community](https://www.getdbt.com/community)! In May we hosted an AMA, presented the Spring 2024 dbt [Community Spotlight](https://docs.getdbt.com/community/spotlight), hosted eight in-person [meetups,](https://www.meetup.com/pro/dbt/) and had a ton of great discussions (and [fire memes](https://www.linkedin.com/feed/update/urn:li:activity:7199072926753570816)) on our Slack channel. Are you ready for the recap? Let’s get started. ## dbt Community Slack AMA Each month we host a live Ask Me Anything event. This month Phoenix Jay, Dave Connors, and Jeremy Cohen hosted the AMA and discussed the importance of collaboration, data quality, and some of the newest features in dbt. Here’s a recap if you missed it. You can also check out the [full video recording](https://www.getdbt.com/resources/community-slack-ama/confirmation). ### Uncovering SQL errors with unit testing One of the standout features discussed in the AMA is [unit testing](https://docs.getdbt.com/docs/build/unit-tests). Jeremy expresses his excitement about how it's been implemented in dbt Cloud and dbt Core v1.8. But why is unit testing such a game-changer? Dave shares a compelling example from the Jaffle Shop project, where unit testing helped uncover a SQL error. By integrating unit testing into your data analytics process, you can catch errors early on, ensuring the accuracy and reliability of your results. ### Exploring the dbt Mesh pattern dbt Cloud customers have been buzzing about the [dbt Mesh](https://www.getdbt.com/blog/dbt-mesh-is-now-generally-available) pattern mentioned by Jeremy. It tackles the challenge of cycles across models and projects. While it sounds complex, Jeremy and Dave find it to be an interesting problem to solve. The mesh pattern provides a solution for teams working on interconnected projects, enabling seamless collaboration and efficient data workflows. ### Collaboration and data quality Collaboration and data quality are the foundation of successful analytics projects. The speakers emphasize the need for well-maintained and documented datasets. Without the right documentation, valuable insights can slip through the cracks. Maintaining data products and ensuring their accuracy is crucial for making informed decisions. Don't overlook the power of good documentation and how it impacts the quality of your data-driven responses. ### Data contracts for seamless integration Data quality and team collaboration are seamlessly integrated into existing workflows through data contracts in dbt. Jeremy highlights their importance as a crucial aspect of implementing data mesh or decentralized structures. By establishing data contracts, you can provide consistent standards across teams, while also emphasizing the significance of data quality and team collaboration. It's a win-win situation for everyone involved. ### Tackling data quality challenges Data quality challenges plague data teams of all sizes. Poor data quality can lead to faulty insights and, ultimately, poor decision-making. The discussion highlights the growing interest in adopting data mesh or decentralized structures to address these challenges. The integration of data contracts allows teams to prioritize data quality and collaboration, making sure that the right data is available when and where it's needed. ### Simplifying data transformation for complex problems Data transformation can be complex, but it also presents an opportunity for tackling more intricate problems. Dave and Jeremy stress the significance of thinking about interfaces and governance early on in a dbt project. By considering these factors from the start, you can streamline your data analytics workflows and make sure that your team is set up for success. Don't let data transformation become a stumbling block. Embrace its potential for solving even the most challenging problems. ### Enabling teams and addressing people problems While tools are essential for data analytics, it's equally important to address people problems and enable other teams within your organization. Dave and Jeremy highlight the value of a strong data team that goes beyond using tools. By fostering collaboration, emphasizing transparent documentation, and providing the necessary support, you can empower other teams to leverage data effectively. Remember, it's not just about the tools, but the people who use them. ### dbt preferences When it comes to dbt preferences, Jeremy emphasizes the importance of explicit configuration and using SQL plus Jinja. Explicit configuration provides clarity and reduces ambiguity, enabling smoother collaboration. SQL plus Jinja offers a powerful combination for manipulating data, making it easier to perform complex transformations. Additionally, unit testing is highlighted as a valuable practice for teams of any size. While the dbt mesh pattern is more relevant for larger teams, it's crucial for all teams to think about interfaces and governance early on in their dbt projects. ### Data contracts as future focus The conversation concludes with a mention of contracts as a potential future topic for an AMA. Contracts hold great promise for streamlining data analytics workflows even further, emphasizing data quality, and facilitating collaboration across teams. Stay tuned for more exciting updates on this front. ### Get ready for another AMA in June [Join us for next month’s AMA](https://www.getdbt.com/resources/webinars/community-ama) on June 27th at 4 pm EST. We have Alex Welch, director of data at dbt Labs, to answer questions and discuss more about BI tools, analytics trends, and the Semantic Layer. Register now to get the link to watch and join the [#dbt-community-merge channel in Slack](https://getdbt.slack.com/archives/C025ZN1L679/p1716308109072499) to participate live! ## Spring 2024 dbt Community Spotlight Every quarter, we highlight community members in the dbt Community Spotlight. These are individuals who have gone above and beyond to contribute to the community in a variety of ways. We're excited to present the Spring 2024 dbt Community Spotlight! This round, we are featuring Johann De Wet, Tyler Rouze, Juan Manuel Perafan, Mariah Rogers, Yasuhisa Yoshida, and Safiyy Momen. Visit the [Community Spotlight](https://docs.getdbt.com/community/spotlight) page to learn about their backgrounds, their plans to grow as leaders, and their experiences—both learning from others and from sharing their own knowledge. If you’re interested in being selected for future rounds of the Community Spotlight, learn more about [becoming a contributor](https://docs.getdbt.com/community/contribute). ### [Johann de Wet](https://docs.getdbt.com/community/spotlight/johann-de-wet) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/c1ef91c5da512761f0ef412b9b61aa44a10015da-1736x1284.png) I'm forever indebted to my manager, John Pienaar, who introduced me to both dbt and its community when I joined his team as an Analytics Engineer at the start of 2022. I often joke about my career before dbt and after dbt. Our stack includes Fivetran, Segment, Airflow, and BigQuery to name a few. Prior to that, I was a business intelligence consultant for 16 years working at big financial corporates. During this time I've had the opportunity to work in many different roles from front end development to data engineering and data warehouse platform development. The only two constants in my career have been SQL en Ralph Kimball's Dimension Modeling methodology...which probably makes me a bit partial to those. ### [Tyler Rouze](https://docs.getdbt.com/community/spotlight/tyler-rouze) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/e310746aeeade737298a55059d568253dc0dd9cc-1736x1284.png) My journey in data started all the way back in college where I studied Industrial Engineering. One of the core topics you learn in this program is mathematical optimization, where we often use data files as inputs to model constraints on these kinds of problems! Since then, I've been a data analyst on both small and large teams, and more recently a consultant shepherding our firm's dbt-based projects towards success. Since joining the dbt Community, I've spoken at the [Chicago dbt Meetup](https://www.meetup.com/chicago-dbt-meetup/), [Coalesce](https://coalesce.getdbt.com/speakers/tyler-rouze) (a milestone for my career!), dbt's Data Leaders Series, and even made open source contributions to `dbt-core`! It has been the joy of my career to be a part of this vibrant community. ### [Juan Manuel Perafan](https://docs.getdbt.com/community/spotlight/juan-manuel-perafan) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/b08e7fd7a8f4c340bcd91d35c714d22b7a0cf0ad-1736x1284.png) Born and raised in Colombia! Living in the Netherlands since 2011. I've been working in the realm of analytics since 2017, focusing on Analytics Engineering, dbt, SQL, data governance, and business intelligence (BI). Besides consultancy work, I am very active in the data community. I co-authored the book *Fundamentals of Analytics Engineering* and have spoken at various conferences and meetups worldwide, including [Coalesce](https://coalesce.getdbt.com/), Linux Foundation OS Summit, Big Data Summit Warsaw, Dutch Big Data Expo, and Developer Week Latin America. I also love meetups! I am the founder of the Analytics Engineering Meetup and co-founder of the [Netherlands dbt Meetup](https://www.meetup.com/amsterdam-dbt-meetup/). ### [Mariah Rogers](https://docs.getdbt.com/community/spotlight/mariah-rogers) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/d9f8f9df4d0ff01860d59ba376283a598ce52c86-1736x1284.png) I got my start in the data world helping create a new major and minor in Data Science at my alma mater. I then became a data engineer, learned a ton, and propelled myself into the clean energy sector. Now I do data things at a clean energy company and geek out on solar energy at work and at home! I attended my first Coalesce virtually in 2021 when my former colleague Emily Ekdahl gave a talk about some cool things we'd been working on. She inspired me to propose a talk the following year, so I submitted two topics and, surprisingly, both were accepted! I ultimately chose to speak about Testing in dbt in New Orleans in 2022, and the community's reception of that talk continues to be a highlight of my career. ### [Yasuhisa Yoshida](https://docs.getdbt.com/community/spotlight/yasuhisa-yoshida) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a6c31f118ddbfe899933041968bd8c87efe76450-1736x1284.png) I currently work as a data engineer at a startup called [10X](https://10x.co.jp/). Specifically, I work with BigQuery to provide data marts for business users. Before using dbt, the queries for creating data marts were overly complex and lengthy, resulting in low data quality. With dbt, we have improved our process by breaking down queries into manageable parts, visualizing data lineage, and enabling easy creation of tests. I am actively involved in the dbt community and share our insights on using dbt at #local-tokyo. Specifically, I shared our experiences with efficient metadata management using dbt-osmosis, and visualizing data quality using elementary. ### [Safiyy Momen](https://docs.getdbt.com/community/spotlight/safiyy-momen) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/5762438845c484e7bcb482dcb489a81e4c035383-1736x1284.png) I've been in the dbt community for ~4 years now. My experience is primarily in leading data teams, previously at a healthcare startup where I migrated the stack. The dbt Community was invaluable during that time. More recently, I've built a product, Aero, that helps Snowflake users optimize costs with a Native extension. I'm exploring ways to automate analytics engineering workflows. I've spoken at various meetups, including the [New York dbt Meetup](https://www.meetup.com/nyc-dbt-meetup/), on data warehouse cost optimization. ## May dbt Meetups In May, we hosted one meetup in North America in New York City. We hosted five meetups in EMEA in Copenhagen, Dubai, Berlin, Stockholm, and the Netherlands. And in APAC, we had two meetups, one in Melbourne and one in Sydney. At our [Stockholm dbt Meetup](https://www.meetup.com/stockholm-dbt-meetup/events/300513833/), organized with our partner Solita with 60 folks in attendance, we ran a Peer Exchange led by dbt Labs’s [Kshitij Aranke](https://www.linkedin.com/in/aranke/). In the peer exchange, a newer format that you'll start to see in more Meetups, we broke out into three groups and discussed the topics of Embracing AI, Data Analytics at Scale, and Analytics Engineering Best Practices. Attendees shared their professional experiences, asked each other good questions, took actionable notes, and forged new relationships. You can see photos from Stockholm, Melbourne, and Copenhagen meetups below! ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/edc1ef2b3a8ca3853a97f5a7f69a1fc1a5c024ef-3866x2577.jpg) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/bd3cf5333bf5b23cc99a1e884bf71ad2529e24c8-4032x3024.jpg) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/1c944ae273d0fa699ae927dbfa16bbf339328067-4080x3072.png) ## dbt Community announcements We’ll wrap up this month’s update with some of the exciting announcements that are regularly posted in our [#announcements](https://getdbt.slack.com/archives/C0VLZM3U2/p1715777392876319) channel on Slack. ### dbt Cloud Launch Showcase On May 14th we had our [dbt Cloud Launch Showcase](https://www.getdbt.com/resources/webinars/dbt-cloud-launch-showcase?utm_medium=internal&utm_source=docs&utm_campaign=q2-2025_dbt-cloud-launch-showcase_aw&utm_content=____&utm_term=all___) virtual event. It was a jam-packed 90 minutes with executive keynotes, new product announcements, and demos that all centered around the theme of how dbt Cloud is helping teams deliver Data That Works. Check out the [recap blog](https://www.getdbt.com/blog/dbt-cloud-launch-showcase-2024) to learn more. ### Snowflake Data Cloud Summit We’ll be at Snowflake’s annual user conference June 3–6! You can connect with the dbt Labs team at booths #1327 and #2503. [Data Cloud Summit](https://www.getdbt.com/events/snowflake-data-cloud-summit-24) has an agenda packed full of interesting talks and workshops, including seven dbt Labs sessions: - Best practices for optimizing dbt models: selecting the correct warehouse size - Welcoming stakeholders to the dbt party - Building customer-facing data products on Snowflake and dbt Cloud - Build, deliver and govern data products at scale using data mesh principles - How Techstyle manages analytics complexity with data contracts and data mesh - How Medtronic optimized and future-proofed pipelines architecture to save $1.4M - Unlocking self-service on unstructured data with Omni ### Databricks Data+AI Summit We'll also be attending the [Databricks Data+AI Summit 2024](https://www.getdbt.com/events/databricks-data-ai-summit) from June 10-13. Join our team for insights on accelerating data workflows and improving reliability. Learn from practical demos and network with data leaders and professionals. Don't miss the chance to advance your data strategy as there are three great ways to connect with us during the conference: - [Book a meeting with our team](https://www.getdbt.com/events/databricks-data-ai-summit#bookmeeting): Schedule a one-on-one session with dbt experts at Data+AI Summit. Gain tailored advice and insights to enhance your data strategy. - Visit our booth at Data+AI Summit (booth #91): Stop by booth #91 and explore live demos, get answers to your data questions, and see our solutions in action. - Looking for a break after the sessions? Enjoy an evening of fun, networking, and entertainment and our [Data on the Rocks event](https://events.montecarlodata.com/dataonrocks/dbt). ### Upcoming events - June 27th - [Register for our next Community AMA](https://www.getdbt.com/resources/webinars/community-ama) featuring Alex Welch, director of data at dbt Labs - Ongoing [Cloud Demo with Experts](https://www.getdbt.com/resources/dbt-cloud-demos-with-experts/) in North America, EMEA, and APAC-friendly times - October 7-10th, 2024, [Coalesce](https://coalesce.getdbt.com/register-2024), by dbt Labs in Las Vegas and Online - Add [dbt Events](https://www.addevent.com/calendar/Tb314369) to your calendar! ### June dbt Meetups We’ve got a busy month coming up with 13 [in-person dbt meetups](https://www.meetup.com/pro/dbt) scheduled. If you’re looking for opportunities to learn with fellow members of the dbt Community, and have fun while doing so, join us at one of the sessions listed below: 🇳🇴 Oslo | Thursday, June 6th, organized by Glitni 🇧🇪 Belgium | Thursday, June 6th, organized by dataroots 🇨🇴 Medellín | Tuesday, June 11th, organized by Factored 🇺🇸 Chicago | Thursday, June 13th, organized by Analytics8 🇺🇸 Atlanta | Tuesday, June 18th, organized by Aimpoint Digital 🇹🇼 Taipei | Wednesday, June 19th, organized by community members Karen Hsieh, Laurence Chen, Allen Wang 🇺🇸 Philadelphia dbt Labs on dbt Meetup | Thursday, June 20th, organized by dbt Labs (from the ‘dbt Labs on dbt’ Meetup series) 🇪🇸 Barcelona | Thursday, June 20th, organized by Spaulding Ridge 🇧🇷 São Paulo | Monday, June 24th, organized by community members Bruno Souza de Lima and Thales Donizeti 🇨🇴 Bogotá | Wednesday, June 26th, organized by Factored 🇨🇦 Halifax | Wednesday, June 26th, organized by community member Esther Fraser 🇪🇸 Madrid | Thursday, June 27th, organized by Astrafy🇮🇪 Dublin | Thursday, June 27th, organized by dbt Labs (from the ‘dbt Labs on dbt’ Meetup series) 🇯🇵 Tokyo | Thursday, June 27th, organized by dbt Labs (from the ‘dbt Labs on dbt’ Meetup series) There are so many exciting things going on in the dbt Community, and we can’t wait to see you all there! If you haven’t yet, [join the community](https://www.getdbt.com/community) today. --- --- title: "dbt Labs Launches dbt for Snowflake, a New Native App on Snowflake Marketplace" description: "Customers can now easily experience the benefits of dbt Cloud while ensuring that the AI apps built on top of Snowflake are built on trusted data" url: "https://www.getdbt.com/blog/dbt-labs-launches-dbt-for-snowflake-a-new-native-app-on-snowflake-marketplace" date: "2024-05-31" authors: ["Alyssa Smrekar"] categories: ["Press"] --- # dbt Labs Launches dbt for Snowflake, a New Native App on Snowflake Marketplace **SAN FRANCISCO – **June 4, 2024: [dbt Labs](https://www.getdbt.com/), the pioneer in analytics engineering, today announced at Snowflake’s annual user conference, [Snowflake Data Cloud Summit 2024](https://www.snowflake.com/summit/), the launch of** **dbt for Snowflake on** **[Snowflake Marketplace](https://www.snowflake.com/en/data-cloud/marketplace/). The new Native App enables Snowflake customers to seamlessly access dbt Cloud’s discovery and semantic capabilities, with built-in integrations to [Snowflake Cortex AI](https://www.snowflake.com/en/data-cloud/cortex/), empowering them to efficiently generate business insights while ensuring trust and transparency in deployed data and AI. “In a fast-paced, AI-driven era where data is paramount to enterprise success, efficiency and quality have never been more important,” said Luis Maldonado, VP of Product at dbt Labs. “With our new Snowflake Native App, anyone, from data analysts to decision-makers, can build meaningful trust and transparency in their data processes, business insights and LLMs.” ## How it works dbt for Snowflake is built to empower users closer to the business. With today’s launch, Snowflake users can access several data and AI capabilities, including: - **dbt Explorer**, a visual and interactive data catalog,** **helps developers understand, troubleshoot, and optimize data pipelines to deliver trustworthy insights. Snowflake users can deploy dbt Explorer directly in the Snowflake UI after installing the app in their account, allowing them to quickly access and monitor the lineage and status of their dbt assets, from source to metric. - **Ask dbt**,** **a new beta experience, is a dbt-assisted chatbot powered by the dbt Semantic Layer and Cortex AI. Data consumers can use natural language to ask company-specific questions without writing a single line of SQL, and unlike traditional chatbots, the dbt Semantic Layer provides critical context about a dbt project to improve the accuracy and transparency of the responses. - **Orchestration observability **provides data analysts with enhanced visibility into each job run, so they can better identify and address data freshness issues at their inception before becoming a much larger problem. Additionally, dbt for Snowflake allows Snowflake customers to get value from data and AI faster through streamlined development processes.  “dbt Labs has been a Snowflake partner for years, and now, we’ve gone deeper through the launch of dbt for Snowflake on Snowflake Marketplace,” said Tarik Dwiek, Head of Technology Alliances at Snowflake. “Data quality is under the spotlight, perhaps more now than at any time in recent memory, and organizations like dbt Labs have an important role to play in the present and future of the industry. Together, we can continue to help customers across industries better activate and unlock their data for large language models and business value.” Snowflake Marketplace is powered by Snowflake’s ground-breaking cross-cloud technology, [Snowgrid](https://www.snowflake.com/news/snowflake-unveils-new-performance-innovations-and-enhanced-cross-cloud-capabilities-for-industry-leading-data-platform/), allowing companies direct access to raw data products and the ability to leverage data, data services, and applications quickly, securely, and cost-effectively. Snowflake Marketplace simplifies discovery, access, and the commercialization of data products, enabling companies to unlock entirely new revenue streams and extended insights across the AI Data Cloud. To learn more about Snowflake Marketplace and how to find, try and buy the data, data services, and applications needed for innovative business solutions, click [here](https://www.snowflake.com/data-marketplace/). ### Additional Snowflake Data Cloud Summit News dbt Labs was named the 2024 Data Integration Partner of the Year award winner by Snowflake. For the second year in a row, dbt Labs has been recognized for its achievements as part of the Snowflake AI Data Cloud, providing joint customers a centralized environment for data transformation and empowering data team collaboration using a shared workflow to model, test, and deploy data sets in Snowflake. Together, dbt Labs and Snowflake operate on the shared principle of making data team’s lives easy by eliminating complexity, optimizing workflows, scaling without limitation, and mobilizing data in end-to-end transformations. For more information, check out dbt for Snowflake via the public listing on Snowflake Marketplace or contact dbt Labs directly at [getdbt.com](http://getdbt.com). ### About dbt Labs Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 40,000 companies using dbt every week. To learn more about dbt Labs, visit [getdbt.com](https://www.getdbt.com/) and follow us on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/). ### Press Contact Method Communications for dbt Labs [dbtlabs@methodcommunications.com](mailto:dbtlabs@methodcommunications.com) --- --- title: "Why managing data source changes in Tableau is challenging (and how dbt can help)" description: "Learn how dbt Cloud and the dbt Semantic Layer simplify managing data source changes in Tableau for consistent, reliable insights." url: "https://www.getdbt.com/blog/managing-data-source-changes-in-tableau" date: "2024-05-30" authors: ["Roxi Pourzand", "Gordon Rose"] categories: ["Learn"] --- # Why managing data source changes in Tableau is challenging (and how dbt can help) [Tableau](https://www.tableau.com/) is considered by many to be best-in-class for data visualization. It’s a popular tool in the world of data analysis and business intelligence, enabling users to create compelling and interactive dashboards and reports. Its user-friendly interface and powerful processing capabilities make it a top choice among professionals for transforming complex data into actionable insights. Published data sources are one of the most widely used and powerful features of Tableau. They're also a source of considerable tech debt in virtually all Tableau Server deployments. Data duplication becomes rampant, extracts fall out of use but continue to consume resources, and the functionality of published data sources is under constant threat from changes to the data they depend on. This amounts to a lack of data consistency, impaired data trust, and burnt-out data teams. That’s the bad news. The good news is there IS a better way, and that dbt Cloud makes it possible to serve up more relevant, efficient, current, and agile data sources. When coupled with the dbt Semantic Layer, these data consistency and velocity issues evaporate. The dbt Semantic Layer gives data teams a governed and scalable way to define metrics and offers a first-class integration into Tableau so downstream consumers can get answers to their questions in an interface that’s familiar, flexible, and accessible. Using the dbt Semantic Layer and Tableau together has numerous benefits, and in this blog post, we'll focus on how these solutions improve how teams manage changes to underlying data. ## What is the dbt Semantic Layer? The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) is a technology that allows you to centralize your metric definitions and make them available to users across an enterprise through a consistent, tool-agnostic interface - in other words, to create the coveted “single source of the truth” for all your organization’s metrics. That “single source of truth” makes it possible to guarantee consistency wherever and however your end users consume data. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/ab2f045eca6e73f7b0578295fa1b970183784149-1128x786.png) The dbt Semantic Layer not only makes the coveted “single source of truth” possible, it also makes it easy to implement. Through a simple, declarative interface, users can model the metrics, dimensions, and data relationships that they ‌wish to expose to their data consumers. The dbt Semantic Layer then leverages that semantic modeling to enable dynamic querying and optimized SQL generation, even the automatic—and correct!—handling of joins to satisfy user requests. Tableau enjoys a [first-class integration](https://docs.getdbt.com/docs/use-dbt-semantic-layer/tableau) with the dbt Semantic Layer through a custom live connector, currently generally available in Tableau Desktop and Tableau Server ([see our Tableau Exchange listing](https://exchange.tableau.com/products/1020)), and in the future will be available in Tableau Cloud. You can connect to your dbt Semantic Layer and guarantee that your Tableau consumers are always getting the same answer to the same question, every time. ## How the dbt Semantic Layer improves data change management As noted above, most data consumed in Tableau dashboards is served up in the form of published data sources. With many users of Tableau building data sources and dashboards that users depend on, the data landscape eventually comes to resemble a “Wild West” of content that's difficult to manage at scale. Duplicate data sources lead to confusion on what is the correct one to build on top of and maintain, and teams will spend unneeded effort ‌updating redundant components. Additionally, the resulting content sprawl means that managing any underlying data change becomes a tedious effort to keep affected Tableau data sources up to date. There must be a better way - and fortunately, there is. Without the dbt Semantic Layer, standard Tableau workflows include content creators connecting to their data and spending time establishing the logic and relationships for how the underlying data relates (e.g.: how `customer` relates to `transactions` through `customer_id`). The logic created from these data sources is often reused, making it increasingly difficult to manage what’s correct and what’s not. Despite this being a fairly common workflow, this isn't how things _should_ work. Users defining joins and creating relationships within their data is error prone and will inevitably lead to inconsistency, duplicate logic, and expensive queries. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/efe27f2cef9d507373c5a8ede25173f9c9e4758f-1644x848.png) And on top of all of that, what happens if the upstream data that’s feeding these Tableau data sources changes? Consider a common scenario in data management where an organization decides to change the structure of a primary data source, such as a sales database. In Tableau, you'd need to track down all instances of this data source and where this changed data was used and completely refactor all of these relationships in Tableau. Such a scenario often leads to significant administrative overhead, where a single change can ripple through the entire analytical framework and require extensive time and effort to return to a consistent state. Furthermore, data changes like the one described above aren’t a matter of “if”; they’re inevitable, and so the question becomes _when_ and _how often_ they’ll occur. Our philosophy with the dbt Semantic Layer is designed to help teams proactively address these inevitabilities head on. It’s architected as a hub-and-spoke model to avoid inefficiencies and problems that arise when managing logic at the edge (where the data is being consumed). By contrast, we believe this logic should be centralized further upstream and then be queried at the consumption layer. ## Consistency at Scale The dbt Semantic Layer acts as a middle layer between your data warehouse and Tableau dashboards that manages all the business logic and relationships across your data models through declarative semantic models. By situating the business logic in the dbt project, teams can centralize data change management. Furthermore, by being an extension of a dbt project, where all of your underlying data modeling logic lies, the dbt Semantic Layer benefits from the automated lineage, documentation, version control, and automated test inherent in the [dbt Cloud platform](https://www.getdbt.com/product/dbt-cloud). When upstream changes occur, updates are made in one place: dbt Cloud. The changes are then run with your scheduled dbt jobs and propagated downstream, without you needing to change anything in your data sources in Tableau. So, a best practice would be not to use Tableau for modeling your data when connected to the dbt Semantic Layer. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/46a6cdb18c16bfe674e76700afc8bbd482828891-1332x648.png) ## Key benefits of using the dbt Semantic Layer and Tableau The benefits of using the dbt Semantic Layer in conjunction with Tableau extend beyond simplified data change management. By centralizing the business logic in the dbt project, users can: 1. **Focus more on analysis and less on the mechanics of data management:** Your analysts can spend their time on what questions they need to answer for the business rather than spending time on preparing the data. 2. **Avoid duplicate work:** Your teams can work more efficiently since they aren't spending time reconciling redundancies across data sources and resolving consistency issues 3. **Foster data trust:** Our recommended approach enhances the overall integrity of the data that’s analyzed in Tableau. With dbt managing the transformations and your metrics being defined in code with a CI process and orchestration, analysts can trust that the data in Tableau reflects the most current and accurate information available, which is crucial for making informed business decisions. 4. **Avoid managing a large number of extracts:** The dbt Semantic Layer offers a live connection and your data will always be up to date when you query, reducing the dependency on extracts. ## Final Thoughts The integration of dbt's Semantic Layer with Tableau makes it far easier to manage upstream changes in data (an inevitable part of any organization’s evolution). Additionally, it alleviates the burdensome task of manual updates in response to updates, simplifies the data workflow, and speeds up time to insights from Tableau. Above all, it delivers that coveted “single source of truth” to you Tableau users, in fact, to everyone in your organization. To learn more about our dbt Semantic Layer and Tableau integration, you can access the [Tableau Exchange](https://exchange.tableau.com/products/1020) or the [documentation](https://docs.getdbt.com/docs/use-dbt-semantic-layer/tableau). --- --- title: "Funnel analytics and AI models for event sequences" description: "Misha Panko, co-founder and CEO of Motif Analytics, on moving from \"what\" questions to \"why\" and \"how\"" url: "https://www.getdbt.com/blog/funnel-analytics-and-ai-models" date: "2024-05-29" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # Funnel analytics and AI models for event sequences _This post first appeared in_ _[The Analytics Engineering Roundup](https://roundup.getdbt.com/p/ep-63-funnel-analytics-and-ai-models)._ Misha Panko has worked in data for a long time, including on high performance data teams at Uber and Google. Today, Misha is the co-founder and CEO of [Motif Analytics](https://motifanalytics.com/), a product focused on helping growth and ops teams understand their event data. In this episode, Tristan and Misha nerd out about the state of the art in computational neuroscience, where Misha got his PhD. They then go deep into event stream data and how it differs from classical fact and dimension data, and why it needs different analytical tools. It's possible to answer specific event-based questions using common tooling and practices today, but it always feels like you're only scratching the surface of the questions that you could ask. Make sure to check out the back half of the episode, where they dive into AI and how Motif is applying breakthroughs in language modeling to train foundation models of event sequences—[check out his team’s blog post on their work](https://www.motifanalytics.com/posts/foundation-models-product-analytics). **Listen and subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Google Podcasts](https://podcasts.google.com/feed/aHR0cHM6Ly9hbmFseXRpY3NlbmdpbmVlcmluZ3JvdW5kdXAubGlic3luLmNvbS9yc3M) - [Stitcher](https://www.stitcher.com/show/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) Key takeaways from this episode: ### Organizations often put a lot of the value from analytics on answering basic questions about the state of business, and even when questions are more exploratory, they are fairly simple. Where do you see differentiation in analytical value between companies like Google and Uber compared to early stage companies? **Misha Panko:** I think you might be surprised how much of Google and Uber are similar to smaller companies when it comes to where the value comes from. Lots of it is, is still about just getting basic facts and getting them to the decision makers when they are asking for it. The very typical loop is product manager working next to an analyst or data scientist. And the product manager keeps asking questions for a deck or to help analyze an experiment or for some dashboard, or just for a custom question. That has been my experience in almost any company ### My understanding is that at Matif, you're significantly less focused on that type of analytical flow and much more focused on a type of exploratory analytics that is very event stream based and helps you optimize the flow of customers through large scale digital systems. I was a product manager on data platform teams serving internally a lot ofdifferent growth teams and product teams. And often I needed to help frame analytics. One of the best lenses that I found was talking about analytics as a pyramid. At the bottom, when you establish analytics, people want to answer “what” questions—this is what I think of reporting. Usually they want to just know the number of users on a platform, the revenue, and they want to build dashboards and monitor that. That's lbread and butter of analytics that we usually think about. But the north star of analytics is to actually guide decisions in business. ### To guide decisions, you need to answer “why” questions and not “what” questions. Exactly. If I didn't have this data, would I make a different decision about launching this feature versus that feature or expanding into this country versus that country. I would probably make a decision some way, based on hunches, based on my mental model of the business. But if I have data coming in can it influence and change my decision? That's where the real value comes now. And when you start thinking about those types of questions, they are rarely just give me a number. They are more about the why. The next level of the pyramid is answering the why questions or root causes. When I'm looking at my metrics, why did it suddenly go down? The top level is where analytics is guiding and answering the “how” question. I don't want to just understand why metrics are moving, I want to understand how I can move my metrics in the direction that I want. It goes from what to why to how, but it's always a pyramid. You can't come into organization and say let’s just start guiding, because first you need to get the bottom layer down. And that's where I see 80 percent of the companies are. They are still building that foundation and that's where a lot of data and data tooling work is. We as a field are making a lot of progress with companies like dbt helping organizations get to the next level to answer why and how questions. ### In most data domains, you actually don't have sufficient data to answer why questions very effectively. That's a good point. I think that you're touching on a lot of things here, including causality. We're after causality. We want to find what drives the business. Now, as you say, causality is very hard. You usually don't have enough data. You also cannot isolate like an experiment very easily. That's why A-B testing is the gold standard, but you can't make every decision based on A-B tests. And so what can you do? I'm a practical guy and I help make practical decisions. PMs very often have to be threading the needle. On one end, you have the total causality of A-B tests. Now that's where you can make very definitive decisions based on causality. It’s rightfully became the bread and butter of how analytics needs to be run. Once you become mature enough, you set up an experimentation infrastructure in your company and run experiments. But even at huge companies like Google, Uber, only a small fraction of decisions, everyday decisions, end up being made on A-B tests. And that's because it is pretty expensive to run one, right? You have to actually build the feature that you are testing. And then usually it takes anywhere from a couple of weeks to a month to run the test, analyze the experiments, make a decision It's a slow process. So what do we have on the other end of the spectrum? It's correlations. We have an idea, A influences B, say, people seeing certain feature affects them subscribing, so let's just plot the correlation right between the two. Unfortunately, the mantra is correlation is not causation. And there's too many correlations that exist in the data. So can we do something better between correlation and this A-B tes? And this is something that we call causal opportunities or practical causality. And what it means is that you take into account the data that you have to try to de-confound, right? Or look at other possibilities of an explaination and still say that my hypothesis of A influences B still holds. It's a much stronger version of correlation. It's not fully causal, but it, what it allows you to do is to really narrow the space of potential ideas or hypothesis to work on. ### Confounders are such a powerful concept. There's some underlying characteristic that you're not able to directly observe, but it's kind of showing up in the data. How do you de-confound that? In general, there are statistical methods to do that. Usually you need to isolate all other variables that you're considering against and consider them together as a group and see if you consider all of them as potential independent variables affecting the dependent variable that you're looking at. If in the presence of this other confounder you still see the effect, then it's probably there. At least it's confounded against that other variable. ### This is the space that the Motif operates in. Broadly, you're trying to be this space between A-B testing, which is extremely expensive but gets you a tremendous amount of confidence around causality, and general reporting, which looks at correlations but doesn't really try to make statements about causality. What is the Motif product experience? The Motif experience is actually very different from traditional tools. We do a lot of education; we still don't allow people to just self onboard. We have a session with them first to explain the main concepts. We wanted to make the trade offs in a different way compared to how traditional reporting tools are doing decisions in analytics. In reporting, you usually pre-think of the questions that could be answered. In exploring what, what we wanted to unlock is you don't know questions a priori, right? You might have starting questions, but you don't know where they're going to take you, depending on what you find. And so what's critical is this fast feedback loop, and ability to touch a lot of data. That's why we go to events, the original raw events, that you have rather than standardized tables. If you're working with a very large dataset, that will be sampled. And then we use rich visualizations to show it to you. Now based on that, you start understanding, okay, well, is it going where I need to go? Do I need to modify my query a little bit? Or maybe I'm ready now to actually query to the end. It’s a three-step process: ask, compute, visualize. ### What types of questions are people asking in this interface? Do they use code to express them? Do they use visual experiences to express them? So you can ask the same questions that you ask with traditional tools. How many users did action A? Or how many searches per session does a user do? You start with those questions, but sequences also allow you matching patterns on sequences. You can start answering questions like, how often does A happen before B. How often do people look at this banner before they subscribe? You start going to these relationships questions. Does A precede B? How far is it before? How many times does it happen? This is becoming more of exploring what affects what. ### My experience doing event stream analytics is that you're sitting in front of this pile of data, and even if you have the ultimate technical capability to traverse this data, it is actually just challenging to know what hypotheses to test for. I would love for you to give your thinking around AI here. Quick stream data or event data could be overwhelming and you’ve got to narrow it down. In fact, that's usually why we see people come to us at Motif. They want to see how to approach this. They want to understand their funnels and figure out whether A affects B, but the space is too big or not well-defined. We start with what we call outcomes. Usually you are pretty good at knowing what you care about in the business. Are you trying to grow subscriptions, hours that people spend on the platform, things like that. That's your outcome. And now what I'm interested is looking at are the predictors. It's still pretty wide because you can look at all the events that happened before, at all the dimensions on that event, but also at any sub-sequences of events. To be able to go through all of that, you have to restrict it to the search space. And usually what people do when they go a little beyond correlations, they start building decision tree models, and then you want to test between them. You're creating a custom model to see if they're affecting the outcome. It takes a lot of time to construct these features. That's usually 80 percent of the work, not running the model, but getting the data in place. And then you're still working with very restricted space. So we're looking at ways you can expand that search, but still keep it under control. One way of doing it is just through exploration. It is giving you access to all of the events before. Where AI comes in, underneath this breakthrough LLMs technology is this transformers. And they happen to encode sequences very well. What we thought is what if we take that and expand it from just using it on words of sequences of words or tokens to using it on sequences of events with dimensions. You train your transformer based model on sequences of events and it understands inherently the structure, which events go together, which dimensions go together or not, and then you fine tune it based on the outcome that you care about. This transformer technology, this technologies behind LLMs, they happen to be good at modeling sequences and sort of compressing the information from sequences as just small embedding vectors.vAnd we are trying to use that for this use case of product analytics. ### What do you hope will be true in the analytics space in five years from now? I hope that we move beyond reporting. Even Google teams, Uber teams, they are still spending most of their time setting up reporting and trying to use data that way. While that's necessary, I think the most interesting part comes after that. So you have your data in place, and now I want to find the best way to find practical insights out of it. I work with, with growth teams. They ask a lot of these questions in terms of why and how do I change things. And I think the answers lie in exploring the relationships between different pieces of the data, rather than just counts. And I hope that the field moves more toward that. Tristan Handy has been curating the [Analytics Engineering Roundup newsletter](https://roundup.getdbt.com/) since 2015, pulling together the internet’s best data science and analytics articles. Tristan now brings the Roundup to real life, with biweekly conversations going deep into the hopes, dreams, motivations, and failures of leading data and analytics practitioners. --- --- title: "More scale, less chaos: dbt Mesh is now generally available" description: "Discover the power of dbt Mesh, now available in dbt Cloud Enterprise. Learn how this deployment pattern enhances data management for organizations, enabling seamless team collaboration at scale." url: "https://www.getdbt.com/blog/dbt-mesh-is-now-generally-available" date: "2024-05-28" authors: ["Azzam Aijazi", "Jeremy Cohen"] categories: ["Product"] --- # More scale, less chaos: dbt Mesh is now generally available At our recent [Product Launch Showcase](https://www.getdbt.com/blog/dbt-cloud-launch-showcase-2024) we announced that support for multi-project collaboration patterns, known as [dbt Mesh](https://www.getdbt.com/product/dbt-mesh), is now **generally available** for customers of dbt Cloud Enterprise. Since introducing dbt Mesh at [Coalesce last year](https://www.getdbt.com/blog/new-dbt-cloud-features-announced-at-coalesce-2023), we’ve seen a rapidly growing number of organizations adopt this deployment pattern across their many data teams and domains. In fact, we're thrilled to share that this number has just crossed one hundred — and it includes some of the largest enterprises in the world. The general availability of dbt Mesh support is a huge milestone in unlocking more seamless collaboration around data, at scale, for more teams and more organizations. Today we want to share why we're so excited about this, a few of the things we’ve learned from customers so far, and a bit about what’s coming next. But first… ## What the heck is dbt Mesh? **dbt Mesh** is a pattern for collaboration across multiple teams and domains of data. It is the culmination of several new dbt capabilities, and was motivated by a key product development priority: _helping data teams handle complexity at scale._ As a dbt project naturally grows in size and scale, navigating it successfully can become challenging. It’s not clear who is responsible for maintaining which models, and stakeholders’ trust in data (and the data team) suffers. Within the team, development bogs down, for fear of breaking downstream uses of dbt models. Onboarding to a monolithic project is too high a barrier for many less-technical contributors. This can be an ugly picture: the data team on the back foot, “shadow” analyses competing with vetted data products. At sufficient scale, one team of data and analytics engineers **cannot** do it all themselves — no matter their talent, no matter the tools on their utility belt. They need a safe, scalable way to invite others into the work of building and maintaining data products. The big ideas introduced by the dbt Mesh pattern are thus: - Teams should own their data, and as such their own dbt projects, from development through deployment. - Those teams must not operate in silos. - dbt models should define the interface for data sharing across teams. But not every dbt model is so worthy — these cross-boundary interfaces must be designated explicitly as such. - The teams maintaining those interfaces should treat them as mature APIs, with a stable set of guarantees that do not break without prior warning — and even then, a careful versioning and deprecation window. - The central data team must retain the ability to set global governance standards, and keep visibility on global lineage across teams. The dbt Mesh pattern allows teams to accomplish these things with a handful of dbt features, rolled out over the past year. Now, data teams can: - define explicit governance rules for dbt models (contracts, versions, access levels) - enable collaborative development with cross-project references - see cross-project lineage in dbt Explorer ![Before and after](https://cdn.sanity.io/images/wl0ndo6t/main/83824de20bb0cfe3ad120049f4e7eb4c37cca6f5-3160x1234.png) > _“dbt Mesh enables us to make data mesh a reality by offering a simple, cohesive way to integrate and manage data pipelines & products across the enterprise using a single platform.” **—Marc Johnson, Data Strategy & Architecture, Fifth Third Bank**_ ## What’s new Alongside the general availability of these capabilities, we announced some new product features that we’re confident will allow even more customers to adopt the dbt Mesh pattern in production. ### (New!) [Cross-project job triggers](https://docs.getdbt.com/docs/deploy/deploy-jobs#trigger-on-job-completion) Now your team can schedule and trigger jobs across projects, so they can own their data models from development through deployment, informed by their upstream dependencies. Beyond this, a major priority for us is to continue building extensions of dbt Cloud’s built-in orchestration, allowing it scale to more complex cross-project topologies. Look out for more on this over the coming months. ![Cross-project job triggers in dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/afe1358f780f6095bb8e45238b4b0bf93917ec91-1554x1156.png) ### (New!) [Staging environments](https://docs.getdbt.com/docs/deploy/deploy-environments#staging-environment) This is a new environment type in dbt Cloud that allows for improved data isolation. It enables developers to contribute in dbt Cloud — with all the cost-saving benefits of [more efficient development](https://docs.getdbt.com/docs/cloud/about-cloud-develop-defer) and cross-project references — but _without_ requiring access to production data. Using it allows organizations to more easily adopt dbt best practices and follow global governance policies. ![Staging environment support in dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/896097dbbb67d10641cb80f168765dec4d3feb59-1960x1036.png) We want to enable customers to adopt scalable, governed approaches to their multi-project deployments in dbt Cloud. To that end, we’re rolling out two related capabilities in the next few months: - **Environment-level permissions (coming soon):** Enable developers to edit and trigger jobs in specific environments (such as Staging), while restricting them from accessing others (Production). - **Environment-level connections (coming soon):** Need to use different authentication, or hit a different database/project/warehouse in your cloud data platform in Development vs. Staging vs. Production? Not a problem. ### (New!) Azure Single Tenant Support The dbt Mesh pattern is now supported for dbt Cloud customers deployed on Microsoft Azure — on Single Tenant deployments today, and Multi-Tenant deployments [later this year](https://www.getdbt.com/blog/dbt-cloud-launch-showcase-2024). We’re glad to offer the same compelling dbt Cloud experiences to more of our largest customers. > _“It’s now easier to ref the models from other projects, and we keep this in one single tool, without having to … identify the database/schema and create the source.” **— Renato Sousa, IT Solutions Expert, Siemens**_ ## What’s next Beyond these enhancements, we have a robust roadmap in place for improving the dbt Mesh experience even further. We’re refining a [bundle of enhancements](https://github.com/dbt-labs/dbt-core/issues/10125) to existing [model governance features](https://docs.getdbt.com/docs/collaborate/govern/about-model-governance) in dbt — making them more intuitive and capable, and making it easier for teams to manage and maintain their project interfaces. We’re exploring ways to provide a first-class cross-project orchestration experience that easily accommodates even more deployment architectures. And we’re thinking now about how to strengthen the connection between the dbt Semantic Layer and dbt Mesh concepts, so that data teams can easily curate and extend semantic objects across projects. As the kids say, “watch this space” for news on these and other product enhancements! ## How should you use dbt Mesh? dbt Mesh is not one feature — it’s a new way of working, and of architecting how data teams should work together. We take seriously our responsibility to provide you with both product experiences and hard-won guidance to **take “data mesh” from theory to practice.** To date, we’ve published a detailed [guide](https://docs.getdbt.com/best-practices/how-we-mesh/mesh-1-intro), [FAQs](https://docs.getdbt.com/best-practices/how-we-mesh/mesh-4-faqs), and a hands-on [quickstart](https://docs.getdbt.com/guides/mesh-qs?step=1). We're excited to update those resources in the coming weeks with lessons learned from successful Mesh implementations at customers large and small. ## Going forward Analytics engineering, at its heart, is about solving data problems and people problems. Tooling is a start, but it’s just a start. Any solution that’s going to be effective at tackling problems of collaboration and complexity around data needs to acknowledge these problems are _socio-technical_. The dbt Mesh pattern is our attempt to give teams a set of tools that they can adapt and employ to reflect _their_ organizational structure, instead of trying to force it to be the other way around. It’s still highly opinionated — there are patterns we encourage and discourage, best practices and anti-patterns — but we firmly believe this is the only kind of solution that will scale. This journey’s not done; it’s only just begun. Alongside our first hundred dbt Mesh customers, we will continue to innovate and to push the boundaries of what's possible with this pattern, supported by the platform we’re building in dbt Cloud. --- --- title: "Analytics engineering and data engineering: Do you need both?" description: "A look at the different problems that analytics and data engineering solve." url: "https://www.getdbt.com/blog/analytics-vs-data-engineering" date: "2024-05-23" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Analytics engineering and data engineering: Do you need both? More companies than ever are adding analytics engineers to their data teams. Should you follow suit? We’ll dive deep into what differentiates data engineering from analytics engineering—and when you should consider developing both practices side-by-side. ## What is data engineering? Data engineering is the engineering practice that designs, builds, and maintains the infrastructure required to store, query, and analyze data. It focuses on a few key responsibilities, including: ### Building new data infrastructure capabilities Data engineers build the data infrastructure on which all users depend. This includes creating data storage facilities such as databases, data warehouses, and data lakes. It also includes self-service tools for managing data, as well as the infrastructure for securing and governing data. ### Enabling and managing data pipelines This involves creating the foundation for all data pipelines—how data transformations are coded, orchestration pipeline execution, handling errors, monitoring, etc. (It may also involve creating the pipelines themselves, especially when an organization’s data practice is just getting started.) ### Building custom data integrations In other words, writing custom code to handle legacy internal systems, currently unsupported file types, or data from third-party APIs. ### Optimizing data storage and query execution Data engineers keep an eye on data read/write performance so that business users can focus on solving business problems instead of fine-tuning their SQL queries. ## What is analytics engineering? By contrast, analytics engineering focuses on providing clean data sets to the end users on their teams. Unlike data analysts, who analyze data, analytics engineers spend their time transforming, testing, deploying, and documenting data. An analytics engineer’s typical tasks include: ### Providing clean, transformed data Analytics engineers primarily create new data pipelines using standardized tools to provide data analysts and other business users with high-quality, relevant data for their business needs. ### Maintaining clean analytics code Managing data transformations at scale requires maintaining a clean codebase. This means [checking code into source control](https://docs.getdbt.com/docs/collaborate/git-version-control), versioning code and data [transformation models](https://docs.getdbt.com/docs/collaborate/govern/model-versions), creating and running [data transformation tests](https://docs.getdbt.com/docs/build/data-tests), ensuring code is [DRY](https://www.getdbt.com/blog/guide-to-dry) by bundling reusable components into [packages](https://docs.getdbt.com/docs/build/packages), and using [Continuous Integration and Continuous Deployment (CI/CD)](https://docs.getdbt.com/docs/build/data-tests) to automatically and safely ship data changes to production. ### Maintaining documentation and data definitions [Documentation](https://docs.getdbt.com/docs/build/documentation) helps data users and engineers understand what purpose data serves, where it comes from, and how it was derived. ### Training business users on tools Analytics engineers bridge the gap between data engineers and business users by using brown bags, one-on-one training, and other teaching modalities to show data analysts and others how to find data and leverage it in reports. ## Do you need both analytics engineering and data engineering? The short answer is: probably. The long answer involves explaining why the analytics engineering role exists in the first place. In the old days (we’re talking pre-dbt here), data engineers weren’t just responsible for data infrastructure. They also took requests from business users to create new data pipelines and address data quality issues. [As data volumes began to spike](https://www.statista.com/statistics/871513/worldwide-data-created/), this strategy proved untenable. Data engineering teams found themselves with months-long backlogs. And business users had to wait weeks or longer to get access to the data they needed. In the old days, this was, perhaps, a necessary evil. Data pipelines were complicated and often required specialized knowledge. Starting in the 2010s, however, we saw an onslaught of new data management technologies. Cloud-based data warehouses like [Snowflake](https://snowflake.com/), data pipeline services and tools like dbt, and easy-to-use Business Intelligence (BI) tools like [Looker](https://cloud.google.com/looker) and [Mode](https://mode.com/) meant more business stakeholders than ever could find, transform, and use data. The analytics engineering practice grew due in large part to this advanced toolset. Tools like dbt Cloud that make it easy to transform data using standard SQL or Python code made data pipelines both less complicated and less fragile. ‌This meant that analytics engineers could take on some of the workload that previously burdened data engineering teams. Introducing an analytics engineering practice provides numerous benefits for everyone in your data ecosystem, including: - Reduces data engineering backlogs - Improves data velocity - Improves data quality - Increases evolution of data infrastructure ### Reduces data engineering backlogs When all requests for new data projects or fixes must go through the data engineering team, that team becomes a bottleneck. We’ve seen this time again as our customers’ data needs skyrocket. Adding analytics engineering means adding a practice focused solely on creating new data pipelines and transformations. Instead of relying on data engineering for data changes, a team can hire an analytics engineer whose sole responsibility is managing the team’s data needs. ### Improves data velocity Data engineering team members juggle multiple responsibilities. That means there’s often a large lag between requests and fulfillment of that request. This can lead to longer-than-expected delays when rework is required. If a business user has an issue with something the central data team delivered, it means making yet another request—and returning to the back of the queue. By contrast, an analytics engineer’s sole focus is developing new data sets for their teams. That means they can ship changes to their users in a shorter timeframe than a central data team—usually days instead of weeks or months. ### Improves data quality Another issue with a centralized data team is that they usually don’t have domain expertise in a specific team’s data. Since they respond to requests from multiple teams, they sometimes have to make their best guesses when it comes to the contents of tables, the format of individual fields, and the calculations required for numeric fields. Analytics engineers, on the other hand, make it their jobs to understand a team’s business model and data needs in detail. This makes them more efficient at implementing their team’s requirements. Analytics engineers embedded with their teams can also get answers to hard questions more easily than members of a centralized data engineering team. ### Increases evolution of data infrastructure Data engineers can absolutely build data pipelines. That doesn’t mean it’s the best use of their time. Most data engineering teams would rather focus on providing new capabilities for analytics engineers and data analysts. When analytics engineers take on more data pipeline work, it frees data engineers up to focus on infra. That means data engineering teams can tackle burning issues such as [creating self-serve tools for initializing new data products](https://www.getdbt.com/blog/key-components-of-data-mesh-self-serve-data-platform), improving the company’s data governance capabilities, improving data query performance, and optimizing data pipelines to achieve the lowest possible cost. Since infrastructure improvements are available to the entire company, this work often has a high return on investment. For example, consider a change that reduces the time it takes to run CI/CD jobs to move data changes to production by 10 minutes. In a company with hundreds of data pipelines, this represents significant time and cost savings. ## Conclusion If your data engineering team is beset with long queues chock full of data pipeline work, it’s time to consider launching an analytics engineering practice. Analytics engineers can produce high-quality data pipelines in less time and with fewer delays. That frees up your data engineers to invest in new data platform capabilities that benefit the entire company. An analytics engineering practice is only possible if you have a solid data platform supporting it. With dbt Cloud as your data control plane, your analytics engineers have a standardized and cost-efficient way to build, test, deploy, and discover analytics code. Meanwhile, data consumers have purpose-built interfaces and integrations to self-serve data that's governed and actionable. Learn more about how to launch an analytics engineering practice with dbt Cloud—[ask us for a demo today](https://www.getdbt.com/contact). --- --- title: "Introducing dbt Assist: a copilot to accelerate dbt development" description: "dbt Assist is a set of AI-enabled workflows for common tasks in dbt. It helps you accelerate (not automate) them, allowing you to ultimately get even more done in less time." url: "https://www.getdbt.com/blog/introducing-dbt-assist" date: "2024-05-22" authors: ["Jason Ganz", "Dave Connors"] categories: ["Product"] --- # Introducing dbt Assist: a copilot to accelerate dbt development _dbt Assist — now in beta — lets you quickly generate documentation and tests in dbt Cloud with the help of AI._ Analytics engineering, at its heart, is about solving problems at the intersection of people and data. Crucially, it’s about [moving up the stack](https://www.getdbt.com/dbt-labs/values): taking care of problems that _used_ to require repetitive manual work for you so that you can focus on delivering even more value to the business. At the annual [dbt Cloud Launch Showcase last week](https://www.getdbt.com/blog/dbt-cloud-launch-showcase-2024) we rolled out a new set of tools for your toolbox that does _exactly_ this: [dbt Assist](https://docs.getdbt.com/docs/cloud/dbt-assist) (now in beta; visit [here](https://docs.google.com/forms/d/1B8txoOrJlfbjmCHTxtmRtNOXKQd86LVz1ajM_tHkRl8/viewform?edit_requested=true) if you’re interested in joining). dbt Assist is a set of AI-enabled workflows for common tasks in dbt. It helps you accelerate (not automate) them, allowing you to ultimately get even more done in less time. dbt Assist allows you to scaffold critical, but time consuming parts of the dbt workflow using AI: - Writing documentation - Generating data tests - Creating Semantic Models and Metrics (coming soon!) ### **Lets see it in action** **Writing documentation** The point of data is for people (and perhaps increasingly LLMs) to be able to use it — and to do that they need to understand it. By [documenting](https://docs.getdbt.com/docs/collaborate/documentation#adding-descriptions-to-your-project) your dbt project, you’re setting it up to be useful to your collaborators, be they human or machine. dbt Assist will generate a first pass attempt at documenting your models, which you can then augment with your own knowledge about your data. [Watch video](https://www.loom.com/share/daeec6a101544d48a8a16cc8e5c939a1?sid=9e6c809f-ec10-4941-8fde-6845c7db75ac) **Generating data tests** dbt [data tests](https://docs.getdbt.com/docs/build/data-tests) help you build confidence in your data. dbt Assist automatically generates a baseline set of tests for your data, making it easier than ever to ensure test coverage across your entire DAG. [Watch video](https://www.loom.com/share/5e33c1a1769c460fa34e8c9ae7e241ee?sid=c8a1fa1d-f526-43ee-ba9f-1cc0e651e227) **Creating semantic models and metrics (coming soon)** The dbt Semantic Layer allows you to standardize your business definitions so that you can query your most important data across any system and get consistent answers. Investing in the Semantic Layer allows you to develop [conversational analytics](https://docs.getdbt.com/blog/semantic-layer-cortex) built on top of it, which we see as a [critical interface](https://www.getdbt.com/blog/semantic-layer-as-the-data-interface-for-llms) for LLM systems. dbt Assist will soon allow you to get up and running with the Semantic Layer by helping build out a first pass of the [semantic models](https://docs.getdbt.com/docs/build/semantic-models) and metrics that power the Semantic Layer. Expect to see support for this in the near future! It’s important to note that as with virtually all AI products, there is an important caveat you should keep in mind: outputs should be verified and validated before you add them to your projects. This is a tool meant to speed up the scaffolding of your dbt projects, and we’ve found internally that it does do that, but it is not a replacement for knowing your data or performing code reviews. ### dbt 🤝 AI This is our first AI-enabled release in dbt Cloud. There are a few guiding principles that we’ve followed in the process of building dbt Assist. We aim to deliver AI-enabled development experiences that are: - **Tailored to the analytics engineering workflow.** The first iteration of dbt Assist is a “Swiss Army Knife” style approach — a number of specialized, vetted tools that accelerate specific parts of the dbt development workflow. - **Built on dbt best practices.** We’ve baked knowledge of [dbt best practices](https://docs.getdbt.com/best-practices) into dbt Assist. That means the code you’re generating is based off of our most up-to-date learnings on how to get the most out of your dbt projects. - **Context-aware.** When using dbt Assist, we dynamically include the relevant metadata about your project in the generation so that the system has access to all of the intelligence that _you_ have put into your projects. The usefulness of LLMs is bounded by the context they have available to them, and dbt Cloud is the place to go to get context about your dbt projects. Going forward, you can anticipate our AI offerings in dbt Cloud will continue to grow. Expect new and more advanced tooling: what you see in dbt Assist today is is the beginning of our efforts here. For a longer-term view of how AI is going to unlock value for data practitioners, read Tristan’s [recent thoughts](https://roundup.getdbt.com/i/144751042/ai-for-dbt) on how dbt Assist + other workflows will shape analytics work moving forward. *** dbt is about solving hard problems once, and then being able to move up the stack. In the olden days, testing and documentation were manual processes managed on an ad-hoc basis. Just like dbt has helped minimize manual work in other areas, dbt Assist lets you quickly create a first pass of tests and documentation, so that you can focus on your most important work, and spend less time writing YAML. This is the first step in our vision toward an AI-assisted development experience in dbt Cloud. We’re looking forward to seeing what it helps you accomplish. --- --- title: "Data products and data mesh: Leveraging both to simplify data management" description: "Used together, data products and data mesh can simplify how you manage data. Learn what they are and how they’re related." url: "https://www.getdbt.com/blog/data-products-data-mesh" date: "2024-05-21" authors: ["Kathryn Chubb"] categories: ["Product"] --- # Data products and data mesh: Leveraging both to simplify data management Many data engineering teams have struggled to keep up with the growing demand for data. Two related solutions have emerged to make it easier for teams to manage data at scale: data mesh and data products. A data mesh is an architectural pattern for scaling data systems and providing domain teams with more local, granular control of their data. Data products are the tangible outputs of this process - curated and managed data assets that serve specific business needs. Used together, the two can make it easier for enterprises to scale, discover, and manage high-quality data sets. In this article, we’ll examine how data mesh and data products work, how the two are closely connected, and the tools available for implementing and maintaining both. ## What are data mesh and data products? Even as software application teams have embraced microservice architectures, many data engineering teams have continued managing data in monolithic data stores. This has resulted in data engineering teams becoming the single point of contact for data requests. This approach has two major downsides. The first is that the data team becomes a chokepoint. Inundated with requests for new data or data pipelines revisions, the team quickly falls behind, unable to devote much - if any - energy to improving the organization’s overall data infrastructure. The second is that the quality of data decreases. A single centralized data engineering team often doesn’t understand the full business context behind a given data set. This can result in a disconnect between data producers and data consumers, resulting in inaccurate data. Consequently, this can propagate downstream to reports and dashboards, leading to misinformed insights and decision-making. Data mesh and data products are approaches to managing and packaging data that aim to make it easier to manage data at scale. ### What is data mesh? A [data mesh](https://www.getdbt.com/blog/what-is-data-mesh-the-definition-and-importance-of-data-mesh) is a decentralized data management architecture comprising domain-specific data. Instead of maintaining a single centralized data repository, data engineering enables teams to own the processes around their own data. A data mesh architecture implements four key design principles. These include: **Data as a product**. In a Data as a Product (DaaP) mindset, data teams think of data, not as a monolith, but as individual products that are built for the benefit of consumers. This means taking the needs of data users across the organization into account when designing data sets, as well as managing changes to avoid breaking existing users. **Domain-oriented data and pipelines**. In the data mesh pattern, business teams own their own data. Each team is responsible for maintaining its data stores, pipelines, and tests. This enables each team to define and enforce its own rules around its data. It can then provide this data to other teams via well-defined interfaces. **Self-service data infrastructure**. Requiring every domain team to build its own data infrastructure would be a waste of resources. It’d also create an impossibly high bar for many teams to hurdle. To enable data ownership, the data engineering team creates tools required to provision data stores, transform and clean data, verify data, and render analytics. **Security and federated governance**. The opposite of a “data monarchy” isn’t “data anarchy.” Federated governance enables teams to secure and classify their data to remain compliant with industry and governmental regulations. It also provides the organizations with the tools required to automatically monitor and ensure compliance in a distributed data environment. ### What is a data product? A [data product](https://www.getdbt.com/blog/key-components-of-data-mesh-creating-and-managing-data-products) is a data container or unit of data that directly solves a customer or business problem. An outgrowth of a Data as a Product mindset, a data product contains a usable unit of data - e.g., a table, a report, a Machine Learning model - as well as the metadata, pipelines, API contracts, and documentation required to produce and use it. Data products are defined by having a core set of attributes, including: **Discoverable**. Other teams can search and find them via some mechanism, such as a data catalog. This eliminates the wasteful overhead involved in finding and using existing data. **Addressable**. Each data set has a unique, labeled location that everyone across the organization can reference to identify it. **Trustworthy & truthful**. The data product contains information about who owns the data set, how often it’s updated, and where it sources its data from. This increases confidence in a data set by providing potential users with self-service answers to fundamental questions about data quality and provenance. **Self-describing**. Data products use metadata, documentation, and contracts to define the data’s format, its intended business purpose, and any other relevant usage information. **Interoperable**. Data products leverage standardized concepts and field across the organization. Additionally, they expose their information via standard query languages, data interchange formats, and Application Programming Interfaces (APIs) to enable data exchange across the org. **Secure & governed**. Data products encrypt their data ar rest and in transit, control access to data via role-based access control, and adhere to organization standards for classifying and protecting sensitive data, such as a customer’s Personally Identifiable Information (PII). ## The connection between data products and data mesh A data product is a key component within a data mesh, designed to make data easy to find, consume, share, and govern. It embodies the architecture's four key design principles: ### Data as a product When working in a data as a product mindset, organizations put data products through a development lifecycle. They view data itself as a valuable product that must be managed, curated, and delivered with the same rigor as software applications. Since high-quality data is fundamental to a successful data product, teams put a greater emphasis on ensuring quality and usability for data consumers. ### Domain-oriented data and pipelines In a data mesh architecture, business domain teams - the teams that truly understand their own needs and requirements - own their own data. This includes owning all of the associate processes, including ingestion, transformation, governance, and serving. In addition, data domain teams are responsible for maintaining data quality, versioning their data sets, and monitoring data usage and storage to reduce costs wherever possible. Packaging data as a data product - with its associated metadata, documentation, contracts, and interoperability mechanisms - makes achieving these objectives easier. Data products give teams a uniform structure for packaging, deploying, and discovering their data sets. ### Self-service data products It doesn’t make sense for every team that owns its own data to reinvent the wheel. Key data and basic data functions—e.g., the tools required to store data, create data pipelines, render analytics, etc.—should still be owned by the data engineering team. In a data mesh paradigm, the difference is that these tools are open and available to all data domain teams who need them. This open data architecture democratizes data by giving every team a consistent and reliable method for creating their own data products. ### Strong security and federated governance Data products are built with discoverability and governance as part of their design. Each team is responsible for associating access controls with its products and classifying and tagging data. Teams are also responsible for publishing their data products to a data catalog, where they can not only be discovered by other teams, but monitored centrally to audit and report on organizational compliance initiatives. ## Data product and data mesh tools Standing up a data mesh architecture and creating data products isn’t something that happens overnight. Fortunately, frameworks like [dbt Mesh](https://docs.getdbt.com/best-practices/how-we-mesh/mesh-1-intro) can make this easier with out-of-the-box tooling support for data mesh and data products. Here’s are the systems and processes that are part of a combined data mesh/data product architecture, along with the support that dbt provides: **Data sources**: The internal and external data storage systems - relational database, NoSQL databases, data warehouses, etc. - that hold your raw, unprocessed data. **Data pipeline and data transformation tools**. dbt enables [using SQL or Python to define models](https://docs.getdbt.com/docs/build/models) that transform source data into transformed and cleaned data for use in a specific data product. dbt also supports versioning model code in source control and automatically [running data quality tests using Continuous Integration (CI)](https://docs.getdbt.com/docs/deploy/continuous-integration). **Data destinations**. The locations - RDBMS databases like [MySQL](https://www.mysql.com/) and [PostgreSQL](https://www.postgresql.org/), data warehouses like [Snowflake](https://www.snowflake.com/) and [Amazon Redshift](https://aws.amazon.com/redshift/), data lakes, object storage, etc. - where you store data after it’s been transformed. **Data management tools**. Your architecture needs a mechanism to version data products, enabling teams to release breaking changes without interrupting the workflows of downstream consumers. dbt enables [explicitly versioning data products](https://docs.getdbt.com/docs/collaborate/govern/model-versions) via dbt models, as well as establishing [contracts](https://docs.getdbt.com/reference/resource-configs/contract) that specify the constraints to which a data product’s data commits. Additionally, dbt supports defining [fine-grained access controls](https://docs.getdbt.com/docs/collaborate/govern/model-access) on data, as well as using metadata to define data sensitivity and classification levels. **Data discovery tools**. Providing tools to find and use data products reduces the risk of data silos and enables organizations to gain maximum business value from their data. [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) gives data consumers a full, 360-degree view of organizational data, along with its corresponding metadata, documentation, and data lineage. **Data analytics tools**. Once exported to a data destination and exposed in a data catalog, a data product’s end users can use their favorite Business Intelligence and analytics tools to consume a team’s work. You can further use the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) to standardize key organizational metrics in a single source. **Provisioning architecture**. The data engineering team, instead of spending all of its time building data pipelines, builds tooling that ties all of these elements together into a self-service data infrastructure. This includes provisioning storage capacity, configuring access controls, and creating data pipeline frameworks and base data models from standardized templates. In terms of provisioning support, you can use dbt to modularize and centralize your analytics code. You can simultaneously provide your data team with the ability to collaborate on data models, version them, and test and document your queries before safely deploying them to production. dbt Cloud plays a critical role in this infrastructure by providing a centralized turn key platform for transformations and modeling. ## Conclusion Data mesh and data products work together to give business teams the ability to own and publish their own data products. That enables greater data scalability and reliability across the organization. By leveraging dbt, you can eliminate a lot of the architectural grunt work involved in standing up the infrastructure required to support this decentralized and democratized approach to data management. For more information on how to get started with dbt Mesh in your company, [contact us for a demo today](https://www.getdbt.com/contact). --- --- title: "Our biggest launch event yet" description: "May 14th saw us announce a ton of new capabilities and provides a window into the future." url: "https://www.getdbt.com/blog/our-biggest-launch-event-yet" date: "2024-05-20" authors: ["Tristan Handy"] categories: ["Product"] --- # Our biggest launch event yet _This post first appeared in_ _[The Analytics Engineering Roundup](https://roundup.getdbt.com/p/our-biggest-launch-event-yet)._ There is a lot to say today. I’m gonna dive right in. On Tuesday, May 14th, we held a launch event. We launched 15-20 new capabilities, depending on how you count. I couldn’t possibly go through all of these here, but you can [read all about them](https://www.getdbt.com/blog/dbt-cloud-launch-showcase-2024). I am incredibly proud of our team—we’re moving quickly innovating on an incredibly broad set of product capabilities, including dbt code authoring, code execution, data quality, semantic layer, enterprise and multi-cloud capabilities, data catalog, and more. Each innovation represents a meaningful extension of dbt’s ability to support the end-to-end analytics engineering workflow for companies from early-stage startups to the largest enterprises in the world. But this event was more to me than just a launch event. It was also a coming out party. In my intro, I talked about our vision of dbt as the “data control plane.” Here’s the slide: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/7c5374a843d24eebec046e48027886afa74d42e6-1720x768.png) If you’ve followed my writing in this space over the past two years, this is likely not a huge surprise to you. I’ve talked a lot about the consolidation trends in our industry, about how the “modern data stack” as originally conceived in all its buy-and-integrate-12-products glory just isn’t how this industry is going to play out over the coming decade. dbt has, for a long time, solved a broad set of customer needs: it has never lived exclusively inside of a single box in an industry diagram. All the way back in 2016 dbt shipped with data quality features. Back in 2018 dbt incorporated a lightweight data catalog. And if you look at the usage data, dbt’s data catalog is one of the most widely-adopted data catalogs in the world with tens of thousands of companies using it. The community intuitively understands that dbt doesn’t live in a single box. Years ago one dbt user replicated much of an entire product category with [a single dbt package](https://hub.getdbt.com/calogica/dbt_expectations/latest/). More recently, a startup [did much the same thing](https://hub.getdbt.com/elementary-data/elementary/latest/) in another category. **dbt created, and exists to facilitate, the analytics engineering workflow. That workflow spans the entire set of underlying activities, regardless of what Gartner decides a particular magic quadrant should be named.** Over the years we’ve extended dbt’s capabilities in orchestration, observability, and cataloging, and customers freaking love them. One of the best things about having an integrated platform is that new capabilities just … magically appear. There is no implementation, no integration, no nothing. No dbt customer had to do any work to get access to dbt Explorer, and now 1,400 companies are actively using it every single week. Given that it just went GA this week, that number is going to grow _fast_. Ultimately, our belief in this single-control-plane vision comes from a small number of first principles: - The analytics engineering workflow spans a broad set of activities, and analytics engineers need tooling that respects that inherent unity. Disparate tooling will never adequately achieve this. In simple / small environments, good enough may be good enough, but in large / complex environments it will not be. - Buyers see the problem solved by the data control plane as a single problem and not as a set of discrete problems. They see the problem as “make data work for my business.” - The 800-pound-gorilla vendors in the space all see the problem this way too and their products / roadmaps very clearly recognize this. - The underlying capabilities of many / most of these product categories are very similar. They are all powered by metadata. The experiences built on top of the metadata may differ, but most of the work is in building the underlying metadata platform. - AI will drive innovation across this entire set of experiences, and AI has a centralizing effect. You don’t want 12 AIs with narrow ranges of capabilities, you want one. This is why ChatGPT is multi-modal. - One of the required capabilities of any mature control plane will be being cross-cloud. When you add all of this up, there is a very clear conclusion. **Welcome to the future of the data control plane.** We are far from the only team who has recognized this, and the race is on to fully live up to this vision. As the [launches from Tuesday](https://www.getdbt.com/blog/dbt-cloud-launch-showcase-2024) hopefully indicate, our team is moving fast. Two other quick notes that are worth making about Tuesday’s launch events. ## Code vs Low-Code: Why not both? We announced a low-code / no-code experience for authoring dbt models. Here’s a screenshot: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0bee3b4b4c818d425f3513d6faaaad202e4d0aa2-1550x926.jpg) Interfaces like this have been around for a long time: I first used one back in 2003. I think there are good reasons why, if you spend a meaningful percentage of your professional life writing code, you should not want to use this type of interface. But! As analytics engineering becomes more pervasive, we should absolutely want more and more stakeholders to participate. And the inherent visual nature of this type of interface makes it extremely accessible and quick to learn. Sure, you can probably learn SQL from scratch in a couple of days, but you can learn how to manipulate data in a visual interface like this in five minutes. Accessibility matters. You can see how much we care about this in all areas of dbt—our mission is to _empower data practitioners to create and disseminate knowledge_, and the more people that can participate in that the better. There is a perception, however, that this type of visual tooling is directly at odds with our vision for analytics engineering. Here’s how one very awesome community member put it to me in a Slack DM: > _In one of the first few blog posts you wrote on the dbt blog and on every job posting on the dbt Labs Careers page, there's this statement:_ > > _We believe that:_ > - > _Code, not graphical user interfaces, is the best abstraction to express complex analytic logic_ - > _Data analysts should adopt similar practices and tools to software developers_ > > _Do these statements not contradict your recent focus on low-code development experiences?_ Certainly, we should update our job postings because I can see where this impression comes from. But I don’t think there’s any inherent tension at all between analytics engineering, mature software practices, and low code. **As long as the low code product writes code!** The most important thing about our just-announced low code editing experience is that **it both reads and writes well-formatted, readable dbt code**. This was a hard requirement for me in order for us to invest in this product capability, and we’ve absolutely lived up to it. - Every model you build using this editor is written to its own model file. - The code it generates is standard dbt-sql. - The code written largely complies with [our SQL style guide](https://docs.getdbt.com/best-practices/how-we-style/0-how-we-style-our-dbt-projects) and we’re continuing to make tweaks to make it even better. - Once ready, model code will get get checked into git and will participate in a standard pull request / CI/CD process just like any other code. - _Our low-code editor can_ _read hand-written dbt code and build an editor experience around it_. So the same model can be both worked on by hand and in the low-code experience by different authors. I am not interested in having us build some part of the dbt experience that totally throws out all of the fundamental principles of the analytics engineering workflow. We built our new low code editor in a way that acts as a good citizen in that workflow, bringing the maturity of analytics engineering into the hands of a brand new set of users. I could not be more excited about where this will take dbt and its community in the coming years. There’s a lot to do. If you’re interested in using the beta once it’s ready, sign up [here](https://docs.google.com/forms/d/1B8txoOrJlfbjmCHTxtmRtNOXKQd86LVz1ajM_tHkRl8). ## AI for dbt Speaking of code: it turns out that language models are quite good at writing code, including all types of dbt code. Documentation, tests, and more. We announced [a new copilot experience for dbt](https://docs.getdbt.com/docs/cloud/dbt-assist) and demoed two initial capabilities: writing tests and writing documentation. There will be plenty more to come in the future, including the ability to create models from scratch. When we originally went all-in on code as the underlying construct from which to build dbt, we didn’t anticipate the innovation in LLMs that we’re seeing today. But dbt’s design could not be more well-aligned with it. Every aspect of a dbt project is already captured in code, and enabling a large language model to accelerate your dbt development does not require major changes to the product or a major research effort. You should anticipate seeing continued progress on this front. There is a worldwide conversation going on about how AI will play out, and one of the dimensions of that conversation is about how it will affect the job market. Here is specifically what I believe about how AI will impact the role of data practitioners: 1. There have never been enough data practitioners to go around. This has always been a huge problem, both for practitioners themselves (who are overworked) and the businesses who employ them (who never feel like they have enough). 2. dbt already breaks down the problem of analytics engineering into very clear components: testing, documentation, modeling, scheduling, etc. Because there is already a clearly-defined framework, it is very possible to teach AI to do each of these tasks with a high degree of both accuracy and testability. 3. Because of (2) and based on early estimates from experiments, we believe it is possible to see efficiency improvements from 50-75% in the core tasks of analytics engineering. 4. Because of (1), these efficiency improvements will not cause businesses to invest _less_ in humans who do this work, but rather _more_ because they are getting so much more value. This is an example of [Jevons Paradox](https://en.wikipedia.org/wiki/Jevons_paradox). There is plenty more to say on this topic, much of which centers around the centrality of the semantic layer to the unfolding AI narrative (more [here](https://roundup.getdbt.com/p/semantic-layer-as-the-data-interface) and [here](https://docs.getdbt.com/blog/semantic-layer-cortex)). But that will have to be a topic for another day. If you’re interested in getting access to the just-announced beta, sign up [here](https://docs.google.com/forms/d/1B8txoOrJlfbjmCHTxtmRtNOXKQd86LVz1ajM_tHkRl8). ## Other stuff I’m paying attention to ### [**Sunsetting open source data-diff**](https://www.datafold.com/blog/sunsetting-open-source-data-diff) This is in itself a non-story: startup releases a subset of its features as OSS, gets some adoption, decides (for whatever reason) that the experiment didn’t work, discontinues support for the OSS distribution. Fine—this is a thing that happens every day, nothing anyone should be surprised about. What’s interesting is the bigger picture and how this tiny little instance points at something larger. OSS is hard. I constantly talk about OSS as “playing on hard mode.” You might get some real advantages out of your OSS strategy but you also create some real challenges. In a world where cash is free and growth rates are high, the advantages seem to outweigh the disadvantages. When that world shifts, that relationship can sometimes flip for an individual company/project. It is not an accident that Datafold launched this OSS project in 2022 and closed it down in 2024. There is a TON of OSS in AI right now. Infra, models, etc. It turns out that cash is cheap and growth rates are off the charts in AI. Awesome! If you love OSS, then these conditions are a great way to see a lot of OSS get built. Finally: open source software without active maintainership is not useful. It quickly gets bricked; software is a living thing. Who maintains OSS is a fascinating and interesting topic, and there are many workable models. But all come with tradeoffs. The truth is though that most OSS projects simply don’t have enough support to continue living on indefinitely and get abandoned. ### [**Sigma Raises $200m**](https://www.sigmacomputing.com/announcements/sigma-raises-200-million-in-series-d-funding) You can read the post yourself, but sounds like the team at Sigma has made some real progress on the business, which is (IMO) freaking fantastic. Great growth, some solid efficiency gains, put Sigma on a path towards a long-lived part of the data ecosystem! The longer any innovative data company stays independent, the longer they get to innovate and define their own road. I think the Sigma product is fantastic—have loved it since I think 2018-19?—and I’m excited to see them build the business and product to the next level. Congrats folks. ### [**Are Emergent Abilities of Large Language Models a Mirage?**](https://arxiv.org/abs/2304.15004) This paper from a year ago is something I just came across and it makes me question one of the things I thought I knew about LLMs: sudden emergence of properties at certain model sizes. Abstract follows. > Recent work claims that large language models display emergent abilities, abilities not present in smaller-scale models that are present in larger-scale models. What makes emergent abilities intriguing is two-fold: their sharpness, transitioning seemingly instantaneously from not present to present, and their unpredictability, appearing at seemingly unforeseeable model scales. Here, we present an alternative explanation for emergent abilities: that for a particular task and model family, when analyzing fixed model outputs, emergent abilities appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale. Specifically, nonlinear or discontinuous metrics produce apparent emergent abilities, whereas linear or continuous metrics produce smooth, continuous predictable changes in model performance. We present our alternative explanation in a simple mathematical model, then test it in three complementary ways: we (1) make, test and confirm three predictions on the effect of metric choice using the InstructGPT/GPT-3 family on tasks with claimed emergent abilities; (2) make, test and confirm two predictions about metric choices in a meta-analysis of emergent abilities on BIG-Bench; and (3) show to choose metrics to produce never-before-seen seemingly emergent abilities in multiple vision tasks across diverse deep networks. Via all three analyses, we provide evidence that alleged emergent abilities evaporate with different metrics or with better statistics, and may not be a fundamental property of scaling AI models. --- --- title: "What is data lineage, and why do you need it?" description: "Data lineage provides an end-to-end view of how data transforms and evolves throughout today’s complex modern data pipelines." url: "https://www.getdbt.com/blog/what-is-data-lineage" date: "2024-05-17" authors: ["Daniel Poppy"] categories: ["Learn"] --- # What is data lineage, and why do you need it? Organizations rely heavily on complex data pipelines to inform critical decisions. However, they also face mounting challenges from data sprawl. As data flows through a complex web of systems, transformations, and dependencies, it becomes increasingly difficult to understand the data's journey from source to destination—and, ultimately, its overall trustworthiness. This is where data lineage comes into play. ## What is data lineage? Data lineage is the process of tracking and documenting the journey of data from its source to its final destination. It provides a clear, end-to-end view of how data moves, transforms, and evolves throughout your organization. Data lineage captures information, metadata about data sources, transformations, and dependencies between data objects. This enables teams to trace errors, assess impact, ensure compliance, and maintain trust in their data assets. ## Why is data lineage important? As data pipelines grow ever more complex, organizations face significant challenges in managing and governing their data effectively. Data lineage is a core tool in the data governance toolkit. However, not every organization prioritizes it as a critical component of their data strategy. Data lineage is important because, without it, you are attempting complex data management with limited insight into your own data lifecycle. This makes life more difficult for everyone in your company that touches data: - [**For data engineers and analytics engineers,**](https://www.getdbt.com/blog/analytics-engineer-vs-data-analyst) the absence of data lineage makes their jobs significantly harder. When issues arise, they often have to spend hours or even days tracing data flows, identifying dependencies, and pinpointing the root cause of problems. This manual process is not only time-consuming and frustrating—it also takes them away from more strategic tasks aimed at delivering value to your business. - **Data analysts and business users** also feel the pain of poor data lineage. They rely on data to make informed decisions. But if your data's trustworthiness is in question, they’re left second-guessing their insights. This uncertainty can lead to a lack of confidence in data-driven initiatives and a reluctance to fully embrace analytics. ## Data lineage fundamentals At its core, it’s about capturing and documenting the life cycle of data as it moves through an organization's systems and processes. This means data lineage can help solve the challenges of complex data pipelines. Data lineage is like being a data detective. It involves following the clues (metadata) left behind as data flows from source to destination, through various transformations and dependencies. There are three key components to data lineage: ### Data origin This could be a database, a flat file, an API, or any other source where data is initially captured or ingested. ### Data transformations As data moves through the pipeline, it often undergoes changes and manipulations such as filtering, aggregating, joining, or applying business logic. Each transformation leaves a trail of metadata that gets mapped by data lineage. ### Data dependencies Data often flows through multiple systems and relies on other datasets or calculations, creating a complex web of relationships. Data lineage tracks these dependencies by showing how data from one part of the pipeline impacts another. By capturing these three components, data lineage creates a comprehensive map of your data's journey—kind of like a GPS for your data, allowing you to navigate the complex landscape of your data pipeline with ease. ## Key benefits of data lineage So why is this map so valuable? Let's consider a few key benefits: ### Transparency With data lineage, you have a clear, end-to-end view of your data flows. This transparency makes it easier to understand how data is being used, where it comes from, and how it's transformed. It's like shining a light into the black box of your data pipeline. ### Traceability When issues arise, data lineage allows you to trace the problem back to its source. It's like having a trail of breadcrumbs that leads you to the root cause of the issue. This traceability saves time and effort in debugging and ensures that problems are resolved quickly and effectively. ### Impact analysis Data lineage assists you in tracking and assessing the impact of changes to your data pipeline. For example, if you need to update a data source or modify a transformation, data lineage shows you exactly which downstream processes and reports will be affected. This impact analysis helps you plan changes more effectively and, maybe more importantly, avoid unintended data disasters. ### Compliance Data lineage is essential for demonstrating compliance with data governance policies mandated by government data handling and privacy regulations like [GDPR](https://gdpr-info.eu/) and [HIPAA](https://www.hhs.gov/hipaa/index.html). It provides an audit trail of how data has been used and transformed, making it easier to meet regulatory requirements and respond to audit requests. ## Implementing data lineage in data governance You can only reap these benefits once you have a solid data lineage practice built into your organization’s data governance. Here are a few key considerations for implementing data lineage: ### Granularity Determine the level of detail you need to capture in your data lineage. Do you need to track every single transformation and dependency, or can you focus on high-level flows? Are you using table-based lineage (which will show additions/deletions of fields at the table level) or column-based lineage (which will show changes to individual fields, such as a data type change)? ### Automation Manual data lineage tracking is time-consuming and error-prone. Look for tools that automatically capture and document data lineage as part of your data pipeline. ### Integration Data lineage should be integrated with your existing data management tools and processes. It should be easy to access and use for all stakeholders, from data engineers to business users. ### Scalability As your data pipelines grow and evolve, your data lineage solution needs to scale with them. Look for tools that can handle large, complex data flows and adapt to changing requirements. ## dbt Cloud for data lineage Gathering metadata and visualizing the relationships between data objects requires a solid set of tools. That’s where dbt Cloud can help. At its core, [dbt is a tool ](https://www.getdbt.com/)that helps data teams transform and manage their data in a more organized, efficient, and collaborative way. One of dbt Cloud’s key features is comprehensive support for data lineage that enables you to visualize and understand the relationships between your data models. dbt’s Cloud [built-in lineage features](https://www.getdbt.com/product/dbt-explorer) provide you with a bird's-eye view of the documentation and lineage of your entire data estate and a way to visualize and understand the relationships between your data models. ### Defining model dependencies with ref() and source() In dbt, you [define your data models](https://docs.getdbt.com/faqs/Models/create-dependencies) using SQL SELECT statements. But instead of referencing tables directly, dbt gives you two special built-in functions: ref() and source(). - The [**ref() function**](https://docs.getdbt.com/reference/dbt-jinja-functions/ref) is used to reference other models within your dbt project, telling you _this model depends on the output of that other model_. When you use the ref function, dbt automatically infers the dependencies between models. - The [**source() function**](https://docs.getdbt.com/reference/dbt-jinja-functions/source) is used to reference raw data sources, such as tables in your data warehouse. It's like acknowledging the starting point of your data's journey. By consistently using ref() and source() throughout your dbt project, you're creating a clear map of how your data flows from source to end result. ### Visualizing the DAG (Directed Acyclic Graph) With your model dependencies defined, dbt can now generate a visual representation of your data lineage. This is known as the [DAG (Directed Acyclic Graph)](https://www.getdbt.com/blog/guide-to-dags) — a diagram of data relationships and connections, displayed as an interactive web page. Each node in the graph represents a model, and the arrows between nodes represent the dependencies between them. This visual representation is incredibly powerful. It allows you to see, at a glance, how your data is transformed and how changes to one model might impact others downstream like the ripple effects of a pebble dropped into a pond. But the DAG is more than just a pretty picture. It's also a valuable tool for debugging and troubleshooting. When you encounter an issue with one of your models, the DAG guides you in tracing the problem back to its source. ### Enhancing data lineage with documentation and tests While the DAG provides a high-level view of your data lineage, dbt offers additional features to enrich and validate your understanding of the data. First, there's [documentation for your dbt models](https://docs.getdbt.com/docs/build/documentation). Good documentation for your dbt models will help downstream consumers discover and understand the datasets you curate for them. dbt provides a way to automatically generate documentation for your dbt project and render it as a website—creating a shared understanding of your data that anyone on your team can reference. It's like publishing a textbook for your data estate, explaining the purpose and functionality of each piece of the pipeline But documentation only goes so far. To truly trust your data, you need to test it. Fortunately, dbt makes it easy to test smarter, not harder, by defining and running tests on your models. dbt provides a simple way to define and run data tests to validate your transformations and catch potential issues early on as a built-in part of your workflow. ## Conclusion Data lineage is an essential component of modern data management. The ripple effects of poor lineage tracking lead to: - Analysts spending hours debugging transformation issues - Teams struggling to understand dependencies - Your organization making decisions based on what might be incomplete or incorrect data Without proper lineage tracking, a minor data discrepancy can cascade into a problem with significant business impact. By providing end-to-end visibility into the data pipeline, data lineage helps your organization overcome the challenges of complexity, ensure data trust, and make informed decisions. Just as a GPS helps navigate unfamiliar roads, data lineage acts as a guide through the complex landscape of data transformations and dependencies. With dbt’s intuitive approach to data transformation and built-in lineage features, you can create a data pipeline that is transparent, traceable, and trustworthy. But the real power of dbt's lineage features is how, by making your data lineage clear and accessible, you're empowering your entire team to work with data more effectively. Learn more about how dbt can help you deliver data products that people trust—ask[ us for a demo](https://www.getdbt.com/contact) today. --- --- title: "From Moneyball to Gen AI" description: "Tristan and Eric Avidon, a journalist at TechTarget, discuss the maturity of AI and analytics and more." url: "https://www.getdbt.com/blog/from-moneyball-to-gen-ai" date: "2024-05-16" authors: ["Kathryn Chubb"] categories: ["Insights"] --- # From Moneyball to Gen AI _This post first appeared in_ _[The Analytics Engineering Roundup](https://roundup.getdbt.com/p/ep-63-from-moneyball-to-gen-ai)._ Eric Avidon is a journalist at [TechTarget](https://www.techtarget.com/contributor/Eric-Avidon) who's interviewed Tristan a few times, and now Tristan gets to flip the script and interview Eric. Eric is a journalist veteran, covering everything from finance to the Boston Red Sox, but now he spends a lot of time with vendors in the data space and has a broad view of what's going on. Eric and Tristan discuss AI and analytics and how mature these features really are today, data quality and its importance, the AI strategies of Snowflake and Databricks, and a lot more. **Listen and subscribe from:** - [Spotify](https://open.spotify.com/show/4BKMMeVXk4jJnAQSqGSJvE) - [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-analytics-engineering-podcast/id1574755368) - [Google Podcasts](https://podcasts.google.com/feed/aHR0cHM6Ly9hbmFseXRpY3NlbmdpbmVlcmluZ3JvdW5kdXAubGlic3luLmNvbS9yc3M) - [Stitcher](https://www.stitcher.com/show/the-analytics-engineering-podcast) - [TuneIn](https://tunein.com/podcasts/Technology-Podcasts/The-Analytics-Engineering-Podcast-p1466362/) - [RSS feed](https://analyticsengineeringroundup.libsyn.com/rss) [Watch video](https://youtu.be/opXe3_UJ3NU) Key takeaways from this episode: ### Compare and contrast writing about analytics and data management versus the Boston Red Sox. I have to imagine that your fans now are far more passionate than those Red Sox fans. **Eric Avidon:** I first heard of analytics through sports really around, around maybe 2000 or so, 2000, 2001. You started hearing about these much more advanced statistics around 2000. ### The Billy Beane era. That's it. They started valuing different things. He created an advantage for his team, was able to create, develop a team on a much smaller budget than teams like the Yankees. The Yankees are spending four times as much as the A's were on payroll, but the A's were able to compete with the Yankees because they were ahead of the curve in discovering that certain things should be valued above other things. ### You've been at TechTarget for five years, covering business analytics and now data management as well. Is that, is that right? Yeah. How would you kind of define the boundaries around those areas? Business analytics is relatively straightforward. I mean, it's really statistical analysis at core and using statistical analysis to drive decisions. Data management is under the hood, it's the deep technology, it is complex stuff that every time I write a story, I'm going and looking up things that I wrote about two weeks earlier to describe some complex technology. When I first started covering analytics, one of the analysts that I frequently spoke with was a Forrester analyst, and he kept saying analytics is a mature technology at that point in 2020. And now I understand what he means as I'm now covering data management. In analytics, there's really not a whole lot new that's happening other than generative AI, of course, but it's more about enhancing what's already there. Whereas data management is always new stuff. And there's all these new vendors that are popping up with specialties and big vendors are buying up those startups to add those capabilities into what they can offer. And that's what I'm seeing the difference between those two. One is really moving fast and the other is enhancing what's there. Everything that leads up to that moment when you're actually visualizing the data and trying to make a decision based off of that data. I would call everything leading up to that data management because it's preparing the data for that moment. ### Generative AI, what's your read on how real this is today? People that log into their Tableau dashboard or their Power BI, whatever, are they mostly using Gen AI capabilities today, or are we not quite there yet? What I hear for the most part is that it's largely still theoretical. There are some tools that are starting to trickle out now. I would say over the last three months, a handful of vendors have made some tools generally available. They're more in the AI assistant realm where you can ask questions without having to write code. Tableau has one Gen AI tool that's GA and one that's now in beta testing. Microstrategy has a has a chat interface that's GA. ### We released our [State of Analytics Engineering](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024), and a shocking percentage of data practitioners surveyed were either currently powering Gen AI capabilities with their pipelines or have active projects in the works. Yeah, I'm hearing the same thing. Among the analytics vendors, I think you're getting the chat bot, the AI assistants that they're rolling out and you're not getting a ton of use yet, but yes, among the data management vendors, I think you're getting a lot of work to facilitate development of AI, genAI applications and models, and that customers are taking advantage of those. ### Are you a Gen AI bull or a bear? Do you think that the way that data people work in five years is going to be completely transformed? Five years, I don't know if I can project that far. I think vendors have to get a handle on accuracy before gen AI can really be everywhere. When I talk to vendors, they’ll have introduced some capability a year ago that's still in preview. And that's obviously not a typical release cycle. A typical release cycle is you release something in preview, and then three months later it's GA. I asked them what's taking so long, and it's about accuracy. It's about making sure that when someone asks a question of their data, they're not getting incorrect answers. ### You can almost leave gen AI aside when you talk about data quality. There's some problems that the data ecosystem over many decades has made progress on, and there are some problems that‌ we haven't made as much progress on. The ability to store and compute large volumes of data, we got way better at that. The ability to scalably build data transformation pipelines on top of that data; we've gotten way better at that. ### [When we survey folks, we find that quality is their number one issue that is a problem today](https://www.getdbt.com/blog/the-2024-state-of-analytics-engineering-report). I'm curious if you see anything in your coverage of the space ‌that makes you optimistic that this might be one of those problems that we solve as an industry. The more you emphasize something, the more it's going to be addressed. People weren't talking about data quality previously; data just gets passed up the line and used. I didn't hear a whole lot about data quality right when I started with Tech Target. I've definitely heard more now, but I think that's tied to AI as AI has exploded over the last year or two. I think there's been a heightened emphasis on data quality because if you’re suddenly relying on an automated system, what's going into that automated system had damn well better be good, or else what comes out of it is going to mess us up. ### What do you hope to be true of the data ecosystem in five years? I hope that all the stuff I've been writing about isn't BS. That I haven't wasted my time covering all this stuff. And that readers haven't wasted their time reading about it. And that vendors haven't wasted people's time promising it. And that it actually comes to fruition. One of the statistics you hear is that in organizations, 25% of people use analytics as part of their job, and that has been stuck there for two decades or so, maybe even more. The hope is that the natural language processing capabilities promised by Gen AI can really break through that. Maybe it's not going to be 100%, but maybe 75% of people don't need to ‌have extensive data literacy training and to use analytics because they can simply type something into a Google-like interface and get responses that can help them. Will we begin over the next tear or two to see that 25% creep up to 40%? Hopefully someone will do a study, whether it's Gartner or someone else, in early 2025, to see where that number is. Tristan Handy has been curating the [Analytics Engineering Roundup newsletter](https://roundup.getdbt.com/) since 2015, pulling together the internet’s best data science and analytics articles. Tristan now brings the Roundup to real life, with biweekly conversations going deep into the hopes, dreams, motivations, and failures of leading data and analytics practitioners. --- --- title: "Delivering data that works: the biggest new dbt Cloud features" description: "Read about the latest features coming to dbt Cloud." url: "https://www.getdbt.com/blog/dbt-cloud-launch-showcase-2024" date: "2024-05-14" authors: ["Tristan Handy"] categories: ["Product"] --- # Delivering data that works: the biggest new dbt Cloud features Today, dbt Labs held our first annual product launch virtual event: [the dbt Cloud Launch Showcase](https://www.getdbt.com/resources/webinars/dbt-cloud-launch-showcase). It was a jam-packed 90 minutes with executive keynotes, new product announcements, and demos that all centered around the theme of how dbt Cloud is helping teams deliver Data That Works. We spent the majority of the time talking through our latest innovations, which fell into one of three categories: - **Quality:** Introducing new and improved ways for dbt developers to build, test, and deploy high quality code, including our new AI copilot experience dbt Assist, unit testing, and more advanced CI capabilities - **Connections:** Building out the context of the dbt DAG to automatically include upstream sources and downstream exposures; additionally, extending our warehouse integrations across the Microsoft ecosystem (Synapse and Fabric) - **Collaboration:** Empowering more people to participate in well-governed data development workflows and the insights they drive via a new low-code UI and improvements to dbt Explorer and the dbt Semantic Layer Supporting all of these innovations is the dbt Cloud platform, which we continue to improve upon in order to make data workloads more reliable, performant, and scalable. With these new and improved ways to test, build, govern, catalog, and democratize data, we’re looking forward to continuing on our journey to help our customers deliver Data that Works. **[Get on demand access to the replay](https://www.getdbt.com/resources/webinars/dbt-cloud-launch-showcase), or keep reading for all the details!** ## ✅ Quality: deliver reliable data without compromising velocity _For data to be useful, it needs to be reliable and it needs to be delivered downstream as quickly as possible. We’re investing in new and improved ways for data developers to build, document, test, and safely deploy data transformations, so that data pipelines continuously hum and business stakeholders have confidence in the data they’re using to make decisions._ ### dbt Assist (AI) [AI is happening.](https://www.getdbt.com/product/ai) You can now use [dbt Assist](https://docs.getdbt.com/docs/cloud/dbt-assist), a new AI-powered co-pilot experience built into dbt Cloud, to boost your productivity and enhance data quality. dbt Assist allows you to quickly generate documentation and tests to augment your dbt models, helping you accomplish more in less time. If you're interested in early access to try dbt Assist in the Cloud IDE, you can register your interest in joining the private beta [here](https://forms.gle/s61CAx38Kjjx29hF8). ![AI assist feature in dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/f5a07ec6e051ef0b7af282ffc2cf82ca9b55f536-2000x1012.png) ### Advanced CI We also announced important enhancements to dbt Cloud’s [Continuous Integration (CI) capabilities](https://docs.getdbt.com/docs/deploy/continuous-integration), so data teams can ensure better data quality at any scale. With the upcoming “compare changes” feature inside of CI jobs, you’ll not only be able to verify that the code in a pull request (PR) _will build_ (which you can do today!) but also ensure that _what is being built meets your expectations_. When enabled, each CI job will include a breakdown of what’s being added, modified, or removed in your underlying data platform as a result of executing the job. ![Advanced CI feature in dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/52b361e685176a8968887e7913066672d5e1bbf3-4067x2190.png) Additionally, the compare changes feature will show a summary of the outcomes of your quality checks _inside_ the actual PR on your Git provider. All of this amounts to more seamless prevention of unexpected changes, a smoother QA process, and a better experience for downstream data product consumers. Stay tuned for the beta coming soon! ### Unit testing You can now use unit tests to validate the behavior of model logic _before_ the model is materialized in production. If a test fails, the model won’t build—saving you from unnecessary data platform spend, while improving data product reliability. To get access to unit testing, as well as other new features in the future, simply select “Keep on latest version” in your dbt Cloud jobs or environments. [Read our docs](https://docs.getdbt.com/docs/build/unit-tests) or check out [this blog post](https://www.getdbt.com/blog/building-a-data-quality-framework-with-dbt-and-dbt-cloud) to learn more about how to approach data testing in dbt. Unit testing is also available in [dbt Core v1.8](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.8). We take our responsibility as stewards of the dbt open source standard very seriously and are excited about this latest release. dbt Core v1.8 highlights include: - A new stable, decoupled adapter interface, giving adapter maintainers more control over when and how they ship updates - Project-level [behavior flags](https://docs.getdbt.com/reference/global-configs/legacy-behaviors): opt into recently introduced changes (which are disabled by default) or opt out of mature changes (which are enabled by default) - A new `--empty` flag for building schema-only dry runs, helping ensure your models will build while avoiding expensive reads of input data ### dbt Cloud CLI We’re pleased to announce that the [dbt Cloud CLI](https://docs.getdbt.com/docs/cloud/cloud-cli-installation), used by hundreds of organizations, is now generally available (GA). Develop anywhere using your code editor of choice, bolstered by the rich features of dbt Cloud including capabilities like [dbt Mesh support](https://www.getdbt.com/product/dbt-mesh), [defer to production](https://docs.getdbt.com/docs/cloud/about-cloud-develop-defer), and improved performance. And if you use VS Code, you can now also use the [Power User for dbt Core and dbt Cloud](https://docs.myaltimate.com/setup/reqdConfigCloud/) extension with the dbt Cloud CLI to bolster your productivity. [This blog post](https://www.getdbt.com/blog/a-closer-look-at-the-newly-launched-dbt-cloud-cli) from our CEO Tristan Handy dives deeper into the features and benefits of building in the Cloud CLI. Get started [using the dbt Cloud CLI](https://docs.getdbt.com/docs/cloud/cloud-cli-installation) today. ## 〰️ Connections: plug into everywhere your data is _For data to be useful, it also needs to be complete. That means you need a holistic view of your entire estate, with the ability to trace and orchestrate your workflow from source to metric to consumer. And to do it seamlessly and automatically. You also need your data workflow to be interoperable with the platforms you’ve invested in—whether that’s Snowflake, Azure, Databricks or anywhere else._ ### Automatic exposures We announced new end-to-end orchestration features that make your data workflow and dbt DAG more automated, context-rich, up-to-date, and cost effective. When you configure automatic exposures, your dbt DAG will automatically reflect the downstream dashboards that your models power. You can also trigger downstream dashboards to automatically refresh when new data is available upstream so your stakeholders are always making decisions from the latest data. These exposures are automatically accounted for throughout dbt Cloud, including in dbt Explorer, scheduled jobs, and CI jobs. ![Automatic exposures in dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/e9c931b5c390437db0de85256fdca7ac1c41a096-1536x1002.png) On our near term roadmap, we’re also incorporating the concept of active sources so that everything downstream _and_ upstream of your dbt workflow is synchronized, automated, and always up-to-date . Automatic exposures for Tableau is expected to go into beta this summer, with PowerBI exposures to follow. ### Microsoft integrations As dbt has become a standard for data transformation on the data warehouse, we have seen significant demand from the Microsoft community for the collaboration and productivity features of dbt Cloud with Azure Synapse and Fabric. We announced our [Microsoft Fabric integration](https://www.getdbt.com/blog/dbt-cloud-is-now-available-for-microsoft-fabric) as GA, and also launched our Synapse adapter in Preview. To get started, create a new dbt project in dbt Cloud and choose Fabric or Synapse as your data platform. ### Databricks OAuth Now generally available, dbt Cloud supports developer [OAuth with Databricks](https://docs.getdbt.com/docs/cloud/manage-access/set-up-databricks-oauth), providing an additional layer of security for dbt Enterprise users. When you enable Databricks OAuth for a dbt Cloud project, all dbt Cloud developers must authenticate directly with Databricks in order to use the dbt Cloud IDE, removing the need to manage credentials in dbt Cloud. With this addition, we now support data platform authentication via OAuth on Snowflake, BigQuery, and Databricks, and look forward to continue adding to the list. ## 👭 Collaboration: empower more people to access, trust, and use data _Data is a means to an end—and that end is always driven by the business. And so, more stakeholders need the freedom and ability to participate in data development workflows. Whether that means building and testing data models from their preferred development environment, querying consistent metrics from their favorite visualization tool, or building a collaborative data mesh architecture—you need a variety of governed inroads to that data to encourage more people to actually use it._ ### Low-code development environment We also unveiled a brand-new development experience we’ve been building in dbt Cloud: a low-code visual editor! The precision and flexibility afforded by code-based development is inarguable. _And,_ we also understand that many potential contributors to dbt development are precluded from doing so if they don’t have SQL expertise. But not anymore. Now, less SQL-savvy analysts will be able to create or edit dbt models through a visual, drag-and-drop experience inside of dbt Cloud. These models compile directly to SQL and are indistinguishable from other dbt models in your projects: they are version controlled, can be accessed across projects in a dbt Mesh, and integrate with dbt Explorer and the Cloud IDE. As part of this visual development experience, users can also take advantage of built-in AI for custom code generation where the need arises. ![Low code editor in dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/53e742642e7a3b6f5bf5b82c376ffd332d10ec32-2936x1660.png) This allows organizations to enjoy the many benefits code-driven development—such as increased precision, ease of debugging, and ease of validation—while retaining the flexibility to have different contributors develop wherever they are most comfortable: via the dbt Cloud CLI, the dbt Cloud IDE, or now, a low-code visual editor. Register your interest in joining the private beta [here](https://forms.gle/s61CAx38Kjjx29hF8). ### dbt Explorer Initially introduced at Coalesce 2023, we’re pleased to share that the foundational [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) experience—including column-level lineage—is now generally available! Over the past few months, we’ve launched [new dbt Explorer functionality](https://www.getdbt.com/blog/proactively-improve-your-dbt-projects-with-new-dbt-explorer-features)—including improved search and lineage, model performance analysis, and project recommendations—designed to help data teams understand, troubleshoot, and improve their data pipelines while also making it easier for stakeholders to discover and analyze trusted data. Today, over 1,400 organizations rely on dbt Explorer as a critical tool for their data workflows. ![Lineage and search in dbt Explorer](https://cdn.sanity.io/images/wl0ndo6t/main/9211804ead305636ed2f1e3d7ccfe4ee4a08c483-1772x1036.png) Looking forward, our vision is for dbt Explorer to help both data producers and consumers improve the ROI of their data and of their valuable time. That is, for dbt Explorer to be an indispensable command center for driving high quality decisions while keeping costs in check. To that end, we pre-announced a few new features coming to dbt Explorer soon: - **Visualize automatic exposures:** Easily (and automatically) enrich your lineage with context into the dashboards, teams, and use cases your models power. We are launching auto-exposures with Tableau, with PowerBI to follow. ![Automatic exposures in dbt Explorer](https://cdn.sanity.io/images/wl0ndo6t/main/f88da2598c84f8e6990778a1e2240c6b14864447-1426x892.png) - **Model query history:** Identify the relative popularity of your models based on usage queries so you can prioritize development and improve data trust. - **Tile embedding:** Surface data health signals, like freshness and quality, wherever stakeholders consume data to build trust and accelerate quality decision making. Also, stay tuned for a revamped dbt Explorer landing page that makes it even easier to find trusted, relevant models or identify pipeline issues. _“dbt Explorer has not only helped us pinpoint areas for code enhancement but also significantly improved our documentation practices. We have effectively mitigated data errors in the bronze/silver layer and can ensure a higher standard of data quality for our end consumers. dbt Explorer is an indispensable ally for any data-driven organization aiming for excellence in their analytics workflows.” **– Shravan Banda, Solutions Architect, World Bank**_ ### dbt Semantic Layer enterprise readiness features In planning our roadmap for the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), a top priority has been delivering what we call “enterprise ready” features—the types of things you’ve come to expect from your SaaS providers that give you the confidence to adopt and embrace a service at scale. We announced a number of enterprise features to the dbt Semantic Layer including: - **Access controls:** Get more granularity, control, and precision to semantic layer permissions with group-level and user-level controls. Soon, dbt Cloud admins will be able to create multiple data platform credentials and map them to service tokens for authentication, meaning that various departments (marketing, data, sales, etc.) have curated access to relevant and governed data. With user-level controls, you can reuse existing developer credentials for more fine-tuned permissions. Both are expected to go into preview by July. - **Caching:** We shipped [result caching](https://docs.getdbt.com/docs/use-dbt-semantic-layer/sl-cache#result-caching) a few months back, which is a useful way to improve load times and reduce compute costs for frequently-queried metrics. With the GA of [declarative caching](https://docs.getdbt.com/docs/use-dbt-semantic-layer/sl-cache#declarative-caching), you have even more power and control to configure the cache that is most relevant to your use case. Now, you’ll be able to “pre-warm” the cache using saved queries and significantly improve the performance of key dashboards or common ad-hoc query requests. Any query requests with the same inputs as the saved query will hit the cache and return much faster. - **SSO & PrivateLink:** We’re further unifying the semantic layer experience with the underlying dbt Cloud experience with the GA of SSO and PrivateLink (both now generally available). You can now develop against and test your dbt Semantic Layer in the Cloud CLI if your developer credential uses SSO. Additionally, dbt Cloud users who deploy with PrivateLink can now use the dbt Semantic Layer. - **[Tableau](https://docs.getdbt.com/docs/use-dbt-semantic-layer/tableau) and [Google Sheets](https://docs.getdbt.com/docs/use-dbt-semantic-layer/gsheets) integrations GA:** With the GA of these popular integrations, you now have the ability to ask better questions, faster. The self-serve GUIs are now on par with the semantic layer API capabilities and you can also save queries to enable collaborative workflows and improve time to insight. The Tableau interface also now supports important constructs like relative dates and parameters. - **MetricFlow improvements:** dbt Semantic Layer is powered by [MetricFlow](https://docs.getdbt.com/best-practices/how-we-build-our-metrics/semantic-layer-1-intro)—a flexible, SQL query generation tool—and we continue to make improvements to MetricFlow to make it even more powerful in helping teams collaborate around metrics. We announced new enhancements to MetricFlow—including metrics as dimensions, sub-day granularity, timezone support, and complex date joins—designed to give teams more flexibility and power as they build and consume metrics with increased velocity and accuracy. _“The dbt Semantic Layer gives our data teams a scalable way to provide accurate, governed data that can be accessed in a variety of ways—an API call, a low-code query builder in a spreadsheet, or automatically embedded in a personalized in-app experience. Centralizing our metrics in dbt gives our data teams a ton of control and flexibility to define and disseminate data, and our business users and customers are happy to have the data they need, when and where they need it.” ** — Hans Nelsen, Chief Data Officer, Brightside Health**_ ### dbt Mesh Support for [dbt Mesh](https://www.getdbt.com/product/dbt-mesh) —a pattern for collaboration at scale in dbt Cloud—is now generally available. The dbt Mesh enables data teams to make use of multiple, inter-connected dbt projects, each aligned to a domain team. Central data teams are able to maintain a view of global lineage and implement governance policies. This pattern improves speed, reliability, and governance relative to a single monolithic project, and it's been adopted to connect thousands of real-world dbt projects, including at some of the largest enterprises in the world. ![Data governance with dbt Mesh ](https://cdn.sanity.io/images/wl0ndo6t/main/9a6899e34b18b72eae23bf1f0b1b03b8d3fb1da1-1804x324.png) _“dbt Mesh enables us to make data mesh a reality by offering a simple, cohesive way to integrate and manage data pipelines & products across the enterprise using a single platform.” **—Marc Johnson, Data Strategy & Architecture, Fifth Third Bank**_ We also announced some new capabilities to streamline dbt Mesh adoption, including support for jobs triggering on completion across projects, as well as for staging environments as a canonical environment type for improved data isolation. Soon, we’ll be rolling out environment-level permissions and warehouse connections, along with support for sharing re-usable macros across projects. ## ⚙️ Platform _On top of the above three innovation themes, we announced a number of improvements to the underlying dbt Cloud platform to make it more reliable, performant, and flexible for the thousands of customers that depend on it every day._ ### “Keep on latest version” in dbt Cloud By providing a fully managed SaaS solution, dbt Cloud allows your team to focus their efforts on _building data products_ instead of maintaining the infrastructure required to run dbt. We recently introduced the ability to “Keep on latest version” in dbt Cloud, allowing you to receive fully vetted new dbt features and fixes in dbt Cloud continuously—without needing to manually upgrade dbt versions—saving your team valuable time and energy. We’re pleased to announce this is now generally available. ![Configuration for Keep on latest version in dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/a9fa4f6457a8a7a33d3533681dda5786c42f96b1-2726x1458.png) And the best news: we’ve also made some optimizations under the hood to significantly improve parse performance in dbt Cloud, cementing it as the most performant way to run dbt. These are available today to everyone running on “Keep on latest version.” To get started, simply select “[Keep on latest version](https://docs.getdbt.com/docs/dbt-versions/upgrade-dbt-version-in-cloud#keep-on-latest-version)” for all your environments and jobs in dbt Cloud, and you’ll be off the version treadmill for good: you’ll never have to upgrade dbt versions again. The engineering team at dbt Labs recently [wrote a blog post](https://docs.getdbt.com/blog/latest-dbt-stability) that delves into the rigorous processes we’ve put in place to ensure a stable, reliable experience for all of our customers that depend on this feature. ### Cell-based architecture Our new cell-based architecture will be the foundation of dbt Cloud going forward. This new architecture offers improved **scalability** (making dbt Cloud maximally performant regardless of the complexity of a customer’s deployment) and improved **reliability** (ensuring we can continuously deliver the same great experiences to _all dbt Cloud customers_ across all regions and deployments without risk to product stability). Our new cell-based architecture will be gradually rolled out to customers over the course of this year. ### Microsoft Azure support We’re striving to bring the same great dbt Cloud experience to all our customers, regardless of on which cloud they choose to deploy it. We announced that dbt Cloud will soon natively support deployment on Microsoft Azure, in addition to [the currently available option](https://docs.getdbt.com/docs/cloud/about-cloud/architecture) of deploying on AWS. Beta is coming soon. ## Thank you We’re really excited about this momentum, and as always, look forward to hearing your feedback! If you want to catch the replay of our launch event (you might even enjoy a Willy Wonka reference or two), you can find it [here](https://www.getdbt.com/resources/webinars/dbt-cloud-launch-showcase). --- --- title: "New dbt Cloud Enhancements Empower Organizations with Trustworthy Data At Scale" description: "Leader in the cloud data ecosystem announces AI and low code development capabilities, among other features, to improve AI and analytics workflows, and enable more people to build and leverage data." url: "https://www.getdbt.com/blog/new-dbt-cloud-enhancements-empower-organizations-with-trustworthy-data-at-scale" date: "2024-05-14" authors: ["Daniel Poppy"] categories: ["Press"] --- # New dbt Cloud Enhancements Empower Organizations with Trustworthy Data At Scale **PHILADELPHIA – **May 14, 2024 – [dbt Labs](https://www.getdbt.com/), the pioneer in analytics engineering, today announced dbt Cloud enhancements designed to help businesses turn data into a competitive advantage. As companies’ data volumes explode and the need for trustworthy, high-quality data increases, dbt Cloud is meeting the market need for streamlined data transformation across pipelines, workflows, and teams. “Accurate and timely data is crucial, which is why we’ve delivered a standardized way to quickly build reliable, holistic, and high-quality data pipelines at scale,” said Luis Maldonado, VP of Product at dbt Labs. “These new features take this even further, significantly improving data workflows and AI workloads, all while empowering more users with powerful business insights.” ## Improve the way data teams build, test, and ship high-quality data pipelines dbt Cloud includes a host of enterprise capabilities for delivering trusted data quickly, securely, and affordably. New features include: - **dbt Assist:** An AI-powered copilot experience, dbt Assist automatically generates documentation and tests to let data developers get more done in less time (beta). - **Advanced CI:** A new “compare changes” view lets teams verify that changes to the codebase meet quality expectations before they are merged into production (beta coming soon). - **Unit testing:** A new feature that allows teams to improve test coverage without driving up data platform spend through earlier validation of modeling logic (generally available). - **dbt Cloud CLI: **Offers developers the flexibility to contribute to projects in dbt Cloud through their terminal or IDE of choice (generally available). ## Seamlessly trace and orchestrate data pipelines in more platforms Data teams rely on dbt Cloud as a control plane to catalog, orchestrate, govern, and observe their end-to-end data workflows. New platform enhancements and integrations give dbt Cloud even more context into the dashboards and decisions that dbt models power. New capabilities include: - **Automatic exposures:** Gives dbt Cloud automatic awareness of Tableau dashboards downstream of dbt models, allowing users to trace and automate end-to-end data lineage to unlock efficiencies, optimize compute costs, and improve data freshness and trust. Auto-exposures are incorporated throughout dbt Cloud including in dbt Explorer, orchestration workflows, and CI jobs (beta coming soon). - **Microsoft integrations: **dbt Cloud now supports Microsoft Azure Synapse (preview) and Microsoft Fabric (generally available). ## Empower more stakeholders to collaborate in the data workflow In order for data to be a true competitive advantage, it needs to be accessible across the organization. Stakeholders—of various technical aptitudes—now have more avenues to engage with dbt Cloud to build, improve, and trust the outputs of data workflows. This is made possible through: - **Low-code development experience: **A drag-and-drop visual editor that generates SQL, which lowers the barrier to entry for more contributors to collaborate on the analytics engineering workflow in dbt Cloud (beta). - **dbt Explorer enhancements: **Introduced at Coalesce 2023, dbt Explorer is an intuitive, interactive catalog for data teams to understand, improve, and troubleshoot their dbt assets across teams and projects (existing capabilities, including column-level lineage, now generally available). Soon, users can do more with dbt Explorer using enriched lineage and auto-exposures, telemetry into model consumption to align development work with business impact, and embedded data health signals in analytics tools for trusted data delivery at scale (all beta coming soon). - **[dbt Semantic Layer](https://www.prnewswire.com/news-releases/dbt-labs-announces-the-next-generation-of-the-dbt-semantic-layer-introduced-alongside-new-integration-with-tableau-301958939.html) enhancements:**_ _Includes enterprise-critical features such as granular access controls and permissions (preview coming soon), Tableau and Google Sheets integrations (generally available), declarative caching (generally available), and improvements to MetricFlow that allow teams to build and consume complex metrics with more velocity and accuracy (generally available). - **Multi-project support: **Allows organizations to manage complexity by supporting multiple inter-connected dbt projects aligned to individual business domains, instead of a single monolithic project. Support for this pattern, known as dbt Mesh, is now generally available. “dbt Cloud allows us to take all the data we’ve collected and actually make it useful to the business,” said Evan Cover, Director BI Engineering & Governance at Klaviyo. “With dbt Cloud at the center of our transformation workflows, we can build data products that represent the reality of our business objectives and model how we go about selling, attracting, marketing, and retaining customers. By enabling more people across the business to collaborate on building trusted data products, dbt Cloud allows us to work faster, more efficiently, and take more advantage of our data.” As a trusted platform for mission-critical data workloads, dbt Labs continues to invest in the performance, reliability, and scalability of dbt Cloud. Recent dbt Cloud platform improvements include: - **Managed dbt versions:** With no manual dbt upgrades required, dbt Cloud users have early and continuous access to vetted features and fixes as they become available (generally available). - **Performance improvements**: Now offers significantly improved parse performance, solidifying dbt Cloud as the most performant way to run dbt (generally available). - **Microsoft Azure support:** dbt Cloud will soon be available on Microsoft Azure (beta coming soon). For more information on dbt Cloud, visit: [https://www.getdbt.com/product/dbt-cloud](https://www.getdbt.com/product/dbt-cloud) ## **About dbt Labs** Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 40,000 companies using dbt every week. To learn more about dbt Labs, visit [getdbt.com](https://www.getdbt.com/) and follow us on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/). --- --- title: "Understanding semantic layer architecture" description: "How a semantic layer provides a single interface that transforms technical data structures into user-friendly business concepts." url: "https://www.getdbt.com/blog/semantic-layer-architecture" date: "2024-05-06" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Understanding semantic layer architecture Today's data-driven business environment brings [exponentially growing data volumes](https://www.statista.com/statistics/871513/worldwide-data-created/) and increasingly complex analytics needs. The semantic layer serves as the crucial bridge between raw data complexity and business usability, providing a single interface that transforms technical data structures into business-friendly concepts. Organizations are pushing for greater data democratization and self-service analytics capabilities. More than ever, they need a standardized and governed way for their teams to access and interpret data. Without a robust semantic layer, companies often struggle with conflicting metric definitions, redundant data transformations, and a bottlenecked data team. When different departments calculate key metrics in multiple different ways, it leads to confusion, inefficient decision-making, and potential business risks. Let’s take a look at the semantic layer's core concepts and key features. We’ll hone in on how it solves these challenges by providing a standardized, governed, and user-friendly way to access and interpret data, no matter where it lives. ## What is a semantic layer? In modern data architectures, the semantic layer is the abstraction layer that sits between your raw data sources (like data warehouses, lakes, or operational databases) and your business intelligence or analytics tools. Functionally, a semantic layer is a standardized framework that organizes and abstracts your organization’s data, whether structured or unstructured, in a single point of access for everyone in your company who uses data in their day-to-day work. It’s the bridge that connects your end users with every type of data asset—both numerical and text-based, as well as media files, videos, presentations, etc. For example, your organization can create a uniform definition for calculating, say, Active Users. You can store this definition in the semantic layer and reuse it across all reports and dashboards. This eradicates the commonly occurring problem scenario where different teams each create and use their own varying definitions for the same metrics, leading to inconsistent and even conflicting reports and decisions. ### What is the purpose of a semantic layer? As the part of your data architecture (or data stack) that provides a consistent, up-to-date, and easily understood representation of your organization’s data, the semantic layer enables self-service analytics while maintaining data governance. Think of it as a universal translator that transforms complex, technical data structures into broadly comprehensible terms and concepts. Instead of dealing with cryptic table names or complex joins, business users can work with logical descriptors like `Customer Lifetime Value` or `Product Margin`. Data users no longer need to think about, much less understand, the underlying data architecture. This is crucial because it allows non-technical stakeholders to easily access and work with data. ### The role of the semantic layer in the modern data stack In the modern data stack architecture, the semantic layer plays an increasingly crucial role as organizations deal with more complex data environments. It sits between your data storage layer (like Snowflake, BigQuery, or Redshift) and your visualization tools (like Tableau, Power BI, or Looker). What makes it particularly valuable in modern architectures is its ability to work with multiple data sources simultaneously, handle real-time and batch data processing, and integrate with various modern data tools through APIs. Whether your data team uses dbt for transformations, Airflow for orchestration, or various BI tools for visualization, the semantic layer can provide a consistent interface for all these tools while maintaining performance and scalability. ## Core components of semantic layer data architecture The semantic layer has five core components. These act as the structural and technical building that define how the system is constructed and how it operates. Think of these as the "infrastructure"—semantic model definitions, metadata management, business logic layer, data access layer, and caching mechanisms. These components control the underlying mechanics of how data is processed, stored, and accessed within the semantic layer. #### Semantic model definitions Creates a logical representation of your business domain, mapping technical database structures to business concepts. Well-designed semantic models significantly reduce the complexity for business users while maintaining the technical rigor needed for accurate reporting. For instance, rather than working with raw tables like `usr_tbl` or `trx_hist`, you define entities like `Customer` or `Order` that encapsulate the underlying complexity. These models also include relationships between entities, like how Customers relate to Orders or Products to Categories. #### Metadata management Essential for maintaining context and understanding within the semantic layer. This component handles information about your data, such as field descriptions, data lineage, update frequencies, and quality metrics. For example, when defining a metric like `Revenue`, the metadata would include not just the calculation logic but also information about which source systems the data comes from, when it was last updated, who owns the definition, and any caveats about its usage. This comprehensive metadata makes the semantic layer self-documenting and helps users understand the context of the data they're working with. #### Business logic layer Where you define calculations, transformations, and business rules that convert raw data into ‌business metrics that are meaningful to your company. This is where you'd implement complex calculations like `Customer Lifetime Value` or `Product Margin` using standardized formulas that can be reused across the organization. The beauty of centralizing this logic is that when business rules change, you only need to update it in one place, and all reports using that calculation will automatically reflect the new logic. #### Data access layer Manages how different users and applications interact with the semantic layer. It handles important technical aspects like query generation, optimization, and security enforcement. When a business user requests information through a BI tool, this layer translates their business-friendly request into optimized database queries, applies appropriate security filters (like limiting access to certain regions or departments), and ensures efficient data retrieval. A well-implemented data access layer is crucial for maintaining performance as your data volume and user base grow. #### Caching mechanisms Vital for maintaining performance and scalability in your semantic layer. These mechanisms store frequently accessed data or pre-calculated metrics to reduce database load and improve response times. For example, if many users are frequently checking the `Monthly Revenue by Region` metric, the semantic layer can cache these results (updating them periodically) rather than recalculating them for each request. Modern caching implementations often include smart invalidation strategies that ensure users always see fresh data when needed while maintaining fast query response times. ### How the semantic layer works These five core components work together to create a robust semantic layer that can scale with your organization's needs. That said, the success of a semantic layer often depends on how well these components are integrated and maintained. Let’s see how these core components work together in practice. When a business user makes a request—let's say they want to see `Monthly Revenue by Region` in their BI tool—these components interact in a choreographed sequence. The **semantic model definitions** first provide the framework for understanding what `Revenue` and `Region` mean in business terms. This triggers the **business logic layer**, which contains the specific calculation rules for revenue. The **metadata management** component provides context about the freshness of the data and any relevant business rules or caveats. As this request flows through the system, the **data access layer** translates this business request into optimized database queries, applying any necessary security filters (like limiting certain regions based on user permissions). Before executing the query, it checks with the caching mechanism to see if this calculation is already available in cache. If it finds a valid cached result, it returns that immediately; if not, it executes the query (and, potentially, caches the result for future use). The interaction between these five components creates a seamless experience where business users can work with familiar concepts while the semantic layer handles all the complex orchestration behind the scenes. At the same time, it also maintains consistency. Whether a metric gets accessed through Tableau, Power BI, or any other tool, these components work together inside the semantic layer, which applies uniform business rules, security policies, and optimizations. ## Key features of the semantic layer Working together, the five core components enable five functional capabilities—the key features of the semantic layer. The key features transform your data into meaningful insights. Features like metric definitions, dimensional modeling, data governance, business glossary integration, and version control are the tangible benefits that business users experience. ### Metric definitions and calculations Represents the standardized way of defining business-critical measurements. Instead of having multiple teams calculate `Customer Acquisition Cost` differently, the semantic layer provides a single, authoritative definition. This means whether a sales analyst in New York or a marketing manager in London runs a report, they'll see the same calculation methodology. These definitions typically include complex logic like time-based filters, weighted averages, or rolling calculations that would be challenging to replicate across multiple tools consistently. ### Dimensional modeling Transforms complex relational data into intuitive, business-friendly structures. This feature creates hierarchical relationships between business entities, allowing users to drill down or roll up data easily. For example, a revenue metric might be viewable at company, region, department, and individual product levels. The dimensional model provides a consistent navigational framework that makes data exploration more intuitive, breaking down complex data relationships into understandable paths. ### Data governance and security Protect sensitive information while maintaining accessibility. The semantic layer acts as a centralized control point that implements where access permissions, data masking, and compliance rules. This means you can define granular access controls—like allowing a regional sales manager to see their region's data but not competitor or corporate-level details—without modifying the underlying database structures. ### Business glossary This feature bridges the communication gap between technical and non-technical team members by creating a common language across the organization. By embedding business definitions directly into the semantic layer, you eliminate ambiguity. You might define `Active Customer` as "A customer who has made a purchase in the last 90 days," and this definition is consistent across all reports and analyses. ### Version control for semantic models Version control treats data definitions like software code, allowing teams to track changes, roll back to previous versions, and collaborate more effectively. This is crucial for maintaining data integrity and understanding how business definitions evolve over time. Developers have long used [Git](https://git-scm.com/) to track code changes. Similarly, data teams can now track how their semantic models have been modified, who made specific changes, and when those changes occurred. Each feature contributes to making the semantic layer more than just a technical component. They elevate it into a strategic asset that improves data quality, accessibility, and organizational understanding. ## Semantic layer architecture in the real world What does semantic layer data architecture look like in the real world? A sample scenario illustrates how these key features work together in a practical implementation. Imagine a global e-commerce company called TechGear that sells electronics across multiple regions. Let's walk through how their semantic layer might be implemented ### Metric definitions at work TechGear’s `Customer Lifetime Value` (CLV) metric is defined with complex logic: total revenue from a customer over their entire relationship, minus acquisition costs, adjusted for inflation and weighted by the recency of purchases. Before implementing a semantic layer data architecture, different teams in the org were calculating CLV in different ways. The marketing team used a three-year window, while the finance team used a three-year window. By centralizing this definition in the semantic layer, they now have a single, consistent calculation that everyone uses. ### Dimensional modeling example Next, with the semantic layer’s dimensional modeling capability, TechGear teams can explore the CLV metric across multiple dimensions: - Geographic: Compare CLV in North America vs. Europe - Product category: Analyze CLV for smartphones vs. laptops - Customer Segments: Break down CLV by new customers, repeat buyers, and enterprise clients. Each of these views uses the same underlying calculation but allows different stakeholders to gain insights relevant to their role. ### Governance and security scenario When TechGear’s regional sales manager for EMEA logs into the semantic layer UI, they automatically see: - Full sales data for European countries - Masked customer personal information - Restricted from viewing global corporate financial details - Access limited to the last three years of data ### Business glossary in action: When definitions for all of TechGear’s business terms are embedded into the semantic layer, a new team member can quickly understand key terms. For example, - "Active Customer" is clearly defined as "Made a purchase in the last 90 days" - "High-Value Customer" is precisely calculated as "Customers with CLV above $5,000" - Terminology is consistent across all reports and dashboards ### Version control demonstration: What if the finance team decides they need to change the CLV calculation? They can work together in the semantic layer to alter the CLV as necessary, while version control lets them track changes and (if things go wrong) roll back to the original version. The process looks like this: - Finance team creates a new branch of the semantic model - Stakeholders can review proposed changes and run side-by-side comparisons with the existing calculation - Once approved, the new CLV calculation goes into production—with full traceability of who, when, and why the modification occurred The result is a powerful, flexible system that transforms raw data into meaningful business insights while maintaining consistency, security, and governance. ## Semantic layer architecture business benefits Implementing a semantic layer into your data stack can solve multiple business challenges. Giving your people a single source of truth while providing self-service analytics and creating a common language across the organization can directly impact both operational efficiency and business decision-making across your entire organization. Here are some of the business benefits: - **A single source of truth** is perhaps the most crucial benefit. With a semantic layer, metrics are defined once and used consistently across all reporting and analytics, eliminating confusion and ensuring everyone makes decisions based on the same information. - **Self-service analytics **becomes much more feasible because business users can access and analyze data using familiar business terms without needing to understand SQL or complex data structures. For example, a sales manager can quickly build reports without requesting help from the data team. This dramatically reduces the time from question to insight and frees up technical resources for more complex work. - **Reduced data redundancy** leads to significant cost savings and improved efficiency. Instead of having multiple teams maintaining similar calculations and data transformations in different tools, everything is centralized in the semantic layer. This reduces storage and computation costs and minimizes the risk of errors and inconsistencies. - **Improved communication.** When everyone speaks the same data language and uses the same definitions, meetings become more productive, cross-functional projects run more smoothly, and decisions can be made faster with greater confidence. - **Faster time to insight **happens when new analytics projects leverage existing, well-defined metrics and data models rather than starting from scratch. For instance, if a new executive dashboard is needed, it can be built quickly using pre-defined metrics and dimensions. This is faster and more reliable than recreating complex calculations and validating data transformations. ## dbt's semantic layer [dbt Cloud’s semantic layer](https://www.getdbt.com/product/semantic-layer) fits naturally into the modern data stack workflow, especially for teams already using dbt. Released in 2022 and continuously evolving, dbt's semantic layer integrates directly with your dbt projects. The core value of dbt's semantic layer lies in its seamless integration with existing dbt workflows and its ability to create a single source of truth for metric definitions. Our semantic later is built on top of [dbt's transformation framework](https://www.getdbt.com/product/dbt-cloud). That means data teams can define, version, and maintain metrics right alongside their data models, using familiar YAML syntax and Git-based version control. It’s a particularly powerful option thanks to dbt's [MetricFlow engine](https://www.getdbt.com/blog/how-the-dbt-semantic-layer-works). MetricFlow handles the complex work of generating optimized queries, managing time-based aggregations, and ensuring consistent metric calculations across all your BI tools and applications. This means whether someone is viewing a revenue metric in [Tableau](https://www.tableau.com/), [Looker](https://cloud.google.com/looker?hl=en), or any other connected platform, they're getting the same number calculated the same way. Plus, since it's tool-agnostic, you're not locked into any particular BI platform, giving your organization flexibility as your needs evolve. Learn more about how dbt Cloud can bring the power of modern data architecture to your organization—schedule[ a demo ](https://www.getdbt.com/contact)today. --- --- title: "DataOps: How to get started" description: "Discover how to apply DataOps with dbt to streamline workflows, improve data quality, and drive better collaboration." url: "https://www.getdbt.com/blog/dataops-dbt-guide" date: "2024-05-06" authors: ["Kathryn Chubb"] categories: ["Product"] --- # DataOps: How to get started [DevOps](https://www.atlassian.com/devops) emerged years ago as a methodology that aimed to break down the silos that had grown between software engineers and IT engineers. The goal was to accelerate the pace of software deployments through shared ownership and automation. DevOps is now recognized across the software development industry as a success story. However, the data industry has remained far behind our software development counterparts in embracing a DevOps mindset. Too often, the people who create data pipelines and data models and those who consume that data don’t have the tools or visibility needed to scalably collaborate with velocity. This results in time-consuming and expensive rework that delays the shipment of new data products and impairs trust in data and data teams. The good news is, it doesn’t have to be this way. In this article, we’ll look at the DataOps development pattern, how it’s patterned after DevOps, the value it offers, and some of the tools you can use to implement it. ## What is DataOps? DataOps is a framework for managing data that removes [silos](https://www.databricks.com/blog/data-silos-explained-problems-they-cause-and-solutions) between data producers (the creators of data products), and data consumers (the users of data products). In DataOps, data producers work closely with data consumers in short, rapid deployment cycles to design, develop, deploy, observe, and maintain new data products that align closely with data consumers’ evolving needs and business goals. DataOps recognizes that a successful data project requires the joint expertise of both data producers and data consumers. Data producers are experts in data technology - integrating and transforming data, storing data efficiently, securing access, etc. Data consumers, on the other hand, are experts in what they _need_ from the data and how they can use it to execute winning business strategies. For example, assume that a Finance team needs a new data set to drive reports tracking sales trends. In a pre-DataOps world, they might log a request to a data engineering team, delivering a final product based on the limited information in the support ticket. The Finance team discovers that the data is missing important inputs or parameters, or varies significantly from the metrics in a similar data pull last month, so they don't trust it or use it and they shoot the request back for fixes. This repeats over days or weeks. Meanwhile, the Finance team lacks the reporting it needs to drive key business decisions. In a DataOps approach, the Finance team and data engineers would meet to discuss the Finance team’s needs in detail. This would include the data required, its shape, the format and correct calculation of fields, allowable values, and importantly, how they intend to use the data outputs to make strategic decisions. The data engineering team would then develop a data product that pulls in all the correct sources, models the data appropriately, and tests and documents the new data models. As in DevOps, the team would use automated deployment tooling such as CI/CD to verify, test, and release the new data product to the Finance team. ## Benefits of DataOps There are many benefits to a DataOps approach—for both data product development and the organization as a whole. **Development value of DataOps**. Like DevOps, DataOps accelerates data development timelines. It does this by focusing on business value from the beginning, aligning teams from across different domains. It also gives data teams the tooling and resources they need to translate raw data into actionable insights in a way that's automated, modular, and tested. It streamlines development cycles by scoping them into smaller cycles that deliver value with every release. **Organizational value of DataOps**. Organizationally, DataOps erodes the silos between data teams and their business stakeholders. This increases collaboration and knowledge sharing, resulting in a better final product and more strategic business decision-making. The rapid releases and automated deployments associated with DataOps reduce bottlenecks, delivering more business value in less time. ## How to implement DataOps So, how does DataOps work? We can think of it as encompassing five phases: - Plan - Build - Deploy - Monitor - Catalog As with DataOps, these phases shouldn’t be considered long, drawn-out projects that take months from conception to completion. Instead, teams often work within an agile development framework, defining a short timespan of work (organized into “sprints”) at the end of which they deliver something of value to stakeholders. The process repeats, with each new iteration adding additional functionality and fixes in response to stakeholder feedback. ### Plan In the planning phase, the data team works with stakeholders to understand what they need, the format in which they need it, how quickly they need it updated, where the data will be sourced from, etc. It’s at this stage that both teams also set various Key Performance Indicators (KPIs) and Service Level Agreements (SLAs) around things such as data quality, data freshness, query performance, etc. ### Build In the build phase, the data team creates the data sets they’ll deliver to stakeholders. This involves building: - Data models - Data transformations - Tests to gauge and certify quality - Documentation describing the data, its business purpose, and how it’s calculated Making tests and documentation a part of the build phase is a hallmark of DevOps and DataOps. Tests ensure that data meets business requirements as well as KPIs and SLAs around quality and performance. Documentation ensures that stakeholders know precisely what business purpose a model serves, where its data comes from, and how it’s calculated. This makes data quality an integral part of the development process rather than an afterthought. ### Deploy In the deploy phase, the data teams push their data set changes out of local development and through a series of environments, testing their changes at each stage to ensure they behave as expected. The goal is to ensure that every change is rigorously tested and validated before it’s pushed to production and made available to data consumers. For example, a team may release a change to a Staging environment, where an automated process runs the data team’s tests on sample data and collects metrics on data quality and query performance. If all tests and metrics checks pass, the process may push to Production, where the team will run the same tests on real-world data. The combination of automation and version control means that teams can both deploy and roll back changes as needed. If a team identifies issues, say, with their Version 2 release in production, they can roll the deployment back to Version 1 easily. This enables stakeholders to continue using the system while the engineering team addresses critical issues. ### Monitor Once deployed into production, data consumers can self-service access to the data, using it in their reports and data products. During this time, the data team will continue to observe metrics, logs, and traces as data updates flow in, responding to any identified data anomalies or performance issues. ### Catalog All of these workflows create extensive metadata that gives stakeholders insight into how data products are built, used, optimized, and debugged. This metadata can be visualized and explored in a catalog, where the data team ensures that its work is discoverable and documented. This might entail: - Making data sets available for discovery in a data catalog - Generating data lineage charts to document a data set's sources, dependencies, owners, and relationships to other data sets - Publishing data models and documentation so other teams can consume their work Cataloging helps reduce data silos and redundant data transformation work. Before embarking on a new data project, a team can search a catalog to discover if another team has already created a high-quality data that they can build. **** ## DataOps with dbt At dbt, we’ve long believed in and supported DataOps as a methodology. That’s why dbt supports DataOps through every step, making it easier to run your business on trusted data. ![DataOps lifecycle diagram illustrating continuous data development, testing, deployment, analysis, and observability in data workflows.](https://cdn.sanity.io/images/wl0ndo6t/main/4a5682d44d9f8b63ca3f5da115bcb613c3e7dc61-1618x854.png) ### Plan: dbt Mesh dbt offers several built-in features that help users plan and align their data projects: - **dbt Mesh: **[dbt Mesh](https://www.getdbt.com/product/dbt-mesh) enables companies to implement a [data mesh](https://www.getdbt.com/blog/what-is-data-mesh-the-definition-and-importance-of-data-mesh) architecture - a decentralized, scalable approach to data management. With dbt Mesh, data teams can design their data workflow architecture to support the unique needs of their downstream stakeholders with domain-specific data, and do so in a governed, automated, simplified way. This enables them to work independently with other teams without sacrificing collaboration, governance, or security. - **SLAs and data freshness: **dbt also provides helpful interfaces for [source data freshness](https://docs.getdbt.com/docs/build/sources#snapshotting-source-data-freshness) calculations. These interfaces are designed to help users determine if source data freshness is meeting pre-defined SLAs. ### Build: Various development environments dbt gained popularity as a tool that allowed developers to build analytics code using SQL. Given the varied skills and preferences of data collaborators within an organization, dbt supports many environments for authoring analytics code: - **Cloud IDE:** [dbt integrated developer environment (IDE)](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/develop-in-the-cloud), a web-hosted IDE that includes SQL syntax highlighting, auto-completing, code linting, documentation, and build/test/run controls to run and debug work on demand. - **CLI: **Developers can build analytics code directly in their [preferred CLI](https://docs.getdbt.com/docs/cloud/cloud-cli-installation) and use dbt command line tools to write, run, and debug model changes. - **Visual editor: **Soon, less SQL-savvy analysts will be able to create or edit dbt models through a visual, drag-and-drop experience inside dbt. These models compile directly to SQL and are indistinguishable from other dbt models in your projects: they are version-controlled, can be accessed across projects in a dbt Mesh, and integrate with [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) and the Studio IDE. As part of this visual development experience, users can also use built-in AI for custom code generation where the need arises. ### Deploy: Scheduling jobs, version control and CI/CD dbt includes an in-app job scheduler to automate how and when you execute dbt jobs. To improve development feedback loops and optimize data platform consumption, users can also “defer to production” for any job run, meaning that when they want to run and test changes to a single model, dbt will build only that changed model (and defer any upstream dependencies to what’s already in prod). A major key to implementing DataOps is tracking changes to work. Without proper change tracking and control, a rogue alteration to source code that causes an error in production could take hours or more to hunt down and fix. dbt Labs' automated testing and version control streamline data pipeline automation, ensuring teams spend less time on manual processes. dbt supports [version control](https://docs.getdbt.com/docs/collaborate/git-version-control) via Git so that every change made to a model is committed, documented, and - if needed - reversible. Using Git, data team members keep their versions of dbt models for development. When they’re ready to commit changes, they create a pull request (PR) that one or more other members of the team can review. This ensures every change receives a second set of eyes before heading towards prod. You can also configure dbt with [Continuous Integration (CI) jobs](https://www.getdbt.com/dbt-learn/lessons/pull-requests#1) to push changes from dev through prod. Once a PR is approved, it triggers a CI job, which runs and tests models in a staging (pre-production) environment before moving them to production. Once a set of changes is verified in staging, dbt can push them to production, a process known as Continuous Deployment (CD). This combined CI/CD process automates code integration, testing, and delivery of updates to production, identifying and eliminating potentially costly errors before they’re shipped to data consumers. The result is reduced manual labor for deployments, resulting in accelerated data development cycles. By incorporating CI/CD for data, dbt helps organizations maintain data quality and consistency across deployments. ### Monitor: Automated testing With dbt, data teams can proactively define assertions—called tests—about their data models. These tests can be designed to validate the behavior of model logic _before_ the model is materialized in production ([unit tests](https://docs.getdbt.com/docs/build/unit-tests)) or about any assertion you want to make about your model (is unique, is non-null, etc). If a test fails, the model won’t build—saving you from unnecessary data platform spend, while improving data product reliability. In addition to setting up tests to proactively catch issues, it’s easy to monitor your production dbt jobs and alert the right people when something goes wrong with Slack or email notifications, logs, run history dashboards, data health tiles, and more. ### Catalog: dbt Catalog Both data developers and consumers alike benefit from having an understanding of data dependencies, freshness, use cases, and other relevant contexts. [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) is an interactive data catalog that represents the metadata created in every dbt run in an intuitive, visual interface. Using dbt Catalog, consumers can find a data asset and view it in context, complete with its metadata, documentation, and data lineage. Data producers can use dbt Catalog to find reusable data assets, as well as to trace lineage to troubleshoot data issues resulting from upstream data defects. The nice thing is that, rather than an afterthought in the DataOps process, the assets viewable by dbt Catalog - models, metadata, documentation, data lineage, security controls, etc. - are all automatically generated during the data development process itself. Updates are pushed automatically to dbt Catalog with every push to production. With dbt, teams can implement collaborative data workflows that reduce bottlenecks and empower faster decision-making across departments. ## Conclusion As a methodology, DataOps can eliminate barriers between data producers and data consumers, resulting in faster data development cycles and higher-quality data. Using dbt, data teams and shareholders can make the DataOps culture part of their daily workflows, building a data framework that combines speed and governance with distributed ownership. See dbt in action and learn how it can support your DataOps journey—[try dbt free today](https://www.getdbt.com/signup) (no credit card required). --- --- title: "Let's talk about AI and its impact on data practitioners, with as little hype as possible" description: "Tristan Handy reflects on AI's evolving role in data analytics, highlighting the current limitations and future potential." url: "https://www.getdbt.com/blog/lets-talk-about-ai" date: "2024-05-02" authors: ["Tristan Handy"] categories: ["Insights"] --- # Let's talk about AI and its impact on data practitioners, with as little hype as possible _This post first appeared in_ _[The Analytics Engineering Roundup](https://roundup.getdbt.com/p/lets-talk-about-ai)._ I’ve been thinking a lot about AI of late. This is a change—I have mostly resisted this for the past ~16 months. Sure, I’ve mused about it in this space from time to time, and I’ve discussed it internally with employees of dbt Labs in many contexts. But AI represents a fairly modest (though growing!) percentage of our product roadmap, and it just isn’t my very biggest near-term priority for the world of data analytics on structured data—the world that you and I live in. I think that will change in the coming years (and there are some fun [seeds getting planted](https://docs.getdbt.com/blog/dbt-models-with-snowflake-cortex)!), but today I think that’s mostly where we are. But first I want to build a foundation and share some thinking on several different topics. I’ll start with customers’ priorities for their data capabilities, hit a bunch of AI developments, and circle back around to connect the threads and how dbt will play into this world. **TL;DR:** the integration of structured data and AI will be driven by metadata. And dbt’s biggest role to play in the AI revolution will be as the source of truth for that metadata. Let’s dive in. ## First: Customer Data Priorities I just spent a full day with some of our biggest customers last week. Here are some things I heard from that conversation that totally check out with me: 1. Their biggest priorities are focused around code quality, data platform spend, end-to-end integrated user experiences and reduction of vendor sprawl, scalability (i.e. how to add more humans to the process of creating and disseminating knowledge), observability, and data trust. These 100% map to my own beliefs about current practitioner priorities throughout the industry. 2. If AI can help with the above priorities that’s great. Currently, it doesn’t seem like there is much of an intersection between these priorities and AI, however, so our customer panel was _mostly_ not involved in current AI projects at their orgs. 3. However, there was universal agreement that these folks would *love* to find AI-powered solutions to their problems, in part because that would give them access to special budgets focused on AI experimentation and deployment. All of this checks out and is healthy. Organizations really do want to push themselves to find nails for this brand new hammer. When done to extremes, this behavior is damaging, but it is important within large companies to push innovation top-down. If you don’t, you will by default get stasis. But the current challenges these folks are facing don’t feel addressable with current AI tooling. The question is: is that a persistent state—i.e. is AI just not that relevant to analytics problems? Is it just too early yet? Or are there some upstream unlocks that need to happen first? ## Second: How Complete is the World Model? Lex Fridman had a [great podcast episode](https://www.youtube.com/watch?v=5t1vTLU7s40) last week with Yann LeCun, the Chief Scientist at Facebook’s AI Research group (FAIR). Yann is probably in the top 5 humans in the world in his knowledge of current AI capabilities and ability to predict future trends; the entire 2+ hour interview is incredible. Ben Thompson at Stratechery [recently wrote about](https://stratechery.com/2024/sora-groq-and-virtual-reality/) Sora, OpenAI’s video-generation capabilities that have been integrated into the ChatGPT experience. Both LeCun and Thompson point out the same thing: the underlying world model created by the transformer architecture is simply not sophisticated enough, not good enough at reasoning through causal chains, to perform certain tasks. Here’s one seemingly [trivial example](https://openai.com/sora) of that: > The current model has weaknesses. It may struggle with accurately simulating the physics of a complex scene, and may not understand specific instances of cause and effect. For example, a person might take a bite out of a cookie, but afterward, the cookie may not have a bite mark. AI can do a lot of things that you can’t do already, but my 4-year-old knows that when you take a bite out of a cookie, the resulting cookie should…have a bite out of it. LeCun points out that what’s going on here is that the transformer architecture, despite the amazing progress it’s created in the past ~6 years, is not really appropriate for this type of world-model-building. And that the reason that we have such fantastic results on top of generative _language_ models instead of on top of models that need to interact directly with the real world (or simulations thereof) is that language is an already-compressed information stream, whereas real-world environments contain dramatically more data in a totally uncompressed state. Models dealing with the open world need to do far more work to compress the vast amounts of input data, create a world model from it, all before attempting to interact in a meaningful way. And we’ve mostly failed at this larger problem so far. Maybe passing the bar (today’s achievement) is something like winning at chess—a task that represents what we have always considered to be a Very Hard Intellectual Thing. In fact, however, it turns out to be far simpler than the things we do without even thinking about them (like reasoning about the state of half-eaten cookies). In this view, transformers and current-generation LLMs have moved very quickly and demonstrated tremendous capability, but aren’t going to be naturally extended to solving other classes of problems. Meaningful net new innovations in model design are required. Now, we can’t say this for certain. One of the ways the industry has come to communicate uncertainty about the path to artificial general intelligence (AGI) is to say “We’re between 0 and X fundamental breakthroughs to get to AGI” where X is lower for optimists and higher for pessimists. (My X is probably 4.) And so maybe the best way to think about the path to AI + analytics is to communicate what we would need to see to get there. My answer to that is: _reasoning about causality_. The core job of a data practitioner is to look through tons of data to answer questions about causality. Not to say “this is true” but rather, “this data suggests that this _could_ be true.” An AI that could do this consistently and effectively would be a massive change to the practice of analytics. Which gets us to… ## Third: MAMBA and JEPA If [attention isn’t all you need](https://arxiv.org/abs/1706.03762), and if transformers are not the only end-state architecture, what else is being worked on? From the [MAMBA abstract](https://arxiv.org/pdf/2312.00752.pdf): > Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module. Many subquadratic-time architectures such as linear attention, gated convolution and recurrent models, and structured state space models (SSMs) have been developed to address Transformers’ computational inefficiency on long sequences, but they have not performed as well as attention on important modalities such as language. We identify that a key weakness of such models is their inability to perform **content-based reasoning**, and make several improvements. And from the same Lex podcast I mentioned earlier: > **Lex Fridman**[(00:28:38)](https://youtube.com/watch?v=5t1vTLU7s40&t=1718) (…) can JEPA take us to that towards that advanced machine intelligence? > > > **Yann LeCun**[(00:29:02)](https://youtube.com/watch?v=5t1vTLU7s40&t=1742) Well, so it’s a first step. Okay, so first of all, what’s the difference with generative architectures like LLMs? So LLMs or vision systems that are trained by reconstruction generate the inputs. They generate the original input that is non-corrupted, non-transformed, so you have to predict all the pixels, and there is a huge amount of resources spent in the system to actually predict all those pixels, all the details. In a JEPA, you’re not trying to predict all the pixels, you’re only trying to predict an abstract representation of the inputs. And that’s much easier in many ways. So what the JEPA system, when it’s being trained, is trying to do is extract as much information as possible from the input, but yet only extract information that is relatively easily predictable. So there’s a lot of things in the world that we cannot predict. For example, if you have a self-driving car driving down the street or road, there may be trees around the road and it could be a windy day. So the leaves on the tree are kind moving in kind semi-chaotic, random ways that you can’t predict and you don’t care, you don’t want to predict. So what you want is your encoder to basically eliminate all those details. It’ll tell you there’s moving leaves, but it’s not going to give the details of exactly what’s going on. And so when you do the prediction in representation space, you’re not going to have to predict every single pixel of every leaf. And that not only is a lot simpler, but also, it allows the system to essentially learn an abstract representation of the world where what can be modeled and predicted is preserved and the rest is viewed as noise and eliminated by the encoder. > > [(00:30:59)](https://youtube.com/watch?v=5t1vTLU7s40&t=1859) So it lifts the level of abstraction of the representation. If you think about this, this is something we do absolutely all the time. Whenever we describe a phenomenon, we describe it at a particular level of abstraction. We don’t always describe every natural phenomenon in terms of quantum field theory. That would be impossible. So we have multiple levels of abstraction to describe what happens in the world, starting from quantum field theory, to atomic theory and molecules and chemistry, materials and all the way up to concrete objects in the real world and things like that. So we can’t just only model everything at the lowest level. And that’s what the idea of JEPA is really about, learn abstract representation in a self-supervised manner, and you can do it hierarchically as well. So that, I think, is an essential component of an intelligent system. And in language, we can get away without doing this because language is already to some level abstract and already has eliminated a lot of information that is not predictable. These are two different model architectures that are both pulling on the same thread: building more sophisticated conceptual representations and reasoning at that level. I certainly don’t want to assert any particular expertise in evaluating novel model architectures, but this is the first time in a while that I’m seeing this topic get any real attention, and it makes me optimistic that we’re going to see some non-linearity in model capability in the coming couple of years. Since GPT 3.5 Turbo came out a couple of years ago, we’ve all gotten at least somewhat familiar with LLM capabilities, their reasoning style, what you can and cannot productively ask of them. While the transformer model when scaled can do a hell of a lot, I still think we’re not done experimenting. I am very very interested in trading off some level of linguistic sophistication of the current generation for models that can reason about complex causal chains that the real world requires. ## Fourth: Groq From [SemiAnalysis](https://www.semianalysis.com/p/groq-inference-tokenomics-speed-but): > Groq, an AI hardware startup, has been making the rounds recently because of their extremely impressive demos showcasing the leading open-source model, [Mistral Mixtral 8x7b on their inference API](https://www.semianalysis.com/p/inference-race-to-the-bottom-make). They are achieving up to 4x the throughput of other inference services while also charging less than 1/3 that of Mistral themselves. ![Throughput vs. Price](https://cdn.sanity.io/images/wl0ndo6t/main/0a61990ecdf910ac6876c251d761ab626456538d-1262x689.webp) [Try it yourself](https://groq.com/)—it’s really freaking fast. Latency and cost both matter a lot as we start to design more complex, and multi-shot, LLM workloads. The fact that we’re seeing compute architectures optimized for LLM inference all the way from the hardware level up is very encouraging. It’s a very open question as to whether or not Groq’s specific approach (which is an interesting one but beyond the scope of this post) is correct or not, but the fact that we are running this experiment and getting these types of results is very good. It means that the idea space has plenty of untapped opportunity. Why would we assume that the same hardware architecture necessarily is optimal for both training and inference? These are two dramatically different workloads. ## Fifth: AI Design Patterns Design patterns are established architectural patterns in software engineering. Understand the most common design patterns and how to implement them and you have an incredibly useful thinking tool as a builder of software systems. The [canonical book](https://en.wikipedia.org/wiki/Design_Patterns) on design patterns, written in 1994, is one of the most widely-taught books on software engineering. The [LLM-as-CPU architecture](https://twitter.com/karpathy/status/1723140519554105733?lang=en) proposed by Andrej Karpathy (see below) begs the question then, what are the design patterns for software developed for the LLM computer? ![LM-as-CPU architecture](https://cdn.sanity.io/images/wl0ndo6t/main/474c0b018d3c0e106929b262e2f5c6ba9a989e11-1456x821.webp) Many folks think in one-shot prompting, because that’s how ChatGPT has trained us all to think. Write a really great prompt and then get a genius LLM-powered answer. But that’s not how most LLM-powered processing will work as we build systems of greater complexity. Tom Tunguz recently wrote [a great post on AI design patterns](https://tomtunguz.com/ai-design-patterns/). Here’s one that’s very easy to reason about: ![AI Router Design Pattern](https://cdn.sanity.io/images/wl0ndo6t/main/6070081d51b706cdfe4c81cdeda38a8ad71bcbe0-1456x974.webp) **From the article:** > The first design pattern is the AI query router. A user inputs a query, that query is sent to a router, which is a classifier that categorizes the input. > > A recognized query routes to small language model, which tends to be more accurate, more responsive, & less expensive to operate. > > If the query is not recognized, a large language model handles it. LLMs much more expensive to operate, but successfully returns answers to a larger variety of queries. > > In this way, an AI product can balance cost, performance, & user experience. One of the design patterns that this article doesn’t discuss but that I think is very central is the [planner-executor-integrator](https://blog.langchain.dev/plan-and-execute-agents/) pattern. Imagine this sequence: 1. Ask an LLM to write an outline for a blog post on a certain topic. 2. Recursively ask the LLM to write content for each bullet of the above outline. 3. Finally, ask the LLM to integrate all of the above-generated content together, edit it to be in a consistent tone and have good transitions, and write an introduction and conclusion that have personality. This is a simple example and some version of this is certainly possible with one-shot prompting, but this approach scales far better to harder tasks and produces consistently better-quality outputs. Imagine this approach for asking an LLM to generate code that accomplishes certain functional requirements, and imagine that each of the “bullet points” (subtasks) has their own unit tests and documentation built out and has been individually iteratively debugged. Overall, I think that AI design patterns are one of the biggest differences between people who understand how to build AI systems and people who just like to talk about AI. Want to build real systems? Start here. ## Finally, the Payoff: A Roadmap for AI in Data Ok, if you’re still with me, let me summarize the above and then draw some conclusions. 1. We’re starting to play with models that are better at developing high-level concepts and reasoning from these concepts. This is what data practitioners do: pull signal from noise and draw conclusions from it. 2. Inference costs and latency are going down and the industry is taking big swings to continue this trend. 3. We are developing better ideas of how to construct AI-powered systems. Cool. But the problem I identified at the beginning of the post was fundamentally that AI systems are not currently solving the problems that sophisticated data teams actually have today…! Will these advances actually change this underlying fact? I think it will, if you pair it with a single trend from the data space: **the advance of the metadata platform**. Frank Slootman is [famous for saying](https://accelerationeconomy.com/cloud-wars/snowflake-ceo-frank-slootman-ai-and-budget-never-in-same-sentence/#:~:text=Near%20the%20top%20of%20his,strategy%20without%20a%20data%20strategy.) “you can’t have an AI strategy without a data strategy.” I think you cannot have an AI strategy—for analytics—**without a metadata strategy**. Certainly, you need underlying data, and you need that data to be consistently accessible to AI systems ([likely via semantic interfaces](https://www.getdbt.com/blog/semantic-layer-as-the-data-interface-for-llms)). But more than that, you need an end-to-end metadata layer that can be fed into your AI systems in order to get reliable, trustworthy, nuanced answers. What does that metadata layer need to include? - Your organization’s **entire** end-to-end DAG - The social graph that overlays the DAG - Human-written documentation - Indicators of trust, including direct human feedback - Usage patterns (i.e. query logs) - System status (i.e. orchestration logs) - Semantic concepts (metrics, entities), the ability to map them to queries, and the ability to map them to each other - Underlying code and git history …and this is just the obvious stuff. For a second, just imagine this idealized metadata platform. It knows everything there is to know about how data flows through your organization and what it means. Now pair that with the advances in AI: higher-level conceptual reasoning, performance, and system design. I think this gets us to a _very_ different place. Today we’re building features like “help me generate documentation for this model.” Tomorrow we will be building features like “help me debug this conversion rate drop.” And that is a VERY different world to be in. Today, AI boosts productivity (which is great!). Tomorrow, AI will generate novel insights that will hit both top and bottom lines of the business. When is “tomorrow”? I have no idea, and honestly, I don’t care that much. It feels incredibly clear to me that this is the world we’re moving towards, and whether it takes 1 or 5 or 10 years to fully get there the payoff is so big that we just have to be building towards it. There will be plenty of interim value created along the way. **dbt’s role in this journey is to be the source of truth for metadata.** To help practitioners build their data systems in a way that generates reliable structured metadata, and then build the universal index of and API for that metadata. We want to incorporate more and more of the above metadata inputs into dbt’s metadata platform that already powers [Explorer](https://www.getdbt.com/product/dbt-explorer) and to consistently develop more and more AI-powered experiences on top of that treasure trove of context. I am very bullish about what this will unlock for data practitioners. --- --- title: "How to hire data engineers" description: "Learn data engineers' evolving responsibilities, how to craft job descriptions, and when to add data engineers to your team." url: "https://www.getdbt.com/blog/hiring-data-engineer" date: "2024-04-30" authors: ["Tristan Handy"] categories: ["Learn"] --- # How to hire data engineers I find myself regularly having conversations with analytics leaders who are structuring the role of their team’s data engineers according to an outdated mental model. This mistake can significantly hinder your entire data team, and I’d like to see more companies avoid that outcome. This post discusses when, how, and why you should hire data engineers as a part of your team. ## What is a data engineer? [Data engineers](https://www.getdbt.com/blog/what-is-data-engineering) are the people who move data from outside of your ecosystem into your ecosystem. They are responsible for your infrastructure and data plumbing. Responsibilities for these folks might include: keeping your dbt instance on the latest version, managing snowflake permissions, managing and writing airflow pipelines, and maintaining the CI/CD pipeline for your repo. ## What does a data engineer do? Even with the availability of new tools that empower data analysts and scientists to build self-service pipelines, data engineers are still a critical part of any high-functioning data team. However, the tasks they should focus on have changed, as has the sequencing in which you hire them. I’ll discuss the “when” question in a later section; for now, let’s talk about what data engineers are responsible for on modern startup data teams. Instead of building ingestion pipelines that are available off-the-shelf and implementing SQL-based data transformations, here’s what your data engineers should be focused on: - managing and optimizing core data infrastructure, - building and maintaining custom ingestion pipelines, - supporting data team resources with design and performance optimization, and - building non-SQL transformation pipelines. ### Managing and optimizing core data infrastructure While data engineers no longer need to manage Hadoop clusters or scale hardware for Vertica at VC-backed startups, there is still real engineering to do in this area. Making sure that your data technology is operating at its peak results in massive improvements to performance, cost, or both. That typically involves: - building monitoring infrastructure to give visibility into the pipeline’s status, - monitoring all jobs for impact on cluster performance, - running maintenance routines regularly, - tuning table schemas (i.e. partitions, compression, distribution) to minimize costs and maximize performance, and - developing custom data infrastructure not available off the shelf. You can get most of your core infrastructure off-the-shelf today, but someone still needs to monitor it and make sure it’s performing. And if you’re truly a cutting-edge data organization, you’ll likely want to push the boundaries on existing tooling. Data engineers can help with both. ### Build and maintain ingestion pipelines While data engineers no longer need to hand-roll Postgres or Salesforce data transport, there are “only” about 100 integrations available off-the-shelf from the modern data integration vendors. Most of the companies we work with have off-the-shelf coverage of between 75 and 90% of the data sources they work with. In practice, integrations are implemented in waves. Typically, the first phase includes core application database and event tracking, with the second phase including marketing systems like an ESP and advertising platforms. These first two phases are available completely off the shelf today. Once you go deeper into your more domain-specific SaaS vendors, you’ll need data engineers to build and maintain these more niche data ingestion pipelines. ### Supporting data team resources with design and performance optimization for SQL transformations One of the shifts we’ve seen in data engineering in the past five years is the rise of [ELT](https://www.getdbt.com/blog/extract-load-transform): the new flavor of [ETL](https://www.getdbt.com/blog/extract-transform-load) that transforms the data after it’s been loaded into the warehouse instead of before. This shift has a tremendous impact on who builds these pipelines. This shift to ELT means that data engineers don’t have to build most [data transformation](https://www.getdbt.com/analytics-engineering/transformation) jobs. It also means that data teams without any data engineers can still get a long way with data transformation tools built for analysts. Data engineers still have a meaningful role to play in building these transformation pipelines, however. There are two key areas where data engineers should get involved: 1. When performance is critical. Sometimes business logic requires some particularly heavyweight transformation, and it’s helpful to have a data engineer involved to assess the performance implications of a particular approach to building a table. Many analysts aren’t deeply experienced with performance optimization within MPP analytic databases and this is a great opportunity for collaboration with someone more technical. 2. When code gets complicated. Analysts are great at answering business questions using data but frequently aren’t trained to think about how to write extensible code. It’s very easy to start building tables in your warehouse and have the entire project get out of hand quickly. Get a data engineer involved thinking through the overall architecture of your warehouse and doing design reviews on particularly pernicious transformations or you’ll find yourself with a spaghetti bowl to clean up. ### Build non-SQL transformation pipelines While SQL can natively accomplish most data transformation needs, it can’t handle everything. One common need is to do geo enrichment by taking a lat/long and assigning a particular region. At the moment, this is not widely supported on modern MPP analytic databases (although this is [starting to change](https://cloud.google.com/bigquery/docs/gis-intro)!), so the best answer is often to write a Python-based pipeline that augments the data in your warehouse with region information. The other obvious use case for Python (or other non-SQL languages) is for algorithm training. If you have a product recommender, demand forecast model, or churn prediction algorithm that takes data from your warehouse and outputs a series of weights, you’ll want to run that as a node at the end of your SQL-based [DAG](https://www.getdbt.com/blog/guide-to-dags). Most companies that are running either of these types of non-SQL workloads today are using Airflow to orchestrate the entire DAG. [dbt](https://www.getdbt.com) is used for the SQL-based portion of the DAG and then non-SQL nodes are added on at the end. This approach gives a best-of-both-worlds outcome where data analysts can still be primarily responsible for the SQL-based transformations while data engineers can be responsible for production-grade ML code. ## Data engineer vs. Machine learning engineer Data engineers build and maintain the infrastructure that allows organizations to collect, process, and store large-scale data efficiently. For example, a data engineer might develop a pipeline that pulls customer interactions from a website, cleans the data, and loads it into a warehouse like Snowflake for analysis. ML engineers use this structured data to develop, train, and deploy machine learning models. Their focus is on turning raw data into predictive insights, such as building a recommendation system that suggests products based on customer behavior. The skill sets for these roles differ. Data engineers work with ETL/ELT workflows, data pipelines, and infrastructure tools like Apache Spark, Airflow, and dbt. ML engineers specialize in model training, feature engineering, and deployment using TensorFlow, PyTorch, and MLFlow. ## When does my team need a Data Engineer? This change in role also informs a rethinking of the sequencing of data engineer hires. The previously accepted wisdom was that you needed data engineers first because data analysts and scientists had nothing to work with if there wasn’t a data platform in place. Today, data analysts and scientists should self-serve and build the first version of their data stack using off-the-shelf tools. Hire data engineers as you start hitting scale points: - Scale point#1: consider hiring your first data engineer when you have 3 data analysts/scientists on your team. - Scale point #2: consider hiring your first data engineer when you have 50 active users of your BI platform. - Scale point #3: consider hiring your first data engineer when the biggest table in your warehouse hits 1 billion rows. - Scale point #4: consider hiring your first data engineer when you know you’ll need to build 3 or more custom data ingestion pipelines over the next few quarters and they’re all mission-critical. The key thing to realize is that data engineers don’t provide direct business value—their value comes in making your data analysts and scientists more productive. Your data analysts and scientists are the ones working with stakeholders, measuring KPIs, and building reports and models—they’re the ones helping your business make better decisions every day. Hire data engineers to act as a multiplier to the broader team: if adding a data engineer will make your four data analysts 33% more effective, that’s probably a good decision. Data engineers deliver business value by making your data analysts and scientists more productive. ## Whom should you hire? As the role of the data engineer changes, so too does the profile of the ideal candidate. My esteemed colleague [Michael Kaminsky](https://github.com/mikekaminsky) put it better than I ever could in an email we exchanged on this topic, so I’ll quote him here: “The way I think about this shift is a change in data engineering’s role on the team. It’s gone from a builder-of-infrastructure to a supporting-the-broader-data-team role. That’s actually a pretty huge shift, and one that some data engineers (who want to focus on building infrastructure) aren’t always excited about. I actually think this is important for startups to appreciate: they need to hire a data engineer who is excited about building tools for the analytics / DS team. If you hire a data engineer who just wants to muck around in the backend and hates working with less technical folks, you’re going to have a bad time. I look for data engineers who are excited to partner with analysts and data scientists and have the eye to say “What you’re doing seems really inefficient, and I want to build something to make it better." I could not agree more with this sentiment. The best data engineers at startups today are support players that are involved in almost everything the data team does. They should be excited about that collaborative role and motivated to make the entire team successful. ## How to write job descriptions for data engineers Hiring, like sales and marketing, is all about the funnel; you are selling candidates on the opportunity of joining your team. Very strong job descriptions are a crucial first step. I recommend job descriptions have five parts: - Background on the role; - Requirements; - Responsibilities; - Hiring process; and - A 30/60/90 day plan (How You’ll Ramp). Many job descriptions don’t have these things, but [candidates really appreciate them](https://youtu.be/T0Z_ibd3Hx0?t=110). For data roles, it’s a job seeker’s market. Investing in thorough job descriptions will help you stand out from the crowd and help ensure a strong candidate pipeline. ### Overview and background Give the candidate an opportunity to understand the company, your goals for this role, and how they will fit into the team. What does the team already have and what is the need that you are filling? Is this someone who’s going to focus on a specific domain or subject area? Let them know up front. This is also a great time to paint a picture of your stack and sell the candidate on your business. ### Requirements What are the hard requirements for your role? (Hint: a college degree shouldn’t be one.) Do you need someone with experience with certain technologies or frameworks? Your requirements list should be as specific as needed but should not be a laundry list. For example, if you use Airflow, you don’t need to have Airflow experience as a requirement, but you might decide that orchestrator experience needs to be, so a candidate who has used Luigi, Prefect, or Dagster is also one you’d consider. If that’s the case, call out “Experience with data orchestration tools” instead of “Experience with Airflow.” Try to keep your list of requirements to 5 to 10 bullet points. Fewer actual requirements is better than a lot of fungible requirements. If you have additional “Nice to Haves,” make that a separate list. [Women are less likely to apply for jobs if they feel they don’t meet all of the requirements.](https://hbr.org/2014/08/why-women-dont-apply-for-jobs-unless-theyre-100-qualified) Help ensure a strong, diverse pipeline by keeping your list of requirements to only requirements. ### Responsibilities What are the things that a candidate will actually do if they move into this role? Try to be as specific as possible. This is your opportunity to paint a picture for a candidate. I always mention in interviews that we are looking for “floor sweepers” — people who are not afraid to pick up a broom and sweep a pile of dust on the floor if it’s in front of them, even though it’s not in their job description. A list of responsibilities is not a list of all the things you will be doing, but this is your opportunity to present what an exciting role this will be. ### Hiring process You should tell candidates upfront exactly what the steps in the interview process are. Let me emphasize: You should tell candidates upfront exactly what the steps in the interview process are. Nobody likes to be in the dark. Tell candidates exactly how many calls they need to do, how long they will be, and who they will be with. Is there a technical assessment? Include that information too. If you cannot write this before posting a role, you have not spent enough time thinking through your hiring process. ### How you'll ramp (30/60/90) Starting a new job is nerve-wracking. Laying out a 90-day plan on how candidates will ramp into a role helps establish standards for performance. It affirms to the candidate that you have thought about what success looks like and helps set their expectations. As in data projects, the time to set clear measures of success is before you invest time and energy, not after. Hiring is no joke and is not a small amount of effort, but investing in the process up front is something that will pay long-term dividends. ## Data engineer vs. ML engineer job descriptions A data engineer primarily focuses on designing, building, and maintaining data pipelines that collect, transform, and store data efficiently for use across an organization. In contrast, an ML engineer specializes in developing, deploying, and optimizing machine learning models, often relying on the data infrastructure established by data engineers to ensure model performance and scalability. The key distinction lies in their objectives: data engineers are responsible for enabling clean, reliable data access, while ML engineers apply that data to build and operationalize machine learning solutions. Additionally, data engineers emphasize skills in tools like Spark, Kafka, and SQL, whereas ML engineers prioritize expertise in algorithms, model training, and frameworks such as TensorFlow or PyTorch. ### Example of data engineer job description Job Overview Included in all roles. This is specific to your company, your team, and your needs. This is your opportunity to sell yourself to the candidate. Requirements - Experience creating production-grade [ELT](https://www.getdbt.com/blog/extract-load-transform) pipelines in Python - Hands-on experience with data orchestrators (we use Airflow) - Excellent written communicator who will enable async work Responsibilities - Evolve our CI/CD strategy on the data team’s code base - Guide and implement architectural improvements to our data infrastructure - Maintain our Airflow infrastructure and ensure efficiency in our orchestration processes Hiring Process 1. Hiring Manager Resume Review 2. Recruiter Screen (1 hr) 3. Hiring Manager Interview (1 hr) 4. Technical Assessment (done on own time, asked to limit to 4 hours) + 1 min 5. Technical Review with peer, scheduled upon submission 5. Peer Interview (1 hr) 6. Executive Interview (30 mins to 1 hr) How you’ll ramp - 30 days: Be in the on-call rotation with named support - 60 days: Be contributing to internal conversations on data organization and structure - 90 days: Rolled out your first pipeline to support your team members with a new data source ### Example of ML engineer job description Job Overview Included in all roles. This is specific to your company, your team, and your needs. This is your opportunity to sell yourself to the candidate. Requirements - Experience with supervised and unsupervised ML and deep learning frameworks like Scikit-learn, TensorFlow, Keras, PyTorch - Strong understanding of data and analytics, including experience with Big Data, real-time, and batch data processing is preferred. Responsibilities - End-to-end ownership of the engineering cycle, including deployment, for ML-based initiatives - Work directly with development teams to guide how features get built from ideation to production Hiring Process 1. Hiring Manager Resume Review 2. Recruiter Screen (1 hr) 3. Hiring Manager Interview (1 hr) 4. Technical Assessment (done on own time, asked to limit to 6 hours) + 1 min Technical Review with peer, scheduled upon submission 5. Peer Interview (1 hr) 6. Executive Interview (30 mins to 1 hr) How you’ll ramp - 30 days: Comfortable deploying data that can be used by developers for feature engineering - 60 days: Helping guide workflows that move ML models from idea to production - 90 days: Have a model you worked on be driving a feature that’s now in production ## How dbt can help Hiring data engineers is a critical step in building a scalable, high-performing data team, and signing up for dbt Cloud is a great first step to learning why they're so valuable. Experience how dbt Cloud empowers data engineers with a collaborative environment for developing, testing, and deploying analytics code. Get a [free 14-day trial](https://www.getdbt.com/signup) for your team today. --- --- title: "How to hire analytics engineers" description: "Learn the analytics engineer's unique value, key skills to look for, and tips for crafting effective job descriptions." url: "https://www.getdbt.com/blog/hiring-analytics-engineer" date: "2024-04-16" authors: ["Erin Vaughan"] categories: ["Learn"] --- # How to hire analytics engineers Modern data warehouses have upended the way that data teams function. [Data warehouse management has entered the realm of analysts](https://blog.getdbt.com/five-principles-that-will-keep-your-data-warehouse-organized/), not just database administrators or data engineers. Modern data warehouses like Snowflake, Redshift, and BigQuery have upended the way that data teams function. Data storage has become cheap and fast; [data transformation](https://www.getdbt.com/analytics-engineering/transformation/) is now done in-warehouse ([ELT](https://www.getdbt.com/blog/extract-load-transform) vs. [ETL](https://www.getdbt.com/blog/extract-transform-load)). These trends are impacting the way that data teams are staffed and organized to bridge the gap between data engineers and data analysts. Many companies still call these people data analysts, but we’ve started to call them “analytics engineers.” Here’s why and how you can hire one too. ## What is an analytics engineer? An [analytics engineer](https://www.getdbt.com/what-is-analytics-engineering) is a technical analyst who applies software engineering best practices to the production and maintenance of analytics code. The analytics engineering workflow cleans and transforms raw data into consumable information and business logic. In the process, the analytics engineering workflow tests data to ensure it is of high quality, documents all business logic and ensures data models are running reliably in a production environment. ### Analytics engineer vs. data analyst In contrast, a data analyst leverages these prepared datasets to generate insights, create visualizations, and drive decision-making by interpreting trends and patterns. Where these roles differ most is in their technical responsibilities and tools: analytics engineers primarily work with coding languages like SQL and Python to design data models, while data analysts often use BI tools and dashboards to deliver actionable insights. Additionally, analytics engineers ensure the infrastructure supports scalability and data quality, whereas data analysts concentrate on storytelling and communicating results to stakeholders. ## Why are companies hiring analytics engineers? In short: analytics engineers can improve teams, data, and infrastructure. ### Increase productivity of the data team It’s still common for data engineers to own 100% of the ETL process in an organization, although this is often a legacy organizational structure from the time when data warehouses weren’t fast enough to allow for data transformation to be done in-warehouse. If you’re using a modern data warehouse, this approach is no longer best practice. For modern data teams, the ideal setup is for analysts, who have a much better understanding of the business logic that goes into data transformation, to own most or all of the transformation process. An analytics engineer can be that seamless bridge that connects data analysis to data engineering. This role is often the difference between analysts being empowered to turn their work around in real time vs. needing to wait in the queue to get data engineering support. It is not uncommon for that difference to result in a 10x in the velocity of the analytical process. ![quotes from data leaders about the need for analytics engineers](https://cdn.sanity.io/images/wl0ndo6t/main/0fb353ea0e666650a6cd9b01ecd90f97ebff21c0-1406x316.jpg) ### Improve data quality Data quality can erode in a few places during the transformation process as an organization matures. Without a tight process, this code can often become full of copy-paste, tables that are no longer used still stick around and create confusion, errors creep in without anyone realizing it, and performance can degrade. All of these issues are simply a byproduct of entropy—the natural state of the world is to degrade towards disorder—and they’ll slow down your team’s productivity significantly. When data engineers own data transformation, quality erodes because they often don’t quite have the depth of understanding of the business needs that data analysts have. Things get lost in translation, and data engineers aren’t actually able to identify what constitutes an “error state” in the data. The analytics engineer improves data quality by bringing a deep understanding of what the business needs into the transformation process, but also by bringing the rigor of software engineering to analytics code. ### Implement specific technology In some cases, the need for an analytics engineer comes from the fact that organizations are invested in specific tools and workflows. [dbt](https://www.getdbt.com) is a tool that is designed to allow analysts to own the entire analytics engineering workflow. Once companies adopt dbt, they start hiring to match that need. ## How do you interview for analytics engineering potential? It’s unlikely that you’re going to find someone who has “analytics engineer” on their resume. Jillian Corkin, Principal Data Analyst at [HubSpot](https://www.hubspot.com/?__hstc=7448941.6e55e52454b48d0dd79baecbb2bc8f35.1731787025216.1734279759362.1734283808517.4&__hssc=7448941.7.1734283808517&__hsfp=3610588553) advises that hiring managers “focus on the skills you need + the aptitude and interest for learning them. There’s a lot of noise in the field and it’s not reliable to look at titles. Also, I suspect there’s a lot of untapped talent who may not even be thinking about analytics as a career in systems analyst-type roles.” In other words, if you’re hiring for this role, you’ll likely be hiring for potential vs. experience. The people we spoke with pointed to some common indicators of analytics engineering potential. ### They have a natural affinity for structure Engineers are structured thinkers, and while data analysts haven’t always been taught that same kind of structure, you can often spot a natural affinity for the engineering mindset. Some more interview questions hiring teams use to evaluate this quality include: - What are the advantages of a columnar data store? And what are the disadvantages? - Walk me through the data stacks you’ve worked with. - You receive feedback that a source of data has been reporting incorrect data for a couple of weeks and that it has been sent out to multiple clients and used for internal analyses in the company. How would you go about identifying the root cause of the problem and how would you communicate your findings to all relevant stakeholders? - What data integrity/governance challenges have you encountered and how did you deal with them? Discuss a time you’ve had someone question your analytic work. How did you explain your process and reasoning? ### They are always learning If you’re hiring for this role, you’re very likely going to be doing some amount of training. Celina Wong from [Simon Data](https://www.simondata.com/), advises that “The novice with potential is a better investment than the expert with arrogance.” Because when you’re hiring for a fundamentally new role, it’s unlikely you’ll find someone who is an expert in every aspect of the work. You’re better off looking for someone who loves learning and has the potential to grow. This approach to hiring isn’t all that different from how organizations approach hiring software engineers. Every engineering team has its preferred languages and tooling, but being an expert in those exact languages and tools is rarely a requirement. Some interview questions hiring teams use to evaluate this quality include: - What resources (blogs, books, newsletters, etc.) do you follow to keep up with data? - If you had built the data team and technology stack at your current organization on day one, what would you have done differently? - When was the last time you learned something new at work just because you were curious about it? What inspired you to take that on? - What is the most difficult data analysis problem that you have solved to date and how did you do it? - Tell me a time when you were caught off guard (eg. fire drill project). What was it and how did you handle it? - Tell me about a project you screwed up and the consequences for the different stakeholders involved. What do you do differently now as a result, and how does that impact each of those stakeholders? ### They care about writing damn good SQL Analytics engineers aren’t afraid of messy data because they can solve it with good SQL. They write SQL in a way that is highly-performant, easy to troubleshoot, and [DRY](https://docs.getdbt.com/terms/dry). They’ll likely write better SQL than either your data analysts or your data engineers. The most popular way to test for technical skills is with a technical test. Tests reveal a lot not just how much SQL a candidate knows, but also open up the opportunity for the hiring team to understand how that candidate thinks about writing analytics code with questions like: - Can you walk me through your thought process here? - What led you to define the problem this way? GitLab asks candidates to rate their SQL skills from 1-5 and then explain why they gave themselves that rating. Bonus points when candidates mention window functions, CTEs, and macros! ## How to write job descriptions for analytics engineers Hiring, like sales and marketing, is all about the funnel; you are selling candidates on the opportunity of joining your team. Very strong job descriptions are a crucial first step. I recommend job descriptions have five parts: - Background on the role; - Requirements; - Responsibilities; - Hiring process; and - A 30/60/90 day plan (How You’ll Ramp). Many job descriptions don’t have these things, but [candidates really appreciate them](https://youtu.be/T0Z_ibd3Hx0?t=110). For data roles, it’s a job seeker’s market. Investing in thorough job descriptions will help you stand out from the crowd and help ensure a strong candidate pipeline. ### Overview and background Give the candidate an opportunity to understand the company, your goals for this role, and how they will fit into the team. What does the team already have and what is the need that you are filling? Is this someone who’s going to focus on a specific domain or subject area? Let them know up front. This is also a great time to paint a picture of your stack and sell the candidate on your business. ### Requirements What are the hard requirements for your role? (Hint: a college degree shouldn’t be one.) Do you need someone with experience with certain technologies or frameworks? Your requirements list should be as specific as needed but should not be a laundry list. For example, if you use Airflow, you don’t need to have Airflow experience as a requirement, but you might decide that orchestrator experience needs to be, so a candidate who has used Luigi, Prefect, or Dagster is also one you’d consider. If that’s the case, call out “Experience with data orchestration tools” instead of “Experience with Airflow.” Try to keep your list of requirements to 5 to 10 bullet points. Fewer actual requirements are better than a lot of fungible requirements. If you have additional “Nice to Haves,” make that a separate list. [Women are less likely to apply for jobs if they feel they don’t meet all of the requirements.](https://hbr.org/2014/08/why-women-dont-apply-for-jobs-unless-theyre-100-qualified) Help ensure a strong, diverse pipeline by keeping your list of requirements to only requirements. ### Responsibilities What are the things that a candidate will actually do if they move into this role? Try to be as specific as possible. This is your opportunity to paint a picture for a candidate. I always mention in interviews that we are looking for “floor sweepers” — people who are not afraid to pick up a broom and sweep a pile of dust on the floor if it’s in front of them, even though it’s not in their job description. A list of responsibilities is not a list of all the things you will be doing, but this is your opportunity to present what an exciting role this will be. ### Hiring process You should tell candidates upfront exactly what the steps in the interview process are. Let me emphasize: You should tell candidates upfront exactly what the steps in the interview process are. Nobody likes to be in the dark. Tell candidates exactly how many calls they need to do, how long they will be, and who they will be with. Is there a technical assessment? Include that information too. If you cannot write this before posting a role, you have not spent enough time thinking through your hiring process. ### How you'll ramp (30/60/90) Starting a new job is nerve-wracking. Laying out a 90-day plan on how candidates will ramp into a role helps establish standards for performance. It affirms to the candidate that you have thought about what success looks like and helps set their expectations. As in data projects, the time to set clear measures of success is before you invest time and energy, not after. Hiring is no joke and is not a small amount of effort, but investing in the process up front is something that will pay long-term dividends. ## Analytics engineer vs. data analyst job descriptions The key distinction lies in the scope of work and technical focus: analytics engineers bridge software engineering and data management by transforming raw data into usable formats, often using tools like SQL and dbt. Data analysts, however, prioritize interpreting this data to identify trends and patterns, focusing on storytelling and decision-making support rather than engineering the data itself. ### Example of analytics engineer job description Job Overview Included in all roles. This is specific to your company, your team, and your needs. This is your opportunity to sell yourself to the candidate. Requirements - Clear and direct communication skills about complex, technical topics - Understanding of SaaS metrics - Track record of working autonomously with organizational and time management skills Responsibilities - Utilize SQL + Git to build new analyses and support existing ones - Communicate findings to a wide range of stakeholders - Help drive a change in the usage of data through the active surfacing of insights to stakeholders Hiring Process 1. Hiring Manager Resume Review 2. Recruiter Screen (1 hr) 3. Hiring Manager Interview (1 hr) 4. Technical Assessment (done on own time, asked to limit to 2 hours) + 30 min Technical Review with peer, scheduled upon submission 5. Peer Interview (1 hr) 6. Cross-functional interview, if supporting specific function 7. Executive Interview (30 mins to 1 hr) How you’ll ramp - 30 days: Working in our business intelligence tool, producing analyses on already modeled data - 60 days: Helping lead metrics alignment conversations across stakeholders - 90 days: Producing proactive insights for the business ### Example of data analyst job description Job Overview Included in all roles. This is specific to your company, your team, and your needs. This is your opportunity to sell yourself to the candidate. Requirements - Comfortable working with Git and the command line - Track record with Python/R and SQL to drive business insights - Clear and direct communication skills about complex, technical topics Responsibilities - Expand our data warehouse with clean data ready for analysis - Help to define and improve our internal standards for style, maintainability, and best practices for a high-scale data infrastructure Hiring Process 1. Hiring Manager Resume Review 2. Recruiter Screen (1 hr) 3. Hiring Manager Interview (1 hr) 4. Technical Assessment (done on own time, asked to limit to 4 hours) + 1 min Technical Review with peer, scheduled upon submission 5. Peer Interview (1 hr) 6. Cross-functional interview, if supporting specific function 7. Executive Interview (30 mins to 1 hr) How you’ll ramp - 30 days: Comfortable working with dbt and our data stack from the command line - 60 days: In the on-call rotation without named support - 90 days: Gathering requirements and scope on projects with little support from more senior members of the team ## Where to find potential analytics engineer candidates Seventeen people shared their expertise with us for this article, and while there are some notable trends even among this small group — the traditional job boards work but communities and meetups are quite popular too — there are less popular ideas with a whole lot of potential — like participating in internship programs. When you’re hiring for a still emerging role, spending 3-6 months showing a group of interns the ropes can be a great way to spot the one or two who have a particular aptitude for analytics engineering. If you currently have an open role, there are 12000+ people in the [dbt Slack](https://www.getdbt.com/community/) #jobs channel. Join the community, share your role, or just pop in there and get inspired by job descriptions from companies that are currently hiring. One word of caution from John Lynch, “Hire one sooner than you think! Analytics engineers are a great way to get data engineers and data analysts/scientists working together more closely.” We couldn’t agree more. ## How dbt can help Hiring analytics engineers is a critical step in building a scalable, high-performing data team, and signing up for dbt Cloud is a great first step to learning why they're invaluable. Experience how dbt Cloud empowers analytics engineers with a collaborative environment for developing, testing, and deploying analytics code. Get a [free 14-day trial](https://www.getdbt.com/signup) for your team today. --- --- title: "The 2024 State of Analytics Engineering report" description: "The annual survey of data practitioners and leaders highlights data quality and data trust issues, team challenges, and AI." url: "https://www.getdbt.com/blog/the-2024-state-of-analytics-engineering-report" date: "2024-04-04" authors: ["Daniel Poppy"] categories: ["Insights"] --- # The 2024 State of Analytics Engineering report The [2024 State of Analytics Engineering](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024) report, sponsored by dbt Labs, is an essential read for anyone involved or interested in the field of data and analytics. This year’s edition, drawing responses from 456 data practitioners and leaders, offers a comprehensive look into the current and future state of the data profession. ## State of Analytics Engineering highlights and key findings - **AI integration in data workflows**: A notable 57% of respondents are either managing data for AI training or plan to do so in the next 12 months, highlighting the accelerated integration of AI into analytics engineering practices. - **Data quality and data trust**: With over 57% of professionals citing poor data quality as a predominant issue—an increase from 41% in 2022—the report illustrates the critical importance of data integrity and data trust across organizations. Respondents identified “increasing data trust” as the leading focus area for data teams. - **Economic impact and investment focus**: The economic landscape has led to 41% of data professionals reporting reductions in their budgets. Despite this, most teams expect to maintain (rather than grow or cut) their investments in data tooling across the board. Decentralized data architectures like data mesh ‌continue to garner investment consideration across the industry, especially at large enterprises. - **Challenges and opportunities for data teams**: The report details the daily responsibilities of data practitioners, with maintaining or organizing data sets (55%) and maintaining platforms or infrastructure (26%) as top tasks. Only 14% of data professionals “strongly agree” that their organization sets clear goals for their team. Respondents cited non-quantitative goals as the predominant measure of their success. 72% indicated that they’re primarily evaluated against either enablement of other teams or project completion, with just a fraction primarily using progress to financial goals or service level agreement (SLA) metrics as their primary measure of success. - **Salary trends across geographies**: In a trend that holds up across regions, analytics engineers can expect to earn considerably more than their data analyst counterparts. In North America, 78% of all analytics engineers earned over $100K per year, compared to 61% of all data analysts and 66% of all data engineers. However, data engineers were the most likely to earn over $200K/year. [Join experts from the analytics engineering space on April 11 for an in-depth discussion](https://www.getdbt.com/resources/webinars/2024-state-of-analytics-engineering-webinar?utm_medium=event&utm_source=webinar&utm_campaign=q1-2025_state-of-ae_aw&utm_content=____&utm_term=all___) on industry benchmarks, macro trends, and strategies for building effective data organizations. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/d049745635c37b0cd0f577b447d0d4cb63a8df00-2400x1260.png) ## **For the Media** This report is invaluable for media outlets covering technology, data science, business, and careers, offering a nuanced view into the challenges and opportunities within the data profession. It is a comprehensive resource for stories on how data practitioners navigate the complexities of data quality, AI adoption, and economic constraints. **Report URL:** https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024 Since 2016, dbt Labs has been on a mission to help analysts create and disseminate organizational knowledge. dbt Labs pioneered the practice of analytics engineering, built the primary tool in the analytics engineering toolbox, and has been fortunate enough to see a fantastic community coalesce to help push the boundaries of the analytics engineering workflow. Today there are 30,000 companies using dbt every week and over 4,100 dbt Cloud customers. To learn more about dbt Labs, visit [https://www.getdbt.com/](https://www.getdbt.com/) and follow us on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/). ### Press Contact For more details, interview requests, or questions, please reach out to [dbtlabs@methodcommunications.com](mailto:dbtlabs@methodcommunications.com) --- --- title: "How to structure your data team" description: "Structure your data team for success. Lessons from SnapTravel, Away, HubSpot on centralized, hybrid, and embedded models." url: "https://www.getdbt.com/blog/data-team-structure" date: "2024-04-02" authors: ["Erin Vaughan"] categories: ["Learn"] --- # How to structure your data team There is no one-size-fits-all way to structure your analytics function. Given the pace of technological change in our industry, it’s fair to assume that your data team structure will need to iterate and evolve over time. The data team is a brand new thing: it’s not “IT”, it’s not finance, it’s not any of the typical business functions within an operating business. So…who does it report to? How does it interact with the rest of the organization? How big is it? These are all questions that are getting answered in real time throughout the industry. And they’re likely questions that you have as you go about constructing, or re-architecting, your data team. As of today, there are no clear answers. Companies are answering these questions in a bunch of different ways, all customized to their particular businesses. This is why this question is such a hard one. ### Overview So rather than opine about what we think the best answers are —and we do have our own opinions! — we figured it would be most useful to collect a bunch of “reference architectures” from amazing companies. Peruse them and see if any of them resonate with the team you’re trying to build. The core themes we picked out that seem to drive the design of a modern data team are: - How centralized / how distributed? Some teams are highly centralized, some teams distribute members to sit and work with organizational units. - Where should [data engineering](https://www.getdbt.com/blog/what-is-data-engineering) sit? Some teams put data engineers on the data team, some draw a dotted line with the engineering organization. - What is the role of the data team? Some teams embrace data as a product, and some teams operationalize data as a service. - What executive does the data team live under? Some teams have VP- or C-level executives leading them who report directly to the CEO. Other times data rolls up to a functional head. ## Why is data team structure so difficult? This topic clearly resonated with a lot of folks, and I think it’s worth considering why that is. It’s something we’re all thinking about right now. My take? Data team structure is difficult because data technology has changed so rapidly over the past five years and this has had a cascading effect on what data people do. Ten years ago, the most challenging problem a data team faced was managing compute and store resources. At this time, data analysts had no choice but to request changes to the data warehouse and patiently wait for the data engineers to deliver. Modern cloud warehousing completely upended this relationship. The challenging problems of managing compute and store resources have largely been solved. The biggest challenges today are around speed: How can we help [data engineers and analysts](https://www.getdbt.com/blog/analytics-engineer-vs-data-analyst) collaborate more effectively? How can we empower analysts to move quickly without sacrificing data quality? How can we empower analysts, engineers, and business users to make sense of the data in our warehouse? These aren’t questions about technology, they’re questions about humans and how we can all work better together. ## Centralized vs decentralized data team structure What I’ve seen working with companies at varying sizes, and what I’ve learned from folks in the [dbt Community](https://www.getdbt.com/community), is that the spectrum of centralized to decentralized, also referred to as “embedded” or “distributed”, is one of the key decisions to make about data team org structure. Here is how David, from SnapTravel, defined the two ends of this spectrum: ![diagram depicting contralized vs decentralized data teams](https://cdn.sanity.io/images/wl0ndo6t/main/83dec2306fc5a7cde33885b5040baa7ed18bfefe-1218x428.jpg) ### Centralized data team structure In the fully centralized data team model, all data resources – people (data analysts, analytics engineers, data engineers, data scientists, etc.) and technology (data warehouse, transform, ingest, BI tools) – are owned by one central data team. If someone from product or finance has a data-related request, they submit it to the data team for prioritization. ![diagram of a centralized data team](https://cdn.sanity.io/images/wl0ndo6t/main/749c446f9c80a1793d75c479afae95c5c12562dd-1180x930.jpg) #### Advantages and disadvantages of a centralized data team structure A few benefits of this model… - Alignment of data resources to company need: When you are a small data team, like Snaptravel was, and are growing, company alignment is particularly important. A small company doesn’t have the bandwidth to do all the things. It’s important to focus data resources on the highest-impact areas of the business. - Knowledge-sharing: By placing analysts and engineers in close alignment, the centralized model prioritizes knowledge-sharing. This makes it easier to build cultural data norms together like naming conventions, syntax, or even how to write and review pull requests. - Mentorship: In the centralized model, analysts get to learn from more senior analysts as well as data engineers. This is incredibly valuable for analysts new to the analytics engineering workflow. The biggest issue with a centralized model is speed. If marketing needs support adjusting their attribution model, it’s likely going to have to wait until the end-of-month reporting is wrapped for the finance team. ### Decentralized data team structure In the decentralized model, you’ll typically see a central core group of data engineers who own the data warehouse with analysts being decentralized, or embedded, within a business function such as finance or product. ![diagram of a decentralized, or embedded, data team](https://cdn.sanity.io/images/wl0ndo6t/main/fec7a40f4031b99d612fae85f87346bd7aa04e42-1214x1026.jpg) #### Advantages and disadvantages of a decentralized data team structure The biggest advantage of the embedded model is speed. Data resources are aligned with department needs (instead of company needs). So if a business user has a request, they don’t need to wait for that request to be prioritized against all of the other needs of the business. Faster time to insights! Speed also comes from having greater context. In a centralized model, work tends to be assigned in a more “round-robin” fashion. In a decentralized model, the marketing analyst owns all marketing requests. They understand the function’s KPIs, know the metric definitions, and are familiar with the quirks of the data. This is often a benefit to both business users (who spend less time explaining themselves) and analysts (who get to go deep into a given function). One of the biggest downsides that we see with the decentralized model is how challenging it can be to keep analysts working closely together and improving their shared knowledge of data analytics. Let’s say your head of finance hires a finance analyst. It’s very possible (likely!) that person will continue to work in the spreadsheets that finance teams are traditionally accustomed to rather than adopt the modern data stack used by your centralized team. ## Data team structure examples Contained in this section are a handful of stories of how data teams have walked this path over the past few years, to help you answer questions like: - How do you divide roles & responsibilities between [analytics engineers](https://www.getdbt.com/what-is-analytics-engineering), analysts + data engineers? - Should our data team be centralized or decentralized? - How are the best teams [interviewing + hiring analytics engineers](https://www.getdbt.com/data-teams/hiring-analytics-engineer)? - When’s the right time to [hire a data engineer](https://www.getdbt.com/data-teams/hiring-data-engineer)? Thanks to the many teams who shared their stories in this section. ### Snaptravel Snaptravel trialed five data team structures over nine months. Five data team structures in nine months is a lot, but the potential efficiency gains for their team felt important enough to make these efforts worthwhile. Each of the five structures that Snaptravel tried was a different mix of centralized vs. decentralized. Ultimately they landed where we see more and more companies land – a hybrid version. The question for data teams is no longer “centralized vs. decentralized?” The question is “What, exactly, should be centralized, and what should be decentralized?” Here are the five structures [Snaptravel’s data team ](https://medium.com/snaptravel/how-should-our-company-structure-our-data-team-e71f6846024d)used[:](https://medium.com/snaptravel/how-should-our-company-structure-our-data-team-e71f6846024d) - Growth Team: When Snaptravel received [Series A funding](https://www.finsmes.com/2018/12/snaptravel-raises-additional-13-2m-closes-series-a-at-21-2m.html), they launched their growth team and began to embed their data analysts to better serve other departments. - Agile: While vacationing in London, England, Nehil discovered dbt, which allowed Snaptravel to keep track of all their data models. To allow the analysts to work together, they quickly centralized analysts onto one team, switching to an agile approach. - Full-Stack: Snaptravel’s agile approach led to a ton of problems within the organization. Data engineers and analysts were not company-level aligned with their priorities and that needed to change. Snaptravel quickly changed this approach and merged four data engineers with four analysts to form a full-stack team. They were finally able to prioritize tasks at a company-level while improving knowledge-sharing between both roles. - Pod: Their full-stack team quickly grew from eight team members to 12 in March 2020, and team meetings became a waste of time for most members, because only one or two people were needed to make a decision. Their solution to this problem was to create multiple pods that specifically owned a full-stack problem in a given area of the business. - Domain Structure: While their pod solution solved an initial problem, it eventually led to a bigger problem that slowed down their team’s progress. The full-stack pod structure lacked ownership over objectives and, at times, there were four to six people all trying to come up with a decision. The last, final change they made to their structure is referred to as a Domain structure. #### Data team structure Finally, after nine months of constant change, Snaptravel landed on a hybrid setup that they call “domain-based” team structure. In this structure, a senior member of the team is labeled “domain lead” for a specific business area in a domain-based structure. They are then responsible for assigning work to other data engineers and analysts on an individual basis to support business priorities ![diagram depicting a "domain-based" team structure](https://cdn.sanity.io/images/wl0ndo6t/main/d068caaa1ac67f5f15a79df23274f4609ad5dd04-681x275.jpg) This filled some critical gaps for them: - Ownership: “One of the reasons domain leaders really really really like this structure is because they have ownership over all the outcomes of a given area of the business,” David said. Data team members aren’t just order takers, they get to see the way their work impacts the results of a given team. - Domain Expertise: David pointed out that this ownership creates something valuable for business users as well – domain expertise. When business users have a data need, they’re always working with the same people and have confidence that this person already knows how their core data sets work and understands the unique nuances of their function. - Collaboration: Data analysts and engineers are able to work on tasks that fit their skill set while sharing best practices with one another. With every analyst and engineer having their own responsibility, they are held accountable to complete tasks in a reasonable amount of time. While this process currently works for Snaptravel, they recognize that it will evolve as their data team and organization grow.“One of the things that I’ve heard from people is that it won’t scale,” David said. “And to be honest, I don’t know. I’ve never worked at a large data organization. What we know is that this works for 10 people, and we think it could probably work for 20 people. Beyond that, we don’t know.” ### Away Travel Mike Berardo Director Data & Strategy Industry: E-commerce Company Size: 250 employees #### The numbers ⚡️3% of the organization focused on analytics 🔎20% of the organization is comfortable using BI tool 🔬1 data scientist (provisions and analyzes data) ⚙️1 data engineer (provisions data) 📊5 analysts (mostly analyze) #### Data team structure Away’s data needs are supported by five people on the analytics team and one person on the data science team, both teams report to the Director of Data & Strategy. The one-person data engineering team works closely with the Data & Strategy team but reports to engineering. Data & Strategy reports to the CEO, though Mike points out that this is an interim setup, long-term, data will report to the CFO. ![diagram depicting Away's data team structure](https://cdn.sanity.io/images/wl0ndo6t/main/96caae321f8ab6a1039c4b263288898dec3a5c51-1600x1282.jpg) The expectation at Away is that business stakeholders can do their own analysis, though customer experience, legal, and people operations all have dedicated analyst support. ### Hubspot Gordon Wong VP of Business Data Industry: SaaS Company Size: ~3000 #### The numbers ⚡️5% of the organization focused on analytics 🔎10% of the organization is comfortable using BI tool ⚙️10 data provisioners 📊120 data analyzers #### Data team structure As the VP of Business Data, Gordon manages data engineering, data warehousing, and analytics enablement. Like many SaaS businesses, HubSpot’s software creates and moves an enormous amount of customer data, which can be challenging to understand. At the same time, data engineering is seen as a skill distinct from full-stack engineering. Because of this, the data engineering team is staffed by skilled software engineers but reports to Business Intelligence. The data engineering cluster on Gordon’s team is focused on building data infrastructure that supports business and product analytics. ![diagram depicting Hubspot's data team structure](https://cdn.sanity.io/images/wl0ndo6t/main/539d50ce87da7b6ae1eac4f3a1aba64c88828e82-1600x1282.jpg) Gordon describes his team today as “engineer-heavy.” They are focused exclusively on enabling analytics, not doing analytics. The analytics function is fully decentralized with each business function hiring its own analysts and data scientists. ## How dbt helps data teams Structuring a data team is an evolving process that depends on your organization’s size, goals, and data maturity. Whether you choose a centralized, decentralized, or hybrid approach, the key to success is ensuring collaboration, data consistency, and scalable processes. dbt helps data teams — no matter their structure — work more efficiently by enabling analytics engineers, data analysts, and data engineers to build, test, and document their transformations in a well-governed, version-controlled environment. With built-in collaboration tools, automated testing, and an intuitive development experience, dbt supports teams in delivering high-quality, trusted data to the business. Get started with a [free 14-day trial of dbt](https://www.getdbt.com/signup) for teams and find out how. --- --- title: "How we think about dbt Core and dbt Cloud" description: "Learn the key distinctions between dbt Core and dbt Cloud, and how dbt Labs balances open-source innovation with commercial growth" url: "https://www.getdbt.com/blog/how-we-think-about-dbt-core-and-dbt-cloud" date: "2024-04-02" authors: ["Jason Ganz", "Jeremy Cohen"] categories: ["Product"] --- # How we think about dbt Core and dbt Cloud Regular readers of this blog will know that dbt Labs maintains two distinct and complementary pieces of software: 1. [**dbt Core**](https://github.com/dbt-labs/dbt-core) is the open source framework that is essential to the work we do, and to our mission of enabling data practitioners to create and disseminate organizational knowledge. 2. [**dbt Cloud**](https://www.getdbt.com/product/dbt-cloud) is a commercial product that extends, operationalizes, and simplifies dbt at scale for thousands of companies. It is state-aware and metadata-rich, powering differentiated experiences across an accessible UI, an extensible CLI, and enterprise-grade APIs. As these two projects have progressed over time, we (Jeremy and Jason) have heard from community members who want to better understand what features go where and why. Our hope in writing this post is to offer stable long-term guidance — to the dbt Community, to open source maintainers and contributors, and to practitioners considering how they will deploy dbt — about our criteria for deciding what functionality goes into dbt Core and what into dbt Cloud. We want you to feel confident building your analytics stack, your everyday workflows, and your career around dbt. **This post is not the announcement of any change.** It’s the distillation of a strategy that we’ve been pursuing for 4+ years. *** First things first: Our commitment to open source is not changing. dbt Core is and will remain licensed under Apache 2.0. Many more people today work full-time contributing open source code to dbt Core than in its early days, and far more people are _using_ dbt Core today than ever before. dbt has become a standard for the industry, and we’re not done extending that standard in important ways. We’re also going to continue investing in making dbt Cloud a world-class product. dbt Cloud allows us to take the vision of dbt so much further, and to so many more people, than would be possible otherwise. It’s worth stating clearly: Building a commercial business around dbt is essential to the long-term sustainability of dbt Labs, and to the sustainability of dbt Core as a well-maintained open source standard. ## What goes where? dbt is the standard for data transformation on cloud-first data platforms. We are serious about treating it as a standard: We believe everyone should work in [the way that dbt makes possible](https://www.getdbt.com/blog/building-a-mature-analytics-workflow). We believe that the work gets easier as more people are doing it, as they engage with the dbt Community and share their hard-won wisdom with the world. To that end, whenever there is an opportunity to standardize _how_ people are defining, executing, testing, and describing their transformations in a data platform, we believe this functionality belongs in open-source dbt Core. This is the basic workflow of analytics engineering. As we built out the standard, over the better part of a decade, we’ve put a lot of that functionality to dbt Core. At the same time, we believe dbt Cloud should be the easiest way to get started using dbt, and the best way to deploy dbt at scale. It should power differentiated experiences that are best built and delivered in stateful, scalable, cloud-first ways. Over the past year, dbt Cloud has delivered on these promises with step-change improvements to development, discovery, and collaboration. At this point, it is our explicit goal to build a superior end-to-end experience for customers of dbt Cloud, while preserving dbt Core’s place as the definitive standard for data transformation. Finding that balance, and finding ways to communicate it clearly, has required us to develop heuristics and to test them in practice. These heuristics are not new, though this our first public post describing them in detail. They’re not carved in stone, but they have remained stable over 4+ years of internal discussions. Paraphrasing [our pal Benn](https://benn.substack.com/p/your-companys-values-will-be-used), a company shouldn’t just talk about what its values are; it ought to show them in the decisions that count. We present three “case studies” over the past year that show how we think about our open source and commercial offerings. ### Case study 1: dbt Explorer Last October, we launched [dbt Explorer](https://docs.getdbt.com/docs/collaborate/explore-projects) as a new platform for discovering, understanding, and optimizing your organization’s data assets, across one project or many. Explorer contains a **lot** of things, and [we’re constantly adding more](https://www.getdbt.com/blog/proactively-improve-your-dbt-projects-with-new-dbt-explorer-features): - Properties and descriptions of production dbt assets (models, sources, etc) - Opinionated recommendations about dbt best practices, such as description + test coverage - Aggregations over historical runs, to analyze model execution timing + test failure rates - Lineage visualizations at multiple levels: - Node-level - Column-level - Project-level (including role-based awareness of who can see what) Explorer is powered by many metadata inputs — including the model and column [descriptions](https://docs.getdbt.com/docs/collaborate/documentation#adding-descriptions-to-your-project) that you define within your dbt projects. **What’s going on here?** There are now many products — built by members of the community, by dbt Labs, and by other companies — that leverage and extend the metadata defined in dbt projects, following the standard spec in dbt Core, in order to deliver an enhanced experience. dbt Explorer is one such product. Where do we draw the line between “standard spec” and “enhanced experience”? We strongly believe that everyone should define descriptions on their dbt models right alongside their transformation logic, in version-controlled code. This means analytics code is documented in one place, and that documentation is updated along with the code it’s documenting, as part of the same code-review flow. This fulfills a key tenet of the original dbt viewpoint. On the other hand, the mechanism for viewing, navigating, and accessing this metadata is not a thing that must be standardized. Put another way, it would be bad if you had to define a dbt model’s description in multiple places — but we expect and encourage multiple ways and places to interact with that metadata. Define once, access everywhere. Back when we added `description` to the dbt standard, we also released [dbt Docs](https://github.com/dbt-labs/dbt-docs), under an open source license, as a lightweight way to visualize that project metadata. dbt Docs fell a bit outside the core (pun intended) functionality of dbt, and with its limited functionality, it left a lot to be desired. At the same time, it has motivated tens of thousands of people to describe their dbt models, to visually consider their DAGs, and to share the fruits of their labor with countless more people. When we set out to build a next-generation experience, we decided it would need a **scalable UI** powered by [**highly available APIs**](https://docs.getdbt.com/docs/dbt-cloud-apis/discovery-api). We needed to rebuild this on real cloud architecture, not static-website-plus-big-JSON-file architecture. That architecture would need the flexibility to blend historical and up-to-date metadata, [logical and applied state](https://docs.getdbt.com/docs/dbt-cloud-apis/project-state#definition-logical-vs-applied-state-of-dbt-nodes). It needed to support end-to-end lineage, across multiple projects, with role-based access baked in from the start — that is to say, _enterprise complexity._ ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f2513bc9d1f23b7ea214a74820b2a72ea4cd3421-684x265.png) For all these reasons, dbt Cloud was the right place to build dbt Explorer. ### Case study 2: dbt Mesh [dbt Mesh](https://docs.getdbt.com/best-practices/how-we-mesh/mesh-1-intro) is a pattern for collaborating across projects and teams, enabled by a suite of new [“model governance” constructs](https://docs.getdbt.com/docs/collaborate/govern/about-model-governance) (groups, access, contracts, versions) and the ability to [resolve references across dbt projects](https://docs.getdbt.com/docs/collaborate/govern/project-dependencies#how-to-write-cross-project-ref). We chose to build a fast and scalable service for resolving cross-project references as a feature of dbt Cloud, while the rest are capabilities in OSS dbt Core. **What’s going on here?** Model contracts, versions, groups, owners, and access levels are entirely new constructs of the core language. They empower teams of any size to treat their models as stable interfaces, and the mechanism for defining those interfaces is part of the dbt standard (one that works across data platforms). **So what about cross-project ref?** Resolving references to models in other projects has long been supported in open source dbt Core — by installing upstream projects as packages. We also improved dbt Core’s mechanisms for [model](https://github.com/dbt-labs/dbt-core/issues/7446), [variable](https://github.com/dbt-labs/dbt-core/issues/6705), and [macro](https://github.com/dbt-labs/dbt-core/issues/7444) namespacing, to improve the scalability of this approach. What our customers really needed, and what we built, was [a strictly better mechanism](https://docs.getdbt.com/docs/collaborate/govern/project-dependencies) for resolving cross-project model references — powered by a state-aware metadata service within dbt Cloud — which enables developers in downstream projects to load up just the needed context about public models in upstream projects. We made that strictly-better mechanism part of the commercial offering, to support organizations with multiple teams collaborating on dbt. This service has enabled the dbt Mesh pattern for our largest customers, and it’s one of many metadata-rich services in the dbt Cloud platform that scales to enterprise complexity. ### Case Study 3: Unit testing This is a big, juicy, [much-discussed](https://github.com/dbt-labs/dbt-core/discussions/8275) feature that we anticipate will drive a ton of value for data teams. And unit testing is coming to dbt Core this spring. **What’s going on here?** Let’s return to the principle above: “Whenever there is an opportunity to standardize _how_ people are defining, testing, and describing their data transformations, we believe this functionality belongs in open-source dbt Core”. We came to the conclusion that unit testing is an important part of the standard for testing data transformations because: 1. Unit testing is an important component of software testing best practices. It has long been dbt’s viewpoint that, whenever possible, analytics workflows should take inspiration from software engineering best practices. 2. We have heard many times over the years from the dbt Community, at organizations large and small, across industries and use cases, that this is something dbt should support 3. There are numerous independent Community implementations of unit testing, from practitioners interested in using this themselves. The need isn’t to make it _possible_, it’s to make it _standard._ This combination of factors — strong overlap with software engineering best practices, durable interest from the Community, and multiple independent practitioner implementations of unit testing in dbt — all in area of data work (testing) that we have long committed as a crucial part of the dbt standard, meant that unit testing was a strong candidate to be added to the open source offering. We look forward to working with the Community to finding more areas like this: big, important problem spaces that are prime candidates for addition to the dbt standard, and the functionality in dbt Core OSS. ## How we’ll keep building it A big new feature is exciting, but even more important is our ongoing commitment to the day-to-day maintenance work of dbt Core. We take its position as a standard seriously. We will continue triaging issues, resolving bugs, and tackling the “paper cuts” that won’t make the marquee but mean better quality of life for the people who use dbt every day. Over the past six months, we’ve also prioritized substantial behind-the-scenes work to decouple and solidify the interfaces between dbt-core and data warehouse adapters — making both easier to develop, test, and maintain going forward. A commitment to “maintainership” means being intentional about what dbt Core is and ought to be; it is not the same as a commitment to “more.” Going forward, you will see us closing issues and pull requests as “out of scope” for OSS contribution, if they fall outside the purview of dbt Core. Our goal is to communicate as quickly and openly as possible what is out of scope. Every single member of the open source community stands to benefit from the investments we’re making in dbt Core, which ensure the stability and rigor of the dbt standard, as well as a clearer definition of what’s included in that standard. At the same time, we are making real and substantial investments in dbt Cloud, with experiences that enhance the analytics engineering workflow and make it accessible to more people than ever. The journey of a mature open source company is figuring out how to make two things true at the same time: supporting a vibrant user community with an ongoing open source roadmap, while also delivering a compelling commercial product that can support the growth of the business. Many companies and open source projects never reach that point; we feel lucky to have the chance to try. We hope this post gives you more clarity on how we’ve been navigating that journey at dbt Labs — and how we’ll keep doing it, for all that’s still to come. --- --- title: "What's new in dbt Cloud - April 2024" description: "Learn about all the newest features and functionality now live in dbt Cloud." url: "https://www.getdbt.com/blog/whats-new-in-dbt-cloud-april-2024" date: "2024-04-02" authors: ["Alexis Jones", "Azzam Aijazi"] categories: ["Product"] --- # What's new in dbt Cloud - April 2024 Our EPD team has been hard at work building new features, user experiences, and integrations for dbt Cloud. From column-level lineage in dbt Explorer to the GA of our Tableau integration for the dbt Semantic Layer to features like unit testing that make analytics engineering more powerful…there’s a lot of newness in our platform! We consolidated it all for you in one easy-to-read blog post, so let’s dive in to what’s new 👇. ## 🔎 **dbt Explorer** Our vision is for [dbt Explorer](https://www.getdbt.com/product/dbt-explorer) to be the best place for data teams to discover, understand, and optimize their dbt Cloud projects so downstream teams can leverage them with confidence. Here’s what’s new: - **🕵️ Model performance evaluator.** Quickly identify models across your dbt projects that could be optimized so you can keep your dbt estate performing as efficiently (and cost effectively) as possible. Check out the [blog post](https://www.getdbt.com/blog/proactively-improve-your-dbt-projects-with-new-dbt-explorer-features) to learn more ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/62eb36ef0a4e9e78700a61b0e22d945420aebd36-1700x1430.png) - **💡 Project recommendations.** Use dbt Cloud metadata to proactively surface ways to improve the test coverage, documentation, and overall project health of your dbt models. Now, data teams can tackle project improvements on their own time to ensure that pipelines are performant and data trust remains solid. [Read the docs](https://docs.getdbt.com/docs/collaborate/project-recommendations) to learn more. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a02149f9ffd4b98e36cfdfc7a040d5cf00cba57f-1990x1436.png) - **🪲 Column-level lineage.** dbt Explorer and the Discovery API now provide column-level lineage for models, sources, and snapshots within a dbt Cloud project. Column-level lineage can be used to improve many data development workflows including auditing, root-cause analysis, and impact analysis. It’s available in beta today for dbt Cloud Enterprise plan customers. [Read our technical blog post](https://docs.getdbt.com/blog/dbt-explorer) to learn more. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/b3d7eb18be3578fb837533c59d09e31d058d8065-2746x1148.png) - **🪄 Better search and lineage experience.** We’ve improved the keyword search interface to make it more intuitive to find the resources and columns you need and filter down results. Lineage is also more performant at scale, providing more navigation options and supporting all common dbt selector methods. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a18eeb47ef098645bb47194bfbf5cfd8a7b05cf3-2440x1484.png) - **🕶️ Lineage lenses.** Think of lineage lenses as map layers for your DAG: they make it easier to understand your project’s contextual metadata at scale. Lenses currently supported include resource type (default), materialization type (e.g. identify incremental model dependencies), last execution status (e.g. diagnose a failed DAG region), and model layer (e.g. discover marts models to analyze). [Read the docs](https://docs.getdbt.com/docs/collaborate/explore-projects#lenses) to learn more. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/fcd825c0d8e04a1e48354c0dd59a185dd0b0b814-2440x1546.png) ## 📈 **dbt Semantic Layer** The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) team has been hard at work since our GA at Coalesce 2023. Keep scrolling for what’s new, [catch the recent webinar](https://www.getdbt.com/resources/webinars/how-to-get-value-from-the-dbt-semantic-layer), or peruse customer FAQs [here](https://docs.getdbt.com/docs/use-dbt-semantic-layer/sl-faqs). - **💾 Saved queries & exports.** Exports allow you to materialize a saved query as a table or view in your data platform. This lets you unify metric definitions in your data platform and query them as you would any other table or view, making it possible to make your metrics available in _any_ analytics platform—even if we don’t yet have a native integration with it. Check out [this video](https://www.youtube.com/watch?v=CNk54JjOhh8) of Jordan demoing the feature, or [read the blog](https://www.getdbt.com/blog/announcing-exports-for-the-dbt-semantic-layer)! - 🗄️ **Result caching.** Improve load times and reduce compute costs by leveraging your data engine’s built-in cache for semantic layer queries. - 🔒 **SSO and PrivateLink support.** You can now develop against and test your dbt Semantic Layer in the Cloud CLI if your developer credential uses SSO 🙌. The dbt Semantic Layer also now supports using PrivateLink. - 🌎 **Integrations:** The Tableau connector is now GA for Tableau Server and Desktop! You can get the latest on all of the dbt Semantic layer integrations [here](https://docs.getdbt.com/docs/use-dbt-semantic-layer/avail-sl-integrations). ## 💻 **Developer experience** - 🚀 **Trigger on job completion.** You can now set up downstream jobs to kick off as soon as an upstream job finishes — even across dbt projects. This lets you orchestrate dbt jobs more flexibly, and easily break up long-running jobs into multiple parts. [This is available](https://docs.getdbt.com/docs/deploy/deploy-jobs#trigger-on-job-completion) on dbt Cloud Team and Enterprise plans today. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/4ea993628781ee0cbf69a6aec63046ff933760f9-1240x640.png) - 🏎️ **Improvements to parse speed.** Every time you type `dbt build` and hit enter, dbt goes through a process we call “parsing.” We’ve rewritten some of the internals of dbt Cloud so that your projects now parse 30% faster than on an M1 Mac. (This requires you to be running “[Keep on latest version](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.8#keep-on-the-latest-version-in-dbt-cloud)” in dbt Cloud.) … But that’s not all! You can also now [enable partial parsing](https://docs.getdbt.com/docs/dbt-cloud-environments#partial-parsing) on your project to parse only _changed_ files_,_ making this process even _faster_. - 🧪 **Unit tests.** Validate your model logic _before_ you materialize your full model in production. This means you can improve your test coverage without driving up data platform spend. Available in Preview for dbt Cloud customers who have opted to “[Keep on latest version](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.8#keep-on-the-latest-version-in-dbt-cloud).” [Read the docs](https://docs.getdbt.com/docs/build/unit-tests) or check out [this blog post](https://www.getdbt.com/blog/building-a-data-quality-framework-with-dbt-and-dbt-cloud) to learn more about how to approach data testing in dbt. ## 🤓 **Platform** - 💫 **Keep on latest dbt version.** dbt Cloud should feel and function like the other SaaS apps your team uses: you shouldn’t have to manually upgrade versions under the hood. Now in Preview, just select “[Keep on latest version](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.8#keep-on-the-latest-version-in-dbt-cloud)” in your environments and jobs to get immediate access to the latest and greatest. Read our [recent blog post](https://www.getdbt.com/blog/seamless-scalability-effortless-upgrades-the-enhanced-dbt-cloud-platform) for more! ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/fa083d89cbf890875562fb7bf1abf723e5895f9b-2664x1412.png) - 📊 **Cell-based architecture.** As your dbt Cloud workloads grow, we’re committed to meeting those increasing demands—with a new architecture designed for improved reliability and scale. We’ll reach out to account admins over the course of 2024 with info on specific migration steps, if any are necessary for your account. More in this [blog post](https://www.getdbt.com/blog/seamless-scalability-effortless-upgrades-the-enhanced-dbt-cloud-platform). - 🦸 **Git repo caching**. Sometimes the things that cause jobs to fail are external to dbt: things like Git provider outages or package dependencies. dbt Cloud can now cache your project’s Git repository, ensuring that your jobs continue to run seamlessly, even if external services are temporarily down. [This is available](https://docs.getdbt.com/docs/deploy/deploy-environments#git-repository-caching) for dbt Cloud Enterprise plans only. - 🤐 **_Account-scoped_ Personal Access Tokens.** These provide a more granular level of access control, scoped specifically at the account level rather than the user profile. This prevents reuse of tokens across accounts, reducing the risk of compromised credentials. Read our [docs page](https://docs.getdbt.com/docs/dbt-cloud-apis/user-tokens#account-scoped-personal-access-tokens) for more information. - 🏃 **New “job runner” permission.** As the name implies, this limited permission-set can kick off a run for a job or cancel an in-flight run. You may find this helpful to assign more fine-grained access controls in larger dbt deployments. See the [docs](https://docs.getdbt.com/docs/cloud/manage-access/enterprise-permissions#account-permissions-for-project-roles) for an updated map of available permissions on the dbt Cloud Enterprise plan. ### 👭 Partnerships and integrations - 👋 **MS Fabric adapter.** Hello Microsoft ecosystem! dbt Cloud is now available for Microsoft Fabric and support for Microsoft Azure Synapse Analytics is coming soon. [Check out the blog](https://www.getdbt.com/blog/dbt-cloud-is-now-available-for-microsoft-fabric) to learn more. - **🔀 Fivetran dbt Cloud integration.** Our friends at Fivetran built an integration into dbt Cloud so you can automate your end-to-end ELT pipelines. Configure the integration to trigger dbt Cloud jobs to run as soon as new data is loaded into your data platform. Read the [blog](https://www.getdbt.com/blog/announcing-the-fivetran-dbt-cloud-integration) to learn more. ## Wrapping up Phew…that was a ton! Looking forward to hearing your feedback on these latest features and continuing to build new functionality that helps you build and deliver trusted data. --- --- title: "Data product examples: What data products look like in practice" description: "Data products make it easier for teams to produce and consume data at scale. Se how data products work in practice." url: "https://www.getdbt.com/blog/data-product-examples" date: "2024-03-25" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Data product examples: What data products look like in practice Data teams use data products as an important method for organizing data development. By shipping new data-driven projects as data products, data teams can make it easier to find, reuse, and version data sets across a company, while also eliminating redundancy and waste. However, even with a detailed definition of data products, it can be hard to understand what a data product _is_. In this article, we’ll demystify the subject by walking through some detailed examples of what data products look like and how they work. ## What is a data product? A data product is more than just a dataset—it’s a structured, curated asset designed to solve a specific business problem. Think dashboards, reports, data tables, or even machine learning models. [Each data product isn’t just raw data](https://www.getdbt.com/blog/data-products-data-mesh); it comes with the context that makes it usable and trustworthy: metadata, documentation, access controls, and clear definitions. Together, these elements ensure that the right people can confidently use the data to make decisions. ### Data products vs. data-as-a-product It’s easy to mix up “data products” with “data-as-a-product,” but they’re not quite the same. Data products are the tangible outputs—what gets delivered: a report, a dataset, a model. [Data-as-a-product](https://www.getdbt.com/blog/data-product-data-as-product) is a mindset. It’s the philosophy of treating data with the same care and intention as a commercial product. When teams adopt a data-as-a-product approach, they: - act more like product teams than service providers. - gather requirements from stakeholders. - prioritize new features and improvements. - manage versions and ensure backward compatibility. Every iteration is focused on making data more useful and valuable to its users. This product thinking is at the heart of data mesh: a model where data is owned at the domain level, with teams accountable for making data accessible, reliable, and self-serve. By treating data as a product, organizations build trust, reduce bottlenecks, and empower teams to work with data on their own terms. ## Attributes of a data product As outlined by the creator data products, Zhamak Dehghani, a domain data product consists of a number of attributes: - **Discoverable**. It should be easy to find - e.g., via registration in a data catalog. - **Addressable**. It has a unique, labeled location from which anyone with the correct authorization can retrieve it. - **Trustworthy & truthful**. It implements metadata, data lineage, and other mechanisms to show data consumers who owns the data, where it came from, and when it was last updated - **Self-describing**. It uses code, metadata, and documentation to describe what the data product is, how it’s structured, its business use, and what each field means. - **Interoperable**. It can work via some mechanism (SQL queries, API calls, etc.) with other data products. - **Secure & governed**. It relies on [federated computational governance](https://www.getdbt.com/blog/key-components-of-data-mesh-federated-computational-governance) to enable teams to manage their own data while also ensuring security, compliance, and high data quality. ## Components of a data product As a deliverable, a data product can be as simple as a spreadsheet or as complex as a machine learning model. However, it’s also more complex than the deliverables themselves. Key components of a data product include: - A contract that specifies what fields the data product contains and any associated constraints - Versioning applied to the contract, so that new releases or fixes don’t break existing data consumers - Any other assets needed to create and maintain the data product, such as APIs, storage, tests, and an orchestration pipeline ## Examples of a data product That leaves the question of what a data product actually looks like in practice, how data producers package and deploy it, and how data consumers access it. To make these concepts more concrete, let’s walk through three examples of a data product: - Cleaned user information table - Lookup view - Joined supertable ### Clean user information table #### Explanation A user information table is a core component that multiple applications need to leverage. Without a single, centralized source for this information, every application or report will end up scraping data from the underlying raw tables in different ways. A cleaned user information table consolidates and standardizes user data from multiple sources, providing a single, reliable source of user information. This data product addresses common issues such as data inconsistency, duplication, and incomplete records by implementing robust data cleaning and transformation processes. To accomplish this, a single team - the one that owns the underlying domain data - takes responsibility for creating a data pipeline that automatically merges and cleans the data. It can do this, for example, by defining a [dbt model](https://docs.getdbt.com/docs/build/models) that defines the structure of the destination tables. It can then run the model continuously [as a Continuous Integration (CI) job](https://docs.getdbt.com/docs/deploy/deployments), pushing changes to the underlying data to the clean model. The team can also define a [contract](https://docs.getdbt.com/reference/resource-configs/contract) for the model to enforce data constraints. Once run, it publishes the contract and model to a data catalog, such as [dbt Explorer](https://www.getdbt.com/product/dbt-explorer). Here, other teams can find the data and incorporate it into their own apps and reports. If the team needs to release breaking changes - e.g., renaming or removing a field - it can do so by [versioning the model](https://docs.getdbt.com/docs/collaborate/govern/model-versions). This will enable data consumers of the current version to continue using it for a period of time until they can upgrade their code to handle the new data format. #### Business value - Provides a single source of truth for critical business data that’s discoverable in a centralized catalog - Avoids wasting computing resources and data engineering time by provide a single, accurate definition of the data set across teams - Prevents sudden downstream work stoppages when releasing a breaking change ### A time-filtered table lookup view #### Explanation Some data sets can grow incredibly large. This makes it harder to query them efficiently for data. If a cluster of use cases only needs data for a limited time window (e.g., the last three months), it makes sense to provide performant access to this critical subset. A time-filtered table lookup view provides efficient access to a subset of large datasets based on time-based criteria. This data product is designed to optimize performance for applications that typically need only the most recent data, reducing computational overhead associated with queries against large tables. A time-filtered table lookup may be as simple as creating a [materialized view](https://docs.getdbt.com/docs/build/materializations) that queries data for the last three months. Most relational databases ([like PostgreSQL](https://www.postgresql.org/docs/current/rules-materializedviews.html)) support the notion of a materialized view, which pre-calculates the result of a complex query and stores it for faster access. Like all models, materialized views in dbt can have associated contracts and versions, providing protection for consumers against breaking changes. #### Business value - Improves query performance by focusing on most relevant data and reduced need for full table query - Reduces compute resources and costs associated with querying large datasets - Can be reused and integrated into multiple applications ### A joined supertable #### Explanation Multiple teams may be using queries that use a large number of JOIN statements to combine data from multiple different tables. For example, a report on customers who have subscribed to insurance plans may pull data from Customer, CustomerDetails, Plan, PlanOption, and PlanOrder tables. Running this query once can be slow and costly - and running it repeatedly, for multiple teams, can make it even more expensive. A joined supertable runs this operation once for multiple teams and makes the results available in a new table containing the most relevant data from each joined table. This reduces resource consumption by providing a single, precomputed location for this data. This data can then be accessed via a simple (and cheap) SQL SELECT query. #### Business value - Eliminate redundant join operations across multiple queries, optimizing resource usage - Simplify data access by providing a consolidated view of related data, making it more accessible to team and team members who aren’t SQL experts - Open the door for exploratory analysis, encouraging users to find relationships in combined data they might not have seen previously ## Data product design process Successfully designing a data product like one of the above examples requires more than just writing a new data transformation. Teams design a data product as a reusable data deliverable that covers a number of use cases and takes into account the needs of multiple users. As such, it requires some forethought and planning to be successful. ### Perform data discovery and define use cases The first step is identifying which data you’ll need and which use cases it fits. Your data products should generally contain all the data required within a [data domain](https://www.getdbt.com/blog/data-domains). Furthermore, that collection of data should be unique across all data products. To identify these data products, look for data that’s consumed by a large number of teams across a diverse set of use cases. For example, in the user table example, teams as diverse as Product Development, Customer Support, Finance, Compliance, Sales, and Marketing will all use this core data in some form or another. Look in particular for data that currently has a lot of duplication and overlap. This is a key signal that the data in question could benefit from rationalization into a data product. ### Set goals and gather requirements Define your goals and any KPIs for your data product. These may include onboarding _x_ teams to the data product within _y _months, eliminating a set amount of redundant computing spend, improving data quality metrics for the underlying data, maintaining a specified average query performance, or similar benchmarks. This is also the point to reach out across impacted teams and gather their requirements. What data do they need to see in the data product for them to use it? What would they like to see in the near future? Failing to gather requirements beforehand could limit the data product’s utility and prevent stakeholders from onboarding. ### Design the data product With the groundwork laid, it’s time to build the architecture of the underlying data product. This includes addressing each of the attributes listed earlier - discoverability, addressability, etc. Other key elements of the data product’s design should include: - **A focus on ensuring data quality.** Teams can accomplish this through a combination of data cleaning, validation, and [testing](https://docs.getdbt.com/docs/build/data-tests). - **Metadata and documentation**. Teams can expose key metadata to make it easier for data consumers to discover, increasing its reach and utility. Solid documentation - such as that enabled by [dbt models](https://docs.getdbt.com/docs/build/documentation) - makes your data products easier to understand and use. - **Performance.** Teams can use optimized queries and make efficient use of resources to help ensure data consumers continue to leverage your data product well into the future. ## The dbt approach to data products When teams think of data as a product, they prioritize usability, trust, and access. A well-built data product doesn’t just deliver information—it reduces redundancy, simplifies access to critical datasets, and prevents breaking changes that disrupt the business. As we’ve explored above, data products can take many forms: reports, dashboards, curated tables, or machine learning models. Each one helps teams discover, understand, and use core data without starting from scratch. With dbt, creating structured, version-controlled data products becomes part of your daily workflow. By focusing on clear contracts, documentation, and metadata, dbt helps teams build trust in the data they share. And once those data products are in place, teams can find, reuse, and build on them—speeding up analysis, reducing duplicate work, and making collaboration easier. ### Related resources Want to dig deeper? Check out these articles to learn more: - [Data products vs. data as a product | dbt Labs](https://www.getdbt.com/blog/data-product-data-as-product) - [Data products and data mesh | dbt Labs](https://www.getdbt.com/blog/data-products-data-mesh) - [How to build a data product | dbt Labs](https://www.getdbt.com/blog/build-data-product) --- --- title: "A closer look at the newly-launched dbt Cloud CLI" description: "The introduction of the dbt Cloud CLI marked a pivotal moment, representing a second way to develop natively within dbt Cloud." url: "https://www.getdbt.com/blog/a-closer-look-at-the-newly-launched-dbt-cloud-cli" date: "2024-03-22" authors: ["Tristan Handy"] categories: ["Product"] --- # A closer look at the newly-launched dbt Cloud CLI Last Coalesce, in October 2023, [we launched](https://www.getdbt.com/blog/new-dbt-cloud-features-announced-at-coalesce-2023) the [dbt Cloud CLI](https://docs.getdbt.com/docs/cloud/cloud-cli-installation) in Preview. In a short period of time, it’s quickly become a critical part of the analytics engineering toolkit for hundreds of organizations. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/b6ba55c9ba32a80ef244d08f9b07eeb707be9cfe-1600x856.png) Today, let’s talk about what the Cloud CLI can help you accomplish, and why users are loving it. ## The developer experience gap today I want to state it as clearly as I can: **dbt Cloud CLI provides a superior developer experience when compared with dbt Core.** Today. And that developer experience gap is going to continue to widen, fast. It’s able to provide this better experience because it takes advantage of the stateful backend services provided by dbt Cloud, because parts of the internals have been rewritten to provide better performance, and because its footprint on your local device is much smaller and easier to manage. Let’s dig into some of the details. ### **Save money and time with deferral** As a user, the biggest win that smacks me in the face as soon as I switched from Core to Cloud CLI was a feature we call “[deferral](https://docs.getdbt.com/docs/cloud/about-cloud-develop-defer#defer-in-dbt-cloud-cli).” Imagine you’re working in a 1000-model project. You’re working on some models in the marts layer to implement a new feature. When you dig into this project, do you start by re-initializing your upstream models so that they’re current with production code? This process can take a _long time_ and be quite cost-intensive (especially if you’re one of tens or hundreds of developers doing this at your company!) With deferral, you can skip that step. dbt Cloud will automatically reference upstream models in production without you even having to think about it. This results in a massively better developer experience (totally eliminates a major source of wait time) and saves real costs simultaneously. You get all of this out of the box with Cloud CLI, and it’s all powered by dbt Cloud’s comprehensive metadata platform and [Discovery API](https://docs.getdbt.com/docs/dbt-cloud-apis/discovery-api). Deferral is the same underlying technology that has powered [Slim CI](https://docs.getdbt.com/best-practices/best-practice-workflows#run-only-modified-models-to-test-changes-slim-ci) for some time; now we’re bringing it to your development workflow too. ### 30% faster parse times vs. your M1-Mac If you’re an enterprise that has made a real investment in dbt, you likely have written thousands of dbt models. Large projects are challenging to operate in for a bunch of reasons; one of these reasons is high parse times. Every time you type `dbt build` and hit enter, dbt goes through a process we call “parsing.” It reads your dbt code into memory and compiles it into SQL that will run on your data platform. For medium-sized projects, this can take tens of seconds, for large projects it can take multiple minutes. Part of the solution to this problem is to move towards a more modular project architecture using dbt Mesh (more on this below!) But part of the solution just needs to be continued investments in faster parse times. And this is where dbt Cloud CLI delivers. Over the past few quarters we’ve rewritten some of the internals of dbt Cloud to optimize performance on exactly the types of very large projects that we see enterprises tend to build. As of today, if you’ve set your environment or job to “[Keep on latest version](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.8#keep-on-the-latest-version-in-dbt-cloud)” in dbt Cloud, parse times are roughly 30% faster on Cloud CLI vs. dbt Core. As we continue to invest in these types of optimizations, we anticipate that gap will continue to grow. Parse times are a major contributing factor to developer experience—writing dbt code should be highly interactive, and at modest project sizes it very much is. dbt Cloud CLI is now a far better developer experience for those folks working in large projects. ### **Build for scale with a multi-project, dbt Mesh pattern** [dbt Mesh](https://www.getdbt.com/product/dbt-mesh) is how we believe sophisticated data organizations should build out their at-scale dbt DAG. It focuses on distributing domain ownership to the edges of an organization while maintaining centralized governance and control over the platform. This allows domain teams to build quickly and autonomously, without compromising governance. If two-pizza teams and microservices were a big part of the answer that has allowed software to scale to the levels of complexity it has achieved today, dbt Mesh will be a big part of that answer on the data side. And dbt Mesh is powered by dbt Cloud. With dbt Mesh, data engineers on a Platform team (as an example) can develop via the Cloud CLI, while BI analysts on the Finance team develop via the browser-based IDE — each operating _within their own dbt projects_, but able to [seamlessly reference](https://docs.getdbt.com/docs/collaborate/govern/project-dependencies) and build on models from each other’s work. ### **Spend less time on maintenance, and onboard new developers faster** Do you support a large dbt Core install base? If you do, you know what it’s like to act as “tech support” to dozens or hundreds of users: - “My local environment got screwed up, I’m getting an error. Can you help?” - “I’m getting a version conflict; it says I need to upgrade. Can you help?” - “I think I installed dbt but maybe it was in the wrong venv because it seems to be gone. Can you help?” …etc. Often times installing, maintaining, and upgrading local software is actually harder than using it! And while dbt Core can be a great user experience, the installation and upgrade experience can be hairy, especially users who don’t go deep on Python. With the dbt Cloud CLI, dbt version upgrades are a thing of the past—[we’re doing away with minor versions for all dbt Cloud surfaces](https://docs.getdbt.com/docs/dbt-versions/upgrade-dbt-version-in-cloud#keep-on-latest-version). This means you’ll always be developing with the most up-to-date version and will no longer have to choose between a) being stuck on a dbt version from 1+ years ago missing out on a ton of new features and fixes, or b) coordinating massive upgrade efforts across your team. Not only that, the dbt Cloud CLI is a _far more lightweight_ application—it’s essentially a wrapper over dbt Cloud APIs—which means its install process is far easier and upgrades are seamless. This install and upgrade process is _already_ easier than with dbt Core, and it will continue to get ever-better as we continue to refine it. dbt Cloud CLI lets you stop acting as tech support and start spending your time focusing on building data products that deliver value to the business. ### **Enjoy improved security** The [dbt Cloud platform](https://www.getdbt.com/security) allows access to your business logic to be centrally managed, audited, and controlled. This reduces the risk associated with distributing sensitive information (such as credentials!) across various local environments or team members, like is often done with dbt Core. This centralized management system allows for easier rotation of credentials and ensures that access can be quickly revoked when necessary, enhancing the security posture of your data projects. > “The dbt Cloud CLI has allowed our core analytics engineering team to continue using the command-line workflows they’re comfortable with. For them, the dbt Cloud CLI has made the transition from dbt Core seamless. Their workflows look and feel the same, even though the mechanics have changed behind the scenes.” > > > **— Katie Claiborne, Staff Analytics Engineer, Cityblock Health** ## Our roadmap for dbt Cloud’s development experience The introduction of the dbt Cloud CLI marked a pivotal moment, representing—for the first time ever—a _second_ way to develop natively within dbt Cloud. Some users love writing code in a browser. There is nothing to set up, no ecosystem of tooling to learn about and deal with, just log in and go. Other users love developing locally—they love the power and control that the local dev tooling environment gives them. There is no wrong answer here—these are just different users expressing their preferences. There are great software development environments in browsers and on the desktop. Our goal is to bring the power of dbt Cloud to your dbt development workflow regardless of which you prefer. And in the future, we expect there will be additional development modalities! More info on this down the road, but you should expect to see a more graphical way to write dbt code as well as a natively-integrated VSCode experience, all powered by dbt Cloud, in the future. Our belief is that there is no wrong way to write code, as long as it goes through mature governance practices (CI/CD, pull request reviews, linting, etc.). So the more options we give people to write dbt code, the more humans will be able to participate in [creating and disseminating knowledge](https://handbook.getdbt.com/docs/mission) at their organizations. ### What’s next for the Cloud CLI? As a user, I already strongly prefer using Cloud CLI for all the many reasons I shared above. I’ve wanted this experience for years now and I’m never going back. But things only get better from here. Over the coming months, here’s what’s on the way: - **Invoke production jobs from the Cloud CLI.** Soon, you will be able to run production jobs on demand from the Cloud CLI in addition to the currently supported development jobs. This will be a very powerful and flexible way to invoke dbt Cloud as a part of your orchestration workloads. - **VS Code extension support.** Get all of the syntax highlighting, linting, and other convenience that you would expect from an IDE language extension. - **Richer integration with dbt Explorer.** Visually interact with your development branch in dbt Explorer. See what you’ve built and easily demo it to others. - **Further improvements to parse performance.** Expect large projects to continue to get snappier. Look out for more announcements over the course of this year related to the dbt Cloud CLI. If you’re already on dbt Cloud, you can [get started using the Cloud CLI today](https://docs.getdbt.com/docs/cloud/configure-cloud-cli). If you’re not using dbt Cloud yet, the Cloud CLI is available to every dbt user for free, forever, on the dbt Cloud Developer tier. [Get started today](https://www.getdbt.com/signup). I’m confident that you, like me, will never want to go back. --- --- title: "Building a data quality framework with dbt and dbt Cloud" description: "Learn about the various testing capabilities in dbt and dbt Cloud to build trust in data." url: "https://www.getdbt.com/blog/building-a-data-quality-framework-with-dbt-and-dbt-cloud" date: "2024-03-21" authors: ["Matt Winkler"] categories: ["Product"] --- # Building a data quality framework with dbt and dbt Cloud We’ve all heard the saying: garbage in / garbage out. Applying that notion to the realm of data has never been more important. Not only are important strategic decisions being made from descriptive data, but [as AI and machine learning continue to grow in importance](https://roundup.getdbt.com/p/lets-talk-about-ai) for organizational decision making, the stakes are higher than ever to manage and optimize data quality. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/573486691df6f30670aba35b1eede8493bbf3c14-1278x612.png) Yet, data pipelines are notoriously hard to manage. Without a solid framework for testing and pipeline maintenance in place, data teams often find themselves with poor outcomes: - They can’t ship fast enough without risking breaking changes, which means - They lose the trust of the business teams that depend on data, and so organizations becomes less data-driven Here are a few examples of how this might manifest: - A data analyst updates a query running live in production, and introduces a bad join, which causes other downstream queries to break. (Not pointing fingers, either! I would venture to say most people working in data have been this data analyst at some point in their career.) - The product team adjusts the order of questions in a questionnaire, which has the unintended consequence of reducing the response rate for questions that have been moved to the end. Meanwhile, executive leadership sees a sudden sharp decrease in a customer satisfaction KPI they use to track customer sentiment, and they sound the alarm to the data team. The data team spends a day tracing their web of SQL scripts back to the original source data to defend their own work. **The testing capabilities in dbt and dbt Cloud enable you to turn this situation on its head, and become _proactive_ about data quality.** With dbt Cloud: - Instead of the data analyst updating a production query directly, they have a workflow to automatically test their changes in lower environments first, before it affects production. - The data team can detect the decrease in the customer satisfaction KPI proactively. They can act on this knowledge by following up with the product team directly and alerting management, increasing trust by putting themselves in the driver’s seat working towards a resolution. > After all, what good is all this data if it takes too long for you to manage, and your downstream consumers can’t derive consistent and actionable answers from it? ## **How to approach testing in dbt** dbt and dbt Cloud offer various testing capabilities, but it can be difficult to know where to start. Below is our recommended framework for how to approach testing in dbt, in sequential order of relative ease of implementation and expected benefit. We’ll go into detail on each of these in this post. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/7d67984ecfca1d44bae842ad8b58872bd7ed5aae-2000x977.png) Depending on the processes you already have in place for data testing, your testing implementation may not follow this exact order. For example, if you already have robust data type checks and constraints in your data loading process, you might skip or deprioritize the third step of implementing source data testing. Here’s an example of how this testing framework is applied across dbt Labs’ Internal Project DAG, focused on revenue data: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/fc563e3cbbd24fdb75c2ae7824b219db1d27aad1-1289x410.png) ## Key concepts > If you need a refresher on dbt tests, keep reading! If not, you can jump straight into the various approaches in the next section. ### What are dbt tests? Generally speaking, [tests](https://docs.getdbt.com/docs/build/data-tests) are operations declared in the dbt codebase to audit the data in your data platform. The dbt ecosystem extends out-of-box dbt test logic with additional packages like the classic [dbt expectations](https://hub.getdbt.com/calogica/dbt_expectations/latest/). There are many other packages out there that you can use to cut down on test development time. Below are a few examples of data tests, unit tests, or source freshness assertions you may make about your data: - A certain column must not contain duplicate values - During a migration, 99% of records from Table A and Table B must be exact matches - Source data must have been loaded within the past 24 hours - Verify a piece of code that formats some wonky timestamp data works correctly For data tests, after each dataset in the dbt pipeline is built, dbt performs the audit, then determines whether it should move on to build the next dataset based on the test result (pass, fail, warn). If a check fails, you can configure dbt to simply raise a warning and still continue building downstream models, or you can dictate that the run should stop and throw an error. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/a708221497cabd6d0c2b97d7881aad1f08f1b6db-346x410.png) ### Running dbt in separate environments When thinking about running tests in this framework, it helps to be aware of dbt’s native capabilities for running data models and their tests in separate [environments](https://docs.getdbt.com/docs/dbt-cloud-environments). dbt reasons about data environments by running your data models in distinct databases and schemas to represent dev, QA, and production. dbt is highly flexible in this manner: your team can choose whether dbt manages separate environments in distinct databases with the same schema, schemas within a database, etc. For example, if I have a dbt model for a table called `fct_orders`, the database and schema structure might look like: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/3e1eda01a42c3854d8c0cc0dfd5bf7cd95b04929-904x748.png) ## Building your testing infrastructure with dbt Let’s jump into the various testing capabilities, how they work, and their benefits. ### CI/CD **What it is** - [Automated testing jobs](https://docs.getdbt.com/docs/deploy/continuous-integration) supported by dbt Cloud, to prevent bad code from making it to production. **How it works** - When analytics engineers make changes to data pipelines, dbt Cloud CI runs ONLY those changes and any related downstream datasets in temporary testing schemas. - This all happens in [parallel](https://docs.getdbt.com/docs/deploy/continuous-integration#concurrent-ci-checks), and builds for outdated commits are canceled automatically. - When the CI process finds a failure, the analytics engineer resolves the issue before testing again. **Benefits** - Automating a “dry run” of your dbt project in a separate pre-production environment greatly reduces the risk of human error. - Users can add branch protection rules to prevent a PR with a failed CI job from being able to be merged into production. - Provides a repeatable process onto which additional quality assurance processes can be built. **Example** If your organization is implementing dbt, CI/CD is a critical component. Without it, you miss the opportunity to verify how new code actually works before it moves to production. You can build CI pipelines in dbt Cloud that run with maximum efficiency. Check out this [guide](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud) to help you get started building your data platform with the power of CI / CD. **Where are models built in CI?** When running your pipelines in CI, dbt Cloud builds into temporary databases and schemas which don’t impact any datasets running in production. It does this by overriding the default and instead building into a schema with a unique name for the pull request that triggered the build. ### Testing DAG outputs **What it is** - Applying dbt tests to the last models in the DAG. These are typically the models that power dashboards, data applications, ML models, etc. **How it works** - Proactively identifies when there is risk that downstream dependencies are broken. **Benefits** - You learn about problems before your stakeholders do. - By tackling issues proactively, you reduce the number of firedrills you’re triaging (and get more sleep!) **Example** After your CI pipelines are in place, the next step is to ensure you have adequate testing on the data models dbt owns and other systems consume. These are the datasets that power things like the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl), BI dashboards, and ML models. For example, in this example DAG, I want to make sure the `fct_order_items` model is well-tested: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/b337e89972184db622407d02d87ec872751c69d8-1600x577.png) I have a few options: I can use [dbt Explorer](https://docs.getdbt.com/docs/collaborate/explore-projects) to get proactive recommendations on when a model isn’t up-to-par (with guidance on what to do next): ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/edc2820b076b2ba45e8b6d6be1cd006714f92635-1265x462.png) I can also add a test and description to this model by updating its configuration, as shown below: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/8b35106c1fab1faac0dacb2a87307f13590890cb-516x184.png) These capabilities allow us to make sure data models have tests applied in key parts of the pipeline, and that they have documentation in place so consumers know how to interpret the data they contain. ### Testing inputs **What it is** - Verifies that datasets at the beginning of the processing pipeline are correct. - Source freshness verifies that the same “raw” datasets are up-to-date with fresh data for downstream processing. **How it works** - Identifies when new records are present in source datasets. - Tests source data to identify any invalid records that should not be processed downstream. **Benefits** - Reduces data platform spend from running data transformations on already-processed data. - Automatically identifies when new raw data is present, so pipelines can start running. **Example** To manage data sources properly, the goal is to make sure they’re both up-to-date and conform to expectations. Going back to our example DAG, the source layer is farthest to the left: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f74addb367b21882b0724c04521ce3fcde171653-1600x776.png) Again, I have choices! I can use [dbt Explorer source details](https://docs.getdbt.com/docs/collaborate/explore-projects#example-of-source-details) to quickly understand the freshness of source data, and debug as needed. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/ce0e0da6c73732b893facbb811db7a895ce30780-1261x519.png) I can also configure both freshness checks and tests on raw sources, as shown below: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/eaeb04be21f3cdfb230089901c71e7eb3a5410cf-718x343.png) Assessing the structure of source data is key in guarding data pipelines against schema drift. Without this capability, data teams remain vulnerable to situations like the example described in the intro, where the product team implemented a breaking change unbeknownst to the rest of the organization. [Freshness testing](https://docs.getdbt.com/docs/deploy/source-freshness) allows you to identify when source data is out-of-date, alert on it, and avoid reprocessing the same records downstream. ### Unit testing > [Now available in dbt Cloud!](https://docs.getdbt.com/docs/build/unit-tests) (And coming to dbt Core with the v1.8 release!) This is a long-requested dbt feature, which we’re thrilled to support. We’re grateful to the many dbt community members who’ve contributed to this capability. **What it is** - Uses static input and expected output datasets to verify dbt models perform expected transformations. **How it works** - Create mock input and output records in .ymls or .csv files. - Enables testing dbt code at the column-level, without needing to specify ALL columns of an input table. **Benefits** - Reduces costs because SQL logic is validated before the entire model is materialized. - Proves that a change implemented by a developer (rather than the underlying data) is the root of a problem. **Example** Unit tests are best applied on the models which you’ve found through experience to be most problematic. They’re a tool suited for making absolutely sure your business logic is correct, before it’s pushed to your data platform. This is most common in situations where there can be many edge cases like regexes or complex case statements. Let’s say you need to deal with some raw orders data with poorly and inconsistently formatted dates from multiple upstream systems. Your dbt SQL might look like: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/f49f52521389be5e70a66c58e5cc1d2eb56c702d-617x119.png) You can apply a unit test to the model like this: ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/bbcda6b040bec295bd21039010852e5f92153c58-560x271.png) This will ensure that the sequence of replacements and date casting in the query actually translates the value ‘01/02/2024T00:00:00Z’ to the simple YYYY-MM-DD date format you want. As your sources evolve and you need to catch additional edge cases, it’s simple to add them to the input rows of your unit test without additional overhead. As a best practice, we recommend to run unit tests in development and CI environments—and to exclude them in production, as static inputs won’t change. ### Linting **What it is** - Determines whether code is formatted according to organizational standards. **How it works** - Using tools like SQLFluff, organizations define rules (e.g. SQL statements must use trailing commas). When developers write code that violates such rules, they are flagged and must be fixed before being able to commit their code. - The dbt Cloud IDE supports linting with SQLFluff natively. **Benefits** - Keeps the codebase readable and standardized. - Readability promotes faster development. **Example** dbt Cloud natively supports both [SQLFluff](https://sqlfluff.com/) and [sqlfmt](https://sqlfmt.com/) in the Cloud IDE. These tools provide an opinion about how code should actually be written and formatted. For SQLFluff, you can stick with the default rules implemented by dbt Cloud, or you can bring your own config file to apply custom rulesets to dbt pipelines. ## Measuring Your Progress dbt Explorer is a fantastic resource for benchmarking the structure of your dbt project(s). It provides intuitive visualizations based on metadata from every dbt run to summarize the structure of what you’ve built in an auditable way. You can use the lineage graph and contextual data in dbt Explorer to: - Verify that all your source data have freshness checks configured, and flag the ones that don’t. - Check that all the ending models of the DAG have tests configured. - Assert rules about the structure of the models and how they connect. And, dbt Cloud provides a number of [out-of-box checks](https://docs.getdbt.com/docs/collaborate/explore-projects#example-of-test-details) to do just that! ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/8978d6d5b44de6ad42ca6e84cd2c90311e4d09df-440x244.png) You can also leverage the [dbt project evaluator](https://docs.getdbt.com/blog/align-with-dbt-project-evaluator) to enforce these best practices. ### Example custom evaluation You can also build on top of the project evaluator in combination with dbt Cloud CI/CD to automatically check that developers are following custom rules for your organization. Say for example, you wanted to make sure your developers don’t add duplicate source mappings. There can be many raw data sources in a dbt project (at enterprise scale, thousands are not uncommon). It can be easy for development teams who want to move fast to create a new mapping, pointing to the same underlying raw data. This redundancy makes it harder to make sure there are good checks on source freshness, and creates the potential for the dbt DAG lineage to show unnecessary information. Downstream dbt models will still run, but in this situation the fidelity of the project documentation is lower. Here we show an example custom test to confirm there are no duplicate source declarations in the project. ```json -- in tests/test_duplicate_sources_exist.sql select * from -- this model exists in the project evaluator package ``` If the test finds any records of sources that point to the same raw table in your cloud data platform, but have separate dbt mappings, they’ll be flagged. Going a step further, you can set up the commands in your dbt Cloud CI / CD jobs to run this custom check each time an analytics engineer creates a pull request. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/5880b1bae8da9b97773aa7f08f5d253b0eb29953-1214x490.png) And with this setup in place, you don’t need to worry about duplicate source mappings! ## Methodically improve your test coverage with dbt Cloud The robust set of tools in the dbt ecosystem give you many options for becoming more proactive about shipping trusted data faster. With so many approaches available, it’s useful to consider how to sequence testing implementation. The test framework we provided (CI/CD pipeline, test outputs, test sources, unit test, lint) gives users a scalable approach to prioritization. Perhaps implicit in all of this, is the value derived from applying data testing best practices in the same tooling used to run transformation logic itself. The simplicity and power of using a single tool makes it much easier to ensure that tests are consistently applied to the right datasets. > If these robust testing capabilities are getting you curious about upgrading to dbt Cloud, [here’s a helpful step-by-step guide](https://docs.getdbt.com/guides/core-to-cloud-1?step=1) on how to get started migrating from dbt Core to dbt Cloud. For all these reasons, we think this information will help analytics engineers sleep better at night, knowing they manage trusted data. To learn more about dbt Cloud, check out our [whitepaper](https://8698602.fs1.hubspotusercontent-na1.net/hubfs/8698602/dbt%20Cloud%20Whitepaper_2024.pdf) or [schedule a demo](https://www.getdbt.com/contact). --- --- title: "Data transformation vs. ETL" description: "A look into the relationship between data transformation and ETL." url: "https://www.getdbt.com/blog/data-transformation-vs-etl" date: "2024-03-05" authors: ["Kathryn Chubb"] categories: ["Learn"] --- # Data transformation vs. ETL “Data transformation” and “ETL” are two terms you’ll come across as a data professional. But what do they mean, and what’s the difference between them? In fact, they aren’t opposing ideas; one is a part of the other. Data transformation is the “T” (“Transform”) in ETL (Extract, Transform, Load). With the introduction of the cloud and the ensuing rapid evolution of the data ecosystem, ETL too has evolved, and a new approach has emerged: ELT (Extract, Load, Transform). In this article, we’ll look at ETL, its relationship to data transformation, and how that relationship is changing for the better with the shift to ELT. ## What is data transformation? [Data transformation](https://www.getdbt.com/blog/data-transformation) is the process engineers use to take data from one state to another. Engineers use a query system like SQL, or a transformation package for a programming language like Python, to write code that takes one table of data and turns it into another table, set of tables, or views on the original table. Data transformation is the heart of data engineering. It allows engineers to clean, normalize, and standardize data to use it for further analysis. Engineers use it to take prepared data and turn it into meaningful aggregates and analytics. Every step of a data processing pipeline is built with data transformation. Data transformation turns raw data into meaningful analytics useful for decision-making. It establishes a quantitative groundwork for understanding an organization's state. It’s the mechanism that turns data into information. ## How data transformation fits into ETL [ETL](https://www.getdbt.com/blog/etl-vs-elt) stands for “Extract, Transform, Load,” and it outlines a sequential pipeline for data processing. Data transformation is the second step of this process. An ETL pipeline for an analytics dashboard is: - **Extract:** A user system connects to a data store and brings it into memory on an instance of computing (e.g., a virtual machine) - **Transform:** The user takes the data and re-formats it - e.g., by converting strings into integer values, or ensuring a customer ID field is using a consistent format across tables - **Load:** The user stores the transformed data in a new location—typically an OLAP database—where a downstream process (e.g., an analytics dashboard) can aggregate and use it to drive business decision-making ETL is an architecture that was designed many decades ago, when client-side storage and bandwidth were limited and costly. All data is in the source database and only select data is loaded into a central repository after it has been transformed. A single transformed table is much smaller than an entire section of a database. Linking those processes together makes sense when you want to limit how much data is moving across your network and sitting on hard drives. But that efficiency comes at a price. In a modern environment, ETL has [serious flaws](https://www.getdbt.com/blog/etl-vs-elt) that get in the way of scaling data operations. The overall problem is inflexibility. Every transformation is defined before data is brought in, so each project is constructed ad hoc, even if it is using similar transformations as another. Similar cleaning and preparation processes are repeated across the organization. Different teams in different branches of the organization develop different metrics when building their dashboards, leading to conflicts in understanding. This chaotic development stack leads to inefficient workstreams, as every pipeline builds from the ground up. Every project that wants to incorporate data must start from the beginning and bring significant data engineering expertise. ETL also increases the time it takes to deliver value. In an ETL process, data consumers don’t have access to the raw underlying data. They have to wait on the ETL process to deliver it to the data warehouse. That means they have to communicate their precise business needs to the data engineering team in detail. This drawn-out process means it can take anywhere from days to months for business analysts and decision-makers to get the data they need. ## How ELT improves transformation workflows ELT (Extract, Load, Transform) is a newer architectural approach that has emerged as data systems have moved onto the cloud. In this paradigm, data transformation is the final step of the process and happens _after_ data has been loaded into the centralized data store. An ELT pipeline takes the following steps: - **Extract: **An analytics engineer connects to a database of raw data - **Load:** The engineer brings relevant data segments into a data platform as is. - **Transform:** The analytics engineer builds a pipeline of transformations in the warehouse on top of the loaded data. These pipelines clean data, prepare it for further analysis, and establish baseline metrics. Data analytics teams build additional pipelines on these transformed tables, producing the final deliverable (a dashboard, an ML or AI application, etc.). In ELT, transformation happens directly inside a data platform, allowing for centralized, organized data processing pipelines. Engineers can pull interesting data from data sources into the data platform, and because storage is so cheap in the cloud they can aggregate any data they think might be useful for analytics. Once loaded, data teams clean (“transform”) the data and create ready-to-use tables for any relevant analysis. Analytics pipelines for one project can be accessed and reused by teams doing similar tasks. Engineers can build standardized metrics so that any team that uses them starts from a single source of truth. A framework of engineered tables to build on allows teams to hook into data without starting from a raw state, lowering the knowledge and effort barriers for all sorts of projects. ## ETL vs ELT: a side-by-side comparison - **Order of operations: ** - ETL transforms data before loading it into the warehouse. - ELT loads data first and then transforms it inside the warehouse. - **System requirements: ** - ETL typically users external tools to transform data before reaching the warehouse. - ELT relies on the power of modern data warehouses to handle transformations, reducing the need for complex pre-processing. - **Efficiency: ** - ETL can be slower for large datasets, as the transformation happens before data is loaded into the warehouse. - ELT can speed up the process, particularly in cloud-based environments, by loading raw data quickly and transforming it later. - **Use cases: ** - ETL is often used in highly strctured environments that require stringent data governance. - ELT is best suited for environments with high volumes of unstructured data, where transformation can occur after the data is already in the warehouse. ## Supporting ELT with dbt data transformation Organizations need tools to support the [potential](https://www.getdbt.com/blog/is-dbt-the-right-tool-for-my-data-transformations) of an ELT architecture. That’s where dbt comes in. [dbt](https://www.getdbt.com/product/what-is-dbt) is a SQL-first data transformation workflow that allows teams to quickly and collaboratively deploy analytics code using software development [best practices](https://www.getdbt.com/blog/data-transformation-best-practices). This gives data teams the agency, control, and visibility needed to deliver world-class data products at scale and promotes data quality and trust. dbt delivers the critical security, governance, automation, and collaboration features required for managing data complexity at scale. dbt also automatically tracks and displays the [dependencies](https://docs.getdbt.com/faqs/Models/create-dependencies) between different transformed tables. That means engineers don’t have to dig through documentation trying to decipher which table depends on which. dbt also promotes data quality and accuracy by allowing users to configure tests that are automatically applied to new datasets. Engineers can define their tests and be confident that their data pipelines meet the required standards to build trust in their outputs. dbt’s debugging capabilities automatically check developed pipelines for errors, highlighting where things have broken down so that engineers know where to look when things go wrong. All of these features add up to a system designed to maximize the benefits of an ELT architecture. Using dbt, engineers can easily organize their efforts, collaborate, share their work, and automate repetitive workflows. The support dbt provides opens up more engineering time to develop new metrics, optimize pipeline efficiency, and explore datasets to find new valuable information. ## Faster time to value with dbt Data transformation is the core of ETL and ELT data processing pipelines. As data systems have shifted to the cloud, ELT has become the standard thanks to the structure it can provide for data engineering projects. dbt offers tools like code-based transformation logic, pipeline diagrams, and automation to maximize the potential of ELT pipelines. Want to learn more about how dbt supports ELT? Check out this [guide](https://www.getdbt.com/resources/how-fivetran-and-dbt-help-with-elt). Want to see if dbt is right for your organization? Create an account and [book a demo](https://www.getdbt.com/contact) today. ## ETL data transformation FAQs **What specific types of data transformations are commonly performed in ETL and ELT pipelines?** **What key advantages does an ELT architecture offer over traditional ETL for modern data environments?** ELT architectures leverage the scalability and processing power of cloud data warehouses, allowing raw data to be loaded first and transformed later within the destination. This approach provides greater flexibility, faster data ingestion, and enables the retention of raw data for future analysis without predefined transformations. ELT also fosters reusability of transformation logic and standardized metrics across teams, leading to more efficient data operations and quicker time to value. **How does dbt specifically support and improve the data transformation process in an ELT setup?** dbt provides a SQL-first workflow that empowers data teams to collaboratively build, test, and deploy analytics code directly within the data warehouse. It automates critical aspects like dependency tracking, ensures data quality through configurable tests, and simplifies debugging by highlighting errors in pipelines. This comprehensive support allows engineers to organize efforts, share work, and automate repetitive tasks, maximizing the benefits of an ELT architecture. **What are common challenges data professionals face when performing data transformation?** Data professionals often encounter several challenges during data transformation, including the complexity of integrating diverse data sources with varying formats and ensuring consistent data quality throughout the process. Scalability can also be an issue when dealing with large datasets, requiring significant computational resources. Additionally, defining and maintaining transformation logic can be inflexible, especially in rapidly evolving business environments, leading to repeated efforts and delayed insights. --- --- title: "Data modeling techniques for more modularity" description: "Learn how modular data modeling with dbt improves data transformation, enhances collaboration, and optimizes analytics workflows." url: "https://www.getdbt.com/blog/modular-data-modeling-techniques" date: "2024-03-04" authors: ["Christine Berger"] categories: ["Learn"] --- # Data modeling techniques for more modularity Back when I started working in the data industry, as part of recruitment you’d get this Army-style pamphlet about all the cool stuff you’re going to do. Then you sit down at your desk, and things get messy. For actual decades, maintaining a productive and reliable data pipeline was an unexciting daily grind of putting out the same fires month after month, restricting access to only the “trusted” builders and watching sadly as those builders moved on to other opportunities and left you with a mess of code that no one truly understands. This is not a fun place to be. With cloud data warehouses and tools like dbt, this reality has changed. While the world is filled with fun things that are terrible for you - like french fries, soda, and fireworks - dbt and the practice of Analytics engineering allow for the rare situation where doing fun thing is also doing the right thing. A large part of this shift is moving from massive, thousand+ line sql scripts and intimidating t-sql procedural nightmares to modular, accessible, and version-controlled [data transformation](https://www.getdbt.com/blog/data-transformation). Work is divided into cleanly separated concerns: - procure raw source data - prepare them as needed in staging models - present them to the end consumer in fact + dimension models and in data marts As analytics engineers, how we prep data for analysis makes all the difference in how trustworthy it is to the rest of our organization. If your end users don’t trust the data, it does not matter how much work you put in to all the previous steps - they will still silo their work in excel sheets and bury important business logic in one off dashboard filters. A few questions can help keep data modeling efforts on the right track: - Are models consistently defined and named? Can someone immediately grok what a given data model does? - Are models easily readable? Are your individual models [DRY](https://www.getdbt.com/blog/dry-principles) enough to be quickly interpretable? - Are models straightforward to debug + optimize? Can anyone identify a _modelneck_ (a long-running model) and fix it? We’ll dig into those three later on, but first, let’s take a step back to the “before times” of non-modular data modeling techniques. ## Traditional, monolithic data modeling techniques Before dbt was released, the most reliable way that I had to model data was SQL scripting. This often looked like writing one 10,000 line SQL file, or if you want to get fancy, you could split that file into a bunch of separate SQL files or stored procedures that are run in order with a Python script. Very few people in the org would be aware of my scripts, so that even if someone else was looking to model data in a similar way, they’d start from source data rather than leveraging what I’d already built. Not that I didn’t want to share! There just wasn’t an easy way to do so. We could call this a _monolithic_ or _traditional_ approach to data modeling: each consumer of data would rebuild their own data transformations from raw source data. Visualizing our data model dependencies as a [DAG](https://docs.getdbt.com/docs/build/models) (a directed acyclic graph), we see a lot of overlapping use of source data: ![messy data model dependency graph](https://cdn.sanity.io/images/wl0ndo6t/main/03b33d53cc9ffa35ba36407b9af9ea2a8229fac6-1999x1124.png) ## What is modular data modeling? With a modular approach, every producer or consumer of data models in an organization could start from the foundational data modeling work that others have done before them, rather than starting from source data every time. When I started using dbt as a data modeling framework, I began to think of data models as _components_ rather than a monolithic whole: _What transformations were shared across data models, that I could extract into foundational models and reference in multiple places?_ _Note: in dbt, one data model can reference another using the_ _[ref function](https://docs.getdbt.com/reference/dbt-jinja-functions/ref)._ When we reference foundational data models in multiple places, rather than starting from scratch every time, our DAG becomes much easier to follow: ![modular data modeling technique](https://cdn.sanity.io/images/wl0ndo6t/main/ec5428e5294d486272186e924d902101e26c6f01-1114x501.png) We see clearly how layers of data modeling logic stack upon each other, and where our dependencies lie. It’s important to note that just using a framework like dbt for data modeling doesn’t guarantee that you’ll produce modular data models and an easy-to-interpret DAG. Your DAG, however you construct it, ultimately is just a reflection of your team’s data modeling ideas and thought processes, and the consistency with which you express them. Let’s get into those 3 tenets of modular data modeling: naming conventions, readability, and ease of debugging + optimization. ## Data model naming conventions A dbt project, at its core, is just a folder structure for organizing your individual SQL models. Within the `/models/` folder of a project, any `.sql` files you publish will be materialized as tables or views to your data warehouse - so you can either make them easy for your team to navigate, or a complete pain - the choice is yours! Without a solid naming convention, our team may end up rebuilding models that had already been reviewed and published, or rejoining data in duplicative + low-performance ways. You could say that a naming convention brings a zero-waste policy to our salad bar, and locks in model reusability. _Note: Even with an established naming convention, we need an equally solid_ _[data model peer review process](https://www.getdbt.com/blog/how-to-review-an-analytics-pull-request), to ensure that it gets followed with each new addition to our transformation logic._ Our data model naming convention defines two things about each of our model layers: 1. The types of models we’ll use (source, staging, marts etc) 2. What types of transformation each of those models types is responsible for My career hit the data space right when cloud warehouses like Databricks, BigQuery, Snowflake et al. were first being adopted - I’ve had the luxury of never worrying too much about compute + storage cost. So when I think about data modeling techniques, I don’t really think in terms of style (Kimball, Data Vault, star schema, etc), although plenty of people find those techniques to be useful. Instead, I generally follow our internal modeling conventions at dbt Labs, which focuses on finding the shortest path between raw source data and the data products that would actually solve a problem for them. Of course your conventions may differ! The important thing is just to have a convention and stick to it. ### Staging models The purpose of staging models (in our convention) is just to clean up and standardize the raw data coming from the warehouse, so we have consistency when we use them in downstream models. In our dbt project, we’ll place them in a staging folder, and prefix filenames with `stg_` for easy identification (so our Zendesk chat log would be `stg_zendesk_chats`, which is based on the raw zendesk.chats source table). ![staging data model](https://cdn.sanity.io/images/wl0ndo6t/main/def70335d60fd4114878f5c231f8d07f95fee2a8-1284x142.png) They’re typically a one-to-one reflection of each of our raw sources, and we do really light transformations at the staging layer. We will very rarely join data models at the staging layer, but instead will perform transformations like: - Field type casting (from FLOAT to INT, STRING to INT, etc), to get columns into the proper type for downstream joins or reporting - Renaming columns for readability - Filtering out deleted or extraneous records Doing these types of base transformations at the staging layer (and the staging layer only) serves as a jumping-off point for our heavier transformation layers downstream. If anything ever changes in the source data, we have a layer of defense, and can be confident that if we fix the staging layer, our changes will flow into downstream models without manual intervention. ### Data mart models The data mart layer is where we start applying business logic, and as a result, data mart models typically have heavier transformations than in staging. The purpose of these models is to build our business’s core data assets that will be used directly in downstream analysis. In our marts project folder our models are generally dimension and fact tables, so we prefix them with `dim_`, or `fct_`. ![project folder of data mart models: facts + dimensions](https://cdn.sanity.io/images/wl0ndo6t/main/d431f4389c78a2453457c43f58c20c42cb858238-520x662.png) The common SQL transformations that you’ll see at the data mart layer are: - Joins of multiple staging models - CASE WHEN logic - Window functions Really nothing’s off limits at the mart layer - this is the space to get as complex as we need to. ### Further model layers to explore (base, intermediate and beyond) At a minimum, you should familiarize yourself with staging, dimension, and fact models. But you can always opt for more layers to better organize your data! Two optional layers we commonly use are base and intermediate: - **Base models** are prefixed with `base_` and live in the staging folder alongside `stg_` models. If a subset of staging models from the same source lack utility on their own, it may make sense to join them together in a base model before moving downstream. - **Intermediate models** are prefixed with `int_` and live within the marts folder. If your marts models are overly nested + complex to read, splitting some of the logic into one or more intermediate models will help with readability down the line. These are two layers that we commonly use internally at dbt Labs, but feel free to make your own conventions! What’s important is that you make a convention for data model layer naming and follow it, but the specifics will vary widely. ## A modular data modeling example Let’s go back to our salad bar example, and go through each layer from the ground up. ### Source + staging models: the raw ingredients We’ll start with our raw ingredients: the raw data that just exists in my data warehouse. Maybe it made it there via an [ETL tool](https://www.getdbt.com/blog/best-elt-tools) or a custom script, but we’ve got raw data flowing. ![raw data sources](https://cdn.sanity.io/images/wl0ndo6t/main/cf02e55cd99762f38e4cb55658f01e64df08b4af-278x384.png) We’ll refer to these as sources, and the first models I’m going to build are my staging models. ![staging and intermediate data models](https://cdn.sanity.io/images/wl0ndo6t/main/8eee14ea7b11a571b3cc7e872c789d9747551995-1416x1020.png) You can see I’m really just prepping each individual ingredient here. If you think of these things as being available on an assembly line, you can see it’ll be really easy to make a salad. Take note of the Italian dressing. I haven’t prepped this yet - that’s because giving cleaned-up versions of the ingredients (vinegar, Italian seasoning, oil, and water) won’t be of much use to anyone on the assembly line. They would have to mix these ingredients every time to create Italian dressing. We need to do something a little bit different with these ingredients - we need to make a join in the staging layer (!). It’s important to note that I want to produce the results of this join in the staging layer because of how **my organization** uses the ingredients - and that’s an important decision to consider. For the purposes of this demonstration, no one will use water, oil, vinegar, or Italian seasoning on their own. In order to keep a 1:1 relationship with the raw ingredients (between raw sources and staging models), I’m going to implement a base layer (shown in purple). ![joining at staging model layer](https://cdn.sanity.io/images/wl0ndo6t/main/6ba214d3684304715c5ab2c2c04e9e5f8c5d7295-1152x402.png) This layer takes over what staging usually does - the reason we provide this is because we always want to have a model which standardizes our data and provides that layer of defense, whether that becomes a base or staging model. Our downstream models will benefit from those transformations and we can start developing consistency in how our data is commonly used. ### Layering in intermediate models: the basic components Now let’s build the intermediate layer, where I’ll conduct some major data transformations. Unlike base or staging models, this layer is completely optional, but it’s especially useful for creating reusable components or breaking up large transformations into more understandable pieces—in our case, those components would be a basic salad + boiled eggs. ![intermediate data models](https://cdn.sanity.io/images/wl0ndo6t/main/e52ec1df4755929b9fb43dad82820d9834241eae-1690x1014.png) ### Fact + dimension tables: the finished product Finally, in order to complete the salad bar, I’m going to join in all of those steps to make our data marts, our fact + dimension tables. A fact or dimension table brings together multiple components to present a unified whole—in our salad example, it’d bring together the basic salad greens, boiled eggs + Italian dressing to form a ready-to-eat salad: ![data mart models facts dimensions](https://cdn.sanity.io/images/wl0ndo6t/main/49460ccc1d08e2c0114407bb49fc0a890b5fc2b2-844x1072.png) This is by no means the end of the data transformation road. A salad (or any individual dish) usually doesn’t make up an entire meal. In an analysis, we’ll usually pull together multiple dimension or fact tables to build a complete picture. But! Every data user has different tastes, and may want to join data mart models in slightly different ways at the point of analysis—so we’ll generally want to leave that last mile of joining together facts + dimensions to the analysis tool itself when possible. ### Where transformation stops and analysis begins It’s really helpful to define where your data modeling effort ends, and where it’ll be picked up by the end user in an analysis tool (a BI dashboard, notebook, or data app builder). You may end up writing SQL in both your data models and downstream tooling, but your decision for where certain transformations live comes down to the features of your tools and the technical knowledge of your end users (BI analysts, ML engineers, business users). If your end user will be writing SQL to pull the data they need, or if they’re using an analysis tool that joins tables for them, then your final data modeling output can be generalized, standardized fact and dimension tables. They can then freely mix and match these to analyze various aspects of the business, without you needing to pre-model the answer to every question in your transformation project. If your end users don’t write SQL or your analysis tooling is limited in terms of self-serve joining, then you’ll probably need to curate datasets which answer specific business questions, by joining together multiple fact and dimension tables into [wide tables](https://www.youtube.com/watch?v=VQklgbZp94I). If that’s the case, we recommend adding an additional ‘report’ model type to your project, with a naming convention of `report_` or `rpt_`. In our data models in dbt, we’re aiming to bring data together and standardize much of the prep work that comes with making an analysis. We are not looking to pre-build in the data warehouse _every_ analysis or complex aggregation that _may_ come up in the future. That’s a better job for downstream tooling, whether that’s at the BI layer, in ML models, operational workflows, or in a notebook app. Sometimes you just can’t avoid pre-aggregating data within data models, and that’s ok. The important thing is just to define (in your model naming conventions + tooling use standards) a line where data mart construction ends and analysis begins. ## Data model readability A solid naming convention will make our data modeling **project** as a whole easy to navigate. But what about our individual data models themselves, the .sql files that will actually be written to our warehouse? I always strive to keep individual model files to roughly 100 lines of code for high readability. Generally data models shorter than 100 lines have avoided doing overly complex joining, either by limiting the raw number of joins, or by joining in simple ways (repeatedly on the same key). That way, anytime someone on the team (or myself!) cracks open the data model, they can understand very quickly what it does and see how they might modify or extend it. _“OK,_ _`dim_intercom_chats`_ _joins together_ _`stg_intercom_chats`_ _with_ _`stg_customers`_ _to map a customer’s plan to their chat log.”_ Problem is, SQL files can be long and tedious to read! If you’ve migrated between 3 different live chat platforms, you’ll have to UNION ALL on those 3 source tables to roll up your full live chat history. If you’re building a date spine to calculate retention, that requires a lot of boilerplate SQL to generate a date spine. That’s where the magic of dbt + Jinja macros comes in - they allow you to invoke [modular blocks of SQL as macros](https://docs.getdbt.com/docs/building-a-dbt-project/jinja-macros) from within your individual SQL files. When I was doing data consulting, my client calls sounded like iPhone ads. _Need to standardize the way you_ _[build date spines](https://hub.getdbt.com/dbt-labs/dbt_utils/latest/#date_spine-source-macros-sql-date_spine-sql-)? “There’s a macro for that.”_ _Need to model_ _[Snowplow events](https:/