/ /
Databricks processes your data. dbt defines what it means

Databricks processes your data. dbt defines what it means

Daniel Poppy

Last edited on Aug 17, 2026

When Databricks announced it was acquiring Neon in May 2025, it published a number from Neon’s own telemetry: over 80% of the databases on the platform were being created by AI agents rather than by people. That number is about Neon. The question underneath it applies to every platform, Databricks included. When an agent writes your revenue model, someone still has to own the definition. When your CFO asks how last quarter’s number was produced, the answer has to live somewhere you can reach.

This is for data leaders whose teams already run on Databricks and are being asked whether dbt is still worth the overhead.

Short answer: Databricks is an excellent platform, and we’re partners. The dbt-databricks adapter is maintained by the Databricks team. If you’re running ML pipelines, building a lakehouse, or doing serious AI feature engineering, it belongs in your stack. Thousands of data teams run dbt and Databricks together in production.

What I want to challenge is a different assumption: that choosing Databricks as your compute platform and choosing Databricks to own your transformation logic are the same decision. Most companies make both at once without noticing. They’re not the same decision.

Databricks is good, and that’s not the point

Lakehouse compute, Delta tables, collaborative notebooks, ML pipelines. These are real capabilities that deliver real value.

Databricks has also been building products that sit squarely in the transformation layer: Lakeflow Declarative Pipelines (formerly Delta Live Tables) for pipeline orchestration, Databricks SQL for query and analytics work, Unity Catalog for governance and lineage. Each is reasonable on its own. Together they represent a set of decisions about where your transformation logic lives.

Choosing Databricks for compute is an infrastructure decision. Choosing Databricks to own your transformation logic is a strategic bet on where your data team’s institutional knowledge will live. Most executives sign off on the first without realizing they’ve also made the second.

What owning the transformation layer actually means

The risks of full consolidation aren’t theoretical. They become concrete the first time you try to leave, audit, or explain.

Start with portability. If your transformation logic lives in Lakeflow pipelines and Databricks notebooks, moving it means rewriting it. That exit cost won’t appear in today’s contract negotiation. It’ll appear as a multi-month engineering project you didn’t budget for, triggered by a pricing change or a strategic pivot that seemed distant when you signed.

Then auditability. When your CFO asks why Q3 revenue was revised, tracing the answer requires Databricks platform access rather than a git diff. Version control of transformation logic depends on the platform, which means it depends on a vendor’s access controls and pricing tier.

And timing. Enterprises don’t feel vendor lock-in when things are working well. They feel it when pricing changes, when they want to run workloads somewhere else, or when the platform pivots. By then the transformation layer is load-bearing infrastructure.

Databricks would counter that Unity Catalog provides lineage and governance. That’s true. But lineage in a proprietary catalog is a record of what happened. Transformation logic in open, version-controlled SQL gives you the ability to change what happens, and to do it independently, anywhere.

What dbt gives you that Databricks can’t replace

dbt solves a different problem than Lakeflow Declarative Pipelines. It puts transformation logic in open, version-controlled SQL that any engineer can read, test, and run, regardless of what’s underneath.

Five things dbt gives you that a Lakeflow pipeline can’t.

One project, any warehouse. The same dbt project runs on Databricks, Snowflake, BigQuery, DuckDB, and a growing list of compatible platforms. The transformation logic is yours. When your warehouse strategy shifts, you port the logic instead of rewriting the pipelines.

Open standards, and the layer above them. On June 1, 2026, dbt Labs open sourced the dbt Fusion engine runtime, releasing it as dbt Core v2.0 in alpha under the Apache 2.0 license. Databricks has made its own open-source moves here too, donating its declarative pipelines framework to Apache Spark in 2025. The difference is what sits on top. An open runtime is table stakes. What your team actually needs to own is the layer above it: the tests, the contracts, the metric definitions, and the lineage that explain what a number means and prove it hasn’t quietly changed.

A semantic layer that travels. Metric definitions in the dbt Semantic Layer are platform-agnostic by design. Define revenue once and the definition holds whether the query lands on Databricks or somewhere else.

A community organized around a standard. More than 100,000 data teams build on dbt. That community coheres around a shared, open standard for transformation rather than around a single vendor’s roadmap, and that’s worth weighing when you decide where your team’s institutional knowledge will live.

Cost control that isn’t tied to a platform. dbt State works as a caching layer for your pipelines: it builds what’s changed and skips what hasn’t. Teams see warehouse compute drop by 30% or more. It runs on dbt Core 1.7+, as a plugin or out of the box in the dbt platform, so the savings don’t depend on consolidating your transformation layer anywhere in particular. This is the shape of the argument in miniature. The efficiency comes from the engine understanding your project, not from the platform owning it.

Five questions executives should ask before consolidating

  1. If we needed to move 30% of our data workloads off Databricks in 90 days, what would that require?
  2. Can we produce a complete audit trail for every transformation that produced last quarter’s revenue number, without Databricks access?
  3. Are our metric definitions in code, or in Databricks notebooks?
  4. Does our data team own the transformation logic, or does it live inside a vendor product?
  5. If the pricing on our transformation tooling changes next year, what’s our fallback?

Each of these describes a situation organizations have already faced at scale. The executives who ended up in difficult positions weren’t naive. They made the two decisions together, without realizing they were separate.

The right boundary: what each layer owns

“dbt vs. Databricks” is the wrong frame. We’re partners, and the combination is production-proven. The risk shows up when companies blur the boundary between the two.

Here’s the boundary that holds.

Layer

What it owns

Databricks

Compute, storage via Delta, ML pipelines, AI feature engineering, collaborative notebooks

dbt

Transformation logic, the semantic layer, data contracts, lineage in code, testing

Databricks processes your data. dbt defines what it means. Those are two different jobs, and conflating them is where the risk comes from.

The executives who regret full consolidation are the ones who approved it while everything was working, before they needed to explain a number to the board, migrate a workload, or watch their data team’s institutional knowledge walk out the door embedded in Databricks notebooks.

dbt is how your organization keeps the ability to reason about its own data, independently of any single vendor.

See how dbt and Databricks work together. Explore the integration

Get started in dbt

Join the analytics engineers building data infrastructure that actually scales.

Install dbt Wizard CLI

Get started with an agent purpose-built for analytics engineering. It knows which tool to call, which context to pull, and checks its own work before surfacing anything to you.

Share this article
The dbt Community

Join the largest community shaping data

The dbt Community is your gateway to best practices, innovation, and direct collaboration with thousands of data leaders and AI practitioners worldwide. Ask questions, share insights, and build better with the experts.

100,000+active members
50k+teams using dbt weekly
50+Community meetups