/ /
With dbt State, stop rebuilding what didn't change

With dbt State, stop rebuilding what didn't change

Daniel Poppy

Last edited on Oct 05, 2026

Assume you have two dbt jobs. They're in the same project and run the exact same models in the same hour.

The difference? One of them rebuilt everything in the data warehouse, while the other produced the same accurate, up-to-date result at far lower cost.

That's the power of dbt State. Now in general availability (GA), dbt State tracks whether each model's logic or upstream data has changed and reuses the previous results if neither has. Across the dbt platform, more than 22 million models build every day — and most of them rebuild even when nothing upstream has changed.

You can see this in practice. We put two hourly jobs side by side, one with dbt State on and one with it off. You could see the difference on screen, with dbt State the clear winner.

Let’s look at what dbt State does, analyze a live run, and run the numbers to show how much it can save you in practice.

The side-by-side, before and after

Compare two jobs that run hourly. The only difference: one has dbt State activated, and one doesn’t.

On the job without dbt State, a run shows that every model and test ran.

On the job with dbt State, however, the dashboard look a lot different. Nearly every model is marked as “reused.” Tests also don’t run for reused models, as those have run and passed previously. Expanding one model, stg_suppliers, explicitly shows that the model is reused because none of its upstreams have data changes.

The best part is that this is model-level. If you had a third job that also used the stg_suppliers model, it would reuse the last valid run from this job, assuming the data sources still hadn't changed. That makes dbt State a powerful feature for optimizing performance across all of your dbt-based data pipelines.

Three actions, not one

dbt State works by comparing each node's compiled SQL logic and its upstream data freshness against the last good build. Based on the result, it takes one of three actions:

Reuse from same schema. If the object already exists in the target schema, the logic hasn't changed, and the upstreams haven't changed within the specified lag_tolerance, dbt leaves the object alone. The job reuses the results from the last build across all jobs.

Clone from a different schema. If dbt finds a matching object with identical logic and fresh data in another schema, such as one built in a development or CI run before the model reached production, it clones that object instead of rebuilding it.

Build. If dbt can't find a model to reuse, it rebuilds the model.

The principle here is that dbt follows a "cheapest valid action" model, prioritizing validity first and cost second.

How to configure and what to watch out for

Another benefit is that dbt State works well with the default configuration, but you can tune it further to your use cases. There are a number of config options you can set on a per-model basis. A few of the key ones include:

  • lag_tolerance: How much time must pass since the last upstream data change before a rebuild happens. You can set this value higher if you know your data needs are pegged to a specific cycle, such as an analytics report that only needs a weekly refresh.
  • require_fresh_data_from: Controls whether a rebuild triggers when any of the sources update, or only once all of the sources have updated.
  • compare_unrendered_code: Whether to compare a model's unrendered code against the last build when deciding whether to rebuild. [Editor's note: the GA launch post references compare_unrendered_code where this draft originally had evaluate_volatile_sql

To see exactly why dbt made the call it did on a given run, use dbt state explain — it shows which nodes were built, skipped, cloned, or deferred, and why.

Moving to dbt State initially takes only minor configuration changes. If you previously used the state-aware orchestration feature, you'll need to map the old build_after value to the new lag_tolerance and require_fresh_data_from values. If build_after doesn't exist, dbt falls back to its defaults of lag_tolerance: 45m and require_fresh_data_from: any.

After this, the only thing to watch is how you estimate pricing. dbt State charges $0.094 for each daily active target table, with every test counted as a separate target table. In other words, if you have a single model with not_null and unique tests on one column, your total cost is $0.282 for each day it's reused, no matter how many runs reuse it that day.

Where dbt State won't help

Not every workload benefits from dbt State. In the following three cases, your models will always be rebuilt, because there's no reliable way to cache the results:

  • Views with select *, because dbt can't resolve the column names without a query. You can bypass this behavior by excluding views from your build with dbt build --exclude config.materialized:view.
  • Non-deterministic Jinja (the template language used by dbt models), such as dbt_utils.get_relations_by_pattern with union_relations, which returns relations in varying order and so changes the hash every run.
  • BigQuery external sources, unless loaded_at_field or loaded_at_query is set.

Fortunately, in most cases, it only takes minor code changes to upgrade these usage patterns to ones that dbt State can cache. If you're not getting the performance benefits you expected by enabling dbt State, check these culprits first.

What teams save

We now have enough customers running dbt State in production to know what companies can‌ save.

The results are substantial.

Parag Shah, vice president of data at CarGurus, for example, reported that, just by turning on dbt State, his company saw a 9% compute reduction, with 35% fewer models built. This resulted in a 15% reduction in Snowflake backfill costs.

At Fanatics Betting and Gaming, senior analytics engineer Alvin Chai said a pilot project went from 0.2% to roughly 15% model reuse, with some projects reaching as high as 25%. The team saw double-digit compute savings on its very first project.

At Virgin Media O2, the head of analytics engineering put it simply: "dbt State created a paradigm shift in how we work. With capacity freed up, freshness codified, and simpler, yet smarter orchestration, we can focus on initiatives that add value to our business." The speed gains show up at the model level, too. At Joe & The Juice, iteration on a 400 million-row fact table dropped from 15-25 minutes to seconds, with manual cloning or deferral setup eliminated entirely.

At RxBenefits, principal data engineer Chris Shepherd reported that since rolling out dbt State, his team has cut costs by 59% on scheduled jobs in the dbt platform running on a Snowflake adaptive warehouse. By reusing more than 700,000 models instead of rebuilding them, the team also saved two weeks of query run time over a 60-day period.

These savings enable teams to move faster. Obie saved so much that they increased their key data pipelines from daily to once every two hours. That means stakeholders get fresh data throughout the day, which allows them to make business decisions faster and with greater precision.

See the bill

A common complaint about cloud platforms is that costs aren't always front and center. Teams are pleased with the technology, but feel sticker-shocked when the first bill lands.

dbt Cost Insights provides full transparency into how much you're paying for dbt State, and how much you're saving. Once you configure your data warehouse costs, Cost Insights will give you a ballpark readout of daily spend.

Cost Insights is GA for Snowflake, BigQuery, and Databricks, and still in preview for Redshift.

Where dbt State runs

Because it's a separate service, dbt State will run no matter how you use dbt: on the dbt platform or via your own orchestration.

dbt State also runs across environments. Most of us will have at least three environments for managing data pipeline changes:

  • A local development environment for each developer.
  • An isolated staging environment, orchestrated via continuous integration (CI), that runs tests and validates functionality in a production-like setting before release.
  • Your production environment where release code and production data live.

dbt State works across all three environments. Your teams will save money and time with every pipeline run:

  • From dev machines while performing ad hoc testing.
  • From isolated CI environments on every approved pull request.
  • From production on scheduled or event-driven runs.

This does more than lower your costs. Because data producers spend less time waiting in every stage, it reduces the time required to surface changes to production.

Stop rebuilding what didn't change

dbt State is a decision engine. With every run, it decides whether to rebuild the model or use the results from the most recent run. Because it makes this decision per model, any job that uses a model can benefit from faster runtimes without sacrificing accuracy.

You can see this in action, along with the other features of the dbt platform that improve developer productivity and reduce costs, by watching the full demo.

Get started in dbt

Join the analytics engineers building data infrastructure that actually scales.

Install dbt Wizard CLI

Get started with an agent purpose-built for analytics engineering. It knows which tool to call, which context to pull, and checks its own work before surfacing anything to you.

Share this article
The dbt Community

Join the largest community shaping data

The dbt Community is your gateway to best practices, innovation, and direct collaboration with thousands of data leaders and AI practitioners worldwide. Ask questions, share insights, and build better with the experts.

100,000+active members
50k+teams using dbt weekly
50+Community meetups