# Section: Case studies --- title: "Boatsetter increases return on ad spend by 100%+ and scales AI" description: "Boatsetter, one of the world's largest peer-to-peer boat rental marketplaces, connects millions of guests with boat owners across more than 700 destinations worldwide, and more than doubled its return on ad spend after unifying its customer data." url: "https://www.getdbt.com/case-studies/boatsetter" date: "2026-08-24" industry: "Transportation & Logistics" --- # Boatsetter increases return on ad spend by 100%+ and scales AI Boatsetter, one of the world's largest peer-to-peer boat rental marketplaces, connects millions of guests with boat owners across more than 700 destinations worldwide, and more than doubled its return on ad spend after unifying its customer data. ### Company details - Headquarters: Florida, United States - Data stack: dbt platform, Fivetran, Snowflake - Data sources: Postgres, Google Ads, HubSpot, Iterable, Facebook Ads ### Results - 100%+ ROAS improvement year over year - 200 hours Data engineering time saved weekly - 6 weeks Getmyboat data integrated post-merger > "The value of Fivetran and dbt isn't just about moving data faster — it's giving every team the same, increasingly in-depth understanding of the customer journey. Gaining that deeper understanding helped Boatsetter to improve return on ad spend by more than 100%, support 2 marketplaces with a lean analytics team, and use AI where it matters most: creating better experiences for guests and boat owners.” > > — Mark Stange-Tregear, SVP of Data & Operations at Boatsetter/Getmyboat Boatsetter and Getmyboat, which merged in January of 2026, form one of the world's largest peer-to-peer boat rental marketplaces, connecting millions of guests with boat owners across more than 700 destinations worldwide. Unlike traditional ecommerce businesses, every booking depends on a complex mix of customer behavior, geography, seasonality, pricing, weather, and boat availability. To continue to grow efficiently, Boatsetter needed to understand not only what customers booked, but how they discovered, evaluated, and chose those experiences. When Mark Stange-Tregear joined Boatsetter as SVP of Data & Operations, there was rich data, and huge potential to gain visibility into customer lifecycles. Customer, marketing, business, and transactional data were siloed, and challenges with pipeline reliability and inconsistent reporting made it difficult to get the data that was wanted to make key business decisions. Without a complete view of the customer journey there were challenges in identifying the marketing investments that were truly driving bookings, understanding where marketplace demand was outpacing supply, and how to put trusted customer insights into the hands of the marketing and operations teams responsible for acting on them. As Boatsetter looked to continue scaling the business — and later integrate Getmyboat following the merger — it needed more than better reporting. It needed a data foundation that could seamlessly connect information across the business, create a consistent view of customers and bookings, and deliver deeper insights so that marketing, operations, and future AI initiatives could all work from the same reliable information. ### Building a complete view of the customer journey The goal wasn't simply to modernize Boatsetter’s data platform — it was to understand the complete customer journey and make those insights available wherever decisions were made. To do that, the company rebuilt its data platform around a cloud-native architecture designed to automate data movement, standardize business definitions, and put customer insights to work. **Fivetran** automatically ingests data from Boatsetter's operational systems, product databases, marketing platforms, customer engagement tools, and third-party applications into **Snowflake**, creating a unified view of customer behavior, booking activity, and business performance. **dbt platform** is** **then used to transform that raw data into governed business models while serving as the operational backbone of Boatsetter's analytics environment. Beyond centralizing business logic, dbt provides [orchestration](https://www.getdbt.com/blog/using-state-aware-orchestration-to-slash-your-data-costs), upstream and downstream testing, lineage, and alerting — eliminating the need for separate tools while giving the team confidence that reporting, marketing, and AI all rely on consistent, validated data. Following the Getmyboat merger, Boatsetter also adopted [dbt Mesh](https://www.getdbt.com/product/dbt-mesh) to manage analytics development across both organizations while maintaining shared business models. Rather than rebuilding audiences and customer attributes in every downstream application, Boatsetter uses Fivetran Activations** **to deliver the governed customer models built in dbt directly into Google Ads, Meta, and HubSpot. This gives business teams access to the same validated customer and marketplace data used for analytics, without maintaining separate models or logic across their day-to-day tools. > “The real breakthrough with Fivetran Activations came when customer journey data became part of every marketing investment decision. We stopped asking, ‘What happened?’ and started asking, ‘Where should we invest next?’” — Mark Stange-Tregear, SVP of Data & Operations at Boatsetter/Getmyboat Teams use Tableau and Hex for reporting and interactive analytics, while the same governed customer models support AI use cases including customer journey analysis, owner recommendations, internal analytics assistants, and marketplace operations. ### Turning customer insights into measurable business growth Understanding the complete customer journey enabled more strategic decisions across marketing, product, and the owner ecosystem. By connecting website behavior, booking history, marketing interactions, and customer communications, Boatsetter gained new visibility into how guests discover, evaluate, and book experiences. With Fivetran Activations delivering those insights into the platforms where campaigns are managed, marketers could more precisely identify which audiences, channels, and moments in the booking journey were most likely to convert — and reallocate spend accordingly. As a result, Boatsetter **improved return on ad spend (ROAS) by more than 100% year over year**. Instead of broadly investing in customer acquisition, the company now uses customer insights to target high-value audiences, reduce ineffective advertising spend, and generate more bookings from every marketing dollar. That understanding now informs owner acquisition, inventory planning, and AI-driven experiences across the entire marketplace. The company has: - **Saved an estimated 200 data engineering hours every week**, allowing analysts to focus on solving business problems instead of maintaining pipelines. - **Integrated Getmyboat's priority data sets in just 6 weeks post-merger **by scaling analytics with dbt Mesh, accelerating post-acquisition reporting and decision-making. - **Improved owner acquisition** by identifying where customer demand was highest. The same customer intelligence that improved marketing performance now powers AI across the business. Boatsetter uses Snowflake AI functions directly within dbt to enrich its data models with customer sentiment analysis and common question detection as part of every data build. [AI-enriched models](https://www.getdbt.com/blog/ai-data-modeling) are starting to power owner pricing, intelligent trip replacement, and emerging agentic capabilities that automate routine work. Rather than replacing people, AI helps owners improve performance, matches guests with the right experience, and helps employees make faster, more informed decisions. > "Because Snowflake AI functions are integrated directly into our dbt models, every recommendation starts from the same business logic that powers reporting. That lets us confidently automate the evaluation of pricing, listings, customer sentiment, and marketplace operations without maintaining separate AI pipelines." — Mark Stange-Tregear, SVP of Data & Operations at Boatsetter/Getmyboat Looking ahead, Boatsetter is working with Big Context & Co, a new player in the space, to manage data context and deliver data agents and a range of other AI-based tools to employees. Boatsetter plans to extend Fivetran Activations beyond today's marketing and HubSpot use cases and sees dbt evolving beyond analytics into the knowledge layer for both employees and AI agents, giving every decision and recommendation the same standardized business context. --- --- title: "How RMIT University evolved its data platform without a disruptive redesign using the dbt Fusion engine" description: "RMIT University, an international university with 100,000+ students, cut data pipeline wait times from minutes to seconds as its shared analytics platform scaled across 20+ teams." url: "https://www.getdbt.com/case-studies/rmit-university" date: "2026-08-24" industry: "Education" --- # How RMIT University evolved its data platform without a disruptive redesign using the dbt Fusion engine RMIT University, an international university with 100,000+ students, cut data pipeline wait times from minutes to seconds as its shared analytics platform scaled across 20+ teams. ### Company details - Headquarters: Melbourne, Victoria, Australia - Data stack: dbt platform, Snowflake, AWS, GitHub ### Results - 3,500 models supported in a single project - 5,500 data tests supported in a single project - 88%+ faster platform performance > “With the dbt Fusion engine, the performance improvements showed up across the entire development experience. Teams spent less time waiting for processes to complete, moved changes through the pipeline faster, and could focus more of their time on building and delivering data products.” > > — Vishesh Jain, Delivery Lead, Data Analytics Platform at RMIT University ### **Building a shared analytics platform trusted across the university** RMIT is an international university of technology, design and enterprise with more than 104,000 students and over 13,000 staff. It has campuses in Australia and Vietnam, a research and innovation hub in Spain, partner-delivered programs across Asia, and research and industry partnerships worldwide. In 2022, RMIT launched a centralized data analytics platform on AWS, Snowflake, and dbt platform to bring previously distributed analytics teams onto a shared foundation with common governance and engineering standards. “When this platform was built, it was built around a principle of bringing people to data, rather than data to people," says Jain. From the outset, dbt gave RMIT University a transformation layer plus lineage, documentation, and catalog capabilities that helped create a shared foundation across the university. This platform sits inside a wider central analytics function. Before that shift, analytics teams were distributed across different parts of the university, each using different tools and processes. Gradually, the platform became part of a broader strategy to create shared standards, reusable engineering patterns, and centrally trusted data products for the university’s data team. As the platform scaled, adoption had grown to 20+ teams, around 70 to 80 dbt users, and roughly 30 to 40 core developers. The main analytics project alone had about 3,500 models and nearly as many tests. That growth showed the platform strategy was working. ### **A platform ready for its next stage** "The platform was doing exactly what we'd designed it to do. More teams were using it, more data products were being built, and adoption was growing. The challenge became how to maintain a great developer experience as everything scaled," says Jain. The friction was cumulative. Parsing, compilation, validation, and pipeline checks each added waiting time. On their own, those delays looked manageable. Combined, they slowed how quickly the team could deliver new work across the university. "If I needed to deploy code, I’d start planning days in advance because I knew the pipelines and processing would take time. Even small changes take longer than they should, and that slowed everything down," says Jain. RMIT's initial response was architectural. The team started planning a project split, with dbt Mesh as the path forward. But this wasn't just reorganizing code. It would have meant moving to a multi-project operating model, reworking Snowflake RBAC, setting up cross-project enablement, and changing how teams developed and collaborated. The platform was already working, and all of that effort was just the cost of keeping performance up at the next scale. ### **From 6,100 migration blockers to Fusion-ready in four days** "We were planning to split the project and redesign parts of the architecture, but when Fusion entered the conversation we asked ourselves: can Fusion solve the problems we're trying to solve? That question changed our approach," says Jain. Fusion addressed parsing and compilation bottlenecks within the existing shared platform, without making that disruptive redesign the immediate next step. At the start of 2026, the team migrated from Bitbucket to GitHub, improving integration with the dbt platform and standardizing the development workflow ahead of the upgrade. The first immediate challenge was technical debt. The project had approximately 6,100 deprecations that needed to be resolved before Fusion could deliver its full benefits. Rather than tackling them manually, the team analyzed common patterns and used dbt-autofix to accelerate the process. All of this happened while the platform was still running on dbt Core, with no disruption to ongoing work. "Within a few hours, we were able to fix more than half of the deprecations. That gave the team confidence that this was actually possible," adds Jain. Four days later, all 6,100 had been resolved, and a quality gate in the CI/CD workflow ensured no new ones could be introduced. To minimize disruption, RMIT University tested Fusion first in a smaller governance project used for telemetry and observability. This surfaced macro behavior that triggered continuous queries in Snowflake, as well as manifest-related issues, allowing the team to resolve problems before expanding deployment across the wider platform.The team took the same careful approach with orchestration, selectively enabling a capability that reuses unchanged assets rather than rebuilding them only where it improved efficiency. ### **Fusion drops parsing to 18 seconds and cuts pipeline runtimes by more than half** The first improvement appeared almost immediately. Parsing times in the university's primary analytics project dropped from approximately three minutes to roughly 18 seconds. For a repository containing thousands of models and tests, that dramatically shortened development feedback cycles. The team then revisited its broader CI/CD workflows, replacing dbt Core-based checks with Fusion-based processes and simplifying the steps developers needed to run before pushing code. As a result, average feature pipeline runtime fell from approximately eight minutes to three minutes before the dbt platform job itself began. Users also began seeing improvements in production environments. Overnight jobs, release processes, and other production workflows completed faster following the rollout. Fusion allowed RMIT to continue scaling its existing environment and improve the developer experience without first undertaking the planned platform redesign. The decision didn't remove the need for future architectural stewardship. It changed the immediate path and avoided layering a broad organizational change on top of a technical migration. The migration delivered benefits beyond performance. By working through the rollout and understanding how Fusion behaved in practice, the engineering team became more confident in identifying and resolving problems when issues occurred. The team also saw approximately 15%–20% reusable assets in workloads where orchestration was enabled. More recently, Fusion combined with model optimization by the engineering team drove approximately 20% lower compute usage through daily pipelines, even as workload continued to increase. > "Fusion gave us a way to achieve the scaling objectives we were aiming for without having to redesign the platform we'd already invested in." - Vishesh Jain, Delivery Lead, Data Analytics Platform at RMIT University The most significant outcome was that day-to-day users experienced very little disruption. Rather than becoming a prolonged change-management program, Fusion simply became part of the platform experience. ### **Creating capacity for the next stage of growth** With the platform performing more efficiently, RMIT is focused on simplifying job design, reducing operational complexity, and identifying further opportunities to optimize compute usage as adoption continues to grow. Longer term, the university wants to bring more teams onto the platform so they can build and manage their own trusted data products within a governed, standardized environment while reusing shared assets and established engineering practices. "People coming on board are not just consumers. They can benefit from the standards the engineering team has set and the quality of models that have already been written," says Jain. The ambition is to extend that pattern across the university, improving engineering maturity beyond the data team. --- --- title: "Lendi builds AI-native future with dbt and trusted data" description: "Lendi, Australia's leading digital-first home loan platform, cut platform outages to zero and freed its engineering team to build instead of firefight." url: "https://www.getdbt.com/case-studies/lendi" date: "2026-07-27" industry: "Banking & Financial Services" --- # Lendi builds AI-native future with dbt and trusted data Lendi, Australia's leading digital-first home loan platform, cut platform outages to zero and freed its engineering team to build instead of firefight. ### Company details - Headquarters: Sydney, Australia - Data stack: dbt platform - Partner: Data Army ### Results - 30 days to migrate from dbt Core to dbt platform - 0 Downtime during the migration - 0 platform outages since go live > “Moving to dbt helped us increase delivery velocity and strengthen trust in our data. By removing the operational burden of self‑hosting, our teams can move faster while maintaining confidence in the analytics we deliver to the business.” > > — Devesh Maheshwari, Chief Technology Officer, Lendi Group ### **Helping buyers make confident home loan decisions, backed by data** Lendi is Australia’s leading digital-first home loan platform, operating under Lendi Group, one of the country’s fastest-growing fintechs. For the millions of Australians navigating the home loan market, the process has long been slow and complex. Lendi set out to change that, combining smart technology with human expertise to help borrowers compare, choose, and apply for mortgages with confidence. Operating alongside the well-known Aussie Home Loans brand, Lendi plays a central role in closing the gaps between borrowers, brokers, and lenders. Its model blends automation, transparency, and expert advice to deliver a streamlined alternative to traditional mortgage broking. As the business has grown, data and technology have become increasingly central to how Lendi operates at scale. Consequently, Lendi has set an ambition to become a fully AI-native organisation in 2026. > “Becoming AI-native is a critical opportunity for us to break away from the pack by embedding AI at the core of how we operate – across our workflows, decisions and customer experiences,” says David Hyman, co-founder of Lendi. **** ### **Growing analytics demand outpaced a self-hosted platform** Lendi is a data-driven business, with strong demand for data insights spanning executive leadership, operations, finance, and marketing. As this demand grew, it put increasing pressure on engineering and platform teams, who were spending time maintaining a self‑hosted analytics environment built on dbt Core and external orchestration, which was slowing Lendi’s ability to deliver new insights to the business and keep pace with growing demand. Platform instability led to failed pipelines and stale reports - preventing access to data when teams needed it to make decisions - while reduced developer velocity caused analytics work to back up. At the same time, complexity in the development environment made it harder for Lendi's data teams (data engineering, analytics, data science, and AI engineering) to contribute to dbt pipelines effectively, limiting the organisation’s ability to scale its analytics function. ### **Moving from self-hosted to dbt platform** As part of a broader effort to simplify the data stack, Lendi evaluated the full cost of continuing to self‑host dbt, including compute, ongoing maintenance, support overhead and opportunity cost. In a detailed cost comparison between self-hosting and dbt platform, dbt came out as the clear winner. dbt offered a way to offload platform management while strengthening governance, collaboration, and development velocity. By moving to a managed platform, engineering teams could spend less time maintaining infrastructure and more time building analytics that deliver value to the business. Lendi implemented dbt in partnership with dbt consultancy [Data Army](https://dataarmy.io/?utm_source=dbt&utm_medium=referral&utm_campaign=lendi-case-study&utm_content=text_link), configuring the platform using Terraform and integrating with Lendi’s identity management platform, Microsoft Entra, to support role‑based access control. Code repositories were integrated with Bitbucket, with deployments co‑managed alongside existing tooling. This approach helped standardise environments and reduce operational overhead as dbt was rolled out across teams. The migration was completed with no downtime, with the second phase finished within a week. The resulting setup supports multiple projects, teams and jobs serving a range of analytical needs across the organisation. > “Data Army consistently delivered exceptional value at Lendi, quickly absorbing context, integrating seamlessly with our squads, and contributing meaningful outcomes from day one,” says Frank Colubriale, Data Product and Enablement Lead, Lendi Group. “Their contribution continues to position us strongly in delivering on our AI native vision and the increasing pace of innovation across the business.” **** ### **Faster development and greater trust in data** By moving to dbt, Lendi reduced the operational overhead associated with maintaining self-hosted analytics infrastructure. Engineering teams now spend less time managing platforms (an average of 16 hours per month saved) and more time working closely with the business to deliver insights. Standardised workflows, access controls and collaboration features increased development velocity and expanded the pool of contributors (by 10 people) to analytics development. dbt features such as the Web IDE, dbt Mesh and dbt Copilot made it easier for multiple teams to safely contribute to shared analytics pipelines and service business demand. At the same time, trust in data improved across the organisation. Tests, documentation and lineage made data quality explicit rather than assumed, giving teams greater confidence in both upstream changes and downstream outputs. As a result, Lendi experienced fewer pipeline failures, less rework, and more reliable reporting to support decision-making. ### **How Lendi evaluated success** Lendi assessed the impact of dbt across delivery velocity, data quality, and cost efficiency. The team look at a combination of indicators to understand how the platform was improving how they worked: These included: - **Time to onboard new data engineers.** Reduced ramp-up time due to consistent project structure, documentation and CI. - **Time to deliver new models or changes.** Faster development cycles with fewer production incidents. - **Test coverage and freshness SLAs. **Increased automated testing and clearer expectations around data availability. - **Reduction in pipeline failures and rework. **Less time spent debugging downstream issues. - **Platform cost optimisation.** More efficient warehouse usage driven by better model design and incremental patterns. - **Data reuse. **Higher reuse of curated models rather than rebuilding logic in downstream tools. ### **What’s next** As Lendi Group works towards its AI‑native ambition, the team is focused on scaling trusted data products across more teams, enhancing automation and data quality, and preparing its data foundations for advanced analytics and AI. dbt remains the core transformation and governance layer supporting this work. --- --- title: "How PetScreening scales a 5-person data team across 6 business domains with dbt" description: "PetScreening supports property managers across the US, now shipping data updates every 15 minutes with dbt." url: "https://www.getdbt.com/case-studies/petscreening" date: "2026-06-16" industry: "Property Technology" --- # How PetScreening scales a 5-person data team across 6 business domains with dbt PetScreening supports property managers across the US, now shipping data updates every 15 minutes with dbt. ### Company details - Headquarters: Mooresville, North Carolina - Data stack: dbt, Snowflake, Airbyte, Hightouch ### Results - 32x increase in data freshness, with data updates every 15 minutes - 90%+ reduction in time-to-insight, from hours (up to a full day) down to near real time - 1000 pet tags/day – up from a backlog stretching months > "I’d definitely encourage people to use dbt. As your environment grows to hundreds of models, dbt keeps you grounded, with orchestration, lineage and documentation that makes it easy to bring others along." > > — Nick Fernandez, Data Analytics Manager ### **Powering property management decisions with quality data** PetScreening helps property managers handle pet policies, assistance animal requests, and related compliance requirements for residential and commercial properties. Its platform connects data across applicants, residents, and pet profiles. As the business scaled, the reliability and freshness of that data became critical not just for analytics, but day-to-day operations and the customer experience. The lost pet alert program alone depends on accurate-real-time data to produce and deliver engraved pet tags to customers. Consequently, the business relies on the data team to connect, transform, and activate large quantities of data so that data is fresh, accurate, and trustworthy. To meet this challenge, PetScreening relies on a data stack that includes Snowflake for warehousing, Airbyte for ingestion, Hightouch for reverse ETL, and dbt platform to transform and orchestrate data across the organization. ### **A data stack that struggled to keep up with the business** As demand grew across the business, stakeholders began to question whether the data team could keep up. With leadership relying on timely operational and KPI reporting, slow turnaround was starting to affect confidence in the team’s ability to support the business at scale. Before dbt, they connected reporting directly to the product’s Postgres database, which was also the live system powering the application. Beyond the technical risk, this setup was very inefficient, straining the production system since every report ran against the same system handling live customer transactions. PetScreening tried a low-code transformation tool. It worked initially, but as data needs grew across multiple business domains, it could not deliver the governance, robustness and delivery speed required. In a business where leadership relies on timely operational and KPI reporting, slow turnaround quickly became a credibility issue. Data requests often had to wait until someone had time, and business stakeholders began questioning whether the data team could ship improvements quickly enough to keep up with demand. ### **Building scalable analytics without the DevOps overhead** For PetScreening, the priority was to build a foundation that would let the team deliver trusted data more quickly across the organization, without adding unnecessary operational overhead. As a small data team, the appeal of dbt platform was straightforward: pipeline orchestration, Snowflake integration, Git-based workflows, and built-in documentation all in one solution. This effectively provided analytics engineering guardrails and operational structure without the need to build or maintain a separate DevOps layer. They team also made a deliberate choice to build on SQL rather than a low-code approach. A SQL-based foundation meant easier hiring, simpler onboarding, and a codebase that more people across the organization could eventually read and contribute to. From day one, dbt helped establish standardization, documentation, and lineage, keeping the system understandable as it scales and reducing technical debt as it grew. PetScreening worked with Aimpoint Digital to architect the right foundation with dbt for long-term scale. ### **What they built first** Initial adoption concentrated on two priorities: - Core reporting across the business, including KPI reporting, financial reporting, reporting for operational teams like customer support and animal assistance (including workflows for reviewing emotional support animals) - High-impact activation in parallel, with pipelines feeding campaigns and initiatives through Hightouch, plus operational workflows like the pet tag pipeline As delivery cycles shortened, teams across the business became more ambitious. Requests that once waited for someone to have bandwidth became fast-turnaround improvements. More PetScreening employees could share SQL with analytics engineers, understand what was happening and collaborate directly. This helped demystify data work and build shared ownership. > “If you’re a data team of one, by using dbt you have a DevOps team built into that, and you’re building correct foundations from the start.” - Will Guicheney, Principal Analytics Engineer, Aimpoint Digital ### **From fixed batch updates to near real-time delivery** Previously, transformation pipelines ran only 3 times per day, which limited when data updates, fixes and new features could be deployed. With dbt platform, pipelines now run every 15 minutes, reducing latency between development and production availability giving teams access to fresher data. This shift reduced the time it took for changes and insights to reach end users from hours or even a full day to near real time — enabling faster decision making and execution across the business. Nowhere was this more visible than in PetScreening’s lost pet alert program, which alerts nearby users if a pet is reported missing, and enables customers to sign up and receive a free engraved pet tag to connect pets back to their owners. The original process relied on CSVs emailed to an external manufacturing partner, often leading to duplicate or missing data and no robust testing. Backlogs stretched into months, creating a poor customer experience. PetScreening treated the workflow as a data problem: ingest order data, transform it into efficient models, and feed clean, deduplicated models into the tag printing process. dbt made it practical to continually adapt the pipeline as requirements change and add matching logic when orders lacked identity data. Up to 1,000 tags per day are now printed by PetScreening, with improved reliability and faster iteration as printing requirements evolve. > “We built ‘gold-level’ models in dbt – almost like a reverse ETL process. Instead of powering a dashboard, the data powered the tag printing process. The backbone is really the lineage of dbt models that allows us to clearly track how data is built at every step and trust the outputs.” — Nick Fernandez, Data Analytics Manager at PetScreening. ### **Scaling impact without scaling headcount** PetScreening’s five-person data team now supports six distinct business domains, more than 30 direct stakeholders, and hundreds of data users across the organization — without adding headcount. The standardization, documentation, and deployment practices built on dbt platform are what make it possible for a small team to operate at this breadth. Without those foundations, supporting this level of demand would have required a substantially larger team. “dbt is what let us scale our footprint without putting pressure on headcount. Without good standardization, documentation, and easy code deployments a team of our size would be immobilized,” adds Fernandez. ### **A foundation for scaling data and AI across the business** As PetScreening looks ahead, its investment in governed, well-modeled data is laying the groundwork for AI adoption. In fact, the team has already put this into practice. Using dbt and Snowflake Cortex, they built a pipeline that takes messy job title data from HubSpot and outputs a clean, standardized field the marketing team can use for segmentation. The team sees two main opportunities: helping developers move faster by building on trusted models, and giving business users easier access to insights that once required technical support. Both rely on the same foundation of reliable, well-documented models and strong semantic definitions. With the right data foundations in place, PetScreening is now in a stronger position to explore AI use cases across the business. “Having that good baseline of data models and well-defined lineage is the enabler for good results from an AI agent looking at the data,” explains Fernandez. With dbt, PetScreening has built the infrastructure that makes that possible. Governed data powers effective AI, and AI provides consistent, trusted outputs across the business. --- --- title: "impact.com scales a $100B+ data platform with dbt to power decision-making and AI readiness" description: "How impact.com scaled a $100B+ data platform, doubled model output, and cut troubleshooting time from two days to one hour." url: "https://www.getdbt.com/case-studies/impact.com" date: "2026-06-15" industry: "Software" --- # impact.com scales a $100B+ data platform with dbt to power decision-making and AI readiness How impact.com scaled a $100B+ data platform, doubled model output, and cut troubleshooting time from two days to one hour. ### Company details - Data stack: dbt platform, BigQuery, Looker, Fivetran - Data & Analytics Hub: Cape Town, South Africa ### Results - 2–3 days/month engineering time reclaimed from pipeline maintenance - 2x output, zero new hires more models shipped without adding headcount - From 1-2 days to 1 hour time to troubleshoot and find the right data > Before dbt, finding the right data could take a day or two of hunting around, but it now takes about an hour. When you’re working across more than $100B in partnership-driven commerce, that makes a real difference. > > — Paul Kotze, Head of Advanced Analytics ### Operating a global partnership platform at massive scale [impact.com ](https://impact.com/)is a global partnership management platform, powering more than one million active partnerships and processing over $100B in partnership-driven commerce each year. The platform connects brands, publishers, and content creators across affiliate, influencer, mobile, B2B, and referral programs – with data at the center of how the business operates. Internally, a central Data and Analytics Group based in Cape Town leads the company’s analytics function, reflecting its role as a growing hub for data and engineering talent. The group supports every value stream across the business, from finance and marketing to customer success and product. ### A data platform that couldn't keep pace with the business As impact.com scaled, so did the complexity and importance of its data. The organization needed a data architecture that could keep pace for reporting, day-to-day decision-making, and a growing set of AI initiatives. Earlier data models had been built in Databricks notebooks with manually managed dependencies. This approach was difficult to maintain and increasingly fragile as the system expanded. Their data team began self-hosting dbt Core. This brought structure and standards, but introduced new challenges as the team scaled. Most analytics engineers came from SQL and BI backgrounds, not software engineering. Local environments made lineage opaque and modeling practices inconsistent — and there was no easy way to onboard someone new without weeks of ramp-up. At the same time, operational overhead continued to slow progress. Self-hosting dbt Core meant orchestration lived in Jenkins, not dbt. Troubleshooting a failed job meant navigating multiple systems. Deployments required other teams so analytics engineers could not operate independently. This ultimately slowed development, limited the team’s ability to iterate and slowed delivery of insights. Data reliability also suffered. Without systematic testing or monitoring, issues often went unnoticed until stakeholders flagged them, sometimes during critical quarterly business review cycles. At scale, this created a deeper problem. Teams did not always trust the data, and time was spent reconciling numbers instead of using them to make decisions. ### From tickets and workarounds to self-service and trusted data impact.com transitioned to dbt platform to improve how its data platform was built and operated, with a clear focus on developer experience and team enablement. “The main problem we were trying to solve was having a team who could write SQL and understood the business, but didn’t yet have the fundamentals of building a data warehouse - dbt helped us close that gap,” explained Kotze. One of the most immediate changes was improved visibility and discoverability with lineage, documentation, and dependencies, all accessible within dbt platform. Analytics engineers could finally see how their models fit together, and monitor and troubleshoot their jobs without filing a ticket, waiting on another team, or relying on external orchestration tools. Most importantly: the time taken to locate data reduced from ~2 days to around an hour, supporting faster decisions across the organization. Another major shift was improved data quality. Before, data quality issues were invisible until they weren't. There was no systematic way to catch problems before they reached the business. That changed with a comprehensive testing framework spanning over 2,800 tests across the warehouse, covering freshness, uniqueness, and integrity. The team now catches issues before they reach stakeholders, and the team spends far less time firefighting and more time building. The third major shift was architectural. impact.com restructured its warehouse from a fragmented collection of reporting datasets built in Databricks notebooks with manually managed dependencies, into a layered, governed architecture that models how the business actually operates. Instead of different teams maintaining their own versions of the same logic, there is now a single source of truth used consistently across teams. ### Empowering the team to build and problem-solve The warehouse grew from ~800 to 1,700 models, without adding headcount. “We doubled the rate at which we could architect, conceptualise and deliver models,” said Kotze, adding, “We were spending two to three days a month maintaining pipelines but moving to dbt freed that time up so the team could focus on building instead. It gave the team capacity to focus on other important work and freed up their headspace, so they didn’t have to worry about running the warehouse.” The visibility and structure included in dbt platform also changed how the team catches and resolves problems. With proactive testing and clear lineage, problems are detected proactively before they surface in reports, reducing rework, cutting technical debt, and rebuilding trust in the numbers. That trust has changed how the business uses data. Leadership operates from shared dashboards built on a single source of truth. Operational teams rely on data as part of their day-to-day workflows, not just occasional analysis. Ad hoc requests have dropped because insight is built directly into the platform. ### What’s next What started as a central data team initiative is now becoming infrastructure for the whole company. Engineering teams are building customer-facing products directly on top of the warehouse. Others are experimenting with the semantic layer to support AI use cases. As more teams start to work with the data, maintaining consistent definitions becomes increasingly important across impact.com. impact.com isn't waiting on AI readiness. They're building for it now. dbt’s Semantic Layer gives every agent and application a single governed interface into the warehouse, where metric definitions and relationships are explicit, not inferred. That means no hallucinated connections between tables, no inconsistent definitions as more teams build on top. Some product and engineering teams are already piloting it, with a broader rollout targeting the end of June 2026. Their ultimate goal is that every team, not just the data team, runs on trusted metrics. --- --- title: "How Kaizen Gaming cut costs and delivers insights faster" description: "Learn how dbt helped Kaizen strengthen its data foundation, increase consistency, and improve operational efficiency" url: "https://www.getdbt.com/case-studies/kaizen-gaming" date: "2026-03-30" industry: "Gaming" --- # How Kaizen Gaming cut costs and delivers insights faster Learn how dbt helped Kaizen strengthen its data foundation, increase consistency, and improve operational efficiency ### Company details - Headquarters: Athens, Greece ### Solution With a standardized framework, testing strategy, and governance in place, Kaizen has established dbt as a core pillar of its analytics capabilities. What began as an effort to improve data workflows has evolved into a stronger data foundation with meaningful improvements in cost efficiency. ### Benefits - Reduced costs - Faster data delivery - Increased productivity ### Results - 60% reduction in pipeline execution times — Downstream teams, such as product, marketing, and commercial functions, are accessing insights sooner and making faster decisions based on fresher data. - 90% fewer workflows required to support analytics use cases — The consolidation of analytics workflows has reduced operational complexity, making it easier to understand, operate, and maintain pipelines while reducing cognitive load for engineers. - 60% decrease in overall pipeline costs — More efficient execution and reduced redundancy have led to a significant reduction in daily pipeline costs while supporting the same data volumes and analytical use cases as before. > “dbt gives us flexibility. By adopting it, we can remain vendor-agnostic. If we decide to expand our data stack in the future, it will be easier to do so.” > > — Stefanos Nikolaou, Principal Analytics Engineer Kaizen Gaming is one of the biggest GameTech companies in the world and the owner of the premium online sports betting and gaming brand, Betano. Today. Kaizen Gaming counts more than 3,000 people across the globe, while it has been consistently recognized for operational excellence and top tier customer experience. In 2024 and 2025 it received the "Operator of the Year” awards in both the EGR Operator Awards and the SBC Awards, the igaming industry’s most prestigious accolades. Kaizen Gaming serves a large and growing customer base across multiple markets in Europe, the Americas and Africa. The company processes a high volume of transactions worldwide and must stay compliant with complex regulatory frameworks across jurisdictions. To support these demands, Kaizen Gaming’s data organization plays a central role across the business, from marketing and engineering to finance, risk, and compliance. As the company’s footprint and ambitions expanded, the data organization has grown in lockstep. In fact, Kaizen Gaming has grown so quickly that the data team itself has more than doubled in less than six months. With that expansion came new expectations for Kaizen Gaming’s analytics environment along with a new stage of maturity: systems and practices that had worked well at earlier stages needed to evolve to support coordination and data reliability at scale. ## New demands for reliability, visibility, and operational efficiency Kaizen Gaming’s broader data organization is structured around several domain-aligned teams. Each team maintains a high degree of autonomy, which allows them to build deep domain expertise and move quickly. As Kaizen Gaming’s analytics organization grew, teams updated their workflows to meet their domain-specific priorities. Over time, this led to a system of well-functioning individual components, but without a shared framework to connect them. For example, pipelines were primarily developed in notebooks using multiple supported languages, including SQL, Python, and Scala. The same flexibility that enabled domains to move quickly also made it challenging for the team to establish consistent coding practices or enforce SQL-first development broadly. As a result, three distinct challenges emerged: 1. **Data reliability. **When Kaizen Gaming’s workflows became more complex and interconnected, the ability to validate data quality earlier in the lifecycle became increasingly important. This highlighted the need for more integrated and proactive validation across downstream reporting and analytics. 2. **Operational efficiency.** Logic duplication across domains made it challenging to maintain consistent metrics: investigating issues could be time-consuming and testing could happen later in the cycle. At times, the team had to address issues later in development rather than earlier in the lifecycle. 3. **Cost and operational overhead.** Managing and maintaining pipelines required a growing level of manual coordination, from updating logic to understanding dependencies across workflows. As data volumes increased, the need for a more scalable and efficient approach became apparent. It was clear that certain aspects of the analytics environment — particularly automation, quality, standardization, and cross-team consistency — needed to change to meet new organizational demands. This was especially evident when more analytics workflows began running during off-peak hours. Without a shared framework to guide development and investigation, troubleshooting issues was painful. Diagnosing failures or subtle inconsistencies took significant effort, and even more resources if the issue was complex. Kaizen Gaming’s analytics practices needed to evolve. As its systems became more interconnected, even small inconsistencies could require increasing effort to identify and explain. Equally difficult was the matter of assessing downstream impact, such as the effects of new features or changes. To maintain trust in the data, the company needed clearer structure, stronger validation, and shared development patterns. To address these challenges, it decided to establish a more unified approach to analytics. Kaizen Gaming’s goals were to strengthen data quality, improve operational efficiency, and provide a consistent foundation across teams. The data team chose dbt as a core framework. ## A unified analytics framework with dbt To optimize the team’s data workflows, Yannis Lazaridis, Analytics Engineering Lead at Kaizen Gaming, began exploring the dbt platform to understand how it compared with other approaches. The team soon found that dbt offered a more structured and transparent alternative to notebook-driven workflows. dbt’s modular modeling, version-controlled development, and explicit dependencies proved particularly valuable in enabling the broader data organization to contribute without needing to navigate Python- or Scala-heavy pipelines. The data organization quickly adopted dbt as a common foundation across domains, establishing shared practices while maintaining flexibility. As a result, Kaizen Gaming saw the following improvements: - **A unified modeling approach. **Kaizen Gaming implemented a layered structure for sources, staging, intermediate models, and marts. Kaizen Gaming’s data engineers now have a predictable way to develop, review, and maintain analytics models. - **Comprehensive, automated testing. **Data quality tests now run automatically with every pull request; they combine built-in dbt tests, dbt-expectations, and custom SQL checks. This proactive approach surfaces issues sooner in the development lifecycle, before changes reach production. - **CI/CD for data quality. **The team introduced Slim CI on every pull request. They also adopted state-based execution, which validates only the models that were affected by a given change during releases. As a result, the team has reduced unnecessary runtime while enabling a more automated deployment process. They now have stronger confidence in data quality while working more efficiently. - **Code-linked documentation. **Models and metadata are documented directly in YAML and kept in sync with code through CI enforcement — a major improvement for data governance. This shared understanding of data structures across domains has sped up onboarding for new team members. - **Infrastructure-as-code for dbt.** Jobs, environments, variables, and connections are managed through Terraform. The team can ensure consistent environment replication and stronger governance across the analytics landscape. “With dbt, we can do data quality checks before a job is completed,” says Thomas Antonakis, Principal Analytics Engineer at Kaizen Gaming. “We have visibility into what goes wrong and can proactively stop a flow without interfering with our production tables.” ## Reduced costs, faster data delivery, and increased productivity With a standardized analytics framework in place, Kaizen Gaming began to observe clear, measurable improvements across its analytics workflows. These gains illustrate how reduced complexity, improved execution efficiency, and more consistent development practices can benefit an organization: ### Workflow simplification With dbt, the number of workflows required to support core use cases decreased by over 90%. The consolidation of analytics workflows has reduced operational complexity, making it easier to understand, operate, and maintain pipelines while reducing cognitive load for engineers. ### Improved runtime and data availability Pipeline execution times have decreased by about 60%. In fact, data is delivered more than an hour earlier than before. Downstream teams, such as product, marketing, and commercial functions, are accessing insights sooner and making faster decisions based on fresher data. “We’ve seen runtimes decrease from over two hours to around 40 minutes,” says Stefanos Nikolaou**, **Principal Analytics Engineer at Kaizen Gaming. ### Lower operational costs More efficient execution and reduced redundancy have led to a significant reduction in daily pipeline costs. Overall processing costs decreased by approximately 60% while supporting the same data volumes and analytical use cases as before. Importantly, these results were achieved within a single team that rebuilt its workflows using dbt. As the broader analytics engineering and BI teams adopt similar shared practices, the data team anticipates further gains in efficiency, consistency, trust, and business value. In the long-term, the shift to dbt has also increased architectural flexibility. “dbt gives us flexibility. By adopting it, we can remain vendor-agnostic,” says Nikolaou. “If we decide to expand our data stack in the future, it will be easier to do so.” ## Looking ahead With a standardized framework, testing strategy, and governance in place, Kaizen Gaming has established dbt as a core pillar of its analytics capabilities. What began as an effort to improve data workflows has evolved into a stronger data foundation with meaningful improvements in cost efficiency. Ever since the data team adopted dbt, other teams at Kaizen Gaming have seen clear advances in data quality and how analytics workflows are developed, validated, and operated. As a result, interest in adopting the same framework has grown, including teams that had not previously worked with dbt. As the analytics organization continues to transform its workflows, dbt will be rolled out more broadly to support earlier data quality checks, greater consistency, and increased trust in data production. Kaizen Gaming is ready to keep growing and views dbt Labs as an instrumental partner in its journey. --- --- title: "From onboarding to AI-readiness: J.Crew’s modernization journey with dbt Labs Services" description: "By working with dbt Labs Training and a Resident Architect, J.Crew’s data team accelerated a complex modernization during the peak holiday season." url: "https://www.getdbt.com/case-studies/j-crew" date: "2025-09-02" industry: "Retail" --- # From onboarding to AI-readiness: J.Crew’s modernization journey with dbt Labs Services By working with dbt Labs Training and a Resident Architect, J.Crew’s data team accelerated a complex modernization during the peak holiday season. ### Company details - Headquarters: New York, NY - Data stack: Snowflake, GitHub and Entra ID ### Results - 50% faster onboarding of the data team - 0 issues for pipelines in production > “We’ve established a solid foundation, thanks to the expertise we got from the get-go. It’s clear to me how dbt fills the gaps that I’ve seen in data engineering my entire career.” > > — Nick Leonard, Director of Data Engineering J.Crew is an iconic American fashion brand known for timeless style and modern classics. Like many established retailers, J.Crew is evolving to meet the needs of today’s consumers and to compete in a digital-first market. But J.Crew’s prior data infrastructure created significant challenges for the business to move both quickly and scale. Built on SAP BW, the 15-year-old infrastructure had become a bottleneck with fragmented workflows that slowed down decision making. J.Crew chose to rebuild its data foundation with dbt Labs. Making matters more complex, the migration was set to occur during the holiday season, at the moment when orders would double or triple. The stakes were high for J.Crew to modernize, fast. To meet its aggressive timelines, J.Crew partnered with the dbt Labs Professional Services team. ### Getting a complex data transformation right J.Crew’s data infrastructure had begun to reach its limits. The infrastructure’s fragmented database structure meant that each engineer worked in a siloed database. This led to frequent permission issues, inconsistent access across environments, and friction that slowed down data insights. When J.Crew decided to work with the dbt Labs Professional Services team, it was just six months before the holiday season. Certain systems, like order processing, would see data volumes multiply. Any issues during the migration would have a significant impact during a critical reporting window. At the same time, J.Crew’s data team was new to using dbt and unfamiliar with workflows like Git version control and pull request reviews. Without a clear strategy for project structure, they risked losing weeks to researching and debating best practices for modern architecture. “We were essentially jumping 15 years forward in time to a really modern stack,” says Nick Leonard, Director of Data Engineering at J.Crew. “Migrating away from a legacy environment that’s been there that long requires a lot of thought to get right. We knew we needed all the help we could get.” ### Onboarding with expert guidance To navigate the transformation, J.Crew partnered with dbt Labs’ Professional Services. The first step of onboarding involved setting up the environment and hands-on training. When it comes to onboarding, dbt Labs takes an applied learning approach. Rather than sit through lengthy lectures, participants work on a real use case within their environment and follow a curriculum designed to match their experience level. With the guidance of a Technical Instructor, J.Crew’s data team learned the fundamentals of dbt while transforming their own data. By the end of the training program, the data team had a production-ready implementation, the skills to migrate other use cases, and the confidence to scale their data transformation. In fact, the success of the initial cohort led J.Crew to invest in a second package to train a new group. “The training exceeded my expectations and helped accelerate enablement across the organization,” Leonard emphasizes. “The dbt Labs Instructor was really fun, personable, and helpful for getting people familiar with the platform.” Following the training, the data team partnered with a dbt Labs Resident Architect to scale their capabilities ahead of the holidays. The Resident Architect acted as a strategic technical partner and integrated directly into the team: joining weekly standups, providing hands-on guidance with Git workflows and code reviews, and advising on which dbt features to adopt (and when to wait). "The Resident Architect showed us what was possible.” Leonard highlights. “Their insights were invaluable for making the most of our investment in dbt.” In short, the Resident Architect helped the team make opinionated, informed architectural decisions during a critical time frame. In just three months, the Resident Architect: - **Streamlined database architecture**, resolving permission errors and simplifying development by consolidating per-developer environments. - **Advised on project structure and scalability**, including when to mesh dbt projects and how to design semantic layers. - **Educated the team on emerging dbt features** like webhooks, contracts, and unit testing, and showed how to use them effectively. - **Provided documentation **to ensure knowledge transfer with current and future team members. - **Constructed decisions with cost in mind**,** **especially as it relates to compute and warehouse uptime**.** “The Resident Architect gave us clear recommendations based on what we had and where we wanted to go,” reflects Leonard. “They helped us think through the tradeoffs for using dbt for certain things, like semantic layer, orchestration, monitoring, alerting, and scheduling.” ### An accelerated implementation with long-term confidence By partnering with dbt Labs Professional Services for onboarding, J.Crew followed a clear, expert-led path. With this guidance, the data team eliminated the guesswork and realized value from their investment faster. Notably, they launched key pipelines during their highest-volume season without disruption. “Overhauling and streamlining our database structures and permission requirements was much more straightforward with the Resident Architect’s guidance,” highlights Leonard. “The impact was immediate. Our setup had been causing missing objects and permission errors, but after consolidating, our developers were unblocked within a week.” The engagement also led to smarter, more cost-effective decisions. The Resident Architect helped the team avoid over-engineering early in the process. Now, the J.Crew data team is clear on what they’re spending and are on a sustainable path, without unexpected compute costs. “We saved at least a couple of months and onboarded 50% faster during an aggressive implementation timeline,” says Leonard. “We had pipelines in production running with no issues or 3am calls that something is broken.” ### Enabling AI-readiness Most importantly, the team is confident in what they’ve built with dbt and their ability to advance the business with timely insights. “We’ve established a solid foundation, thanks to the expertise we got from the get-go,” Leonard affirms. “It’s clear to me how dbt fills the gaps that I’ve seen in data engineering my entire career.” Modernizing J.Crew’s data foundation was just the beginning. With cleaner data models, streamlined processes, and greater cost transparency, the data team is excited to drive J.Crew’s next era with AI. “From day one, working with dbt meant we were making our data AI-ready,” concludes Leonard. “We’re building an ambitious roadmap around AI, and it’s backed by a foundation we can trust with dbt Labs.” --- --- title: "CHG Healthcare gets their data modernization right the first time" description: "When CHG Healthcare needed to move to the cloud, they partnered with a dbt Labs Resident Architect. Now they’re ready to scale for years to come." url: "https://www.getdbt.com/case-studies/chg-healthcare" date: "2025-04-16" industry: "Healthcare" --- # CHG Healthcare gets their data modernization right the first time When CHG Healthcare needed to move to the cloud, they partnered with a dbt Labs Resident Architect. Now they’re ready to scale for years to come. ### Company details - Headquarters: Midvale, Utah, USA - Data stack: dbt Cloud, Snowflake ### Results - $20k Saved in time and resources on martech - 0 Migration redos - #1 highest-rated project in satisfaction survey > “dbt Labs is a critical infrastructure partner. It’s not enough to just know how dbt works. We needed strategic direction from dbt Labs.” > > — Mark Menatti, Director II, Data and Engineering ### Streamlining healthcare resources to deliver better patient care CHG Healthcare is an organization that helps healthcare facilities fill staffing gaps by connecting them with providers like doctors, nurses, and technicians. In the United States, [access to timely healthcare](https://www.mercer.com/en-us/about/newsroom/future-of-the-us-healthcare-industry-labor-market-projections-by-2028/) is a critical issue. To meet patient demand, many facilities work with CHG Healthcare to quickly staff up during peak periods, find interim practitioners, and hire for high-demand medical specialties. Behind the scenes, CHG Healthcare handles a lot of data. Until a year ago, that was stored using an on-premises SQL server data warehouse. In order to scale, CHG Healthcare migrated to Snowflake and dbt Cloud. “Moving from legacy architecture to the cloud was a big jump for our team,” says Mark Menatti, Director II, Data and Engineering at CHG Healthcare. “An experienced partner could guide us through a seamless migration and ensure business continuity.” To meet their data requirements, CHG Healthcare chose to work with a Resident Architect from dbt Labs’ professional services team. ### Legacy infrastructure with a complex environment For larger organizations like CHG Healthcare, cloud migrations are complex with little room for error. “Getting your data transformation right the first time is essential,” says Menatti. “If you have to redo your migration or constantly fix issues afterwards, it’s painful for the business—and costly.” What’s more, CHG Healthcare’s data engineering team was accustomed to monolithic systems. To architect a successful migration with dbt, they needed to learn entirely new best practices: concepts like modular data pipelines and [DRY](https://www.getdbt.com/blog/guide-to-dry). It’s a lot to manage—and a cost-efficient modernization would require expert guidance. ### Strategic guidance for a smooth cloud migration By working with a dbt Labs Resident Architect, the data engineering team understood exactly how to architect a new system from the start. The Resident Architect helped them: - **Get CI/CD right the first time**. The team quickly implemented best practices for CI/CD. “The Resident Architect brought a depth of expertise that bridged our knowledge gaps,” says Menatti. “Instead of debating if a tactic was right or wrong, we could talk to an expert who always steered us toward the best solution.” - **Develop best practices.** The Resident Architect provided training sessions, clear documentation, and tailored examples for macros, project structure and dbt mesh. Now every team member understands how to approach modular architecture. - **Set up infrastructure that’s ready to scale.** The Resident Architect ensured that CHG Healthcare designed their data models in a modular, scalable way. As a result, the organization can handle changes to their data systems without disruption to data access. For instance, if a team changes to a new CRM system, they can still track the same metrics for reporting—even though the underlying platform is different. The dbt Labs’ Resident Architect also supported the marketing team with their own migration. The team needed to modernize their data stack and create a net new data assets—an implementation that the Resident Architect helped them complete seamlessly. “We accelerated our martech roadmap by an entire quarter,” says Luke Stenis, Principal Product Manager, Data and Engineering at CHG Healthcare. “That amounts to $20,000 in saved time and resources.” ### A powerful data foundation, built to scale By partnering with a dbt Labs Resident Architect, CHG Healthcare implemented an efficient data transformation. “We got our foundation right from the start,” says Menatti. “There’s little infrastructure work for us to redo, which significantly reduces costs.” Notably, the data transformation has improved organizational trust and buy-in. In fact, the transformation received the highest ratings in a recent business satisfaction survey—a testament to the work done with dbt Labs. Now that the modernization is complete, the data engineering team is focused on optimization. For example, they’re scaling up testing and developing more nuanced scenarios for using macros. On the marketing side, CHG Healthcare is building a marketing operations team, which will use their centralized data sets for marketing campaign orchestration. “We’re confident in our data models and long-term scalability, thanks to dbt Labs’ Resident Architect,” concludes Menatti. --- --- title: "WHOOP improves efficiency by implementing dbt Core and migrating to dbt Cloud" description: "At WHOOP, every decision starts with data. Migrating from dbt Core to dbt Cloud was critical for improving data integrity, accuracy, and governance at scale" url: "https://www.getdbt.com/case-studies/whoop" date: "2025-04-15" industry: "Health & Fitness" --- # WHOOP improves efficiency by implementing dbt Core and migrating to dbt Cloud At WHOOP, every decision starts with data. Migrating from dbt Core to dbt Cloud was critical for improving data integrity, accuracy, and governance at scale ### Company details - Headquarters: Boston, Massachusetts, USA - Data stack: dbt Cloud, Snowflake ### Results - 1 Number of days to migrate from dbt Core to dbt - 3 Number of months it took to to migrate from Redshift to Snowflake — - 32+ Number of hours saved per month resolving data errors and issues — > “Access to accurate data is critical, it allows us to improve the customer experience and increase retention, lifetime value, and profitability.” > > — Matt Luizzi, Senior Director of Analytics for WHOOP ### Tracking data to power performance and health [WHOOP](https://www.whoop.com/) is a wearable fitness tracker that helps people monitor their sleep, activity levels, and health. That also makes WHOOP a data company. Internally, its algorithms capture and analyze biometric data for people all over the world, 24/7. Across the business, every decision starts with data. “Access to accurate data is critical,” says Matt Luizzi, Senior Director of Analytics for WHOOP. “It allows us to improve the customer experience and increase retention, lifetime value, and profitability.” To manage the data, the WHOOP data analytics team initially adopted dbt Core. They quickly stood up a single, centralized layer for orchestration and transformation to increase visibility. For a lean and technical team, dbt Core was effective. But as the team grew, they required a more scalable solution. They chose to migrate from dbt Core to dbt Cloud. ### Data quality issues made it difficult to trust the data While dbt Core provided a single, systematic approach to data transformation, the team encountered a few challenges as they grew. For one, there was no centralized governance. Analysts could independently create nearly identical dbt models for the same data, without visibility into each other’s work. For another, they lacked built-in scheduling and orchestration. Since dbt Core doesn’t itself allow them to deploy their dbt models, the team relied on external tooling for orchestration. “As we grew, we experienced bottlenecks from relying on other engineers just to do our work,” says Luizzi. “That meant stakeholders weren’t getting consistent answers to their questions or able to make decisions as quickly as they should.” This setup also made it difficult to trust their data. If a model failed or a downstream model was skipped, it was difficult to troubleshoot the issue or vouch for the data’s integrity. To address these errors, Luizzi spent as much as a day every week addressing and resolving data quality issues—an unsustainable dedication of resources. “The business was keen to invest in AI, but our models could only be as good as the data they're trained on,” says Luizzi. “To drive the business forward with AI, we needed data quality. We couldn’t be scrambling to fix inconsistent data.” Finally, as the team prepared to migrate from AWS Redshift to Snowflake, they needed to ensure that the incoming data was clean and well-governed. That required strategic management; simply “lifting and shifting” their Redshift database wouldn’t cut it. “I was already running into problems where the data wasn’t accurate or even available,” says Luizzi. “For our migration, trust in our data was our most important measure of success. When people lose trust, it's hard to win it back.” ### Establishing a clean database structure on dbt Cloud For the migration, data accuracy and integrity were the team’s foremost priorities. Given the technical debt they had accumulated, they decided to build a new project from scratch in dbt Cloud. The team started with first principles: they identified key metrics to track, and then determined the tables and DAGs required to support these metrics. As a result, the actual migration from dbt Core to dbt Cloud took just one day—while ensuring they had all of the data they needed to create a single source of truth. By migrating to dbt Cloud, the team also achieved the following: - **A deliberate, incremental migration.** To make the most out of their prior investment in dbt Core, the team brought new workloads onto dbt Cloud—while preserving existing infrastructure on dbt Core, like CI/CD processes, dev/test environments, and custom macros, tools, and command-line utilities. - **Data governance best practices.** The team established a weekly production release cadence based on software development practices: every Friday, they fork their dbt code off of production. An analytics engineer then implements the week’s QA changes into a single release, gets code-owner review, and runs unit tests to ensure data integrity. After pushing the update to production, they share release notes with the business detailing the changes. “All of this is possible because of the dbt ecosystem and its developer framework,” says Luizzi. - **Breaking down data silos.** The analytics team created a WHOOP Commons dbt project containing company-wide data and reusable code (e.g., macros) for use across data projects. With [dbt Mesh](https://www.getdbt.com/product/dbt-mesh), every team can access these common assets via dbt Explorer and reference them in their own projects. Now, they no longer need workarounds to bring models in from different projects. It’s all native within dbt Cloud and all of their models are transparent. “dbt Cloud simplified our data transformation,” says Luizzi. “Now we only need a single analytics engineer to maintain dbt Cloud, which would have been unsustainable with dbt Core.” Another benefit of moving to dbt Cloud has been achieving 99% documentation coverage. To maintain that, they use dbt Copilot. "dbt Copilot has streamlined our process by cutting PR review times from thirty minutes to five—making it easy to maintain our 99% documentation coverage inertia,” says William Tsu, Senior Analytics Engineer at WHOOP. “This efficiency allows our analysts to focus on developing robust data models. It’s a significant step forward in our data warehouse strategy." ### Scaling data transformation and improving business efficiency As a result of the migration to dbt Cloud, the data analytics team is operating more efficiently. “Now that we’ve implemented a weekly release process, we don’t spend time addressing ad-hoc data-update requests,” says Luizzi. “Our team is able to focus on data analysis, not building a data pipeline.” It has resulted in greater trust, too. Previously, the team experienced at least one issue per week. Today, they experience no unexpected changes to historical data, no dbt production job failures, and no accidental errors in production code. The data team is confident in the accuracy of their data; most importantly, the business trusts their data ecosystem and integrity. “I can sleep at night knowing that I'm not going to wake up to a Slack message from dbt Cloud, saying that a production job failed,” says Luizzi. “Even when we make changes to the database, our stakeholders can always trust the data they pull to make decisions for the business,” says Luizzi. The migration to dbt Cloud also set the team up to adopt AI quickly and effectively. “We've been one of the fastest companies in the world to adopt some of the most cutting-edge AI technologies and actually put them into production,” says Luizzi. “That’s a direct result of solving our data quality problem with dbt Cloud.” As the team looks ahead, they’re exploring how dbt Semantic Layer can enable them to support AI use cases with natural language chatbots. If a machine can understand the semantics of their models, a user could query the data without knowing SQL or any type of coding language. “dbt Semantic Layer will be a huge player in AI,” says Luizzi. “We’ve seen exciting new features from dbt Labs, and their roadmap is aligned with what we want to build next.” --- --- title: "Bilt Rewards builds scalable incremental models quickly" description: "Bilt Rewards needed creative solutions to handle its most complex datasets. By partnering with a dbt Labs Resident Architect, Bilt unlocked greater value—and dramatically reduced data warehouse costs." url: "https://www.getdbt.com/case-studies/bilt-rewards-professional-services" date: "2025-02-11" industry: "Banking & Financial Services" --- # Bilt Rewards builds scalable incremental models quickly Bilt Rewards needed creative solutions to handle its most complex datasets. By partnering with a dbt Labs Resident Architect, Bilt unlocked greater value—and dramatically reduced data warehouse costs. ### Company details - Headquarters: New York, NY - Data stack: dbt Cloud, BigQuery, Github ### Results - $20K per month cost savings in BigQuery cost - 10x faster implementation delivering cost savings in hours instead of months > “The [dbt] Resident Architect helped us reduce the volume of data to be scanned on our key datasets by 99%.” > > — Ben Kramer, Senior Director of Data Analytics ### A rewards program designed for renters, built on data [Bilt Rewards](https://www.biltrewards.com/) is a platform that enables individuals to earn rewards on rent payments and in their neighborhood. For more than 100 million renters in the United States, housing is their largest monthly expense. Yet unlike credit-card payments, rent doesn’t typically help renters build credit, let alone earn rewards. Most landlords don’t accept credit cards for rent; those that do often charge steep fees. In 2021, Bilt set out to make rent more rewarding. The result has been extraordinary growth. Today, Bilt processes over $40 billion in annual rent and HOA payments—more than 3x the previous year. Bilt oversees a complex ecosystem of data, including tenant information, card transactions, payment processing, and first party data. To manage it, Bilt’s data analytics team adopted dbt Cloud for data modeling, transformation, and orchestration. During onboarding, they quickly realized they needed a custom approach to handling their massive data. “Our team was already proficient with dbt Cloud, so we were able to set up basic best practices smoothly,” says Ben Kramer, Senior Director of Data Analytics at Bilt. “But for our complex use cases, we needed a strategic expert and partner who would help us achieve value quickly.” To meet its business requirements, Bilt chose to work with a Resident Architect from dbt Labs’ professional services team. **** ### Building incremental models for complex use cases Every day, Bilt processes millions of transactions across its users, merchants, and partners. Unsurprisingly, the cost of reprocessing the entire dataset for every update is substantial. However, due to the unique shape of its data sources, Bilt's initial attempts to make its models incremental proved technically challenging and remained expensive. “Our team didn’t have the time to figure out complicated incremental models,” says Kramer. “We needed a highly customized approach, which would have taken days to weeks to do on our own.” For Bilt—a fast-growing company pioneering a new market segment—speed is paramount. Stakeholders need timely access to data to set strategies, build the right features, and address customer questions. And, the business needs to get answers at scale. “Keeping data warehouse costs low was critical for our leadership team,” says Kramer. “By bringing on a dbt Labs Resident Architect, we would save months of trial and error. Without their expertise, our costs would have continued increasing while we figured out incrementality.” ### A strategic partner dedicated to creative solutions The Resident Architect developed a customized approach for Bilt—one based on deep expertise in both dbt modeling and Bilt’s data warehouse platform. First, they reviewed Bilt’s data sources, data warehouse logs, performance problems, and existing workflows. Next, they introduced a scalable and repeatable modification to Bilt’s initial implementation. “The Resident Architect helped us reduce the volume of data to be scanned on our key datasets by 99%,” says Kramer. The Resident Architect also advised the data team on best practices, like modeling conventions, when to utilize multiple projects, and optimizing job schedules in the data pipeline. Once they established a solid maintenance framework, the team quickly pivoted to long-term scalability—exploring custom tests and testing packages, dbt Mesh, and multi-threaded projects. Throughout the partnership, the Resident Architect provided the following: #### Trusted expertise The data team relied on the Resident Architect as a sounding board for their big-picture questions and ideas. From there, the Architect helped the team prioritize and determine the best place to start. “The Resident Architect was a trusted partner from the beginning,” says Kramer. “They enabled us to dive deeper into every aspect of dbt and unlock its advanced potential early on.” #### Strategic guidance and recommendations The data team always knew the right next step to take, thanks to the Resident Architect’s direction. For example, when the team considered creating multiple dbt projects, the Resident Architect identified the optimal timing and business threshold. When it came to data testing, they provided numerous examples, training, documentation, and recommendations for how to make the most of their custom tests. #### Dedicated support for challenging use cases As the team tackled increasingly advanced use cases, like optimizing authorized view permission grants in the data warehouse, the Resident Architect was a dedicated partner. They researched the best approaches and developed custom strategies—tasks that would have distracted the data team from business priorities. “Having an expert in the room gave us confidence in our decisions,” says Kramer. “The Resident Architect has seen a variety of projects and best practices across many organizations. Instead of spending days piecing together best practices and debating internally, we could get clarity in a single one-hour meeting.” ### Significant cost savings with a roadmap for innovation As a result of its work with dbt Labs’ Resident Architect, Bilt saved tens of thousands of dollars. Building performant incremental models took just hours to complete, rather than months. “Our work with dbt Labs’ Resident Architect has driven the business forward, faster,” says Kramer. “We were able to maintain our focus on key business priorities without sacrificing weeks to ensure our warehouse costs were in order.” With a solid structure and clear documentation in place, Bilt is shifting to a more proactive approach with data. The organization, structure, and best practices have freed the team from constantly answering ad-hoc questions. It’s even eliminated the need to hire an analyst dedicated solely to handling queries. Looking ahead, the data team is considering implementing business intelligence—something that’s easier to do with dbt. They’re also exploring generative AI and enabling AI-driven recommendations for both internal and external use. Because both business and technical users can use dbt, the team can focus on hiring the right talent to support their initiatives (rather than only the most technical). “We’re confident that our investment in dbt is delivering enormous value,” concludes Kramer. “Thanks to dbt Labs’ Resident Architect, we understand precisely what we want to do next and how to position ourselves to grow.” --- --- title: "Enpal fuels data efficiency with dbt and saves 70% on data costs" url: "https://www.getdbt.com/case-studies/enpal" date: "2024-12-18" industry: "Renewable Energy" --- # Enpal fuels data efficiency with dbt and saves 70% on data costs ### Company details - Headquarters: Berlin, Germany - Data stack: dbt Cloud, Fivetran, Stitch, Airflow, Azure DevOps, Snowflake, Sled, Tableau, Excel ### Results - 1,400+ models migrated and maintained on dbt Cloud - 70% monthly cost savings from migrating to a modern data stack - 30x speed increase in execution for the heaviest queries > “We used to have outages on a regular basis where the organization wouldn’t have updated data for a whole day. This year, we’ve only had three or four minor hiccups that we could fix in a few hours.” > > — Alexander Novikov, Director of Data and BI ### Overcoming data bottlenecks to power smarter DataOps [Enpal](https://www.enpal.com/), Germany’s leading solar and heat pump installer, faced challenges with their legacy data infrastructure. Bottlenecks, poor data quality, and slow processing times were holding back their vision of empowering employees with reliable, real-time insights. Operating across Europe with a team of several thousand employees, Enpal knew their legacy systems couldn’t keep up with the growing demands of their data-driven operations. To overcome these hurdles, they set out to transform their data stack and workflows, ensuring scalability, efficiency, and cost-effectiveness. ### A vision for centralized analytics across the organization Enpal’s operations—from customer acquisition to energy generation and supply chain management—depend on seamless access to trustworthy, reliable data. Dashboards and integrations with tools like Salesforce and Pipedrive ensure that business users can make data-driven decisions in their daily workflows. To support this, Enpal’s central data team (~ 30 data professionals) focuses on providing the infrastructure for raw data and integrations. Collaborating with three groups of domain expert analysts, they develop data models designed to empower employees across business and fulfillment functions. While the central data team currently supports 1,000 employees, their long-term vision is to scale these capabilities and deliver actionable insights to the entire organization. ### Tackling poor data quality and low velocity Due to Enpal’s rapid organisational growth, their data infrastructure struggled to keep pace. The legacy setup, built on five separate databases dependent on a single larger one, lacked the consistency and scalability needed to meet growing demands. This fragmented design created l issues: - **Scattered transformations and data silos:** Scripts were spread across multiple Microsoft SQL Server clusters and Azure Data Factory databases, creating duplicate datasets that were hard to maintain and inconsistent thus eroding trust of the users . - **Slow performance:** Queries were inefficient due to the fragmented structure, hindering analytical workflows. - **Low data velocity:** Dependencies on different teams with varying permissions caused delays - **Data quality issues:** Frequent raw data changes and slow processing times caused pipeline malfunctions, affecting metrics and reports. With their aging data warehouse, the data team took a critical decision: investing in a scalable, modern infrastructure. “Mornings would start with ‘Alex, can you fix this?’,” said Alexander Novikov, Director of Data and BI at Enpal. “. We had to choose between firefighting and investing in a data structure that would meet our growing needs”. ### Choosing simplicity and scalability To address these challenges, Enpal took a bold step toward transformation. They opted to move away from disconnected platforms and adopt dbt’s unified data control plane. This approach centralized all transformations in a single, collaborative environment where analysts and engineers could efficiently work together, eliminating redundancy, and ensuring consistent data quality. Leveraging CI/CD and adhering to software development best practices, the team streamlined operations, improved collaboration, and established a scalable foundation for modern data analytics. “Work before dbt was a rollercoaster. We had multiple places where transformation was happening”, shared Alex. ### Shedding legacy systems: seamless migration with dbt Cloud Recognized as the "state of the art data collaboration tool," dbt quickly became the solution of choice for Enpal’s data team. After testing dbt Core, they adopted dbt Cloud for its user-friendly interface and minimal onboarding effort, which was especially critical as they transitioned from their reliance on Azure Data Factory. With dbt Cloud in place, Enpal set an ambitious goal: decommission their legacy server within three months while continuing to meet all business data needs. This required a strategic and phased migration process: - **Introducing analytics engineering concepts:** Business In analysts, were onboarded to dbt Cloud. They learned the fundamentals of code management and CI/CD, enabling them to adopt modern data engineering workflows. - **Cleaning and structuring data:** The migration provided an opportunity to organize thousands of legacy models into a streamlined data warehousing structure. Layers were created for raw data, cleaned data, core models, and reporting. dbt’s macros significantly reduced duplication, enhancing efficiency. - **Ensuring a seamless transition:** Over 1,400 models were successfully migrated to the new infrastructure. Dashboards and reporting tools continued functioning smoothly, even as the legacy server was decommissioned. This migration not only met Enpal’s immediate needs but also set the foundation for a scalable and collaborative data ecosystem, empowering the team to achieve their long-term vision. ### Achieving data efficiency, stability, and 70% cost savings Enpal’s data transformation delivered tangible results, driving operational improvements and cost reductions: - **Enhanced collaboration and sustainability:** dbt’s built-in best practices, combined with Enpal’s new ground rules for naming conventions and pipelines, established a strong foundation for scalable, sustainable data development. - **Dramatically faster processing:** Engineering optimizations such as modularity, macros, and the migration to Snowflake reduced job processing times from up to 36 hours to just one hour, significantly accelerating insights and decision-making. - **Improved system stability:** Pipeline breaks became rare. What were once daily disruptions transformed into infrequent, quickly resolvable issues, ensuring reliable operations. - **Major cost savings:** Migrating to Snowflake, dbt, and Fivetran reduced infrastructure costs by 70%. Enpal now spends just 30% of their legacy system’s six-digit monthly costs, with additional savings from decreased maintenance requirements. Through this transformation, Enpal achieved not only cost-efficiency but also the operational resilience needed to scale their data capabilities across the organization. ### Paving the way for broader data accessibility with dbt Semantic Layer Enpal is looking ahead to make data more accessible across the organization. While dbt Cloud is currently used by the Data and BI team, plans are underway to extend access to domain experts, empowering them to leverage governed data without compromising compliance. The dbt Semantic Layer and MetricFlow will play a critical role, enabling stakeholders to interact with accurate, reliable data through tools they already use, such as Google Sheets and Excel. “The path to data accessibility isn’t just dashboards but meeting people where they are, in the tools they already use,” said Alexander Novikov, Director of Data and BI at Enpal. By building a modern, streamlined data stack, Enpal has positioned itself for scalable, efficient analytics that not only enhance internal operations but also support its mission of accelerating the adoption of renewable energy. The future holds even greater potential as Enpal continues to democratize data access and foster collaboration across its growing organization. --- --- title: "Bilt Rewards saves 80% in analytics costs with the dbt Semantic Layer" description: "How Bilt Rewards leverages the dbt Semantic Layer to deliver reliable, personalized embedded analytics to their partners and customers." url: "https://www.getdbt.com/case-studies/bilt-rewards" date: "2024-12-11" industry: "Banking & Financial Services" --- # Bilt Rewards saves 80% in analytics costs with the dbt Semantic Layer How Bilt Rewards leverages the dbt Semantic Layer to deliver reliable, personalized embedded analytics to their partners and customers. ### Company details - Headquarters: New York, NY - Data stack: dbt Cloud, BigQuery, Github ### Results - 80% cost savings from migrating off BI tool embeds - Increase in trust, reliability, and scalability - Positive reviews from partners on new portal and the unprecedented level of insight > “By centralizing our entity relationships in the dbt Semantic Layer, where all of our data transformations already live, we could easily create visualizations in our B2B product. We delivered an improved data experience for our B2B partners by eliminating a step in our process, decreasing our data costs by 80%, and increasing reliability and trust.” > > — Ben Kramer, Senior Director of Analytics ### Serving diverse analytics needs with a lean data team Founded in 2021, [Bilt Rewards](https://www.biltrewards.com/) is a platform that lets members earn points through rent payments, neighborhood dining, and travel. Bilt faced the challenge of managing massive datasets while delivering reliable, tailored insights to partners (merchants, property owners, and more), members, and their internal teams. With a lean data team of just three analysts and over 10,000 external data consumers—primarily B2B partners who depend on accurate reporting—Bilt needed a scalable, cost-effective solution to streamline their analytics for every data consumer. #### A complex data infrastructure with billion rows Bilt’s data infrastructure is vast and intricate. Hosted in BigQuery with over 200 schemas, hundreds of billions of rows, ingested from 50+ application databases—from financial partners (Mastercard, Wells Fargo) to property managers, to merchants, and to travel partners. The company has hundreds of different transaction logs, with changing dimensions and various join patterns. Data serves three main types of data consumers: internal teams, business partners, and members: - **Internal teams**: Employees of Bilt who leverage data (via Reverse ETL & BI tools) to deliver personalized user experiences and improve business performance across CRM, website, app, and paid campaigns. - **Members**: Those using Bilt to pay for their Home, Neighborhood, and Travels purchase and earn best in class rewards. Members can access their past usage and rewards through Bilt’s mobile app and website, as well as personalized benefit recommendations. - **Business partners**: Financial institutions, property managers, merchants, and other partners who access critical data on Bilt’s B2B portals, where trust and accuracy are paramount. ### Growing pains across data operations and costs Although the Bilt team was doing several data transformations for multiple use cases, they had not yet invested in data engineering best practices across the analytics development lifecycle (ADLC)—such as documentation, data lineage, version control, CI/CD, and automated testing. With only transformation in place, these gaps left the team vulnerable to inefficiencies, errors, and eroding trust with both internal and external stakeholders. #### Rising analytics costs and lack of personalization for 10,000+ business partners Bilt initially used a BI embed to deliver data reports and charts to their business partners. While the solution served the initial use case, the per-user pricing model quickly became unsustainable as costs scaled linearly with each new merchant or partner onboarded. Additionally, this pattern limited the types of visualizations Bilt could offer, further constraining their ability to meet partner expectations. Last, data transformations lived in both the data warehouse and in the BI tool causing troubleshooting headaches. For a rapidly growing company, this approach was not viable. ### Implementing dbt Cloud to increase the velocity of a lean data team Managing such a large and diverse data ecosystem as Bilt’s required an efficient and scalable solution. Some data team members had been long-term dbt users before joining the company and identified the tool as a solution for their DataOps problems. Bilt Rewards initially considered dbt Core but opted for dbt Cloud to meet their speed and scalability needs. dbt Cloud provided tools to centralize business logic, streamline transformations, and implement practices like [documentation](https://docs.getdbt.com/docs/build/documentation), [lineage tracking](https://www.getdbt.com/blog/guide-to-data-lineage), and [governance](https://www.getdbt.com/product/governance). These out-of-box features were essential for their lean data team of just three individuals. Together, this created a transparent and reliable analytics framework, setting the stage for broader use cases and sustainable growth. ### Cost savings and improved data quality #### Simplifying data transformation with Semantic Layer Bilt implemented the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) to centralize their metrics and dimensions in the same place as all of their data transformations in dbt Cloud. This allowed them to efficiently deliver data for key use cases like embedded analytics for their B2B partner portals and deliver personalized member experiences on their mobile app and website. > “What makes us most excited about Semantic Layer is that we only need to define metrics once,” said Ben Kramer, Senior Director of Analytics at Bilt Rewards. “We don’t need to define relationships and definitions in all of our downstream tools. We write it once, store it in GitHub, and now we have a true and trusted model we can use everywhere”. #### 80% costs saved and improved user experience from switching to dbt Semantic Layer By moving transformation logic from their BI tool into dbt, Bilt transitioned off their BI Embed solution, significantly reducing costs tied to its per-user pricing model. The dbt Semantic Layer decouples business logic—metrics, dimensions, and calculations—from front-end visualizations, enabling a more cost-efficient, headless BI approach through dbt’s GraphQL endpoint. The shift to the dbt Semantic Layer lowered query costs and streamlined the delivery of accurate and customizable data visualization. > “All we really needed to leverage for our customer-facing visualizations in our B2B portal was the data. Once we centralized data transformations with dbt and their Semantic Layer, we could easily create the visualizations to our front end,” said Ben. “It became super simple. We migrated quickly to sending data via the graphQL endpoint, and our data costs decreased significantly by 80%.” #### Superior data quality with DataOps best practices Standardizing metric transformation and entity relationships in the Semantic Layer was one of the ways dbt Cloud helped Bilt improve data governance and efficiency across their data ecosystem. It built trust with Bilt’s partners and encouraged them to develop new campaigns and complementary offerings. > “We had a lot of data quality issues, but dbt Cloud really solved most of them. And for the hardest problems, the dbt team collaborates with us to create solutions,” said James Dorado, VP of Data Analytics at Bilt Rewards. ### Leveraging Semantic Layer for AI and ML use cases The B2B portal was just the beginning use case for the dbt Semantic Layer at Bilt Rewards. They’re now exploring how to use semantic models to ensure consistent and reliable metrics power their Machine Learning and AI LLM initiatives: “Our B2B portal demonstrated one powerful way to leverage the Semantic Layer,” said Ben. “It helped us model our business in a way that’s easy for both end users and B2B partners to understand while querying efficiently. It was an excellent starting point for our Semantic Layer journey with AI and BI as our next steps!” --- --- title: "Symend implements a robust data foundation fit for scale with dbt Cloud" description: "This is the story of how dbt Cloud enabled Symend to improve data quality and decrease costs" url: "https://www.getdbt.com/case-studies/symend" date: "2024-07-23" industry: "Software" --- # Symend implements a robust data foundation fit for scale with dbt Cloud This is the story of how dbt Cloud enabled Symend to improve data quality and decrease costs ### Company details - Headquarters: Calgary - Solution: Conscious Engagement delivers industry-leading conversion rates, positive customer experiences and enhanced brand affinity through deeply personalized outreaches. - Data stack: dbt Cloud, Snowflake, Microsoft Azure, Looker, PowerBI, Sisense, Github ### Results - 70% decrease in daily warehouse credit consumption - 90% reduction in time to debug - 20x improvement in pipeline velocity > “I fell in love with dbt when we were decreasing data latency. We were previously doing full loads, twice a day, and runs took a full 12 hours. We implemented an incremental solution with dbt that significantly increased Snowflake's speed so we could get latency down to two hours." > > — David Petiot, Principal Data Architect, Enterprise Analytics ### Leveraging data-driven behavioral science #### **Identifying and reaching financially at-risk customers** Founded in 2016, [Symend](https://symend.com/) helps companies, like Canadian telco Telus, with digital customer engagement “treatment strategies” to collect debt from customers. Millions of consumers have been reached by Symend, helping them avoid negative credit outcomes. #### **A new vision with data at its core** Although Symend started with a single channel—emails—over the years, data’s role expanded from results reporting to shaping the product. Today, the company uses Artificial Intelligence, and machine learning to identify, segment and create content and journeys aligned to drive desired outcomes. The longer a customer is with Symend, the more value they receive thanks to a deeper understanding of their data and their behavior, which allows for further personalization. [Watch video](https://www.youtube.com/embed/p6wGs8hvPVM) ### A legacy data stack unaligned with the business’ new data vision Symend’s “V1” data warehouse was originally built on a tier-1 cloud data warehouse SQL offering to run weekly reporting. However, the company’s growing needs were leading to issues, such as: - **Poor data quality**: with more time-sensitive data needs, the pipeline was pushed from weekly to daily operations, which resulted in regular incidents. - **Limited accessibility:** more team members needed access to data— to expose it, the data team had to create and maintain a complex third-tier system as a workaround for technological limitations. - **High data latency**: although not an issue historically, Symend’s new data products required lower data latency. - **High cost**: the operating cost of the “V1” data warehousing solution proved to be much pricier than initially estimated by the data team. With the legacy stack, Symend was building its advanced analytics on top of a shaky foundation—a risk to scaling further. Since data had become mission-critical, it became clear to the team that data needed a better home. #### The search for a new data stack, designed for engineers #### **Back to the drawing board, starting with Snowflake** Symend settled on Snowflake as the anchor of their data system. During their evaluation phase, one of the companies’ core considerations was that the new stack should fit their workflow as an engineering organization: the software development lifecycle (SDLC). Code shouldn’t be written in production and large SQL stored procedures should be avoided, replaced by CI/CD and modular development instead. #### **Evaluating, selecting, and migrating to dbt Cloud** dbt Cloud plugged directly into Symend’s SDLC, with a direct connection to GitHub. The product also offered critical features important to Symend: [Jinja and macros](https://docs.getdbt.com/docs/build/jinja-macros) support for modular models, built-in observability, and concurrent pipeline runs. The migration to dbt Cloud and Snowflake was completed in 1.5 quarters, marked by the first data set going live. The team went through [dbt’s on-demand training](https://www.getdbt.com/dbt-learn) alongside support from a consultant to understand the product and build comprehensive documentation. ### Delivering value faster, at a lower cost #### **Improved accessibility** Symend’s first live bronze data set on dbt Cloud and Snowflake was made immediately available to analysts. In under 2 weeks, analysts with no previous dbt experience were onboarded and exploring the data sets. Whereas previously analysts had to query production directly, they now had unlimited (yet governed) freedom in the Snowflake environment to uncover valuable insights for the business. #### **Productivity gains and data quality improvements** The ease of use of dbt Cloud, paired with out-of-box features like the [job scheduler](https://docs.getdbt.com/docs/deploy/job-scheduler), data testing, modularity, and [documentation](https://docs.getdbt.com/docs/collaborate/documentation), increased the efficiency of the data team: > “The way dbt is designed and documented makes it very easy to follow and comprehend. If something goes wrong, it’s easy to debug,” said Raman Singh, Tech Lead at Symend. > > “With dbt at the heart of our data transformations, we were able to do the work of 8 people with a team of 4 FTEs.” added Ziko Rajabali, VP of Engineering at Symend. The built-in development best practices provided scalability to Symend’s data infrastructure; the solution now seamlessly loads tens of millions of records daily. #### **Decreased cost for higher frequency data** As part of the migration from their legacy stack, Symend decreased data latency from one week to 12 hours. As part of the migration from their legacy stack, Symend decreased data latency from one week to 12 hours. With dbt, the team was able to create very modular systems, because they could choose how they materialize the data—not everything needed to be materialized to be reused. dbt allowed the team to embrace engineering best practices to build a system that was scalable, version-controlled, and accessible to more stakeholders. All in a way that didn’t break the bank. At a higher level, the paradigm shift to the cloud allowed the team to develop a more streamlined solution, increase data velocity with a simpler architecture, and allowed them to improve costs with more efficient executions of the full data pipelines. This amounted to drastically reduced costs compared to what they were paying for their legacy stack, totaling 89% less: > “Even though we are loading more data, at a higher frequency, higher quality, and in a more accessible manner, our new costs were only a fraction of the previous price. There was no loss in the equation,” shared Ziko. #### **Lower data latency opens new product development capabilities** Symend’s improved 12-hour latency from moving to the cloud was still deemed too lengthy for product analytics needs. They were ultimately able to reduce latency even further, to only two hours, by revamping the design of their dbt models to make them run incrementally. They leveraged dbt features like built-in macros and dynamic queries, enabling them to run 200 models every 2 hours, instead of every 12 hours. All while setting themselves up for future scale, and notably, _without_ driving up compute costs. In fact, they were able to save 70% of daily credits on Snowflake. This improved latency opened opportunities to build new capabilities based on data. With data pipelines refreshing 12 times a day, they were able to ship three new data-driven product features, improving both customer retention and Symend’s competitive advantage. > “I fell in love with dbt when we were decreasing data latency,” said David Petiot, Senior Manager of Enterprise Analytics. “We were previously doing full loads, twice a day. With dbt’s built-in macros, we implemented an incremental solution that significantly increased Snowflake's speed so we could get latency down to two hours.” ### Looking ahead: BI migration and a semantic layer Symend will soon be incorporating a new BI component, Sisense, into their platform—a move simplified by all data models and logic now living in dbt Cloud. The team is also exploring how to best use additional dbt Cloud features, such as leveraging [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) to ensure metric governance and increased data trust. --- --- title: "Siemens implements a data mesh architecture at scale with dbt Cloud" description: "How Siemens increased data velocity with a decentralized and accessible data infrastructure, used by hundreds of thousands of employees" url: "https://www.getdbt.com/case-studies/siemens" date: "2024-07-10" industry: "Industrial Automation" --- # Siemens implements a data mesh architecture at scale with dbt Cloud How Siemens increased data velocity with a decentralized and accessible data infrastructure, used by hundreds of thousands of employees ### Company details - Headquarters: Munich, Germany - Data stack: Snowflake, dbt Cloud, Gitlab, Tableau, Qlik Sense, PowerBI, AWS ### Results - 93% reduction in daily load time, from 6 hours to 25 minutes - 90% reduction in costs to maintain certain dashboards - 35 ERP systems used for one dashboard > “Already in our first dbt Cloud project we were amazed by the seamless collaboration dbt Cloud offers, allowing us to effortlessly work together on the same Snowflake project. With built-in tests, simple job scheduling, and easy deployment, dbt Cloud enabled us to immediately focus on the business case rather than spending time on our data architecture setup.” > > — Rebecca Funk, IT Business Partner ### Leading innovation across many, many industries **Employing 300,000+ people across the world** For the last 180 years, Siemens has led innovation across diverse industries—from healthcare to infrastructure. In the mobility business alone, Siemens produces the software and machinery to design and build cars, as well as manufacturing car batteries and charging stations. **An immense, complex, on-premises data infrastructure** Having operated for so long, Siemens found itself internally dealing with the consequences of a legacy, monolithic data infrastructure. The company used an on-prem HANA system, which was difficult to use, maintain, and grow. [Watch video](https://www.youtube.com/embed/Li1G7Xyz85E?si=Kjt-pfsydJjiOpuB) ### The Siemens Data Cloud (SDC) project Siemens started by defining the requirements for their data infrastructure. After reviewing the existing architecture, they landed on a new direction: an open ecosystem to unify all data products—from business intelligence to machine learning—and share those data products across the 70,000 internal data consumers. #### **Building for scale with data mesh** Given Siemens’s scale, the [data mesh framework](https://www.getdbt.com/blog/what-is-data-mesh-the-definition-and-importance-of-data-mesh) was chosen as a compass to help the company achieve its data vision. It would change how teams access data assets and participate in data development—leveraging the thousands of Siemens’ domain experts to increase data velocity. ##### Previous way of working - Internal users locate the correct owners and manually request access to needed data assets—a process that can take weeks. - BI and AI data products are created in silos and do not share governance, sources, or metrics. - Data products are created for 1 team and often used only once. - Data products are created in silos by data engineering and analysts, producing inconsistencies in metrics, duplicated logic, and code few can understand. - Data pipelines are centralized and exported consecutively. This leads to slow load times and unstable pipelines where one break can affect all data products. ##### Data mesh way of working - **Internal users self-serve data needs** with a semi-automative process; access to raw and modeled data assets is granted within minutes. - BI and AI data products are created within the same data stack and **can share data sets**. - **Data products** are stored in accessible locations and **repurposed** by multiple teams. - Data products are created with **testing, documentation, continuous integration (CI), and observability**—improving data quality and decreasing duplicated work. - Data pipelines are loaded in parallel. There is **workload isolation** and, therefore, **improved stability**. Siemens first assessed if it could use the existing central stack for a data mesh setup. However, the team realized a “lift and shift” approach would fall short due to: - **Hardware capacity**: Siemens could not keep up with the hardware needed to run the data mesh infrastructure on-premises. - **Expensive fixed costs**: Costs for HANA were federated with limited consumption-based billing—leaving little opportunity for efficiency gains. - **Not built for purpose:** Unlike on-premises ETL solutions, a decentralized cloud-based stack is built for expansive organizations like Siemens, offering better capabilities for managing high-volume data. ### Delivering the SDC vision with dbt Cloud and Snowflake To achieve its data mesh goals, Siemens needed an intuitive interface for internal users to discover, access, and model data while ensuring data governance. dbt Cloud, with its accessible browser-based IDE and shallow learning curve, was picked as the UI for data modeling in SDC. #### **Migrating from on-prem to the cloud with Accenture and dbt Labs** The data stack for SDC needed to be an open architecture, to enable and facilitate any future migrations or modernization projects. The two tools chosen as the backbone of the SDC enabled this flexibility: “With Snowflake and dbt, you’re agnostic to cloud providers or BI tools,” explained Tobi Humpert, Product Owner of SDC. To set up their dbt project following best practices, the Siemens team recruited the help of Accenture and dbt Labs’ services. Data teams completed dbt Labs-led onboarding sessions and group training and leaned on the support of a dedicated dbt Labs Resident Architect. Siemens also leveraged their internal Learning Management System (LMS) and dbt Labs-produced content to encourage dbt training—all while tracking dbt learning and development across unique company divisions. #### **Setting up the infrastructure for data mesh with dbt Mesh** As a customer with large scale and complexity, Siemens participated in the closed beta of [dbt Mesh](https://www.getdbt.com/product/dbt-mesh): a set of features that enable companies to implement and maintain a data mesh infrastructure. This empowered Siemens to bring their new way of working to life with: - **Multi-project discovery**: dbt Explorer provided the data team with complete lineage across hundreds of decentralized projects. - **Data contracts**: dbt model contracts offer an additional validation layer for production code to govern data quality and prevent downstream issues. - **Federated governance**: dbt’s built-in governance features enable the SDC team to define which datasets should be shared with the wider Siemens data community, and which should only remain within a certain project. ### Reaping the benefits of a modern cloud-based Data Mesh infrastructure #### **Democratized data access decreases time-to-value** To increase data participation across the organization, Siemens launched the _Siemens Data Cloud Marketplace_ where all employees can browse and purchase data products, such as platform components and ML models. The Marketplace enables stakeholders to reutilize assets created by other teams, reduce duplicated work, and transform data teams into profit centers. All data products live in one Snowflake account, managed by the central IT department. Employees request access to a Snowflake project, which is a bundle of Snowflake, dbt Cloud and git access—a process automated end-to-end. Within minutes, analysts can self-serve and access the data they need to build data products on top of it: “Already in our first dbt Cloud project we were amazed by the seamless collaboration dbt Cloud offers, allowing us to effortlessly work together on the same Snowflake project. With built-in tests, simple job scheduling, and easy deployment, dbt Cloud enabled us to immediately focus on the business case rather than spending time on our data architecture setup,” said Rebecca Funk, IT Business Partner at Siemens. The new workflow eliminated the corporate IT bottleneck causing delays in data deliveries. And, most importantly, it decentralized data production by enabling domain experts to own a much larger portion of the data development process: “All teams are now working in the fields of their expertise, which empowers them to experiment within their data domains and innovate,” explained Tobi. “This set-up enables teams to scale independently, with faster development cycles and decreased risk of failure.” #### **Centralized governance complies with security and privacy standards** dbt Cloud’s federated governance works behind the scenes to enable Siemens’ domain teams to self-serve and participate in data transformation: “The barrier to entry is low because end users don’t need to worry about contracts, control, or security. That’s already all defined centrally in dbt,” explained Tobi. The Central IT team increased data velocity without compromising on security by: - **Defining contracts and access in dbt: **User access is managed with row-level security and by publishing datasets into a distribution layer within Snowflake, all implemented on dbt. Meanwhile, contracts guarantee new code doesn’t affect downstream data. - **Using a singular Snowflake dbt account**: All stakeholders’ data sources and transformations live in a central place where IT can set guardrails, audit, and visualize the full lineage with [dbt Explorer](https://www.getdbt.com/product/dbt-explorer). #### **Increased efficiencies and 90% cost savings** The migration to the cloud enabled Siemens to move away from a fixed-cost structure to transparent usage-based pricing. In particular, dbt’s incremental materializations led to substantial efficiency gains, as only the latest data available is now loaded, as opposed to full table loads. “In our on-premises stack, data for our business analytics dashboard would take six hours to load every day,” explained Nuno Pinela, Data Engineer at Siemens. All 35 ERP systems that fed the dashboard had to load sequentially. “With dbt, if a model is not dependent on another, they automatically run in parallel.” Today, that same data is loaded in 25 minutes, leading to a 90% reduction in costs. ### A successful migration As of February 2024, over 700 projects have been migrated to the Siemens Data Cloud and dbt Cloud within a year and a half. Siemens has over 600 developers onboarded to dbt Cloud, using a single dbt instance, who are maintaining 5,000+ active dbt models. The company celebrated phasing out their legacy SAP HANA system with a global party spanning 5 physical locations. --- --- title: "AXS delivers business value with an analytics engineering workflow" description: "This is the story of how the ticketing industry leader improved data velocity with dbt Cloud" url: "https://www.getdbt.com/case-studies/axs" date: "2024-05-15" industry: "eCommerce" --- # AXS delivers business value with an analytics engineering workflow This is the story of how the ticketing industry leader improved data velocity with dbt Cloud ### Company details - Headquarters: Los Angeles - Solution: ticketing technology / live entertainment - Data stack: dbt Cloud, Snowflake, Hightouch, Fivetran, Looker ### Results - 50% faster to deploy new models, using incremental models - 10% faster troubleshooting, using lineage graph - 40% saved work hours on maintenance > “dbt Cloud is great because it’s so approachable. You can log in on your browser, select a project, and start building models right away. You can build things with native SQL that you’d think were only possible with Python or a more object-oriented language. The SQL-first approach creates commonality among the tech teams, improving communication and collaboration.” > > — Michael Colella, Senior Director of Data & Analytics ### AXS: selling over 65 million tickets a year [AXS](https://www.axs.com/) is a global leader in ticketing, serving North America, Europe, Asia, Australia, and New Zealand. They are the ticketing platform of choice for major festivals (like Coachella), over 500 venues (including the O2 Arena), and popular sports teams (LA Galaxy, LA Kings, and more). ### Providing data for diverse stakeholders The AXS data team is responsible for complying, transforming, and sharing data with their customers (venues and promoters), internal users (from marketing to business intelligence teams), and operationally, end users (fans). Like any e-commerce platform, data plays a pivotal role in driving revenue across the business, by enabling activities such as: - **Pricing recommendations: **Calculating optimal ticket prices to increase average revenue per ticket and events’ sell-through rate. - **Conversion optimization: **Measuring customer journey to uncover opportunities to drive conversion rates up. - **Marketing: **Reporting on sources and channels leading to purchases, and sharing data with marketing channels to improve campaign performance. - **Predictive analytics: **Forecasting the company’s growth trajectory and events’ performance (pre-sales metrics). - **Fraud detection: **Reducing fraud—a huge problem in the ticketing industry. - **Customer experience:** Improving the fan experience with relevant content, event recommendations, and more. ### Identifying a gap: analytics engineering best practices To turn raw data into actionable insights faster than ever before, the data team needed to build a new way of working with [analytics engineering best practices](https://www.getdbt.com/blog/analytics-engineering-six-best-practices). The team identified the following areas of focus:: - **Observability: **improving visibility into workflows and errors, including clear data lineage, version control, governance, and alerting. - **Data consistency:** Metrics should be consistent across data sets to prevent confusion and data quality issues. - **Automated ETL process: **shifting data ingestion and transformation from a a manual task to an automated, fit-for-purpose workflow less prone to errors and long hours of troubleshooting. To achieve their analytics engineering vision, the team searched for new tools that would enable the workflow and technological shift. ### Choosing dbt Cloud, starting with dbt Core The must-have requirement that led AXS’ search was [version control](https://docs.getdbt.com/docs/collaborate/git/version-control-basics). To align development practices across teams, the new stack would have to align with internal users’ (data engineers and analytics engineers) preferences—git integration, and support for Python and SQL. The AXS team started with a two-week proof of concept on [dbt Core, before choosing dbt Cloud](https://www.getdbt.com/product/dbt-core-vs-dbt-cloud) for lower complexity and ease of onboarding: “[dbt Cloud](https://www.getdbt.com/product/dbt) is great because it’s so approachable. You can log in with your browser, select a project, and start building models right away,” said Michael Colella, Senior Director of Data & Analytics at AXS. ### Reaping the benefits of an analytics engineering workflow #### Improved collaboration with a common workflow Today, all data projects across AXS’ global regions are centralized on dbt Cloud. Teams can check on existing jobs, write new jobs, rerun models, and start new development all in the same place. “dbt Cloud helped us remove complexities and gave us a [framework in analytics engineering](https://docs.getdbt.com/best-practices/best-practice-workflows) for organizing our work,” explained Michael. “By centralizing data work in one tool with one language (SQL), it created a commonality among our tech teams and facilitated communication. That was very powerful.” #### Driving data quality with testing and documentation The dbt workflow involves building staging models, using a governance layer, and testing before rolling out any production changes—directly improving data quality. dbt Cloud’s [built-in documentation](https://docs.getdbt.com/docs/build/documentation), now updated whenever engineers perform code changes, maintains traceability so future team members can uphold that data quality too. “Three years after an engineer has left, you can see what they did and why they did it,” said Michael. “I can’t overstate enough how convenient it is to have the documentation right in dbt.” #### Easier debugging with lineage and automated alerts When issues do get by the data governance guardrails, dbt makes it easy to figure out what went wrong. AXS has leveraged dbt Cloud to facilitate root-cause analysis, using [lineage graphs](https://www.getdbt.com/product/dbt-catalog): “One of the great things with dbt is easier debugging. To debug faster, you need visibility into your pipelines and solution architecture,” explained Michael. “dbt’s lineage graphs tie your models together to visualize downstream impact and see how things are working overall.” The data team also takes advantage of [automated job scheduling and alerting](https://www.getdbt.com/product/deploy). Jobs are set up once and if they fail, the team is automatically notified via Slack. This decreases the amount of time dedicated to maintenance and enables data practitioners to spot bugs much earlier. #### 50% faster deploy times and lower time-to-value Since migrating to a modern data stack and implementing an analytics engineering workflow, AXS has achieved efficiency gains. With the help of features like [incremental models](https://docs.getdbt.com/docs/build/incremental-models)—where only modified data is loaded—the team cut deployment time by 50%. This has led to faster processing times and decreased warehouse costs. Data velocity‌ has improved; [macros](https://docs.getdbt.com/docs/build/jinja-macros) enable engineers to reuse pieces of code, reducing the net new code required. And since the team now spends less time on data munging, they can better focus on delivering value. --- --- title: "Rocket Money modernizes financial reporting with dbt Cloud" description: "How Rocket Money transformed financial processes and cross-functional collaboration with a dbt-powered Quote-to-Cash system" url: "https://www.getdbt.com/case-studies/rocket-money" date: "2024-05-15" industry: "Banking & Financial Services" --- # Rocket Money modernizes financial reporting with dbt Cloud How Rocket Money transformed financial processes and cross-functional collaboration with a dbt-powered Quote-to-Cash system ### Company details - Headquarters: Silver Spring, MD - Solution: Financial products and services - Data stack: BigQuery, dbt, Looker, Shortcut - Description: Rocket Money is a personal finance app that provides tools to help users take control of their finances. Features include tracking spending, budgeting, monitoring credit scores, and automating savings. Rocket Money also offers bill negotiation and subscription cancellation services. ### Results - 3,000 Tests implemented to ensure data quality - 0 deficiencies in SOX audit after modernizing system > "Having this automated Quote-to-Cash system run in dbt with our test suite allows us to confidently and quickly close our books each month." > > — Amber Oar, Staff Analytics Engineer ### Data-driven digital finance Founded in 2015 Rocket Money has gained rapid recognition for its personal finance app. Through an array of features such as budgeting, spending tracking, bill negotiation, and subscription cancellation, Rocket Money allows its users to precisely control their spending. Like many other financial organizations, Rocket Money’s business model leans heavily on reliable data and analytics, with processes such as revenue recognition, cash forecasting, and monthly close depending on high-quality data. ### The need for a quote-to-cash system The Rocket Money team quickly found that as the business expanded, keeping track of high-volume information wasn’t a simple task. The company has many different revenue streams and works with several different payment providers, which led to the requirement of a single system to serve as a source of truth to combine these disparate data sources. Amber Oar, Staff Analytics Engineer, and her team quickly realized that if the business was to marshal the data needed to keep decision-makers up-to-date with the latest information, it needed a single source of truth for financial reporting - one that was GAAP (Generally Accepted Accounting Principles) compliant and easy to audit. _“We didn’t want different metrics being calculated in different ways by different teams when they’re all supposed to be the same thing,” _explained Amber._ “Whenever we are calculating revenue for a given month for our company, we want to make sure everyone is pulling it the same way and ending up with the same number.”_ The team needed a quote-to-cash system to help them ensure their financial data was as accurate—and accessible—as possible. To comply with GAAP, the new system needed to track users and transactions through the entire accounting lifecycle, from purchase orders and invoices to revenue, cash, and accounts receivable. ### Mapping complex revenue streams Given the business’s complex financial workflow, the Rocket Money team knew they needed support from specialized data tools. They singled out dbt Cloud to support the new system. _“The previous process was very time consuming. We needed a system that would allow us to scale and also quickly close the financials each month.” _said Amber. One of the first challenges was mapping Rockey Money’s complex web of revenue streams. This was a time-consuming task that stretched across multiple teams, which included product, engineering, and accounting all needing to coordinate efforts. Fortunately, dbt Cloud significantly assisted the team in managing this complexity. The tool’s modular architecture allowed them to reduce masses of complex financial logic into more manageable, maintainable layers. By breaking down intricate processes and data flows into simpler components, the team could more efficiently manage and update their financial systems. ### Building trust with rigorous testing One of the cornerstones of any financial system is reliability; when cash is on the line, you don’t just need accurate numbers, you need to be able to prove they are accurate. The Rocket Money team ensures data quality with a rigorous testing scheme. Before introducing the quote-to-cash system, a lot of this testing was conducted by the finance and accounting team. As the product scaled, the team realized that this would not be sustainable long term. Now, it runs 3,000 tests (that test both quote-to-cash and other production data models) daily, which include business logic and accounting logic checks to ensure that everything is operating as expected. These tests create accountability for the data produced by the new system and allow for easy collaboration across teams. This comprehensive testing gives the Rocket Money team confidence that upstream changes will not cause unintended regressions. ### Guaranteeing compliance with auditability In 2021, Rocket Money was acquired by a publicly traded company. Now that the business was part of a publicly traded company, the company’s accounts needed to be consistently easy to audit. This means having not just clear data but also solid lineage and documentation. dbt Cloud set up the team for scaling success by supplying access to auto-generated documentation and audit logging. In addition, [data diffs](https://docs.datafold.com/data_diff/what_is_data_diff/#:~:text=%E2%80%8B,data%20and%20guarantee%20data%20quality.), a tool that compares datasets, furthered auditability. _“We wanted to make sure that we had an audit-worthy system that could pass any audit even before we needed it,” _said Amber. This careful planning paid off when Rocket Money underwent its first SOX audit—an annual requirement under the US Sarbanes-Oxley Act. _“We passed with zero deficiencies,” _smiled Amber. “_That's great. This system is exactly the kind of thing auditors look at during SOX audits.”_ ### Streamlining financial processes Introducing an automated system sped up the team’s month-end close tasks and provided a single source of truth to power the business’ enterprise reporting. _“Introducing dbt and the Quote-to-Cash process has made close a lot less stressful because we have a testing suite that runs every single day,” _explained Amber._ “If there is a bug that is going to impact our accounting team, we know immediately, and we're working with the development team to fix it before the month closes.”_ The extra visibility has also improved communication between teams. Rocket Money’s engineering team is now more aware of the accounting impacts of their decisions—such as releasing products at the very end of the month. The product team can confidently assess the impacts of experiments because the system supports production of key financial KPIs at the granularity of a purchase. ### Driving impact across Rocket Money Introducing the dbt Cloud-powered quote-to-cash system has transformed Rocket Money’s financial processes and unlocked cross-team alignment. _"Now the accounting team feels like they're shipping like the larger product and engineering team,” _said Amber. “_They're aware of changes before they are made that may impact the revenue streams."_ Looking ahead, the team is building on the scalable foundation to implement advanced analytics across the business as changes are made to existing revenue streams and new ones are rolled out. --- --- title: "DISH Digital Solutions scales data operations with dbt Cloud" description: "How dbt Cloud’s data development workflow helped DISH Digital Solutions increase data quality, velocity, and governance" url: "https://www.getdbt.com/case-studies/dish-digital-solutions" date: "2024-03-20" industry: "Industrial Automation" --- # DISH Digital Solutions scales data operations with dbt Cloud How dbt Cloud’s data development workflow helped DISH Digital Solutions increase data quality, velocity, and governance ### Company details - Headquarters: Düsseldorf, Germany - Data stack: GCP BigQuery, Gitlab, dbt Cloud, Airflow, Tableau ### Results - 30% fewer bugs with improved data quality - 15% less time spent troubleshooting with faster root cause analysis - 270+ models built on dbt Cloud > “Our process is totally different than it was before dbt Cloud, where people worked in silos without visibility on who was doing what. The new workflow increased our data quality while improving transparency and collaboration.” > > — Ramon Marrero, Senior Head of Data & ML Operations ### Digitalizing the hospitality industry under Metro AG [DISH Digital Solutions (DISH)](https://www.dish.digital/en/about-us) provides a suite of digital products to the hospitality industry: from streamlining operations to improving guest service. Based in Germany, the company belongs to the [METRO AG ](https://www.metroag.de/en)conglomerate which employs over 90,000 people across Europe and Asia. The Data Insights team at DISH is responsible for delivering value directly to the customer, by leveraging Machine Learning models to help customers optimize menus and manage inventory. Their goal is to enable restaurants to thrive in an increasingly competitive landscape. The team also works closely with other departments, such as Sales and Finance to deliver accurate and timely data for decision making. ### Siloed analysis and wavering data trust The Data Insights team was facing challenges, such as data quality issues, which impacted the reliability and accuracy of their datasets: - **Duplicated logic and discrepancies:** There was no central place for data transformation, which led to the duplicate, differing metrics. - **Friction on the BI team:** The BI team sits the closest to operations, embedded directly into business units. As such, stakeholders blamed data quality issues on the BI team despite the root cause lying further upstream. - **Siloed analysis:** The quality issues were causing teams to mistrust the insights delivered by the Data Insights team. Instead, teams were performing their analysis. DISH's existing data stack—using Google Cloud Platform ETL tools like Pub/Sub and DataProc—no longer served their requirements to improve data quality. ### A transformation layer to improve data quality Upon joining, Senior Head of Data & ML Operations Ramon Marrero identified the need for a tool that’d sit between the raw data and the reporting layer. This tool would enable the Data Insights department to perform data transformations and quality tests in a staging environment before delivering models to reporting or machine learning. “We had no quality assurance, tests, documentation, or lineage tracking. We started searching for a tool that’d help us fill the gaps,” explained Ramon. “We needed to have tests embedded in our processes to substantially improve our data quality.” #### Evaluating and advocating for dbt Cloud The Data Insights department, led by Chief Data Officer Dr. Olaf Maecker, embarked on an initiative to modernize its integrations and data processes. Ramon had successfully used [dbt Core](https://www.getdbt.com/product/dbt-core-vs-dbt-cloud) before—the open-source version of [dbt](https://www.getdbt.com/product/what-is-dbt). He had seen the tool “bring more transparency, improving data quality and collaboration.” However, using dbt Core was resource-intensive. Its open-source nature required teams to set up and maintain the infrastructure themselves: “If you want a scalable solution that easily integrates with your version control tools, then [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) is the best option,” said Ramon. Set on the fit of dbt Cloud, the Data Insights department presented the business case to stakeholders. They explained that resolving data quality issues required the appropriate tooling—the costs vs. benefits resonated, and the team received the go-ahead from leadership. ### Improved data quality, security, and velocity ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/bb6c276f507ce08d9aa839113cbfd1f74d556bfe-2044x1444.png) #### A transparent and collaborative data development workflow DISH Digital Solutions integrated dbt with their version control tool, Gitlab. This enabled the Data Insights department to move towards an increasingly transparent workflow, based on branches: “Before, people just shared SQL statements through Jira tickets. It was chaos. Now everything's centralized. There's more transparency and much more visibility,” said Ramon. “dbt has improved our data quality, as well as enabled safe collaboration.” The new setup has enabled all data engineers to work independently on features, perform reviews, and push their code to the different branches. Without the bottleneck of centralized approvals and re-working code errors, the Data Insights department can deliver data products and fulfill data requests faster. #### Improved data quality, fewer bugs, and increased trust The transparent [engineering workflow](https://www.getdbt.com/resources/the-analytics-development-lifecycle) decreased the occurrence of data issues by 30%. Not only are there fewer incidents, but with dbt features like Data Lineage, the team can fix remaining bugs faster. With fewer data quality issues, business stakeholders trust the data more, leaning on the Data Insights department to help with business-critical analytics. “Analysts and business users can investigate transformations themselves on [dbt Cloud](https://www.getdbt.com/product/dbt-cloud),” said Ramon. “They only need to involve the engineering team if there’s something abnormal with the data, which now happens less and less. That means engineers can dedicate much of their time to providing value, instead of just fixing bugs.” #### A secure and compliant data environment After migrating to dbt Cloud, all models are now stored in the same geolocation as their data warehouse. This reduced the need for cross-border data transfers, simplifying compliance efforts. The company also leveraged dbt Cloud’s multi-tenant environment, ensuring that logs and metadata around their data remain within the EU. ### Scaling data products within dbt Cloud The Data Insights department is continuing to explore dbt features, including incorporating Python and model contracts. With contracts, the team will be able to enforce model logic to further increase both collaboration and governance. In certain use cases, like machine learning, Python is a better fit than SQL. The team can bring their new engineering workflow—with data lineage, version control, documentation, and modularity—to ML models by consolidating their code from Jupyter notebooks with dbt. --- --- title: "Retool builds scalable, self-serve analytics with dbt Cloud and Databricks" description: "Discover how Retool streamlined data workflows with dbt Cloud and Databricks, enabling scalable and easy-to-maintain analytics." url: "https://www.getdbt.com/case-studies/retool" date: "2024-02-29" industry: "Industrial Automation" --- # Retool builds scalable, self-serve analytics with dbt Cloud and Databricks Discover how Retool streamlined data workflows with dbt Cloud and Databricks, enabling scalable and easy-to-maintain analytics. ### Company details - Headquarters: San Francisco, CA - Solution: Software Development Platform - Data stack: dbt Cloud, Databricks Data Intelligence Platform ### Results - 25% decrease in run time - 50% cost decrease in production jobs > “dbt on Databricks allowed Retool to get value out of data really quickly. We were generating large quantities of product usage data and we needed insights without having to hire a data team first.” > > — Samuel Garfield, Analytics Engineer at Retool ### A fast-growing developer platform with over half a million apps built [Retool](https://retool.com/?ref=dbt-case-study) enables businesses to quickly build software by connecting to an existing database or API. Companies like Zappos and Doordash use Retool to create new web applications, mobile apps, and even AI / ML logic workflows. The data team’s focus at Retool is straightforward: create a scalable data system that enables Retool and its customers to get value from their data both today and in the future. To do so, they hire data-savvy team members and provide them with the necessary, governed data for stakeholders—from marketing to HR to product—to self-serve their data needs. #### Empowering both internal business users and Retool customers with data activation Business users leverage the shared data to create reports and insights. Retool’s data is also used beyond reporting to automate processes, improve efficiency, and increase productivity. The same data activation approach is also reflected for Retool customers. For example, customer Zappos rebuilt their entire enterprise demand planning system using Retool. In less than 3 months they connected their myriad of data sources, addressed years of backlogged feature requests, saved thousands of dollars by consolidating tools, and made the relevant data analytics team 20% more efficient. #### Providing holistic views to different teams The Retool data team also creates and maintains key dashboards. One of them, a Customer Success (CS) “mega dashboard,” ties together data from multiple sources—including Salesforce and machine learning models—to calculate customer health. In this central dashboard, CS can leverage the insights to drive customer satisfaction, improve retention, and identify expansion opportunities. ### A simple, stable data structure with dbt Cloud and Databricks Retool is a fast-growing and fast-moving company with new feature launches, expanding sales motions, increasing data volume, and growing teams. This means the data team has to stay nimble so they can adapt to the company’s evolving data needs. With this in mind, Retool prioritized ease of set-up, minimal maintenance, and scalability for their data infrastructure. #### Setting up dbt Cloud to empower a small team Retool started using [dbt](https://www.getdbt.com/product/dbt) several years ago when the company was much smaller and employees wore many hats. The tool brought an efficient data workflow with built-in collaboration and [continuous deployment](https://docs.getdbt.com/docs/deploy/continuous-deployment). It also decreased maintenance and troubleshooting time with out-of-the-box automated testing and data quality alerting features. “dbt Cloud allowed Retool to get value out of data really quickly,” shared Samuel Garfield, Analytics Engineer at Retool. “We were generating large quantities of user data on product usage and needed insights without hiring a dedicated data team first. Engineering, growth, and CS could build out and maintain data models on their own.” #### Integrating Databricks Data Intelligence Platform Databricks' Data Intelligence Platform also offered the same ease of use, flexibility, and scalability Retool was after. Their lakehouse architecture brought AI and BI development under a single roof and came with “killer” features like [Databricks SQL](https://www.databricks.com/product/databricks-sql), serverless data warehouse, and [Unity Catalog—](https://www.databricks.com/product/unity-catalog)a unified governance solution. It’s easy to use Databricks SQL to control compute costs and ensure a given dbt job uses the appropriate type and amount of resources. Databricks SQL powers most of the analytics use cases at Retool. The team used the dbt Databricks adaptor to set up and integrate their data. “The dbt-Databricks adaptor allowed us to automatically integrate our Databricks Unity Catalog to dbt so we didn’t need any extra configuration,” explained Samuel. “All the security and permissions are automatically derived from dbt models, so it is easy to maintain a compliant, end-to-end view of all data and AI assets. It also automatically does [Photon](https://www.databricks.com/product/photon) optimization, which has been a big contributor to our performance and cost improvements.” #### SQL or Python, whenever wherever Both [dbt and Databricks](https://www.getdbt.com/data-platforms/databricks) support SQL and Python within the same project. This flexibility enables Retool’s data analysts, business users, and engineers to use whichever language is most accessible to them. This joint workflow incentivizes cross-team collaboration while enabling unified governance. ### Moving forward with AI Retool has migrated existing AI models to Databricks and plans to explore Databricks’ AI & ML features in the future. “I have to shout out the [Databricks Assistant](https://www.databricks.com/blog/introducing-databricks-assistant) which has made data much more accessible at Retool. Many of our stakeholders are answering their questions by asking the Assistant about our data model. In many cases, the assistant pulls its answers from the dbt documentation itself. This seamless combination of [Databricks and dbt Cloud](https://www.getdbt.com/data-platforms/databricks) will make it easy to deliver value from the data in ways that our business and stakeholders don’t even imagine,” concluded Samuel. --- --- title: "Purple builds data trust with dbt Cloud" description: "Learn how Purple used dbt Cloud to boost data quality and optimize costs during a company acquisition." url: "https://www.getdbt.com/case-studies/purple" date: "2024-01-25" industry: "Retail" --- # Purple builds data trust with dbt Cloud Learn how Purple used dbt Cloud to boost data quality and optimize costs during a company acquisition. ### Company details - Headquarters: Lehi, Utah - Solution: Mattresses - Data stack: Fivetran, AWS Glue, Snowflake, dbt Cloud, Looker, Atlan ### Results - 20% inconsistency reduction in financial data - 60% fewer support requirements - 80% time reduction to respond to data requests > “dbt Cloud has been a game-changer in the way we handle data, bringing efficiency and reliability to our processes.” > > — Bryan Kerr, Analytics Engineering Manager Launched in 2015, [Purple](https://purple.com/) has quickly become a leading mattress provider in the US, driven by its unique "Hyper-Elastic Polymer" technology. After several expansions, the business now sells its products through several channels, including retailers and direct to customers. This combination of a broad product portfolio and various distribution channels generates a high volume of data that Purple’s analysts use to measure the business’ performance and spot new opportunities. However, as the amount of data grew, the company’s legacy data systems struggled to scale accordingly. “When I joined, we had too many interpretations of data, which led to huge debates in meetings over simple things like dates or gross sales definitions,” explained Matt Meads, Purple’s Senior Director of Analytics “This became a big problem, especially with changes in our C-suite. New executives questioned the discrepancies in data across different dashboards and systems—where was the right number, Looker or our financial tools like NetSuite and Adaptive?” These inconsistencies led team members across the business to conduct their analyses using various technologies and systems, creating what Matt described as “a sort of freelance data analytics across the company.” This ad-hoc approach to data delivered results for individual teams in the short term, but as the months passed, the lack of a central source of truth created more and more confusion across the wider organization. ### Leadership pushes for reliability The team recognized that to understand how it was performing, Purple needed a new data system. “In senior leadership meetings, our CEO would express frustration about inconsistent data and the inability to agree on accurate numbers,” said Matt. “These instances made it clear that we needed to make changes quickly.” Bryan Kerr, Analytics Engineering Manager, added: “If there’s one single event that kicked off the change, it was when the CEO came to analytics with a request to move BI tools.” The team soon began exploring options for a new data stack to support their current use cases and allow for future migrations. While attending a manufacturing user group meeting, some data team members noticed many presenters were using [dbt](https://www.getdbt.com/product/what-is-dbt). As they began analyzing different architectures, dbt consistently emerged as a best practice due to its simplicity and cost-effectiveness. “Essentially,” said Bryan, “word of mouth led us to dbt.” The team soon created a new database for the structured data model within the business, shifting their data ingestion from Matillion & Fivetran to AWS Glue & Fivetran while still channeling data into Snowflake. They also integrated [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) as the modeling layer, utilizing Looker for data visualization and Atlan for data governance. ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/bb2024a68268850d0ea229b4658ea4bcfcbfd168-2170x1178.png) ### Modularizing complex business logic Once set up, the new tools quickly delivered results. The team soon found the new setup simplified handling the complex business logic underpinning Purple’s core sales model. “Previously, any modifications we made to the model would inevitably disrupt other elements, so we were in constant firefighting mode,” explained Matt. “With dbt, we've established a development and QA process before pushing changes to production. Now, we have a more structured, safer approach to making changes.” This governed approach reduced the time spent on bug-hunting and maintenance while improving trust in the data across the business. “The value of a good data team is reflected in the quality of data provided,” noted Bryan. “You can’t put a price on trustworthy data.” Rather than having to hunt through multiple definitions for the same metric, [decision-makers can now have confidence in the data they work with—allowing them to act more decisively and in alignment with the data team.](https://www.getdbt.com/product/build-trust-in-data-and-data-teams) ### Minimizing on-Call support Another major success of Purple’s new architecture was reducing the time and energy devoted by the data team to support duties. The previous ad-hoc approach to deployment meant the data team would run into issues every single day. The on-call staff always had to be ready to firefight issues like broken dashboards and data inconsistencies. With the new, safer development process, the team now deploys changes twice a week, only needing to monitor for problems for two hours immediately following deployment. “There are always humans in the process, so there will always be a need for support,” said Bryan. “But with dbt’s monitoring and Atlan’s ability to detect downstream issues, our support demands have reduced significantly. They’ve gone from being a constant, 24/7 requirement to just a few hours a week.” Matt explained that dbt’s proactive [monitoring](https://docs.getdbt.com/docs/deploy/monitor-jobs) and [testing abilities](https://docs.getdbt.com/docs/build/data-tests) have been “game-changers.” “We can often address issues before they even result in tickets,” he continued. “The data source freshness tests are especially helpful, allowing us to identify upstream problems early. We've also implemented various built-in and custom tests to enhance our monitoring.” dbt has reduced the time the data team spends on support rather than development and improved morale by virtually eliminating one of the most stressful aspects of the data team’s jobs: the Monday morning support rush. “Monday mornings were just the worst,” winced Matt. “If we let somebody push to Snowflake on a Friday, without fail, on Monday morning, something was broken. Every Monday was dedicated to support for at least two or three people.” “The transition has made Mondays far less stressful,” added Bryan. “It’s almost eliminated Sunday night stress!” ### Optimizing costs as data expands As the volume and complexity of data models increase, it tends to drive up computation costs. However, despite the implementation of dbt coinciding with a significant corporate acquisition by Purple—and the added strain associated with managing data from the new business—the company found that its costs have remained stable. Some of this stability is thanks to the data team’s careful management, but Matt also attributes much of the [cost management success](https://www.getdbt.com/product/cost-optimization) to the increased efficiency of the new system. “Moving our business logic from LookML to Snowflake and utilizing dbt Cloud for modeling has significantly reduced our Looker warehouse compute costs,” he explained. “Our latest consumption reports show stable computing levels, which is notable given the additional systems integrated since adding IntelliBed (the acquired company). Overall, we believe our tools and actions are effectively managing our compute costs.” ### Managing a major data re-architecture With the new data stack in place, the data team finally emerged from firefighting mode and began planning for the future. These plans have combined into a major re-architecture effort driven by dbt. The project's primary goal is to supply consistent and reliable data across all departments. “We need to ensure that our data is uniform and trustworthy across various analytics teams. This project aims to build confidence in our data, enabling better business decisions and reducing skepticism and deflection at the top level.” Bryan added: “dbt offers simplicity. We want to efficiently handle data requests, aiming for a turnaround of a few days rather than weeks or months. It's about increasing efficiency and response time for our data needs.” The extensive project is ongoing, with the team’s current focus on retiring Matillion. The team is transitioning Python scripts and queries from Matillion to AWS Glue for source data and rewriting transformations in dbt for data marts and similar tasks. “With Matillion, it was difficult to separate what was a source transformation versus what was a transformation getting used in Looker, he said. “With the project structure we have in place now in dbt, it's all very clear. And we’re working on extending that clarity across the company.” --- --- title: "One NZ unifies customer data with dbt Cloud and Data Domain" description: "Learn how New Zealand’s largest telecom streamlined customer tracking and cross-sell enablement with dbt Cloud and Data Domain." url: "https://www.getdbt.com/case-studies/one-nz" date: "2024-01-19" industry: "Telecommunications" --- # One NZ unifies customer data with dbt Cloud and Data Domain Learn how New Zealand’s largest telecom streamlined customer tracking and cross-sell enablement with dbt Cloud and Data Domain. ### Company details - Headquarters: Auckland, New Zealand - Solution: Telecommunications, wireless carrier - Data stack: Snowflake, dbt Cloud ### Results - 3.91K unique dashboard views within the first month of launch - 75% reduction in time to collect and research data > “By mastering the data we have on complex and often bespoke customer solutions, our single view of the customer capability has empowered our sales and customer success teams to more efficiently and effectively serve our customers." > > — David Redmore, Head of Enterprise Product and Commercial [One New Zealand](https://one.nz/) (formerly Vodafone New Zealand) was part of the global Vodafone business network. It ceased to be a subsidiary of the UK-based parent Vodafone Plc in 2019. One New Zealand (One NZ) is the largest wireless carrier in New Zealand, accounting for 38% of the country's mobile market share in 2021 with 2.4 million customers. In 2022, One NZ launched a business initiative aimed at elevating customer engagement and marketing for its business customers. Partnering with data consultancy [Data Domain](https://www.datadomain.co/), the initiative focused on the following business value: - **Cost reduction**: Reducing reliance on external teams to access customer information and reporting, along with automation and consolidation of existing reports. - **Improving customer experience:** Reducing lead time to call customers, and enabling more proactive conversations through easier access to relevant data - **Increasing revenue and reducing churn:** Using business data to create cross-sell opportunities, targeted campaigns, and customized solutions. ### Problem Statements - The data and analytics for business customers were disparate and inconsistent. - No single view of business customers for One NZ to drive strategic sales and campaigns. - The team lacked a single source of truth in the data to drive growth and retention initiatives, resulting in manual duplicated effort and analytics errors. - To build a view of accounts and engage with customers, the sales teams needed to manually collate customer information, searching across multiple systems. - Existing data sources, structures, and quality were complicated and poor. The solution required automated testing to resolve core data issues and expedite delivery velocity. To address these issues, the team created data products to consolidate service and account data for business customers. With new tools and processes, the team could ensure the build quality met the standard required to support customer communication, campaigning, and in-application capabilities. ### The Solution One NZ implemented an agile squad, consisting of a Product Owner and a team with multi-disciplinary technical capabilities. Data Domain led the discovery, design, analysis, build, and testing dimensions. The team aimed to provide value quickly, iteratively, and collaboratively. Due to the complexity of the source data, the approach was to split the delivery into two phases: Proof of Concept and Productionisation. The solution had to accommodate core requirements and enable small incremental changes, modifications, or additions to be included in future phases. The requirements were initiated with all stakeholders, including end users, to ensure the data product would be fit for purpose. The build was carried out with input from sales teams to identify changes early and incorporate them into the development phase. Development occurred collaboratively with other vendors and internal stakeholders, while delivery management, across multiple workstreams, ensured value promptly. The development cycle ensured end users could immediately use the data product, understand the scope of upcoming development cycles, and receive comprehensive training to achieve high satisfaction. ### The Outcomes - The Data Domain team deployed a pioneering solution for One NZ using [dbt Cloud and Snowflake Data Cloud](https://www.getdbt.com/data-platforms/snowflake)—now the build standard for the Tribe. - The team introduced new ways of working, with a new template-driven dbt development framework, and baseline build. The dbt’s automatic code generation framework increased developer productivity and enforced best practices across data engineering squads within One NZ. - The project’s success initiated a roadmap for continuous funding of similar initiatives and the establishment of an ongoing dedicated squad for Business, with the potential to expand into Machine Learning and Artificial Intelligence programs of work. - The data products became valuable assets, providing customer-facing teams with an easily accessible, consolidated view of information previously challenging to obtain or entirely missing. The consolidated customer view reduced CS data collection and research from 20 minutes to less than 5 minutes. - The sales teams at One NZ use the customer dashboard to make timely decisions and inform conversations, yielding 3.91K unique dashboard views within one month. ![outcomes](https://cdn.sanity.io/images/wl0ndo6t/main/08b6feed0ac31d23439f658028e10732ac90e760-2126x1312.png) > _“By mastering the data we have on complex and often bespoke customer solutions, our single view of the customer capability has empowered our sales and customer success teams to more efficiently and effectively serve our customers."_ > > — David Redmore, Head of Enterprise Product and Commercial, One NZ - The Power Bl dashboard saw frequent and high usage, with positive feedback from stakeholders, making it a business-as-usual (BAU) tool. Customer-facing teams now have easy access to consolidated views with the top users consistently interacting with the dashboards hundreds of times per day. - The successful implementation of the One NZ Business Customer Initiative not only improved operational efficiency and customer engagement but also set a benchmark for future data initiatives within the organisation. The project's results, including cost reduction, enhanced customer experiences, increased revenue, and reduced churn, underscored its importance as a strategic move towards data-driven growth and retention. > _"This has been a game changer for our teams! We have everything we need to understand the services our customers have at a click of a button."_ > > — Simone Cuthbert-Scott, Head of Customer Success, One NZ The initial success of the data platform led to new data products leveraging the same workflow and solutions. For example, customer loyalty program dashboards and enterprise account management. The team is now working on retiring legacy data products. --- --- title: "Plentific implements a robust, scalable data workflow with dbt Cloud" description: "Learn how Plentific leverages dbt Cloud to ensure accuracy and stability across data pipelines and reporting." url: "https://www.getdbt.com/case-studies/plentific" date: "2023-11-16" industry: "Real Estate" --- # Plentific implements a robust, scalable data workflow with dbt Cloud Learn how Plentific leverages dbt Cloud to ensure accuracy and stability across data pipelines and reporting. ### Company details - Headquarters: London, UK - Solution: Property management platform - Data stack: dbt Cloud, Looker, Snowflake, AWS, GCP, PostgreSQL, Stitch, Hevo Data, Github ### Results - 99% decline in data pipeline breaks since implementing automated end-to-end testing - 70% data team growth over the last 3 years - 20 reporting views modeled from 1000s of raw tables > “Instead of relying on product engineering to manually communicate changes, we leveraged dbt Cloud to automatically prompt us before code goes to production. We’re very proud of this system because it protects the integrity of our downstream data products like advanced analytics and machine learning recommendation engines.” > > — Raúl Aviles Poblador, Head of Data Engineering ### A real-time property management orchestration platform [Plentific](https://www.plentific.com/en-us/) is pioneering real-time property operations for real-world impact. Its software as a service (SaaS) platform seamlessly connects owners, operators, service providers, and residents in one place, making operations simpler, faster, and more efficient. Working with clients to streamline operations, unlock revenue, enhance resident experience, and remain compliant, Plentific empowers clients with data-driven insights that drive action. Plentific is dedicated to building stronger communities where people can thrive, with its growing network of 1.5M+ properties and 25,000+ service providers worldwide. #### Providing recommendation models and customer-facing data products To recommend the best contractors for each service, Plentific uses machine learning models. Data not only impacts revenue analysis but also generates direct revenue for Plentific. With their flagship data product, Advanced Analytics, customers can view in real-time everything about their property’s repairs, maintenance, inspections, and compliance jobs. ### A lack of data development standardization Like many companies, in the first years of existence, Plentific was a start-up that operated without a data warehouse and data infrastructure in general and instead used ad-hoc SQL scripts, which ran against their production databases. This led to: - **Siloed data knowledge:** Business users needed help from SQL-proficient technical stakeholders to extract or interpret data. - **Pipeline instability:** Without pull requests or version control, new code from the product would break data pipelines. There were no alerts or testing in place. - **Lack of observability:** The data team could not visualize the status of their cron jobs or easily perform root cause analysis. As Plentific grew, these issues hindering data productivity and quality were remediated by reassessing the data stack and implementing a scalable infrastructure. ### Selecting the right tools Plentific’s data stack reassessment began with selecting a new data warehouse ([Snowflake](https://www.snowflake.com/en/)) and business intelligence tool ([Looker](https://cloud.google.com/looker-studio)). But they were still missing a piece in their architecture puzzle: a middle layer for data transformation. “We needed to transform approximately 1000 tables in a normalized form to 15-20 tables consumable by non-technical users,” explained Raúl Aviles Poblador, Head of Data Engineering at Plentific. “We searched for tools for this use case and found there wasn’t any real competition to [dbt](https://www.getdbt.com/product/what-is-dbt). Airflow was doing something similar, but it wasn’t SQL-specific. Looker had a semantic layer, but it was at the time more focused on business intelligence.” To assess whether dbt could meet their transformation requirements, the Plentific data team got started with dbt’s open source offering, [dbt Core](https://www.getdbt.com/product/dbt-cloud), before later migrating to dbt Cloud. ### Reaping the data quality benefits of a modern data stack #### Improving data stability with automated testing The team implemented an automated testing system to prevent pipeline breaks caused by product changes. Now, whenever there’s a new pull request on Github, a Jenkins job checks the JSON generated by dbt to assess whether the new code could impact existing data pipelines. Since implementing these new processes, Plentific hasn’t yet had a single broken data pipeline in months. The new system boosts the product team’s confidence, as they know their changes won’t break anything downstream. “We’re very proud of this system because it protects the integrity of our data products and obviously, both we and our clients highly value that," emphasized Raúl. “dbt brought this new structure into place. It made it easier for us to test our changes across machine learning and advanced analytics,” added Bruno Lopes, Principal Data Engineer at Plentific. #### Consolidating revenue reporting Plentific’s revenue reports, built on Looker, are widely used—accessed by 1 in 3 employees across the company. “Each financial management vendor reports data differently so we need to consolidate all revenue data into a single, accurate model that can be consumed and understood by all stakeholders,” explained Raúl. Data from Netsuite and Xero is transformed and joined in dbt, simplifying the creation, management, and maintenance of the revenue reports. #### Honoring Advanced Analytics SLAs with increased stability Beyond reporting, dbt is used to create the consumer-facing models in Plentific’s flagship data product: Advanced Analytics. These models are updated in almost real-time and accessed by power users at large property manager organizations in the UK, Germany, and the US. “dbt professionalized all transformation steps and enabled us to scale, minimize risk, and increase stability,” said Raúl. #### Automated documentation with dbt Cloud, Github, and Confluence Plentific set up an automated workflow that creates a Confluence page whenever new code is published to Github. Business and data users can now leverage this [automatic documentation ](https://docs.getdbt.com/docs/build/documentation)to answer their questions and reduce dependency on data and engineering stakeholders. “It’s very quick for everyone to see what tables and sources are being used in the Confluence docs. And, because now we centralize all our data on dbt, this workflow was easy to implement,” shared Bruno. ### What’s next for Plentific: enhancing the platform with more machine learning and AI Moving forward, Plentific will keep using its data stack and workflow to support the development of all data products. Next on the list: building and improving new revenue-impacting machine learning models, like pricing recommendation algorithms, and generative AI that enhances work order accuracy and diagnostics. --- --- title: "Secret Escapes modernizes web analytics with dbt Cloud" description: "Discover how Secret Escapes empowers analysts to build business-critical data products safely with dbt Cloud." url: "https://www.getdbt.com/case-studies/secret-escapes" date: "2023-11-06" industry: "Travel" --- # Secret Escapes modernizes web analytics with dbt Cloud Discover how Secret Escapes empowers analysts to build business-critical data products safely with dbt Cloud. ### Company details - Headquarters: London, United Kingdom - Solution: Luxury Travel - Data stack: Snowflake, dbt Cloud, Tableau, Github, Airflow, AWS, Snowplow ### Results - 14 Analysts developing in dbt Cloud - 5 analyst-powered production datasets - 3-Month cross-team project reduced to 3 weeks > “With dbt Cloud, you can give analysts autonomy, while maintaining data governance. Our analysts have taken to it like ducks to water." > > — Robin Patel, Head of Data & Analytics Engineering [Secret Escapes](https://www.secretescapes.com/) is a UK-based luxury travel company that specializes in curated trips to locations across the world. The company acts as an agent for hotels and operators and markets deals to its members. As a digital travel company, Secret Escapes generates a huge amount of data. This data is used to build key insights—such as how users interact with the company website, or the status of relationships with hotels and operators—and improve core business areas and customers’ experience. ### Foundations of Secret Escapes’ Modern Data Stack In 2018, Secret Escapes invested in modernizing its data stack to ensure resilience through trade volatility. “When I first joined, we only had Snowflake and a foundational infrastructure that allowed our pipeline to ingest data into our warehouse,” said Robin Patel, Secret Escapes’ Head of Data & Analytics Engineering. Over the next few years, the Data Platform team worked to build out the business’ data warehouse, with a vision of easier, more efficient reporting. Forever an evolving product, the data warehouse soon proved to the business that it was capable of serving data for multiple use cases—whether analytics or feeding external systems. However, its initial success led to more data requests, and a new problem emerged; the Data Platform team became a bottleneck for creating new datasets. Secret Escapes needed a solution that would empower analysts to provide data products, not just data engineers within the Data Platform team. The team demoed and ultimately chose [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) with this goal in mind: providing autonomy to analysts around the business, while also instilling data governance and development principles. ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/fd18caa6a46a995c5da6db0dd888ce0d1cc739c6-1006x680.png) ### Democratizing and de-risking legacy data systems using dbt Cloud One of the first datasets the team set their sights on improving using dbt was an existing system for modeling and comparing the return on investment (RoI) of Secret Escapes’ marketing initiatives. Before introducing dbt Cloud, the model was functional but cumbersome. 5 years ago, Secret Escapes commissioned an agency to create their custom marketing RoI solution; the system connected to external marketing systems to ingest data into an old MySQL database and then into an Access database where costs were married with revenue by ingesting a CSV mapping file. “It was not only complex but also isolated and risky,” explained Robin. The team knew there were improvements to be made, and took on the project as their dbt POV and evaluation: “Rather than incurring migration costs from external contractors to build more custom connectors in our warehouse, getting our engineers to model the data, and then working with business stakeholders to understand if a model is fit for purpose, we instead taught our analysts how to use dbt Cloud, gave them the data, and they did the rest,” said Robin. Before dbt, similar requests would have sat in a queue for months. “It wouldn't have met prioritization, so it would have taken a couple of months for us to pick it up—and then any iteration needed would require a new ticket, which takes time to implement. With dbt, analysts got the data product delivered quickly, and the marketing analysts own all the logic,” explained Gianni Raftis, Head of Data Products & Strategy at Secret Escapes. ### Discovering dbt Cloud: More than Documentation Gianni and the team first heard of dbt in [peer discussions](https://www.getdbt.com/community) on documentation. “A lot of data professionals are mentioning dbt, and I’d heard great things about its documentation features,” said Gianni. Secret Escapes runs a monthly hackday: “On the last two days of each month, engineers can work on whatever they want, free of business requests,” explained Gianni. “We started to experiment with dbt and realized that it could do so much for us. And it just snowballed from there.” Data Platform then tested dbt in their hack day and they soon realized that the solution could provide much more than they expected. Since implementing dbt Cloud, the data team at Secret Escapes has found that they’re able to give autonomy to users outside of the Data Platform team to [model and structure their data and develop pipelines](https://www.getdbt.com/product/develop). ### Uniting Data across Subsidiaries One of the most significant changes brought about by using dbt Cloud is the integration of data across subsidiaries. “Our business incorporates a few subsidiary brands, acquired over the years,” explained Robin. “From a group-level perspective, we've got several technology stacks with different ways of reporting metrics. Because they all have slightly different products, it's not as simple as just rolling them up together.” Using dbt, one single [analyst](https://www.getdbt.com/product/analyst) was able to take the different data sets and combine them together for group-level reporting—removing significant amounts of manual labor compared to their previous process, which was heavily reliant on Excel and not automated. "dbt has empowered me to build a unified data model that incorporates all trading metrics for each entity with Secret Escapes, a model previously nonexistent in the business. This has streamlined our reporting processes, enabling the automated creation and distribution of daily and monthly KPI reports to the business and board of directors," said Dharmita Bhanderi, Senior FPA Analyst. ### Increased C-Suite Understanding: Business Predictability Another area where dbt Cloud proved itself invaluable was in communicating information to stakeholders. “We had a high-level need to better understand our user base,” explained Robin. “Using dbt Cloud, we were able to very quickly build something that would have previously taken an exorbitant amount of time.” The solution’s utility, combined with the team’s ability to iterate further leveraging dbt, helped sell it to other departments at Secret Escapes. “The model delivered great value, helping the wider business and the C-suite understand a lot more about our user base—with metrics such as how often a user drops from our service flow and becomes inactive,” Robin added. “Now we can use our model to better predict how our business will perform based on how much we spend.” ### Rapid Iteration: Modeling Email A/B Tests dbt’s use cases at Secret Escapes extend beyond reporting. Building on top of dbt’s models, the team implemented a series of A/B tests on their email and web recommendations to test different approaches and strategies and judge their performance. “We can maintain and quickly adjust code in dbt Cloud—without restarting the whole development process within our warehouse,” said Robin. Because Secret Escapes can track the performances of the A/B tests within dbt Cloud, it’s easier and quicker to output results and iterate. “I managed to build a whole model module around A/B testing with loads of different data sets and create a scalable product,” Gianni added. “Now, every test I run is all in dbt—and it runs every day. Even the CRM (customer relationship management) executives are using the model, which is just brilliant.” ### Growing Analyst Confidence With dbt Cloud’s assistance, the Secret Escapes data platform team is rapidly developing its capabilities. Robin is careful to dedicate time and resources to the people using the data tools as well to ensure the technology meets users where they are. “We can't expect our analysts—not to mention everyone else in the business that wants to use the data—to be advancing at the pace of our technology. It's just not sustainable. Instead, we try to empower our end users.” Rather than serve raw data, the Secret Escapes data team provides other teams access to a semantic layer of approved core data products that have already undergone modeling & transformation. “If we had just given everyone the full raw datasets to play with, most individuals would feel overwhelmed,” explained Robin. “They would be hesitant to leverage our centralized metrics, resulting in redundant data, redundant processes, and disengagement from our products.” By tailoring which datasets users interact with through dbt Cloud, Secret Escapes maintains team engagement and thoughtful collaboration. “It stops data being misinterpreted downstream, and [gives analysts greater confidence](https://www.getdbt.com/product/analyst) by serving as a sort of safety net.” ### Guard Rails: Enabling Risk-free Data Development Echoing the spirit of safe data usage, Secret Escapes used dbt Cloud to create guard rails that ensure data modeling occurs in a safe, governed manner. “Some of the guard rails are there to ensure that the analysts’ work never affects any of the core data sets,” Robin explained. “Anything they do rests on top of the core data, appearing in a completely separate database. This keeps things from getting messy, letting the analysts get to work without accidentally disrupting our core data. And because the development layer sits within the same tool, people can still navigate it easily.” The team has also used dbt Cloud to separate certain data into departmental layers. “We have sub-sections inside our dbt database and schemas specific to departments,” Robin added. “If I’m looking for CRM data, for instance, I can just go to the CRM schema.” Far from hindering analysts, these guard rails ensure that dbt can be navigated with confidence. “The guard rails reduce risk and help analysts boost their knowledge and confidence within dbt,” said Robin. ### Looking to the Future: Training and Expansion Going forward, the data team is hoping to further push the benefits of dbt Cloud out to the wider organization. “We’re going to be introducing targeted sessions around documentation and testing, as well as more community-oriented initiatives like drop-ins,” said Robin. Meanwhile, with analysts now safely and confidently managing the data they need, Secret Escapes is looking to make the most of its modernized web analytics capabilities. “The next steps will be polishing our insights, and helping people unlock dbt’s complementary features like documentation and freshness.” --- --- title: "Watercare builds a scalable, resilient data warehouse with dbt Cloud" description: "Discover how Watercare modernized its analytics with dbt Cloud, moving from manual processes to a scalable, resilient data warehouse." url: "https://www.getdbt.com/case-studies/watercare" date: "2023-11-06" industry: "Utilities" --- # Watercare builds a scalable, resilient data warehouse with dbt Cloud Discover how Watercare modernized its analytics with dbt Cloud, moving from manual processes to a scalable, resilient data warehouse. ### Company details - Headquarters: Auckland, New Zealand - Solution: Water Utility Provider - Data stack: Snowflake, dbt Cloud, Power BI ### Results - 30,000 smart meters tracked in 1 dashboard - 7x increase in the number of data models produced in a year > “All of the improvements driven by dbt Cloud are starting to add up, and the quality of our data has exponentially grown. We can put testing in place and enforce rules to ensure that data going to users is correct without having to spend time manually checking it.” > > — Diego Morales, Analytics & Insights Practice Lead ### From Water Pipelines to Data Pipelines [Watercare](https://www.watercare.co.nz/) provides essential water and wastewater services to over 1.7 million people across the greater Auckland region in New Zealand. As a utility company responsible for critical water infrastructure worth billions of dollars, Watercare relies heavily on data to optimize its operations and assets. However, Watercare's legacy architecture made it difficult to get value from the vast quantity of data it was generating. Siloed systems led to unclear data lineage, hindering troubleshooting, while a lack of data modeling and transformation capabilities limited the team’s ability to perform analytics. Diego Morales, Ex-Analytics and Insights Practice Lead, said: “Watercare has terabytes and terabytes of data to manage. We easily fall into the 'big data' category. But when I stepped into the role of Analytics and Insights Practice Lead, there wasn't a solid strategy on how to harness this data for insights.” This lack of insight increased in severity as the business handled growing volumes of data from customers, facilities, and users. With no way to easily manage this information, it became clear that Watercare needed a modern data solution. “We couldn't drive trends, and we couldn't understand when something changed in our system,” Diego explained. “We didn't have any data lineage. The only way we could track data was by looking directly at the code...which made it hard to understand and trust the information it was producing." “It was clear we weren’t going to get much out of our existing set-up. We needed a tool that allows us to build reliable data models.” ### Leveraging high-volume smart meter data The team began exploring options for a modern data stack. After receiving a recommendation from a contact in the data sector, Diego evaluated [dbt Cloud](https://www.getdbt.com/product/dbt-cloud). “After doing the research into the best tools for the job, it became obvious that dbt Cloud was the standard” he noted. “If you speak to people in the industry, everywhere everyone's talking about dbt.” Despite their enthusiasm for the tool, however, the team still needed to demonstrate its effectiveness to the rest of the business. To do this, they used dbt to power a proof of concept (POC) project to develop dashboards drawing on data from the business’ smart meters. These meters are intelligent, networked devices that digitally measure and record water consumption in real-time. Diego explained that the project turned out to be extremely useful, as the support team is regularly contacted by customers querying unusual billing patterns. When these unusual patterns include a sudden increase in usage, the Watercare team investigates potential leaks. “With our new data model, we can offer a clear visual representation of their consumption trends and easily spot any spikes,” Diego said. “This data means we can swiftly notify our smart meter team about potential issues, such as leaks, tampering, or changes in pipe pressure. This is extremely valuable for teams that are deployed in the field.” The project was, in Diego’s words, “incredibly successful.” Despite starting as a small proof of concept, the dashboard quickly became Watercare's most frequently used data tool. “Our customer service team uses it every day,” he said. “Customers might ask them to confirm they don't have a leak, and with a quick check of the dashboard, they can provide answers. Out of the 600 dashboards deployed within the organization, it’s the one people access the most. That's big." ### Building Trust and Speed with dbt Cloud The success of the POC project encouraged the organization to realize the benefits of trusted analytics and responsive dashboards. The data team was entrusted with NZ$2.5 million in funding to expand the use of dbt Cloud. Over the next year, Watercare went from using just 25 data models across its business to being able to produce 175 with the help of dbt Cloud. As the use of this data grew, a solid foundation of testing and documentation ensured analysts and other users that they could trust the numbers they were seeing. For example, the team was able to integrate its CI/CD pipelines with the dbt Cloud API, which made both collaboration and governance much easier. “We have started documenting our dbt models in the open so different teams can discover and reuse existing logic. And we can work together on new transformations that meet multiple needs. On the data team side, adopting dbt has helped us work more efficiently and align us more tightly with the goals of the analytics users we aim to empower,” said Diego. “It's a complete game-changer.” At the same time, Watercare has been making use of dbt’s [community hub](https://www.getdbt.com/community/). The open-source community enables the team to accelerate their development using ready-made packages supplied and documented by members of the dbt community and adapt them to suit their own needs. “All of the improvements driven by dbt Cloud are starting to add up,” explained Diego. “We can put testing in place and enforce different rules to ensure that the data going to users is going to be correct without having to spend time checking them. “ ### Delivering Value Across the Business: Machine Learning and Custom Analytics Watercare successfully transitioned from a series of disjointed legacy data silos to an agile, analytics-focused data architecture. This modern foundation has prepared Watercare to support advanced use cases like working with IoT sensors and has allowed more and more areas of the business to start making the most of its data. The business’ critical financial and work order reporting is now built on top of dbt models, and users are being encouraged to self-serve—using dbt Cloud to explore the data without the need for an expert to walk them through it. The risk-free accessibility is helping Watercare enable teams on the value of data as an asset rather than a static means to an end. “We're bridging that gap because we're providing self-service analytics,” he explained. ”The POC project created a model and used it to generate reports, but we’ve also used it to provide capacity for self-service analytics. “We are using dbt Cloud to produce and create many data models, but now people can take all this highly organized data in dbt and use it to create new use cases for themselves.” Watercare’s trusted [dbt models](https://docs.getdbt.com/docs/build/models) also feed machine learning pipelines, used to create tools that run automatically and spot potential errors. "We managed to produce such a high-quality data model that now consuming it into machine learning became easy,” Diego explained. “Immediately after that, because the data was in good condition, we started a process to build an ML model to detect any anomalies in the data.” ### Building for Better Water Quality Looking ahead, Watercare has plans to further enhance analytics by integrating [Databricks](https://www.databricks.com/), a lakehouse platform for machine learning workloads. The team is also looking to integrate new data sources into its pipeline, pulling in information from sensors able to measure parameters like temperature and pollution levels. "That data is, as you can imagine, incredibly fundamental for a company that works with water. We're going to take all of that information and move it into dbt Cloud,” concluded Diego. --- --- title: "Nitrogen accelerates data velocity with dbt Cloud" description: "Learn how Nitrogen streamlined data workflows with dbt Cloud, cutting integration time in half and reducing operational costs." url: "https://www.getdbt.com/case-studies/nitrogen" date: "2023-11-06" industry: "Banking & Financial Services" --- # Nitrogen accelerates data velocity with dbt Cloud Learn how Nitrogen streamlined data workflows with dbt Cloud, cutting integration time in half and reducing operational costs. ### Company details - Headquarters: Auburn, California - Solution: Wealth Management - Data stack: Snowflake, dbt Cloud, DOMO, Snowpipe, Github ### Results - 50% time reduction to build and deliver data integrations - 30% faster to deliver data sets - 2-4x fewer operational costs than running ETL pipelines > “Even though we had a smaller team than all of our past integration projects, we were able to deliver the integration in just six weeks. That’s exactly half of the time that projects with similar scope took previously.” > > — Andrew Waters, Director of Platform Engineering Founded in 2011 in California as “Riskalyze,” [Nitrogen](https://nitrogenwealth.com/) started in the wealth management industry with a portfolio risk analytics tool. Today, under the Nitrogen brand, the company delivers firm-wide analytics, compliance solutions, and dozens of data integrations with custodians, banks, and other fintech vendors. Nitrogen enables wealth managers to get a full overview of their customers’ portfolios on different custodians, such as Fidelity or Charles Schwab. They can use the product to design new investment strategies or prepare pitches for prospective clients. “Without current accounts and investment positions data landing early in the morning every day like clockwork, our customers can’t rely on the data to get results for their clients,” said Andrew Waters, Director of Platform Engineering at Nitrogen. Timely, accurate data is essential for the success of the product and, therefore, the business: “For us, data is not just for internal teams or business intelligence. It’s a key value of the product we offer,” explained Andrew. ### The role of data integrations at Nitrogen Data integrations, in particular, are the backbone of Nitrogen’s product. The company has data integrations across dozens of firms, including banks and other tech providers. “Our financial advisors (customers) need to see their data flowing in both directions,” explained Andrew. “We have to work seamlessly with their other tools.” “Establishing those data feeds is core to anything they do across our product. A key value that we're offering is powered by the data that we have available,” emphasized Andrew. ### A unique opportunity to move upmarket comes with a data caveat In recent years, Nitrogen’s target audience has expanded from SMBs—the “mom-and-pop shops” of financial advisors—to also include mid-sized and large firms. Earlier this year, a unique opportunity arose to land a mid-sized customer. The prospect wanted to purchase the product for their entire firm of hundreds of advisors. However, it came with a caveat: a new data integration had to be delivered within six weeks before their contract was due to start. “Our sales team did an incredible job landing this customer,” said Andrew. “We wanted to put every effort to meet their requirements.” The data team investigated whether they could deliver this new data integration with their existing workflow: “The deadline was impossibly tight, so we evaluated all options. We quickly ruled our existing process of building a new interface in our legacy data pipeline, which typically took over a quarter to complete,” said Andrew. “We knew we needed to come up with an innovative solution as more of these larger wealth management firms would have similar requirements.” ### Legacy infrastructure hinders upmarket expansion #### The limits of the existing data infrastructure The question of delivering on the data integration request promptly reopened discussions on Nitrogen's legacy infrastructure—a mix of in-house software data transformation applications and traditional relational databases such as MySQL and Postgres. “Various components of our stack weren’t scalable, and took patience to develop in,” explained Andrew. “We knew the way we were currently doing things wouldn’t scale to where we want to go in the future...we were pushing the limits of what we could do with our old systems.” #### The search for a new data infrastructure Faced with these challenges, the Nitrogen data team started the search for alternative solutions: “We saw numerous opportunities if we moved to a central data lake. We could leverage machine learning to improve retention, and also decrease costs by moving out of expensive relational databases,” said Andrew. ### Evaluating dbt Cloud and modern data tools #### Scoping requirements: including the business & investors The Nitrogen data team began by interviewing departments across the company to identify data use case opportunities and priorities, and involve other teams in the evaluation process. HG, a majority investor at Nitrogen, also participated in the evaluation process. #### Recommending dbt Cloud: past success and ease-of-use Based on past successes with other portfolio SaaS companies, HG put forward [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) as their recommendation, kicking off Nitrogen’s evaluation. “We saw dbt Cloud as an opportunity because of its simple learning curve. It’s an approachable SQL-based skillset, which made it easy to adopt,” said Andrew. “It also looked like an efficient and cost-effective solution that fit our business needs.” “After a quick evaluation, we determined dbt Cloud suited our requirements and was our only option that could deliver the integration on time; so we greenlit the project.” #### Getting started with a hackathon With [dbt Cloud and Snowflake](https://www.getdbt.com/data-platforms/snowflake), the team hit the ground running: “We got a group together from different teams that were interested in the new data stack and did an early hackathon,” shared Andrew. “Half of the team explored the transformation side of things on dbt Cloud, and the other focused on Snowflake.” Although this was a small-scale project, it was a successful first effort. Armed with product knowledge and confidence in their new solution, the team could move on to their production use case: building the requested data integration for the prospective customer. ### Successfully tackling the integration use case, and landing the customer Unfortunately, most of the data engineering team at Nitrogen was tied up with other commitments and could not assist with the data integration work: “Normally, I don’t get into development but we only had two people with enough capacity,” explained Andrew. “I paired up with our principal data engineer and together, the two of us dove into learning dbt and building out the first integration use case.” Despite the compact team size, Nitrogen successfully delivered the integration in time: “Even though we had a smaller team than all of our past integration projects, we were able to deliver the integration in just six weeks. That’s exactly half of the time that projects with similar scope took previously,” noted Andrew. To speed up the onboarding process, Nitrogen had a dbt Labs trainer take them through dbt capabilities in a 3-week, [6-course series](https://www.getdbt.com/dbt-learn): “[dbt Learn](https://www.getdbt.com/dbt-learn) was a fantastic introduction,” said Andrew. “Having access to those session recordings helped us rapidly onboard the rest of our team. We still use them to onboard newcomers and other teams that want to get started with dbt Cloud.” ### Reaping the benefits of migrating to a modern data stack ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/d357ba010dff06a4033df41cd3124f7d71125212-1600x791.png) #### A substantial decrease in the cost of data maintenance The new data structure driven by dbt Cloud simplified the set of systems needed to maintain the Nitrogen’s data infrastructure: “We don't have as many steps in the process, such as going through different queues and then staging the database,” explained Andrew. “The whole architecture has been streamlined. The new workflow on dbt Cloud and Snowflake runs a lot more efficiently.” One of the sources of increased efficiency was the move away from large and cumbersome relational databases. “We don’t have a final measurement on saved costs yet, but it’s already very clear that it’s substantially less. We’re probably spending 2 to 4 times less,” said Andrew. #### Faster data delivery and improved customer experience The streamlined infrastructure also had a positive impact on data delivery: “Our daily imports now run under 4 minutes, which is 30% faster than similar-sized data sets in our old process,” said Andrew. The speedier data delivery has a direct impact on customer experience, with wealth managers receiving their customers’ and market data earlier: “Data lands in customers’ hands earlier in the day, and enables us to hit our objective of delivering all account data before financial markets open,” explained Andrew. #### Increased visibility and collaboration Incentivizing collaboration wasn’t an objective for Nitrogen when they formed their new stack, but it’s proven to be a positive consequence of the migration: “Collaboration is just now becoming a visible benefit of having a central repository and clear data lineage,” said Andrew. “Sometimes knowledge is locked into one person’s head, but now it’s obvious—anybody can explore where the data came from and understand how it was transformed. The dbt lineage feature and how data is linked explicitly in the product is perfect for that.” #### A simpler workflow with SQL at the forefront With SQL-first dbt Cloud in place, the [data transformation](https://www.getdbt.com/blog/successful-data-transformation) process is magnitudes simpler: “Before, the logic and custom applications, such as using JSON objects, for data translation were far too complex,” shared Andrew. “SQL-based development is inherently faster and a better fit for a proper data architecture.” “Today, to make a simple transformation layer, we only have to worry about three things: getting the data into Snowflake, transforming the data in dbt Cloud, and then exporting to a destination.” ### Looking ahead for Nitrogen #### Phase out legacy technologies The first project replacing the legacy data infrastructure was a success for Nitrogen. Moving forward, the data team will continue working on the transition from the legacy stack and the depreciation of their existing, expensive relational databases. “We’ve proven with the first project that this new stack is better in performance on all factors we care about. It beats the baseline, which is very positive from an early prototype,” said Andrew. “Now over the next year, we’ll continue expanding that logic to handle all our current and future data use cases.” #### Spread dbt adoption to other teams, including compliance Business users and executives have already started using the new dbt data views, but other teams are still being onboarded to the platform. Compliance officers, in specific, see a big opportunity in using the new stack to prove they’re meeting fiduciary responsibilities: “That’s the second big use case we’re delivering, which is crucial in our industry,” explained Andrew. “We’re building a new data warehouse to serve customer-facing BI in our application. This will assist wealth managers in assessing and visualizing risk for their different customers.” #### Improve governance with dbt Cloud’s automated testing Now past the immediate goal of taking their first use case live, fast, the Nitrogen team will continue to focus on data governance and data quality over the next year. “We want to explore dbt Cloud’s data quality features, such as alerts and data validation,” said Andrew. “We’re still early in our journey; we see a big opportunity in improved data governance, and we’re excited to layer on standards, oversight, and everything else dbt Cloud has to offer.” --- --- title: "Fullscript doubles data sources with dbt Cloud and Secoda" description: "Learn how Fullscript used dbt Cloud and Secoda to scale data workflows and modernize its stack during a company merger." url: "https://www.getdbt.com/case-studies/fullscript" date: "2023-10-31" industry: "Healthcare" --- # Fullscript doubles data sources with dbt Cloud and Secoda Learn how Fullscript used dbt Cloud and Secoda to scale data workflows and modernize its stack during a company merger. ### Company details - Headquarters: Ottawa, Canada - Solution: Health supplement platform - Data stack: Snowflake, dbt Cloud, Fivetran, Secoda ### Results - 300% enhancement in data pipeline project delivery efficiency - 10x increase in reporting dashboard performance - 100 new sources added following a major company merger in just 1 month >  “I can't imagine how we would have been able to scale so easily or double in size without dbt Cloud and Secoda. Looking back, I feel great that we went with this set of tooling.” > > — Amit Jain, Data Team Technical Director ### Connecting Health Practitioners with Prescriptions Founded in Ottawa, Canada, in 2011, [Fullscript](http://fullscript.com) works to help more than 90,000 health practitioners deliver better care to their patients through its online prescription platform. With Fullscript, health practitioners can create personalized plans for their patients and access over 20,000 practitioner-grade products from more than 300 brands. Any operation of this scale generates a vast amount of information, and Fullscript is no exception. “Data is fundamental to Fullscript,” said Amit Jain, Data Team Technical Director. “We rely on data to report on both the financial and operational aspects of our business.” The team tracks customer engagement through clickstream data on its front-end applications as well as back-end operations for warehousing and shipping. These insights boost customer service and allow the business to make decisions based on accurate, reliable numbers. As Fullscrip grew through 2021, the organization's development structure presented challenges when it came to effective collaboration between data engineering and data analysts, who used separate workflows and tools. “Data engineers were inclined to use software engineering development practices, which meant they didn't speak the same language as data analysts,” explained Amit. Over time, the initial team members with a deep understanding of the legacy data engineering systems underwent role transitions, leaving the data warehouse team devoid of crucial historical knowledge. ### “Almost at a Standstill”: The Need to Upgrade a Legacy Data Stack Fullscript’s legacy data stack consisted of Postgres, SQL, and Python scripts embedded in Argo—a Kubernetes-based orchestration tool. However, it soon became clear that the organization needed to update these systems to a more accessible, reliable data stack in order to scale. “From a data analytics point of view, our legacy systems were tough to manage when things went wrong, and the data didn’t match between source and target,” explained Jain. “We couldn’t identify faulty sources to apply the appropriate fixes.” The fact that all of Fullscript’s tools were custom-developed further complicated matters. “We were almost at a standstill,” said Jain. “The data pipeline was written using complex orchestration mechanisms and had SQL and Python wrapped inside the code. Our team was spending all its time just making things run.” This massive drain of time and resources meant the team didn’t have the capacity to improve the systems or build new products. Fullscript wanted to bring in data from their CRM, NetSuite ERP, Google Analytics, and Facebook Ads, but there was no time to add new sources. “Our teams were at capacity just managing issues as they came in. We didn’t have time to add new requirements or sources.” ### Modernizing Online Prescriptions with dbt Cloud and Secoda Jain and his team knew their legacy stack was limiting the team's potential. They needed to find reliable, scalable, accessible tools to transform and catalog their data. “We estimated that our modernization project would take almost nine months,” said Jain. “We wanted a way to track how the migration was progressing because we had to have both systems running in parallel.” After researching and experimenting with several potential tools, the Fullscript team decided that [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) was the best candidate for data transformation, while [Secoda](https://www.secoda.co/)—an AI-powered data search, cataloging, lineage, and documentation platform—could help them keep track of information as it moved through their systems. According to Jain, “Fullscript chose [dbt](https://www.getdbt.com/product/what-is-dbt)—and dbt Cloud in particular—because it offers excellent integration of data engineering and analytics. It’s also very scalable, has great data analytics capabilities, and is easy to adopt and ramp up quickly.” “At the same time, we found Secoda was the most flexible cataloging tool available. We looked at many other options, but there were always some gaps or source types they didn't support. I can’t imagine anyone running a modern data stack without Secoda.” Soon, the team was able to run a proof-of-concept project during one of the company’s regular hackathons. Almost immediately, they delivered impressive results. “I worked with a few data analysts on our team and used dbt Cloud to create a report within a couple of days,” said Jain. “In the old setup, where everything was custom developed, this would have taken months.” #### Starting Small and Scaling Up Transforming an organization's data stack isn’t a small or simple project. That’s why Fullstack’s team took advantage of dbt's ability to easily scale up as required. “You could start with the open source version and then explore options to add dbt Cloud and then dbt Enterprise for the added support,” said Jain. “It was a very flexible model. You could easily start small and prove your concept; this was especially useful when getting executive buy-in.” With no engineers left from the team that had designed Fullstack’s legacy tools, the business needed a tool that was easy to onboard and adopt. “That's what we found with dbt—a solution easy to teach and use,” said Jain. “Using just SQL, somebody with a data analytics background could use the tool fairly easily. For the initial setup, much of the tools just ran out of the box.” #### Solid Support Leads To An Easy Implementation Clear documentation and accessible support are vital in getting projects off the ground. The Fullscript team found dbt Cloud could deliver on both fronts when [modernizing their data stack](https://www.getdbt.com/product/data-modernization). “We used dbt Cloud's documentation to create a reference architecture for the initial implementation,” said Jain. “From there, we were able to work with the dbt enterprise team to refine it according to our needs." “The support and training capabilities dbt Cloud offered were almost as important as the tool itself.” #### Ease of Use Delivers Cost Savings While it’s easy to think of ease-of-use and well-designed support packages as fringe benefits, the combination of features allowed dbt to deliver when it came to Fullscript’s bottom line. Jain noted that even though the business had moved from free software to a paid product with dbt Cloud, the switch still helped save on costs. “A Kubernetes-based solution required a lot of engineer support and a wide array of specialized engineers,” he explained. “Once we moved to dbt Cloud, we didn't need engineers with very high Docker and Kubernetes expertise. Knowledge of SQL and analytics engineering capabilities were enough to use the platform.” ### Tackling an Acquisition Mid-Platform-Migration The flexibility of dbt Cloud, Secoda, and Fullscript’s new data stack was tested mid-implementation. The team was only a few months into the migration when the business merged with Natural Partners—a supplements provider operating on roughly the same scale as Fullscript. “Overnight, our data sources doubled,” emphasized Jain, “The acquisition plans had been under wraps, so teams weren't even aware it was coming.” Fortunately, the Fullscript team had already started making full use of Secoda, which was able to help them integrate the metadata from Natural Partners’ data warehouse. “Secoda was very helpful,” said Jain. “We were able to quickly connect to these new sources and bring in metadata from their new source databases.” While the acquisition represented a significant challenge for the data team, it also provided an opportunity to showcase the power of their modern data stack. Leveraging dbt Cloud, the team could quickly extend its existing data model and effectively map the data. “Mapping the data took a lot of upfront work, but once it was done, we could extend our existing models,” said Jain. “All of that was powered by dbt Cloud, and it proved the scalability of the modern data stack. We did it quickly and finished the migration within the planned nine-month period.” “I can't imagine how we would have been able to scale so easily or double in size without dbt Cloud and Secoda. I know we made the right choice with this set of tooling, and if I were to go back in time, I would choose it again.” ### Looking Ahead: AI and Expansion With a modern, powerful data stack now in place, Fullscript plans to continue building its capabilities over the coming months and years. One of the key goals is to extend its data models to support the new requirements expected to come in from across the business. For example, the ability to support more Artificial Intelligence Markup Language (AIML) use cases, anticipating increased use of AI across both Fullscript and the broader data ecosystem. “We’d like to use AI/ML around the recommendations engine,” explained Jain. “Specifically, for product recommendations using large language models, ChatGPT-style.” “For example, the ability to develop guidelines for product recommendations would greatly help our practitioners. Something like that would have a major impact on their day-to-day and we now have to tools and foundation to make it happen.” --- --- title: "Rebtel improves data quality and data team productivity by migrating from Matillion to dbt Cloud" description: "Discover how Rebtel streamlined data workflows and adopted engineering best practices with dbt Cloud." url: "https://www.getdbt.com/case-studies/rebtel" date: "2023-10-12" industry: "Telecommunications" --- # Rebtel improves data quality and data team productivity by migrating from Matillion to dbt Cloud Discover how Rebtel streamlined data workflows and adopted engineering best practices with dbt Cloud. ### Company details - Headquarters: Stockholm, Sweden - Solution: Fintech, Telecom - Data stack: Stitch, Snowpipe, Snowflake, Looker, AWS CodeCommit, dbt Cloud ### Results - 80% reduction in licensing costs by migrating from Matillion to dbt Cloud + Stitch - 6 months to implement and launch new data stack - 3-4 weeks per year dedicated to maintenance saved by the data team > “Before, we couldn’t kick off any machine learning projects because incidents would happen. We had little visibility on our data lineage and code on Matillion, so we never really knew if the data was correct or complete. Now we’re working on an ML fraud prevention use case which will have a direct impact on our profit margins.” > > — Chandan Singh, Head of Data ### A one-stop shop for migrants and international nomads #### Starting as the “rebel of telecommunications” Founded in Sweden in 2006, [Rebtel](https://www.rebtel.com/) started in the telecommunications industry helping people make cheaper international calls. Nowadays, the company has pivoted towards offering migrants and international nomads a diversity of services, such as money transfers and mobile top-ups. ### Rebtel’s diverse data types Like any telecom, Rebtel stores and analyzes a large quantity of call data: “We have a lot of data to monitor our call volumes, call quality, delivery, success rates and the likes,” explained Chandan Singh, Head of Data at Rebtel. “Because we are a paid service, we also have payments and subscription data.” As Rebtel’s business expands further into new use cases, new types of data, data structures, and data deliveries arrive: “Now that the business has launched the international money transferring product, we also have data on send volumes,” added Chandan. ### Legacy data stack leads to heavy maintenance work and lack of trust #### Challenges delivering on business requirements Rebtel’s business teams needed timely, accurate data to understand how their expansion was performing and confidently make decisions accordingly. “Stakeholders were requesting different deliverables from the data team,” said Chandan. “But our legacy stack prevented us from addressing as many as we’d like. We had a lot of infrastructure incidents happening on a day-to-day and week-to-week basis.” Because of all of this firefighting, the data team was unable to dedicate the necessary amount of time to business requests. #### Managing the legacy infrastructure, full-time Rebtel was already using [dbt](https://www.getdbt.com/product/dbt) (in the form of open source [dbt Core](https://www.getdbt.com/product/dbt-core-vs-dbt-cloud)) and Snowflake, alongside GUI ETL tool Matillion. They had migrated to this stack a few years ago from a Microsoft SQL Server stack. Although some issues had improved with the migration, the maintenance load had not: “A lot of our time wasn't spent actually working on data modeling or data transformation, but on infrastructure,” explained Chandan. “We had our jobs running in Airflow, which was hosted in a Kubernetes cluster within Azure Cloud. The data team was maintaining all of this, which derailed our focus on building data products." ### Lack of engineering best practices on Matillion At the time, Rebtel’s data team operated without [engineering best practices](https://www.getdbt.com/resources/the-analytics-development-lifecycle)—such as documentation, versioning, and CI/CD. This made bug fixing, infrastructure maintenance, and root cause analysis an even greater task. “In our Matillion setup, we didn’t have a place to create documentation or testing,” said Quentin Coviaux, Data Engineer at Rebtel. “Without these standards, people could write whatever code they wanted. It was more like every engineer had their own best practices,” said Chandan. ”We couldn’t enforce rules, like using SQL Fluff.” Because the team hadn’t integrated Matillion with Git, Rebtel also didn’t have code versioning in place: “You need to upgrade Matillion to have code versioning, but that takes time because you need to spin up new servers, among other things,” said Chandan. “So even though code versioning is essential, we didn’t have it.” #### Discrepancies in data and lack of trust As the number of incidents kept increasing, the trust in Rebtel’s data went the opposite direction. Business stakeholders were no longer certain if the data they were using for decision-making was accurate. “We were frequently getting questions like ‘Can you tell me if this number is correct?’” shared Chandan. “Although it’s easy to ask the question, for the data person, it would take hours of effort to make sure nothing was broken.” With the time required to maintain the infrastructure and constantly double-check the data’s accuracy, Rebtel’s data team had little bandwidth to deliver business value. ### The last straw: incident leads to loss of historical data #### A 2-hour ingestion job failure Rebtel’s legacy data stack was leading to more and more pain points, which the team flagged and discussed with the business. But it was a 2-hour ingestion job failure that broke the camel’s back: “The ingestion process was written ages ago and no one in the team knew how it worked. There was no documentation on Matillion we could rely on,” said Chandan. #### Challenges rerunning jobs and recouping historical data with Matillion Around 70 tables were affected in the incident. One of them was the accounting table, critical for the business: “We spent two weeks fixing just this one table,” shared Chandan “And we realized it was not possible to rerun the transformation in Matillion without spending another few weeks or months to get the data up to date.” #### Decision to depreciate Matillion Given the complexity of retrieving this historical data, Rebtel had to “accept the data loss”: “That’s when we reevaluated and said ‘Okay, we have to migrate away from Matillion because we cannot afford to be in this situation again,’” said Chandan. ### Reflecting on pain points and starting on a holistic data journey #### Mission to add value For the data team, the incident was proof their previous data stack could not support their use cases, and highlighted where they were spending their time. They were working on activities that weren’t adding business value nor improving employee satisfaction or retention: “As data engineers, we want to create value with data. We don’t really want to build infrastructure or upgrade Airflow versions,” said Chandan. With this mission in mind, the data team started putting together an architecture diagram to understand the problems to fix in their legacy structure. #### Visualizing data stack pain points in a diagram Rebtel’s data team put all their architecture in one diagram to visualize their pain points. The exercise helped them get the final buy-in from leadership on their digital transformation project. It also assisted them in identifying specific issues to fix on their existing stack: “We had different redundancy points in our stack, such as doing data transformation on both dbt and Matillion,” explained Quentin. ![challenges](https://cdn.sanity.io/images/wl0ndo6t/main/33fa20aac0868b30a22e2ae6486df30f7b03a885-1628x923.png) #### “Drag and drop” over SQL wasn’t a fit for the team [Data transformations](https://www.getdbt.com/blog/successful-data-transformation) were one of the key problem areas in the legacy data stack. Although Matillion is a “drag and drop” (GUI) tool, the data team was still starting and finishing all transformations on the platform by writing SQL code. “I’ve worked with Matillion previously and was a big supporter of the tool. But when I came in, I quickly changed my mind,” said Chandan. “We were using Matillion as a scheduler only, not like the ETL tool it is.” “It’s a good tool if you find it easier to work with a visual tool instead of writing SQL. But we were very comfortable with using SQL for all our data transformations. It made no sense to pay for an expensive ETL tool without using its intended functionality.” Establishing a SQL-based workflow as the standard, rather than a specialized Matillion skillset, would also facilitate the future onboarding of new data employees. ### Depreciating Matillion As a continuation of their “pain points” diagram exercise, the team started to come up with an architecture that’d best suit their needs: “Little by little, we removed everything until we had a clean architectural design,” explained Quentin. “Our goal was to remove everything from Matillion—both the transformation and ingestion.” ![tech stack](https://cdn.sanity.io/images/wl0ndo6t/main/67717cd3d1d2e50f9b28da6e824708971f0e7bca-1600x846.png) “Instead of having one do-it-all tool that is a big black box, we separated the extraction and loading from the transformation. That really helped us,” said Quentin. During this process, Rebtel also migrated from [dbt Core to dbt Cloud](https://www.getdbt.com/product/dbt-core-vs-dbt-cloud): “With these moves, we took away all the infrastructure maintenance overhead we were spending a lot of our time on,” said Chandan. ### Reaping the benefits of a transparent, governed data infrastructure #### Enforcing engineering best practices with dbt Cloud Certain [engineering best practices](https://www.getdbt.com/resources/the-analytics-development-lifecycle)—such as documentation and CI/CD—are out-of-box for [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) and are automatically enforced. With the depreciation of Matillion, Rebtel’s team saw their workflow improve: “As we’ve been doing the migration from Matillion to dbt Cloud, we have really raised our set of standards for development. dbt made it so much easier to create and maintain documentation,” said Quentin. “The new stack makes perfect sense for us, which is why we’ve decided to invest more in dbt instead of consolidating in Matillion,” explained Quentin. “Before, we didn’t have any documentation or testing. Now, we’ve set up about 80 tests with dbt Cloud and growing.” #### [Improved data quality and trust](https://www.getdbt.com/product/build-trust-in-data-and-data-teams) With the better-governed workflow, data quality has improved and, as a consequence, the business’ trust in data has increased: “We’ve been focused on making sure people can trust the data we are presenting,” said Quentin. “dbt is really helping in governing our data; it has improved our data quality, which is amazing.” The data loss incident that kickstarted the stack migration in the first place would also have been easily avoided with the new infrastructure: “One of the things that works really well on dbt is that you can do a full refresh and catch up on all the transformations in your data pipeline,” said Chandan. #### [Saving costs and time with a simpler workflow](https://www.getdbt.com/product/cost-optimization) With the new stack, the amount of time the team devotes to maintenance has dramatically decreased: “Before, we dedicated three to four weeks per year in maintaining Airflow and the different components of our stack,” said Chandan. “Now we have near zero maintenance work. We don’t have issues with dbt or Stitch.” By moving from Matillion to dbt and Stitch, the team has saved on licensing costs as well as time. They estimate an 80% reduction in data infrastructure costs. ### Bandwidth to focus on non-maintenance activities With the migration from Matillion to the new stack nearly complete, the team has achieved “data stability.” Instead of focusing on maintenance and infrastructure, they can double down on complex data work: “Before, we couldn’t kick off any machine learning projects because incidents would get in the way. We had little visibility on our data lineage and code on Matillion, so we never really knew if the data was correct or complete,” explained Chandan. “Now we’re working on an ML fraud prevention use case, which will have a direct impact on our profit margins.” ### Moving forward with further automation and advanced data analysis #### Automated pipelines and documentation Although Rebtel has already seen big improvements in data governance, they’re still early in their journey of implementing and enforcing engineering best practices. “Right now, we have a set of standards that are written in documentation that developers are supposed to follow,” said Chandan. “There are packages that prevent you from pushing code if you don’t have documentation, but we’re not there yet.” “We’re keen to work on CI/CD pipelines to make the development flow more secure, reliable, and faster,” added Quentin. #### Investigate the dbt Semantic Layer and data catalog The data team has yet to explore dbt Cloud’s features that support a [Semantic Layer and data catalog](https://www.getdbt.com/product/dbt-catalog). Together, these two functionalities can help Rebtel improve data accessibility while maintaining data quality: “We want to work on our data catalog to increase data availability even more for newcomers,” said Quentin. “In the future, we’re also interested in exploring the dbt Semantic Layer and Model Contracts.” --- --- title: "Fundrise increases data velocity and stakeholder trust with dbt Cloud" description: "This is the story of how Fundrise leverages dbt Cloud to scale its data infrastructure and unlock time for high-impact analysis" url: "https://www.getdbt.com/case-studies/fundrise" date: "2023-10-10" industry: "Banking & Financial Services" --- # Fundrise increases data velocity and stakeholder trust with dbt Cloud This is the story of how Fundrise leverages dbt Cloud to scale its data infrastructure and unlock time for high-impact analysis ### Company details - Headquarters: Washington D.C. - Solution: Alternative investment platform - Data stack: Mixpanel, Snowflake, Fivetran, dbt Cloud, Tableau ### Results - >50% reduction in time to introduce new data sources - 0 time spent on data pipeline maintenance > “We can develop much faster than we could before, and our stakeholders are confident that the numbers they’re working with are correct. Those two factors really resonate together.” > > — Jack Ploshnick, Analytics Manager ### A Rapidly Growing Alternative Investment Platform [Fundrise](https://fundrise.com/) is an investment platform that provides direct access to alternative investments. It has experienced rapid growth over recent years, managing over $3 billion on behalf of more than 400,000 investors. The company relies heavily on financial data about its investments, assets, and transactions, as well as customer data from its website and app. Charles Wood, VP of Analytics, explained that their legacy analytics stack had been designed “to be as scrappy as possible and to get the best results we could with minimal overhead.” However, while this bare-bones set-up was sufficient in the company’s early years, it was never designed to meet the needs of a large organization. With Fundrise’s success came an exponential growth in users, which translated into similar growth in data volume and product complexity. As the small data team struggled to keep the system functioning under the strain of incoming information, they had to spend an excessive amount of time on maintenance and troubleshooting. It was challenging to innovate or improve, and soon, errors began to creep in. “Our data wasn't able to be correct all of the time because of the complexity we added,” explained Charles. “For a little while, we just tried patching things up, but at a certain point, it was clear that our set-up just wasn't viable.” ### Modernizing with a New Data Stack In dealing with a steady stream of minor errors, the team realized they needed to make changes to ensure consistent and reliable business metrics. Fundrise determined investing in new tools would be more efficient than hiring and training engineers unfamiliar with the company’s unique data needs. The team chose [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) to shape data sources in Snowflake, which would then be funneled into Tableau for visualization. As part of the tooling change, the team took advantage of the opportunity the migration offered to reorganize their data warehouse. “We made everything unified and organized,” explained Charles. “Doing so using dbt Cloud was much easier than if we had attempted a clean-up without its capabilities.” As they worked to reorganize the warehouse, the team quickly experienced the benefits of dbt Cloud’s [testing capabilities](https://docs.getdbt.com/docs/build/data-tests), [automatic documentation](https://docs.getdbt.com/docs/build/documentation), [lineage tracking](https://docs.getdbt.com/docs/explore/dbt-explorer-faqs#column-level-lineage), and [error alerting](https://docs.getdbt.com/docs/deploy/monitor-jobs)—all helpful in freeing up the data team’s time and reducing the amount of troubleshooting needed. The team’s newfound confidence in the data accuracy and lineage has rebuilt trust and collaboration with users across the business, who previously questioned the quality of the data they were provided. “On a day-to-day basis, we're now confident that we’re delivering more data than ever before, that it’s on time, and that it's correct,” said Charles. Jack Ploshnick, Fundrise Analytics Manager, continued: “We can do things much faster than we could before, and our stakeholders are confident that the numbers they’re working with are correct. Those two factors really resonate together.” ### Time to Focus on Strategic Analysis Now that Fundrise’s data team spends far less time maintaining pipelines, its analysts can now focus on high-value analysis to answer strategic questions. The business’ data velocity has increased, enabling it to increasingly transition from reactive to proactive work. “dbt allows us to create new business intelligence dashboards that never existed before,” explained Jack. He continued: “More importantly, it allows us to spend much less time creating those dashboards and more time on in-depth analysis. We have more time to answer the harder, impactful questions.” “That benefit can't be understated,” added Charles. “There was a huge opportunity cost associated with losing out on the highest-value work that we could be doing just to keep the system going. The amount of time we spend on maintenance is almost zero now.” The marketing team was the first to benefit from the new system and the data team’s new capacity, with a data tool that provided insights into the efficiency of their ads in terms of cost vs. revenue. While this was something the team was able to provide on its previous data stack, the new configuration was much more reliable and time-efficient. “In the past, we had to check in every day to make sure the marketing attribution numbers were correct and that all the data was flowing through as it was supposed to,” said Jack. “If the data source says that a particular campaign was efficient, we would need to go in and spend time making sure that number was actually right. “That is now zero work. We know the dashboard is correct. The question now becomes: Why is the marketing program efficient or why is it inefficient? We can finally focus on those kinds of questions.” ### The Power of Self-Serve Analytics One of the major advantages of the modern stack lies in empowering users to work without direct input from the data team at all. Since the introduction of dbt Cloud, the entire business has seen the ability for users to self-serve their data needs dramatically increase. More and more of Fundrise’s business stakeholders are building their own dashboards without the need for direct supervision or aid from the analytics team. “We expose the documentation to the stakeholders, and they're able to confidently do their own analysis,” said Jack. “They weren't able to before.” This move towards self-service analytics hasn’t just made Fundrise’s day-to-day operations smoother. Giving analysts access to simple, user-friendly tools has also made hiring and growing the team much easier. Jack explained: “We don't have a job title called analytics engineer at Fundrise, because it's so easy to work with dbt Cloud. If analysts have SQL skills, they can learn dbt. We no longer have to prioritize technical expertise over business savvy, and that’s a huge unlock.” ### A Scalable Foundation for the Future of Fundrise The introduction of dbt Cloud and a modern data stack has allowed Fundrise to scale up and navigate evolving business goals. For example, while Fundrise initially began as a platform dedicated to real estate investment, over time, the company diversified its offerings. As a result, they introduced a variety of alternative assets and other investment products to their portfolio. This expansion hasn't been without its challenges, particularly in terms of data management. The original data model was based on simple one-to-one relationships. However, with the inclusion of multiple new investment products, the data model needed to evolve to accommodate more complex one-to-many relationships. “Using dbt, we were able to adapt to those changes very quickly. Many of the stakeholders expected a lag between the product change being rolled out and being able to understand the data that was coming in,” explained Jack. “However, there really was no lag; we were able to adapt almost instantaneously because dbt makes it so easy to make changes to the data model and understand how something upstream affects something downstream.” ### What's Next: Expanding Usage and Governance No longer swamped by maintenance and troubleshooting, the Fundrise data team is planning for the future. An important part of this vision includes expanding the use of dbt Cloud across more of the company’s teams. “dbt is currently used by two analytics teams at Fundrise—the product & marketing team and the real estate investing team,” explained Jack. “Our next step is unifying the data landscape across all the different teams in the organization.” The data team also plans to expand the use of dbt Cloud and leverage its more advanced features for governance such as extended data lineage and access controls. --- --- title: "LendInvest speeds loan processing and reduces risk with dbt Cloud and Synq" description: "Learn how LendInvest uses dbt Cloud and Synq to proactively detect and resolve data issues before they impact customers." url: "https://www.getdbt.com/case-studies/lendinvest" date: "2023-10-09" industry: "Banking & Financial Services" --- # LendInvest speeds loan processing and reduces risk with dbt Cloud and Synq Learn how LendInvest uses dbt Cloud and Synq to proactively detect and resolve data issues before they impact customers. ### Company details - Headquarters: London, UK - Solution: Mortgage Lending - Data stack: AWS, dbt Cloud, Synq, Tableau, Metabase ### Results - 11x reduction in data runtime - 10s of issues detected in near real-time - 125% reduction in time to identify business-critical data issues > “We now catch issues proactively— the relevant teams outside of the data team are responsible for fixing breaking changes caused by data from systems they own. This frees up the data team and reduces the time to resolution.” > > — Rupert Arup, Data Team Lead ### Connecting Unconventional Mortgage Borrowers and Lenders Founded in London in 2008, [LendInvest](https://www.lendinvest.com/) is a fintech firm that acts as one of the largest non-bank mortgage lenders in the UK. It provides service lending by positioning itself between large institutional lenders and individual borrowers who—for a variety of reasons—may not meet mainstream lending criteria. "Our focus used to be on buy-to-let and specialist lending, but we've just entered the residential mortgages market. With that business expansion comes a lot more pressures, more regulated reporting, and the need for more testing," explained Rupert Arup, LendInvest's Data Team Lead. LendInvest handles the full lending journey—from underwriting through origination. Their platform is designed to equip underwriters with the data and customer insights needed to make informed decisions. This information is vital when clients are dealing with unconventional borrowers. "Our primary customers are brokers,” said Rupert “The market is intermediated in the UK, so customers don't directly choose who they're going to get their financing from." “Therefore, a lot of our work is around creating that relationship with those brokers. Being able to present them with clear, concise, accurate information is vital to providing a good customer experience." ### Legacy Systems Creating Major Data Risks Supplying the brokers with the data they needed, however, has not always been easy. Until a few years ago, LendInvest relied on a legacy MySQL data warehouse, Python, and SQL scripts to handle its data needs. While the setup had been sufficient when the business was smaller, the system had several systemic issues that grew as the years passed. Perhaps the most risky of these issues was that the legacy system struggled to handle the thousands of data models and dashboards LendInvest needed to meet the demands of its customer base. The team struggled with testing and troubleshooting, making it hard to surface and reconcile data mismatches. "We bring in data from three different systems—our origination platform, servicing platform, and internal loan engine,” said Rupert. “If we write a loan, we expect the same loan amount, interest rate, reversion rate, etc. on the offer letter to match up with all of the systems. "However, we were finding small errors where someone changed data in one place, and it hadn’t synced through to other systems, or they retrospectively changed the data after we had already produced reports or offers." For a business operating in the heavily regulated financial services sector, this kind of mismatch can be much more than a mere annoyance. Making errors on loan offers and failing to provide pristinely audited data pose serious issues; even relatively small mistakes can result in severe reputational and regulatory consequences. "If we end up submitting something incorrectly to the Financial Conduct Authority (FCA), we can get fined," Rupert explained. "If you get it wrong, and you find out later, there's a real risk that you incur a huge charge.” “If you keep making errors for a sustained period of time, you could be fined by the FCA or regulated out of existence." ### An Opaque Legacy System At the same time, the tooling barriers to easily surface information made [data lineage](https://www.getdbt.com/blog/what-is-data-lineage) tracking difficult, hampering the team's ability to reproduce and debug failures. "We were unable to recreate the past,” Rupert winced. “Whenever we wanted to run something end to end, we had to know the order that things ran in, but our legacy stack just couldn't do that reliably. Things became unscalable." Without clear data lineage and reproducibility, the team struggled to resolve bugs or inaccurate data buried in the pipeline. The lack of clarity of the legacy orchestration scripts made troubleshooting difficult and prevented innovation. The team spent most of its time "just making things run" rather than delivering new value. Processing speed degraded over time and the team was investing all their time into just meeting service level agreements (SLAs) for timely data delivery. Failure to meet SLAs, together with frequent complaints about stale or incorrect data from business users, provided the impetus to upgrade to a modern cloud-based data stack. ### Automating Data Reconciliation with dbt Cloud and Synq To address their data consistency challenges, LendInvest implemented a powerful solution leveraging dbt Cloud and [Synq](https://www.synq.io/), a data reliability platform. First, the data team built [dbt models](https://www.getdbt.com/product/dbt-cloud/) to extract and transform loan data from the organization’s three core systems before normalizing it into comparable formats. Custom dbt assertion tests then check for discrepancies across the systems, such as differences in loan amount, interest rate, or other key fields. These tests validate any new loans against the downstream data models in near real-time. When discrepancies are detected, Synq automatically routes alerts to the responsible teams based on severity. For example, a high-severity alert is sent immediately to a senior stakeholder who can escalate if needed. Lower-priority issues are routed to operational staff. This development process creates a scalable workflow and audit trail. Alert recipients can dig into the specific data differences in [Metabase](https://www.metabase.com/) and mark issues as acknowledged, fixed, or acceptable one-off differences. "We get these alerts on an hourly basis rather than a quarterly basis,” said Rupert. It means that we can alert on different levels. We have integrity alerts, source freshness alerts, exception alerting, and more recently we've been adding anomaly alerting." This automated, rapid reconciliation between core systems ensures data accuracy and prevents regulatory mishaps or negative customer experiences. ![processes](https://cdn.sanity.io/images/wl0ndo6t/main/2aa930276d7f5e73b2efd3ab3d153afaae747a2c-1600x1567.png) ### Rapid Debugging Frees Up Valuable Time Previously, the data team had to dig into the data, find issues, then track down who was accountable. This consumed hours of data team time. With Synq and dbt Cloud, debugging accelerated as the team can reuse and run SQL scripts against the database within 10 minutes rather than an hour to isolate the problem: "It speeds up everything by at least 100%," Rupert emphasized. By setting granular alert routing based on severity and groups of models failing, resolution times are also faster. This means the data team spends more time innovating versus wrangling tedious debugging. Whereas previously data was processed in 12-hour intervals, the team can now provide near real-time processing. This gives the entire business—and the brokers it works with—confidence the data is accurate. ### Next Steps: Expanding Tests to Cover More Use Cases Looking ahead, LendInvest has plans to continue expanding its use of [automated testing](https://docs.getdbt.com/docs/build/data-tests) and reconciliation powered by dbt Cloud and Synq. ![use cases](https://cdn.sanity.io/images/wl0ndo6t/main/f84b6ab0af8c5d60b08f2555c065dff136526240-1600x931.png) One priority is implementing more anomaly detection tests to uncover unusual spikes, dips, or other unexpected patterns in the data. This will provide another layer of monitoring to catch potential issues early. The team recently took the first steps by monitoring their 50 most important tables with Synq volume monitor checks that look at table row count. It’s important for the team to catch issues in near-real time so the checks run every 30 minutes. “Last week, we were notified of a sudden drop in the number of rows in our postcodes table we may otherwise have missed. This was caused by a change in the categorization that the Office for National Statistics uses, which would have gone unnoticed otherwise, potentially exposing us to missing postcodes in our database. As a property business, this is critical data, and Synq helped us detect this, rectify the problem, and mitigate any future risks," said Rupert. LendInvest also aims to involve more business analysts outside the core data team in the dbt development process. By giving analysts self-service access to curated dbt models, their domain expertise can be leveraged to build even more insightful tests tailored to their specific business needs. "Moving to a less centralized data team model will be key over the next six months. We’re looking to see how we can scale the insights we're getting out of our platform without scaling our team," explained Rupert. Ongoing investment in testing coverage will reduce reconciliation errors and improve data accuracy across the board. The ultimate goal is continuous validation so customers—whether borrowers, brokers, or auditors—always see pristine, trustworthy data. --- --- title: "Lundbeck centralizes marketing data with dbt Cloud to drive omnichannel strategy" description: "Discover how Lundbeck used dbt Cloud and Snowflake to unify marketing data and boost top-line impact." url: "https://www.getdbt.com/case-studies/lundbeck" date: "2023-10-09" industry: "Healthcare" --- # Lundbeck centralizes marketing data with dbt Cloud to drive omnichannel strategy Discover how Lundbeck used dbt Cloud and Snowflake to unify marketing data and boost top-line impact. ### Company details - Headquarters: Copenhagen, Denmark - Solution: Pharmaceuticals - Data stack: Fivetran, AWS, Snowflake, Snowpark, dbt Cloud, Qlik Sense, Github, Dagster, Hightouch ### Results - 7 omnichannel marketing campaigns launched - 5 data sources centralized in the data warehouse - 2-3% target top-line impact for brand, customers, and markets > “This is the first time we're collecting data from all these sources and uniting them to get a holistic view. Our data stack has enabled us to launch seven omnichannel campaigns so far, with more revenue-impacting campaigns to come.” > > — William Møller, Data Engineer ### A vertically integrated pharmaceutical company: from research to commercialization Founded in 1915 in Copenhagen, Denmark, [Lundbeck](https://www.lundbeck.com/us) is a pharmaceutical company specializing in neurological medicine. Employing 6,000 people, they develop and sell treatments for diseases such as depression, Alzheimer and Parkinson's. Unlike most pharma organizations, they own their entire value chain: from research, to trials, to production and, eventually, commercialization. Lundbeck’s value chain ownership gives them a unique opportunity to leverage data. The company’s wide-reaching visibility—from research trials results to sales conversations with medical professionals—can be harnessed as a competitive advantage for the commercialization of their products. “Since we have a lot of unique systems running, we have a lot of data with different formatting, from different sources,” said Daniel Thoren, Data Engineer at Lundbeck. ### The vision: enabling omnichannel marketing campaigns with data #### A personal approach to acquisition and upselling Lundbeck’s marketing team had a vision for how to harvest this data: adopting an omnichannel approach to acquisition and upselling. In this new approach, data from multiple marketing sources—such as marketing analytics and lifecycle campaigns—would be accessible and actionable by the marketing team. “For example, if a contact doesn’t open our newsletters, marketing could send a personal email or arrange a face-to-face meeting instead,” explained William Møller, Data Engineer at Lundbeck. Sales would also benefit from a comprehensive view of their leads and customers. “By enriching CRM contacts with data points such as website analytics, the team could understand what content interests a given audience and adjust their conversations with medical providers accordingly.” Marketing’s long-term vision was to use this holistic approach to further enable personalization. Sales reps would know at what time doctors liked to read content, on what topics, and when they are interested in ordering, among other information. #### The requirements for data But, to enable and support marketing’s vision, Lundbeck needed to take their data stack back to the drawing board: “The omnichannel project quickly became very data-focused,” said William. “It meant that we needed to combine data from many different sources and create a holistic 360 view of our audience—something our old data stack couldn’t really support.” ### Bringing Omnichannel Marketing to Life #### Better collaboration by moving from on-premises to the cloud Several teams were trying to crack omnichannel marketing; in the past, Lundbeck had outsourced related projects, but the team was confident that they could accomplish it in-house with the right tools. “It ignited a desire to liberate our data for digital initiatives from our on-premises systems,” said Lars Schöning, Tech Lead of Lundbeck’s Data Platform Team. The company leverages technology based on the SAP suite, with most of the maintenance outsourced. Although sufficient for corporate purposes, it was not able to support digital omnichannel initiatives. “We also had other needs that weren’t being met,” added Lars. “The system in place wasn’t accessible to the average developer at the organization.” #### A capable but still too-complex infrastructure Lundbeck’s data team settled on building a data lake on AWS. The new setup gave the team freedom to configure their infrastructure, but problems arose with both accessibility and observability. “There were more maintainability requirements with AWS than expected. We needed more skilled developers across the board,” said William. “We didn’t have a front-end that was accessible to non-developers. We also didn’t use a friendly language like SQL that would enable the team to work effectively and collaboratively.” ### Onboarding dbt Cloud, starting with Snowflake #### Sharper data requirements: collaboration, stability, and flexibility With the learnings from their AWS experience, Lundbeck landed on a third data infrastructure: Snowflake as the data warehouse, with SQL-first [dbt Cloud](https://www.getdbt.com/product/dbt) to transform and deploy the data. “We had a clear vision of what we needed,” said Lars. “A capable solution for all analytical use cases that’s also easy to understand. We also required a resilient and flexible architecture where we’d be able to swap out components later.” #### Centralizing marketing data on Snowflake with dbt Cloud and Fivetran Working alongside the marketing team, the data team got started on their Snowflake instance. With their new data stack, different marketing sources—such as Google Analytics, Salesforce Marketing Cloud, and Veeva CRM—are ingested into Snowflake via Fivetran. This raw data is transformed in dbt Cloud before being loaded back into the CRM with Hightouch. The new set-up allowed commercial leadership to access up-to-date marketing data via the team’s BI tool (Qlik Sense). And, it also enabled the marketing omnichannel vision. ![diagram](https://cdn.sanity.io/images/wl0ndo6t/main/914d2418ee336387bdb750f6b89e81ff0ec912f2-1384x722.png) ### The benefits of a centralized marketing data infrastructure #### Omnichannel campaigns Since Lundbeck debuted its new data infrastructure, they’ve launched seven omnichannel marketing campaigns leveraging their improved data access and workflow. “This is the first time we're collecting data from all these sources and united them to get a holistic view,” shared William. “It’s one of the first times we’re actually doing lead generation, with our data stack enriching our CRM with insights sales teams can use to close deals.” #### Faster turnaround times for sales representatives Lundbeck’s various marketing sources are loaded into Snowflake, and after performing data transformations in dbt Cloud, are then passed through Snowpark for advanced analytics and finally are made accessible to business stakeholders to leverage. [Snowpark](https://www.snowflake.com/en/data-cloud/snowpark/) is a Python development framework within Snowflake. It meets developers where they are and allows data engineers, data scientists, and data developers to code in a familiar way while executing data pipelines, ML algorithms, and data apps faster and more securely in a single platform inside Snowflake. “We’re now providing our sales representatives with actionable, current data,” shared Lars. “We have many use cases for dbt Cloud paired with Snowpark. One of our sales representatives received 300 ML powered suggestions across their accounts based on the data,” said William. “For example, it can suggest sales arrange meetings with doctors who ordered samples on our website.” This type of integration used to require multiple components, but now [Snowflake and Snowpark orchestrated through dbt Cloud](https://www.getdbt.com/data-platforms/snowflake), have created new opportunities for Lundbeck. #### Centralized data across APAC, EMEA, and LatAm During the data centralization process, the data team also tackled centralizing international sales data. Lundbeck sells their products across 100 countries, with nearly 50 sales representatives operating in total. “We combined APAC, EMEA, and LatAm data in one holistic view,” said William. “It’s now possible to do all these new things because we don’t need to create net-new data models for each market or combination of markets.” “Instead of having 50 different data science teams, we have one data team that has rolled out a solution that works everywhere,” added Lars. “This only became possible with our new data stack and workflow. dbt Cloud has helped us tremendously in harmonizing our global data.” #### Increased independence and speed for data stakeholders Today, all data used commercially by Lundbeck lives in Snowflake. While the data is centralized, data access is decentralized. “The centralization comes from being able to access all our data in one place, not from the centralization of ownership,” explained Lars, “Different teams have the responsibility of loading and processing this data to meet their needs.” The workflow, paired with the low barrier to entry for using dbt Cloud, has led to more autonomy for data stakeholders: “Our data bottleneck problem stemmed from having one central data team review data requests,” said William. “As we mature, we’ve leveraged our tools to give embedded data teams more and more autonomy so we can scale effectively.” --- --- title: "Code42 increased data team productivity by switching from dbt Core to dbt Cloud" description: "This is the story of how Code42's data team delivers revenue-impacting insights by switching from dbt Core to dbt Cloudto the whole organization" url: "https://www.getdbt.com/case-studies/code42" date: "2023-08-09" industry: "Industrial Automation" --- # Code42 increased data team productivity by switching from dbt Core to dbt Cloud This is the story of how Code42's data team delivers revenue-impacting insights by switching from dbt Core to dbt Cloudto the whole organization ### Company details - Headquarters: Minneapolis, MN - Solution: Cybersecurity software - Data stack: Snowflake, Fivetran, dbt Cloud, Tableau ### Results - 40 hours saved weekly on maintaining models - 10% decrease in warehouse computing costs - 15% saved in ingestion costs with dbt snapshots > “Before dbt Cloud, we were happy if we could restore something the same day it broke. Now, we can preview results before merging, and my team no longer needs to switch between Snowflake and VS Code. The time saving there alone is just amazing.” > > — Josh Carlson, Director of Analytics ### Helping companies secure their IP Founded in 2001, [Code42](https://www.code42.com/) helps employers secure their data while still harvesting a collaborative environment. The organization has developed technology that enables companies to identify leaks, source code, or IP theft by their workforce. Based in Minneapolis, they employ over 350 people. ### Data as the compass for decision-making Data has always been at the forefront of how Code42 makes decisions: “How do we make good decisions based on data? That’s what we’re looking for,” said Josh Carlson, Director of Analytics at Code42. “ We have a small team of six, so we’re critical about where we allocate time. Our job is to build data models that help the whole business.” ### Early analytics at Code42 #### Starting with Salesforce data on Tableau The analytics team at Code42 got started seven years ago with a sales use case. Their goal: understand sales performance and uncover new revenue-generating opportunities. “Our first request was to build sales dashboards, so our database consisted of just Salesforce data,” said Josh. “We started out with Tableau with a bunch of custom SQL.” Although the nimble set-up provided analytical value in the short term, it quickly hit issues. First, SQL code needed to be repeated multiple times, leading to errors and discrepancies on how metrics were calculated. “It got overwhelming quickly,” explained Josh. “It was really messy. We ended up with the classic ‘Why does this report say this number and this other report say another number?’ We’d have one different character in a query and it’d take hours to find.” Second, maintaining the Salesforce API integration was a complex activity. At one point, one of the six members of Code42’s data team worked exclusively on keeping the integration running. #### Moving to Snowflake with Fivetran The next step for Code42 was to upgrade their Postgres internal database to a cloud warehouse ([Snowflake](https://www.snowflake.com/)) and to purchase an ingestion tool ([Fivetran](https://www.fivetran.com/)). With the new tools, importing Salesforce data got a lot simpler, but the issues with querying and modeling the data weren’t resolved. “With less time spent on the API integration, we could develop more views. We’d create them in Snowflake and query in Tableau,” explained Josh. “But it was getting harder to chain these things together smartly.” ### Enter dbt: bringing governance and interconnectivity to SQL models The last piece brought into Code42’s data stack was [dbt](https://www.getdbt.com/product/dbt). The company started with [dbt Core](https://www.getdbt.com/product/dbt-core-vs-dbt-cloud) to interconnect the different Snowflake views. This enabled Code42’s data team to reuse and piece together different SQL models. “We could build one SQL query and easily reference it in another query. You didn’t have to repeat yourself,” emphasized Josh. ### Implementing a collaborative data development workflow with dbt Core #### Enabling ROI-driven, cross-team collaboration Code42’s data team supplies data and insights to the whole organization. With the number of net-new, custom models greatly reduced, they built a source of truth for the business. Different teams could look at the same metrics and collectively work to reach their revenue goals. “Our marketing team and sales teams now communicate and collaborate because they share common sales funnel metrics,” said Josh. “Before, we focused just on sales for dashboarding. But now that we’ve expanded to the full sales funnel, teams work together from lead generation to signed contract.” “Our data has also created a partnership between our product, operations, and finance teams to help manage costs and improve margins. That's some of our most impactful work since it's directly tied to profit." ### Outgrowing dbt Core #### Difficulties scaling data infrastructure without CI/CD The successful adoption of Code42’s data infrastructure dramatically increased the complexity of data maintenance. As the team built new queries on top of existing views, their models had more and more dependencies. “dbt enables you to have [dependencies](https://docs.getdbt.com/faqs/Models/create-dependencies), but running on Core, we lost sight of what breaks might occur if we push new code,” said Josh. #### A data engineering workflow—based on CI/CD—provides this oversight: “We didn’t have CI/CD, nor did we have the time and resources to burn to build out a CI/CD framework on dbt Core,” said Josh. #### Limitations of an on-premises system Additionally, data developers working in dbt Core on their local machines lacked visibility on how their code would affect existing data infrastructure. Publishing code to production was therefore error-prone and difficult to troubleshoot. “[Analysts](https://www.getdbt.com/product/analyst) would go to Snowflake to run exploratory queries. After they had a good data source, they’d head to their local developer environment and commit to our on-prem Git. Then they’d hope it didn’t break anything,” explained Josh. “If it did, we’d have to dive into AWS to see the logs and try to understand what happened.” This debugging process often took hours: “I was happy if we managed to restore something the same day it broke,” winced Josh. ### Deciding to move from dbt Core to dbt Cloud Code42’s initial investment in dbt Core proved a success. But dashboard and view errors were increasing, and restoring them took longer each week: “We kept experiencing quality issues because we'd make changes with unexpected downstream impacts,” said Josh. “Dashboard users were letting us know something was offline or incorrect before we even noticed. So we decided to move to [dbt Cloud](https://www.getdbt.com/product/dbt)”. Since dbt Core had already proved its value, the argument for investing in dbt Cloud was simple; with [higher quality data](https://www.getdbt.com/product/build-trust-in-data-and-data-teams) and [fewer maintenance costs](https://www.getdbt.com/product/cost-optimization), the data team would be more productive and deliver better revenue-impacting insights. ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/255cbc1268707cf95ea9cc460fdf5362a9f7dd7c-2346x1532.png) ### Time savings and higher quality data with dbt Cloud Leveraging dbt Cloud’s Github-integrated IDE and out-of-the-box [CI/CD](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud), Josh’s team establish a new, scalable process for data development. The outcome: increased productivity and uptime. “We got back 40 hours a week from time that was being spent maintaining models,” said Josh. “Now the team can provide higher quality dashboards, with way less effort—our dashboard uptime is now at 95%, whereas before the move to dbt Cloud, it hovered at 80%.” #### Code stability and repeatability By adding dbt Cloud to the stack, Code42’s team improved the consistency and stability of their data. Armed with observability tools—such as scheduling, logging and alerting—the team could spend less time on the maintenance of their data infrastructure. “I no longer needed to worry about ‘When was this table last refreshed?’” shared Josh. “I could go look in my dbt runs and see.” The issue of metric discrepancy from repeating the same SQL queries for different views was also resolved: “Before, I had to copy and paste the same code into whole new queries. Now I could use the same views for product, sales, and customer success—an immense time saver.” With these changes, the team could instead use their time to deliver more value: “By focusing less on maintenance, we decrease mental load and increase our mental capacity,” explained Josh. “The team can focus on what they love and we can all be more productive.” #### Insights for reducing churn and increasing expansion With increased time and capacity, Code42’s data team built more dashboards with a direct impact on the company’s bottom line. As a SaaS business, their “life and breath” lies in understanding user engagement in-product: “We built dashboards that identified patterns and recommended corresponding tactics to the sales and customer success teams,” said Josh. “For example, we predict customers likely to churn and those primed for expansion." “It took us multiple iterations and data products for us to build these sophisticated dashboards,” said Josh. “We need to schedule multiple runs from different data sources in a designated order, join the outputs, and define the metrics. That is not something that’d be possible without dbt’s data development framework.“ Today, this is the product usage data that Code42’s product management team relies on to shape their roadmap. ### Next for Code42: semantic layer and model contracts #### Standardizing metrics with the dbt Semantic Layer Post successful migration to dbt Cloud, Code42’s data team is now exploring more capabilities, like the [dbt Semantic Layer,](https://www.getdbt.com/product/semantic-layer) to help the data team predict what’s happening with leads: “We’re starting to leverage metrics in machine learning and it’s all starting to come together,” said Josh. “We want to predict the win probability of a certain outcome and, then, use[ standardized metrics](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) around marketing and sales to help inform the ML model." With this approach, the data team would deepen its direct impact on revenue. #### Leveraging model contracts for governed data dbt Cloud enables developers to define a set of upfront “guarantees” that define the shape of your model, called “[contracts](https://docs.getdbt.com/docs/collaborate/govern/model-contracts)”. This allows for further governance, by identifying whether a model’s transformation produces a dataset matching up with its contract. Code42 wants to explore this functionality to make it easier and safer to leverage shared models. “Our centralized team manages many models, each with different contexts. If we can define a contract once, I can free my future self from having to remember the logic the next time I use that model. I’m guaranteed to build a better data product with less effort via CI/CD and contract enforcement.” #### Advancing revenue use cases The combination of a semantic layer, model contracts, and [multi-project querying](https://www.getdbt.com/blog/analytics-engineering-next-step-forwards/) will enable the team to split their work into contexts—effectively defining data services so the team can deliver insights just-in-time for highly actionable, revenue-impacting metrics such as product usage and engagement data. “We believe we can improve our sales funnel by increasing data velocity,” explained Josh. “There are a number of sales dashboards that are not real-time today. We’ve begun to test how to decrease delivery time, and it’s already working really well.” --- --- title: "Inventa Builds Confidence in Data Through the dbt Semantic Layer" description: "This is the story of how Inventa uses dbt Cloud to boost accuracy and domain ownership of its data" url: "https://www.getdbt.com/case-studies/inventa" date: "2023-07-26" industry: "Industrial Automation" --- # Inventa Builds Confidence in Data Through the dbt Semantic Layer This is the story of how Inventa uses dbt Cloud to boost accuracy and domain ownership of its data ### Company details - Headquarters: Sao Paulo, Brazil - Solution: B2B marketplace - Data stack: Snowflake, Hex, dbt Cloud, Fivetran, Rudderstack, Eppo ### Results - 90% reduction in data maintenance time - 13 contributors  to data models, increased from 2 - 83 metrics centralized and implemented in the dbt Semantic Layer > “One thing I used to hear a lot was 'can I trust this data?' Now everyone knows: if it's there in the warehouse, you can trust it because it's been tested and centralized.” > > — Gabriel Marinho, Lead Analytics Engineer ### A booming business Based out of Sao Paulo, Brazil, [Inventa](https://www.inventa.com.br/) is one of South America’s fastest-growing prospects. The company functions as a B2B marketplace designed to quickly and efficiently connect countless small businesses with major suppliers, allowing them to stock up on everything from snack foods to cosmetics. Unlike businesses with similar portfolios operating in North America and Europe, Inventa offers a service tailored to the unique demands of the South American marketplace. “In Brazil, it's not easy for small businesses to open a credit line,” explained Gabriel Marinho, Lead Analytics Engineer at Inventa. “They also have low leverage with suppliers, so they pay huge shipping fees. They can’t get discounts, and oftentimes the minimum order quantities suppliers require far surpass what small businesses can—or need to—buy. Many businesses, therefore, don’t use traditional purchasing software. Instead, they work through messaging programs and place orders by directly communicating with sellers—a painful process of bouncing between WhatsApp messages and PDF catalogs. Inventa provides the large market of small businesses and suppliers with a safe, simple platform to complete these transactions. ### The need for good data Connecting a large number of clients and customers means Inventa must handle a lot of information. As the business grew in both reach and ambition, so did its need for reliable data. “For us, data is almost everything,” said Gabriel. “We want to offer credits to people without a bank history, and it's not easy to extend a credit line without data. So we have our own credit risk model." This model allows Inventa to evaluate credit applications and help customers improve their access to cash flow, without taking on unwanted risk.” On the other side, Inveta needs to provide suppliers that sell on their platform with visibility on how their store performs. “With in-depth reporting, we can help suppliers that are struggling or striving for better results,” shared Gabriel. “Our data assists them in identifying how to improve their products and increase sales.” ### An MVP system to start When Inventa began operations, the company used a Postgres database directly managed by Gabriel. “At first,” he explained, “we were mainly sourcing free or cheap ways to manage our data. So when I got to the company, my first job was to list the cheapest way we could build a data stack.” As the information volume Inventa handled started to increase, it quickly became apparent that this low-budget approach would start holding the company back and restrict its ability to grow. The data team spent too much time on maintenance and patching systems together rather than delivering value to the rest of the business. “I was duct-taping our platform together so it could cope with our growth,” Gabriel said. “This was a problem. Once we received investment funds, we decided we should start fresh—not only for the data team but for the company, too.” Gabriel and his team started evaluating solutions and soon encountered dbt. “We wanted to use SQL,” he said. “And that's a methodology [dbt](https://www.getdbt.com/product/dbt) leads the market for.” ### dbt Cloud: Investing in data the right way After experiencing the difficulties of working with an unstable software stack, Inventa’s data team sought to ensure that their next setup was as reliable and reputable as possible. “We didn't want to go to a new tool that was not mature enough and would potentially leave us needing to migrate again or change our workflow in the next few years,” Gabriel explained. “We didn't need to cut costs for data infrastructure anymore because we knew data would be the pillar of the business for the next few years of development. It was no longer perceived to be a cost center but rather an investment in the organization's future." Another of dbt’s major draws was the wealth of information and documentation associated with the solution. With a rapidly growing team, it was crucial to have easy access to resources that detailed best practices and helped the company upskill. “One of the reasons we went with dbt Cloud was so the rest of the team could try to build machine learning solutions, build better dashboards and BI systems for stakeholders, and create new models themselves,” he explained. ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/9c291843e7aa82ffebdcadbef7a7aab0f0ac9f2a-1600x1020.png) ### Benefits #### Reducing maintenance time and costs While building and managing their own data stack in Inventa’s early months had its benefits, keeping everything online and running as expected was time-consuming and frustrating. Once the company switched to dbt Cloud, Gabriel and his team freed up a lot of time—time that could be spent developing the business rather than fighting fires. “Before, maintenance was taking up 80% of my time,” he said. “Now, it's almost 0.” Gabriel noted that when something fails in dbt, “99% of the time, it's a problem with the source data rather than the system.” Before, it could have been a problem in the pipeline, an accessibility issue, or one of the dozens of other potential errors. “Finding out the cause,” he added, “used to be a lot of work. A lot of frustrating, time-consuming work.” With dbt, Inventa’s data pipelines are much easier to understand, assess, and maintain thanks to dbt Cloud's reliability and clear data lineage. “We now have a huge amount of documentation,” he explained. “With dbt, we can add Git pre-commits hooks and use [project evaluator](https://docs.getdbt.com/blog/align-with-dbt-project-evaluator) to enforce the CI/CD for documentation. “I'm setting the guidelines, but I’m also accountable for following them rather than pushing code without documenting it.” #### Easy testing Another significant advantage of moving to dbt Cloud was the ease of testing. Inventa had previously used DynamoDB, which came with its challenges. “We used to have a lot of bad data,” he explained. “With dbt Cloud, I can easily test the columns that I care about, and when we have an actual problem, we can quickly flag issues to the engineering team.” With Inventa’s previous data stack, the engineering team's first indication that something went wrong was when the production pipeline broke. Now, the team is alerted the moment a test fails, even if it's not breaking the data pipeline. “Things are way more robust now,” he said. “This helps the product team avoid any issues reaching customers and helps the data team identify any issues during ingestion.” #### Automated reporting to suppliers with the dbt Semantic Layer The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) enables teams to create standardized metrics that return the same consistent and precise data across tools. Inventa was already using this feature to power their business analytics—it [ensured different teams were looking at the same calculation of metrics](https://www.getdbt.com/product/build-trust-in-data-and-data-teams), such as supply revenue: “When you put everything on dbt, you ensure everyone is seeing the same number,” said Gabriel. “You don't get that message saying, ‘oh, my director got this GMV number and I'm getting this different one.’” Supplier analytics—which had shown in customer interviews to be crucial for overall supply and demand performance—was one of their dbt Semantic Layer powered data products. “We had a lot of MVPs that needed to be worked on and supplier analytics was no exception,” shared Daniel McAuley, Data Lead at Inventa. “It was a manual, error-prone process. We imported a .CSV from a Hex dashboard locally, turned those results into different PDFs, and then uploaded them to Google Drive to send to suppliers." “It would sometimes take an entire day of work for our business analysts to generate these reports,” winced Daniel. Once built out, dbt’s Semantic Layer powered all of Inventa’s internal dashboards; next, they saw an opportunity to also use the feature externally: “We realized the dbt Semantic Layer could take the same metric we used for business analytics and reuse it for supplier reporting,” shared Gabriel. Inventa used the dbt Semantic Layer to power automated Hex reports for suppliers. Each supplier had a dedicated dashboard, filtered based on their unique supplier ID. The data team could now ditch their former manual process that was taking a full day of work. “The initial goal was to get this day back. We succeeded, and the feedback from the suppliers has also been really positive,” said Gabriel. “We can now see which suppliers and accounts are looking at the data and have our account managers reach out proactively to help them succeed”. #### Distributing ownership through collaboration One of the less immediately obvious benefits of dbt Cloud was its impact on the broader data team. Inventa’s previous data stack had been too complex for most team members to handle easily, but the relative simplicity of dbt Cloud allowed for much more collaboration—and, with that, ownership. “Before, I was the only person creating models,” explained Gabriel. “We had five or six data scientists, but they didn't know how to build the pipeline, integrate with Airflow, or deploy a lambda function to AWS directly." “With dbt Cloud, however, they don't need to know any of that, and they can easily collaborate.” As the team began collaborating, members adopted best practices on how to model different data, what to discard, and how best to display their insights to end users. This granted the team ownership over their data, enhanced by their new ability to easily check [lineage and documentation](https://www.getdbt.com/product/dbt-catalog) with dbt. “I can now say that I spend only about 1% of my time on maintenance,” Gabriel said, “because I now have 12 people maintaining their own data.” ### Company-wide value Once the new data stack was in place, it didn’t take long for the benefits delivered by dbt to filter out from the data team to the rest of the company. The [seamless integration between dbt Cloud and Inventa’s Hex and Snowflake systems](https://docs.getdbt.com/docs/cloud-integrations/avail-sl-integrations) allowed the data team to meet stakeholders where they are. “The whole company benefits,” Gabriel explained. “A lot more people across the business are using the data because they can see more value. As a result, the whole company has started to do more analysis and be more data-oriented. “And for hesitating teams, we can say: ‘Look, we have this trusted data set that you can use to test your assumptions and hypotheses. Why don’t you try and use it?' And they can.” This new approach has helped Inventa to run more efficiently and reliably than ever. Where the data pipeline was once confusing and prone to errors, it’s now viewed as a trustworthy source of valuable insight. “One thing I used to hear a lot was “can I trust this data?” Gabriel said. “Now they know; if it's there in the warehouse, you can trust it because it's been tested and centralized." “That is incredibly valuable.” --- --- title: "Tempo builds a virtual personal trainer with dbt Cloud and Stemma" description: "Discover how Tempo uses dbt Cloud and Stemma to enhance user experience and e-commerce operations with collaborative data development." url: "https://www.getdbt.com/case-studies/tempo" date: "2023-06-30" industry: "Health & Fitness" --- # Tempo builds a virtual personal trainer with dbt Cloud and Stemma Discover how Tempo uses dbt Cloud and Stemma to enhance user experience and e-commerce operations with collaborative data development. ### Company details - Headquarters: San Francisco, CA - Solution: Smart home gym - Data stack: Fivetran, Snowflake, dbt Cloud, Stemma ### Results - 30-40% efficiency improvement from dbt Cloud and Stemma's workflow > “Growing from five to 300 people meant we needed more consistent data, version control, testing, discoverability—all the features that we could get with dbt Cloud and Stemma” > > — Eric Ganbat, Senior Data Engineer, Tempo [Tempo](https://tempo.fit/) is a rapidly growing smart home fitness company based in San Francisco, California. Offering an AI-powered home gym, the company’s technology uses 3D sensors and AI to track user motions and provide personalized form corrections and custom workout plans based on the data. Founded in 2015, Tempo’s aim is to democratize the personal trainer experience for customers so they can better enjoy their home fitness experience. Data underpins all of the company’s activities, from user behavior to supply chain metrics. ### Powering digital fitness, e-commerce, and operations with data This data becomes truly valuable when transformed into advanced insights that help power business functions. “We provide a lot of recommendations in real-time,” explained Chong Sun, Senior Director of Engineering at Tempo. “Let's say a user is working out—we can say: maybe you should increase your weight or adjust your form to be more productive and effective.” Data is similarly useful in analyzing Tempo’s supply chain, to help better understand how the business is running and optimize the e-commerce website. “In order for our primary marketing channel to be more effective, we need to understand where we should land users on our website and how we can optimize web pages,” Chong explained. “From there, we can also accurately understand how much money we're spending to attract each user.” ### Choosing tools for scale: testing, version control, and lineage With its former data set-up, however, Tempo struggled to marshal the data it needed to power its various functions. “With our previous solution, we had a big pain point where the data modeling was not well structured.,” said Chong. “There was no way for us to maintain it and produce high-quality data with a small team.” Replacing the old solution, then, was an opportunity to start fresh while adding important new functionality. “There were a few key components that we were previously lacking,” explained Eric Ganbat, Senior Analytics Engineer at Tempo. “One was a way of testing data to make sure everything coming in is accurate. Second, we previously did not have a good way of doing version control, reviewing each other's code, and establishing lineage.” With those capabilities in mind, Tempo searched for a new solution. While the team was set on using Snowflake and Fivetran in their data tech stack, “We needed something that could clean, test, and transform data in Snowflake,” Eric said. Tempo's decision to upgrade its data stack was also prompted by the company’s growth. “Our previous solution was not fit for us anymore,” Eric added. “We reached a state where we had multiple data engineers and analysts together as a team, and we knew it was time for us to think about how to organize, optimize, and ultimately build better data models. That’s when we discovered dbt Cloud.” ### Modernizing the data stack with dbt Cloud and Stemma: In June of 2022, Tempo decided to bring [dbt](https://www.getdbt.com/product/what-is-dbt) on board. The team already had Stemma in place as part of their goal to democratize data. “From my past experience at Uber, I knew how important a data catalog is to a data-driven culture,” said Chong. “We needed one place where we can easily identify which tables are servicing other tables. We also wanted to democratize data so everybody could access it, not just the data team, so the catalog was a central effort for us.” By having their data cataloged in Stemma the team was able to make a smooth upgrade to [dbt Cloud](https://www.getdbt.com/product/dbt-cloud). “Growing from five to 300 people meant we needed to rethink how we would maintain consistent data, version control, testing, discoverability—all the capabilities that we could scale up with dbt Cloud and Stemma,” said Eric. “Once we set up development and production environments, made sure dbt Cloud was talking to Snowflake, and that credentials were set up correctly, everything was smooth sailing.” Using a combination of internal training and consultants with additional expertise, Tempo quickly onboarded and ramped up their use of dbt and Stemma. “Stemma provided a huge value by helping us clearly understand what's going on with the current table transformations and set up our migration plan and priorities,” said Chong. “Without Stemma I'm not sure how we could have easily finished the migration.” ### Building an enterprise data stack: data unification, testing, and velocity With dbt Cloud and Stemma in place, the team at Tempo was able to realize a range of new benefits from their new data stack. #### Unifying data for improved insights “Using dbt, we are now transforming all our data models with proper [testing](https://docs.getdbt.com/docs/build/data-tests) and [documentation](https://docs.getdbt.com/docs/build/documentation),” said Eric. “We’ve covered a wide range of data sets from different products, generating events through user activity; additionally, we have our financial subscription, product usage data models, as well as marketing and supply chain data.” With such a large amount of data now at their disposal, Tempo can fully benefit from comprehensive insights drawn from different dashboards spanning product usage, customer profiles, web traffic, and inventory. “We needed a solution to streamline our data from all sources, and dbt Cloud has truly offered that,” said Eric. “It fits our needs in terms of data collection, understanding customer needs, and generating comprehensive reports and tables for internal stakeholders.” #### Catching issues with unit testing One of the most noticeable benefits of moving to dbt has been the ability to fix pre-existing issues with data, as Frank Wang, Data Engineer at Tempo, explained: “The biggest thing that I noticed moving to dbt is the range of unit testing functionality. A lot of our supply chain data is pretty fragile in terms of the underlying data quality. In the past, stakeholders would reach out to us and we would have to address issues manually.” With dbt features like regularly scheduled unit testing, the team now has the ability to simply and automatically react to issues before they reach customers or stakeholders. “The first week we launched dbt within a portion of the supply chain data, we caught some pretty glaring data quality issues that we could easily fix,” said Frank, “and we were able to resolve them before any stakeholder noticed on the downstream dashboards.” #### Velocity: shipping data faster Thanks to better quality and downstream reporting, Tempo’s data team now has the capacity to respond to the business’ demands faster and more effectively. “We're able to work much quicker to respond to ad hoc requests now,” said Frank. “That has been immediately obvious to our product analysts—we're able to get more done in the same amount of time.” Analysts at Tempo are now able to work smarter, rather than harder, significantly increasing productivity and decreasing delivery times. “Now we’re able to change just one table or make one small transformation, and our changes will flow downstream,” added Chong. “I’d say it’s improved our efficiency by about 30 or 40 percent for certain data engineering tasks.” ### Collaborative data development with dbt Cloud Much of the efficiency gain came from dbt Cloud’s built-in workflow which allowed the team to collaborate while leveraging their existing skillset. “None of our team members had extensive expertise in command lines. So it’s really helped us to work faster and elevate our team’s capabilities. dbt Cloud is able to quickly compile and run everything we need,” said Eric. dbt’s embedded software development best practices introduced code reviews and collaboration to the data team’s previously siloed workflow. “Eric and I have been able to review each other's code in a way that we just weren't doing beforehand. I went from not looking at anything to reviewing everything he touches,” explained Frank. “It’s gone from 0 to 100. Collaboration has been a huge win for us.” ### Modeling business-critical data With months of dbt experience now under their belt, the data team at Tempo has restructured the models providing new levels of insight into the workings of the business. “Our orders table is particularly important because it tells us what products we’ve sold,” said Eric. “That's one of the key financial metrics the finance and marketing teams use, which in turn provide our growth metrics.” Another foundational model combines all of Tempo’s product insights, collecting information such as who is doing what workouts. “The unified workout table is a big one,” added Frank. “Whenever we're doing any sort of product testing or feature development, that's the table that always gets referenced.” ### Managing the data development lifecycle with Stemma All of Tempo’s tables and dashboards are documented and observable through [Stemma](https://www.stemma.ai/). With its strong lineage capabilities, Stemma supports common operational use cases in data engineering. “Stemma has been a very powerful tool for us to quickly identify work that needs to be done to address new requests. If we need to change a data field, Stemma helps look upstream to find dependencies and trace back to the source so we can make requests of engineering to rev data quickly,” said Chong “Lineage tracking also plays a role in avoiding breaking changes and understanding the downstream impacts of changing columns and tables.” “I've been using a column lineage quite often because column naming wasn’t consistent when we migrated. But with Stemma I can easily see when one column is derived from another column with a different name,” said Eric. “It’s helping me go back and clean up names after the change.” Stemma has continued to help the team manage the development lifecycle of their data assets well after the migration. “Data table usage is another key feature of Stemma,” said Eric. “We had a case where we knew there were many tables not being frequently used. By looking at lineage and dashboard usage in Stemma, we were able to easily trim down about half of our tables that weren’t being frequently used.” ### Looking to the future: better models and real-time data After a six-month migration, the team is pleased with the possibilities of their new data infrastructure. “We’re very happy with our current data setup,” said Chong. “Overall, we’re confident we have the right tech stack with dbt at the center, combined with Stemma and Snowflake. It satisfies all our current needs.” Now that the data infrastructure is in place, Tempo’s Q2 plan is focused on optimizing and expanding use cases. With solid foundations in place, they are now looking to extend the benefits to improve data usage across other teams within the organization. #### Refining Tempo’s data architecture The first step towards making that a reality is ensuring they are doing as much as possible with the data Tempo collects. “The question now is: how can we leverage the current data stack to improve the overall architecture?” said Chong. “We are now able to build much better models. And by building better data models, we can bring more value to our products and to the business.” #### Boosting the customer experience through real-time data Looking further ahead, Tempo is interested in using the power afforded by dbt to improve its real-time data capabilities. “We’re looking into a number of real-time use cases,” explained Chong. “Once we get the data from our client side, we can rapidly move that information into a data warehouse, another third-party service, or downstream where real-time decisions can be made.” Tempo is particularly interested in unlocking additional value for its customers: “Tempo is not just a smart home gym. It is a data-driven platform that helps its customers reach their fitness potential, enhance their well-being, and have fun while working out. By improving the platform, we can improve the way we use our customer’s data to personalize their training, optimize their performance, motivate their progress, and connect them with others.” --- --- title: "Wellthy automates real-time data alerts with dbt Cloud, capturing $80K in revenue" description: "Learn how Wellthy used Snowflake, dbt Cloud, and Hightouch to power real-time alerts and notify teams of data issues." url: "https://www.getdbt.com/case-studies/wellthy" date: "2023-06-23" industry: "Healthcare" --- # Wellthy automates real-time data alerts with dbt Cloud, capturing $80K in revenue Learn how Wellthy used Snowflake, dbt Cloud, and Hightouch to power real-time alerts and notify teams of data issues. ### Company details - Headquarters: New York City, NY - Solution: Digital Healthcare - Data stack: Fivetran, Snowplow, Snowflake, dbt Cloud, Hightouch ### Results - 27 hours of manual data work saved - 2500 test implemented over 300 dbt models - $80,000 in revenue captured from two alerts > “I’ve been working in the data space for over ten years, and Snowflake, dbt Cloud, and Hightouch have solved 95% of the problems that I’ve run into. Together, these tools power everything we do at Wellthy.” > > — Kelly Nelson Pook, Sr. Analytics Engineer ### A data-driven care concierge platform [Wellthy](https://wellthy.com/) is a care concierge platform that leverages technology to provide personalized and comprehensive support to individuals and families navigating the complexities of caregiving. A dedicated care team works closely with caregivers to understand their unique needs, preferences, and goals. All of this is facilitated on the Wellthy platform with seamless communication and collaboration between the care team, the caregiver, and other healthcare providers involved in their loved one’s care. Wellthy covers nearly 2 million lives with comprehensive care infrastructure and benefits, and in 2023 was named one of Fast Company magazine's "Most Innovative Companies" and a "Top 10 Most Innovative Workplace" company. ### The challenge: tracking ever-changing member data Wellthy works with some of the best-known companies in the world and their employees, offering caregiving support to thousands of unique caregiving situations and millions of sponsored lives. The company creates leading health plans for their members so delivering a consistent and smooth experience is a constant operational challenge. For Wellthy, member eligibility changes every single day, so it’s vital that these changes are surfaced to the appropriate teams. Inaccurate data can completely derail the onboarding process, leading to fewer signups, conversions, and lost revenue. In order to combat this problem, the Care Team was spending hours hopping between systems trying to manually reconcile users with the necessary information, creating significant operational inefficiencies. With no centralized view of members and minimal resources to model, transform, and track member data, Wellthy needed to: - Build out a new data platform and establish a single source of truth - Implement a robust data modeling framework to rectify and enrich member data from many sources - Create automated alerting around the member onboarding process and experience To get to where they wanted to be, the company turned to Snowflake, dbt Cloud, and Hightouch to address each step simply with the least operational cost and fastest time-to-value. > “There’s nothing more stressful than caring for a loved one. We need to be able to provide personalized support when and where they need it. Being armed with this information ahead of time makes for a better member experience,” said Kelly Nelson Pook, Sr. Analytics Engineer. ### Wellthy's modern data stack ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/e584039abea4d7fee61caa82fb0114a7ad528c09-1534x651.jpg) #### Snowflake To solve for building a new source of truth, Wellthy implemented Snowflake. Before Snowflake, the data team was operating out of an application Postgres database. This was causing the team to run into both concurrency and scalability issues because production databases simply aren’t built to handle analytical workloads. Additionally, it was very difficult to pull in data from other sources. > “We’re trying to position ourselves to serve as many people as possible with the best caregiving assistance possible, and Snowflake powers everything we do from analytics, to operations, to activation," explained Kelly. Since adopting Snowflake, the data team has been able to easily and reliably pull data in from any source to aggregate and merge it into one clear, coherent member record. Wellthy no longer has to spend time manually scaling resources for various analytics jobs thanks to Snowflake’s fully managed platform and auto-scaling capabilities. #### dbt Cloud The second step was to build a robust data modeling framework without complex processes or manual builds. Enter [dbt Cloud](https://www.getdbt.com/product/dbt-cloud). To fully unlock the value of Snowflake and better serve the needs of the business, the data team needed a reliable and scalable way to automate, orchestrate and manage all of their data models within Snowflake. This quest quickly led the data team to implement and standardize dbt Cloud across the entire organization. Since adopting dbt, the data team has built over 325 models and 2,500 different automated tests. > “Previously, data teams would try to shove all of their transformation logic into one gigantic query. dbt Cloud lets us modularize the different stages of our transformation layers, standardize everything with version control, and easily use DAG visualizations to see how our models relate to one another," added Kelly. #### Hightouch Once Wellthy had their data centralized and modeled, Hightouch unlocked activation and alerting. With so many dbt models running in the background of Snowflake, Wellthy is constantly trying to identify possible exceptions to these models that are indicative of operational issues. However, there was added complexity because Wellthy needed to identify these exceptions without outright canceling or erroring existing dbt jobs that were running. The data team had built a ton of tests in the transformation layer, and stopping these dbt jobs would interrupt downstream data flows, and ultimately operational reporting. In search of a solution, Wellthy quickly landed on Hightouch to power two core use cases: - Exception-based alerting: User and account-level data is constantly changing for Wellthy, and delivering a good member experience means proactively identifying potential problems before they affect the member. The Care Team needs to know when an account or user is missing key information because this can halt the entire onboarding process. - Look-forward alerting: Wellthy works with Fortune 100 companies, so it’s important to notify internal teams of recent changes to client files. The Client Success team needs to know when company X increases (or decreases) the number of eligible users from 1,000 to 2,000, so that they can proactively confirm this new number, rule out upload errors, can allocate appropriate resources. ![Hightouch](https://cdn.sanity.io/images/wl0ndo6t/main/570398ac27a6fcea974f926db3cb7958b674e326-739x351.png) > “We have a ton of dbt jobs and tests running at all times, and we need to be able to identify potential problems without canceling or stopping our data flows. Hightouch lets us proactively send alerts to Slack so our Care Team can take action immediately.” Previously, identifying and rectifying these eligibility issues could take as much as six hours to pull data on individual users manually–and in many cases, key stakeholders were even forced to contact the client or member directly. With Hightouch, this manual process is eliminated, saving Wellthy over 27 hours per month and generating huge cost savings because key employees can now redirect their efforts to more important tasks. Wellthy has just scratched the surface with alerting, but the company has already been able to capture $80,000 in annual lost revenue from these two alerts. The data team is looking to implement this at a wider scale to proactively identify even more problems. ### Results Leveraging their new data stack, Wellthy: - Established a single source of truth using Snowflake - Captured an additional $80,000 in revenue from two alerts - Saved 27 hours of manual data work, leading to huge cost savings - Built over 300 models and 2,500 tests using dbt Cloud ### What’s next? As Wellthy continues to build out a modern data stack, the company is focused on activating the rich member data living in Snowflake. Wellthy is in the process of using Hightouch to move data out of Snowflake and enrich various downstream business applications like HubSpot and Salesforce to power hyper-personalization and ensure that every business team is working off of the same view of the member. > “I’ve been working in the data space for over ten years, and Snowflake, dbt Cloud, and Hightouch have solved 95% of the problems that I’ve run into. Together, these tools power everything we do at Wellthy.” --- --- title: "Vivian Health Connects Healthcare Professionals to their Dream Job with dbt Cloud and Metaplane" description: "Discover how Vivian Health uses dbt Cloud and Metaplane to deliver a seamless job search experience for healthcare professionals." url: "https://www.getdbt.com/case-studies/vivian-health" date: "2023-06-14" industry: "Healthcare" --- # Vivian Health Connects Healthcare Professionals to their Dream Job with dbt Cloud and Metaplane Discover how Vivian Health uses dbt Cloud and Metaplane to deliver a seamless job search experience for healthcare professionals. ### Company details - Headquarters: San Francisco, CA - Solution: Healthcare Staffing - Data stack: Snowflake, dbt Cloud, Metaplane, Census ### Results - 3x increase in data pipeline contributors - 1 OKR achieved, around data quality goals - 100% of data models created in dbt > “dbt is a critical part of our infrastructure, and Metaplane allows us to ensure that it is running smoothly around the clock." > > — Max Calehuff, Data Engineer At [Vivian Health](https://www.vivian.com/), data is everything—it’s what underpins their core platform that connects healthcare professionals with the right job opportunities. Max Calehuff, Data Engineer at Vivian Health, sheds light on the “transformative” initiatives undertaken by the company's data team, who handles both 1) the machine learning training for their product, and 2) data modeling for the company's internal dashboards. To meet the need for data, Max sits on a team of 10, including data engineers, analysts, and machine learning engineers. Between him and one additional data engineer, they tackle everything from: - Creating and maintaining pipelines - Data cleansing - Source integration research - Ensuring data integrity stays high - Translating business asks into data models ### Vivian’s Data Stack ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/ea17282edcb8f86421ecbecfd345f6b10bf4402a-1492x842.png) To serve Vivian's vast data needs, the team set up a best-in-class data stack. - **Snowflake:** the data warehouse solution at the core of their infrastructure - **Various data ingestion solutions:** Vivian combines Stitch, Airflow, Kinesis Firehose, and in-house builds, to funnel data into Snowflake. Upstream sources include multiple PostgreSQL databases that serve different workloads, various business applications, and events captured by Segment. - **Looker:** the business intelligence tool that empowers business teams with visualizations and reporting capabilities. - [**dbt Cloud**](https://www.getdbt.com/product/dbt-cloud): the transformation solution to enhance Vivian's data modeling processes. - **Census**: reverse ETL to pipe curated data back into business applications such as Salesforce or Amplitude. - **Metaplane**: data observability with integrations to Postgres, dbt, Snowflake, and Looker ### How Vivian uses dbt to expand modeling capabilities The Vivan team knew they needed a powerful and accessible transformation tool to align with the capabilities of Snowflake's Data Cloud and power their customer experience. Everything from analytics tables to machine learning models is transformed and modeled in dbt. For example, the training data for ML models used to show the most accurate recommendations to job seekers in-product. The data team models data from various upstream sources—their CRM, web traffic, and application events—in dbt, using [DRY code](https://docs.getdbt.com/terms/dry#:~:text=DRY%20code%20means%20you%20get,and%20energy%20on%20tedious%20syntax.) to create repeatable data models fit for the scale needed to ensure that their job recommendations continue to improve. Vivian’s data team also relies on dbt to model data for internal use cases including: - **Analytics:** the same web and application data models feed into internal product performance dashboards, as well as customer success and marketing dashboards. These business-focused visualizations are used to increase customer retention and usage rates, as well as inform organic search strategies to attract more customers. - **In-App Decision Making:** Beyond Looker dashboards, the team funnels their dbt models further downstream using Census to send modeled data feeds back into business applications themselves, making it even more convenient for internal stakeholders to draw insights in the tools they’re familiar with. Across use cases, the team benefits from dbt Cloud’s testing so data practitioners can write and test analytics code all in one place. “The ability to load sample data into dbt allows us to verify our models work as intended before deployment," said Max. #### A Transformative Solution Creating data models with both SQL and Python in dbt has allowed the data team's analysts to quickly self-serve the data they need without having to wait for the infrastructure-focused data engineers to assist them. In the past, Max has been part of teams that relied on performing transformations exclusively within business intelligence (BI) tools like Tableau and Looker. Max recalls that when using Persistent Derived Tables in Looker to hold transformed data, the approach not only required specialized LookML knowledge to set up, but also led to more fragile code, resulting in wasted time on constant maintenance. With dbt, the sheer amount of models and supported dashboards speaks for itself. By having one tool where the data team can write analytics code in SQL or Python, they’ve tripled the number of people on the team who are able to create models. Going forward, “we want to use the Python functionality more. It’s already given us results,” said Max. The increase in development velocity frees up time for Max and other engineers on the team to focus on more strategic initiatives, such as improving the machine learning models and writing advanced tests. “Everyone on the team knows how to create a [dbt model](https://docs.getdbt.com/category/models), and with only 2 data engineers that traditionally were responsible for modeling, that’s pretty cool. I can make a new dbt model as fast as it takes me to write a SQL query. It’s so much faster than how we did things in my previous roles. I could not go back, ever," emphasized Max. ### How Vivian uses Metaplane to improve data quality coverage With customer careers and by extension, healthcare patient experiences, at stake, Vivian’s data team knew that data quality needed to be prioritized. The search for a data observability tool came shortly after incorporating other components into their data stack, prompted by a series of data incidents, with downstream impacts that were difficult to identify and equally as difficult to fix. The usual solution would be to deploy unit tests on their pipeline, but there were several issues with this workflow: - **Time-consuming:** Setting up tests for a single pipeline could take anywhere from an hour to an entire day. - **Test accuracy:** Test thresholds needed to be updated over time as their data and expectations changed. - **Scale for new objects:** consistent product growth meant ingesting net new sources and creating new models, which weren't monitored. All of this led to the need to look for a more scalable solution to data quality monitoring. Luckily at that time, the team was already using dbt and started off by implementing dbt tests. They were able to capture incidents with causes such as stale data but wanted to continue their success with more test types and additional object coverage. #### Higher data quality, better patient care With the implementation of Metaplane, Vivian Health now has a comprehensive solution to monitor incidents across their entire data pipeline. Data quality is monitored from the upstream transactional PostgreSQL database to dbt-modeled tables in Snowflake, with the ability to see how incidents impact dashboards in Looker. “If I get a Metaplane alert, it’s always something that’s gone wrong, which is what we want. We want to ensure that we’re catching everything without over-alerting,” said Max. The practical application of Metaplane within Vivian Health's operations spans across a wide variety of use cases such as: - **SEO Analytics:** For web analytics, Metaplane plays a role in verifying the stream of events, facilitating proactive conversations where the data team can alert the web team to adjust their ingestion pipeline. - **Product development:** To support the Vivian Health platform, Metaplane helps ensure that healthcare professionals receive the most relevant job postings by monitoring training data for the recommendation models, in addition to other user experience improvements, such as models that indicate when a posting might be stale or inaccurate. After researching and evaluating other data observability solutions, Vivian Health found Metaplane to be the most user-friendly, cost-effective, and, most importantly, capable of identifying data quality issues. Their implementation improved their ability to monitor data integrity for all of their critical objects, resulting in improved data reliability and more efficient incident resolution, while saving time spent setting up and maintaining acceptable thresholds for data quality tests. Outside of the product, Max also added: “We've been using Metaplane for at least a year and a half...I've seen the many improvements that Metaplane's made. It's pretty cool to talk directly to developers and 90% of the time, when I bring up a feature, it's either being worked on, or, they'll just tell me: 'Oh we fixed that already. I'll turn it on for you now'." ### Using dbt with Metaplane Today, Vivian Health still uses dbt Cloud and Metaplane side by side. Not only do they continue to rely on dbt tests, but also use Metaplane’s monitors to track the outputs of those models, while deploying similar monitor types on their tables. As any good data engineer knows, it’s a good practice to set up comprehensive data quality solutions, particularly when data is of such high importance to the success of the company. ![dbt with Metaplane](https://cdn.sanity.io/images/wl0ndo6t/main/57baaa1c724d2736cf02f134e04b64c838f2d62e-1600x570.png) One use case where both tools shine together is in their relatively new use of Snowpark. The team wanted to use Snowpark’s ability to execute more complex Python transformations to generate training data weekly. In the process of testing Snowpark, Vivian set up a new compute warehouse and infrastructure efficiency monitoring to track warehouse usage. Rather than undergo the arduous process of calling Snowflake’s API and parsing the JSON into usable tables, they elected to use dbt to model the usage data, and placed Metaplane on top of that output table to alert whenever costs spiked. In addition to monitoring the outputs of models, Metaplane also monitors the job runtimes themselves, to alert the team to additional latency that might impact their downstream usage. dbt has allowed both analysts and engineers to [generate models ](https://www.getdbt.com/product/develop)faster than ever before, accelerating work directly impacting customer satisfaction and retention. And across all data pipelines, Metaplane provides full data quality coverage, saving weeks of effort in creating & updating custom tests. "We had an OKR last quarter related to data quality that Metaplane helped us achieve,” said Max. “dbt is a critical part of our infrastructure, and Metaplane allows us to ensure that it is running smoothly around the clock," Max concluded. --- --- title: "Vida Health uses Fivetran and dbt Cloud to provide personalized healthcare" description: "How a well-integrated modern data stack empowered Vida Health’s data team to deliver fast, reliable insights and operationalize data for care delivery" url: "https://www.getdbt.com/case-studies/vida-health" date: "2023-06-05" industry: "Healthcare" --- # Vida Health uses Fivetran and dbt Cloud to provide personalized healthcare How a well-integrated modern data stack empowered Vida Health’s data team to deliver fast, reliable insights and operationalize data for care delivery ### Company details - Headquarters: San Francisco, CA - Solution: Digital healthcare - Data stack: dbt Cloud, Fivetran Enterprise, Google Cloud, BigQuery, Looker ### Results - 6 months to implement and launch a new data stack - 80% of work to create new data products can be self-served by the analytics team - 1 day to trace the root cause of issues, down from two weeks > "With Fivetran and dbt Cloud, more teams are able to get the pipelines they need. We have more efficient collaboration between teams that previously did not exist. Our teams are able to share insights that serve the customer the best healthcare experience possible." > > — Trenton Huey, Senior Director of Data ### Data-driven, digital healthcare [Vida Health](https://www.vida.com/) is a digital health company that provides personalized virtual care, via employer-sponsored plans, to individuals with chronic conditions such as diabetes, obesity, and depression. To tailor support and resources to individual needs, Vida Health collects data on the customer's medical history, past insurance claims, lab test results, and log data from health-tech devices such as fitness trackers and digital scales. “With data, you get a really good health profile of the member before they even sign up,” said Trenton Huey, Senior Director of Data at Vida Health. “It helps us personalize their onboarding and identify the right programs for them.” The data enables Vida Health to monitor user progress, adjust treatment plans as needed, and give users visibility into results via the Vida Health mobile app. “The data is not just used for reporting and analytics, but also to directly guide the end-user experience in our application,” said Trenton. ### Moving away from inefficient custom pipelines Before moving to [Fivetran and dbt Cloud](https://www.getdbt.com/data-platforms/fivetran), the Vida Health team relied on a custom-built solution, using Python scripts and cron jobs to load and transform data in BigQuery. This approach presented several problems: - The solution was not scalable: The homegrown pipeline struggled and often failed when data volume spiked. - The pipeline was opaque and fragile: Components were poorly documented and understood by just a few people on the data team. When issues arose it sometimes took weeks to find and fix the root cause. These issues sometimes resulted in reporting downtime of 2-3 days. Data was not reliable or accessible to the teams that needed it to serve their customers best. “Non-engineering teams were struggling to get anything done,” shared Trenton. “They’d end up building their own individual pipelines locally. But we weren’t able to reuse that work or feed it into any other systems”. ### Starting afresh: The needs and targets for a new data stack #### Increasing collaboration in a newly unified data team Vida Health had recently consolidated its data engineering, data science, and data analytics functions into one team—seeking to break down silos and improve collaboration across these roles. “It was more than just the tech components. We were asking ourselves: 'How can all of these teams work better together?'” explained Trenton. With accessible user interfaces in [Fivetran and dbt Cloud](https://www.getdbt.com/data-platforms/fivetran), Vida Health's data analysts could take greater ownership of the data loading and transformation process and move their work forward more autonomously. They could consult with data engineers for advice, instead of leaning on them to create and maintain every change to the pipeline. “We have opened new lines of collaboration by using dbt and Fivetran,” said Trenton.” For example, research done by clinical researchers can now be reflected quicker and more granularly in the product.” And this new collaboration helped the team achieve its goal: they had less than six months to onboard more than ten new clients. This meant they needed a pipeline that could ingest, transform, and deploy eligibility, claims, and lab data—at scale. They knew they could tackle the challenge with Fivetran + dbt. ### Achieving data movement at scale with Fivetran The first step in onboarding these new clients was centralizing their data across SaaS and proprietary data sources into one source-of-truth warehouse. They needed a way to do this at scale, with built-in automation and security—especially when handling health data. That’s why they turned to Fivetran. “Fivetran’s reputation spoke for itself. We knew it could move our large—and growing—volumes of data from all of our sources securely. We didn’t even evaluate other solutions.” But, for Trenton, “The fact that claims data stored as flat files in GCS could be moved was the clincher. Everyone from data engineers to analysts was bought into Fivetran.” With fully-managed Fivetran connectors, Vida Health moves data from all their SaaS applications—like Salesforce, Stripe, Zendesk, and Google Analytics. This SaaS data, covering sales, marketing, and product, powers business reporting and helps drive the company’s growth. They also move proprietary Protected Health Information (PHI)—including laboratory data, past medical claims, and pharmacy orders, among others—from multiple sources, such as PostgreSQL, MySQL, and Google Cloud Storage. They can do so in different formats in the same BigQuery instance. With Fivetran’s Enterprise plan, they also get maximum security, ensuring this data is safe and HIPAA compliant. This secure, centralized data serves as a reliable foundation for Vida Health to provide personalized care to customers. Today, Vida Health uses Fivetran to centralize data from 148 different connectors into one BigQuery. “Fivetran helped modernize infrastructure and democratize who ingest data. Previously data centralization was a slow process because it required custom Python jobs. But now many teams set up pipelines and it’s easier to maintain. When we start new projects with new systems or data sources, we often say ‘it all starts with Fivetran.’” ### Accelerating product development with dbt Cloud Once data was loaded into BigQuery with Fivetran, the team used [dbt Cloud ](https://www.getdbt.com/product/dbt-cloud)to refactor old, unwieldy pipeline code into modular SQL queries that could be more easily read, reused, and maintained. “We could cut thousands of lines of code," said Trenton. “At my last company, we needed several Spark jobs to transform data. With dbt, we need far fewer resources to do the same work.” Prior to their migration to dbt Cloud, Vida Health's research and product teams had used separate pipelines to process data, often relying on local notebooks. This kept team knowledge siloed and sometimes hampered their ability to represent clinical data with appropriate nuance in their product. With its user-friendly IDE and auto-updating data lineage, dbt Cloud made the data transformation process accessible to Vida Health employees across clinical research, product, and data teams. It has enabled the teams to share knowledge, add more precision to their data insights and product—and move faster. "dbt has enabled things that were just not happening at all before. It's like zero to one on certain things," said Trenton. “Analysis on the treatment of a certain condition used to take months, even quarters. And after that was ready, we’d still have to wait months until those insights were fed into the product,” said Trenton. “Now we get access to the data faster with Fivetran and build the logic directly on dbt, so we’ve reduced that time from quarters to weeks or, sometimes, days.” ### A more resilient and maintainable pipeline with Fivetran and dbt Cloud ![A more resilient and maintainable pipeline with Fivetran and dbt Cloud](https://cdn.sanity.io/images/wl0ndo6t/main/7fb4ab465b1477d43c52158fe5cb1eaa33231198-1590x1028.png) Vida Health now uses built-in features from dbt Cloud and Fivetran to prevent issues that could create pipeline downtime and ensure quick recovery when incidents (rarely) occur. At the top of the pipeline, “having access to pipeline creation and monitoring (with alerting) all in one place with Fivetran is helpful. We can easily pinpoint issues with data ingestion as they come up,” said Trenton. "We also appreciate having fully managed connectors with schema migration and automation built in, for all data sources that power our technology.” This ensures data is always available. The team uses [dbt Cloud's data lineage](https://www.getdbt.com/product/dbt-explorer) to trace bugs, snapshots to detect data changes over time, and built-in CI support to standardize testing and improve data quality. "We have CI/CD testing on all pipelines—that didn't exist for many of our pipelines in the past," said Trenton. ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/4ff8181a08e6c1d456d9319d767bc8bf72c8b1af-1600x617.png) ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/5d7f8728035a7703f02147900b50ef882a1b34cb-1600x693.png) “Now we have much more precision on what we’re creating,” emphasized Trenton. “We can trace issues down. Before, we had issues where it took two weeks to find out what was wrong; now we have an answer within a day.” On top of less time troubleshooting, Vida's new data stack has also increased the team's velocity through significantly reduced maintenance requirements. With a modern data stack, Vida Health has the foundation needed to provide best-in-class care to its clients—driven by data. --- --- title: "McDonald’s Nordics establishes a nimble and standardized Data Vault structure with dbt Cloud" description: "This is the story of how dbt Cloud enabled McDonald’s Nordics to implement Data Vault 2.0" url: "https://www.getdbt.com/case-studies/mcdonalds-nordics" date: "2023-06-01" industry: "Food Services" --- # McDonald’s Nordics establishes a nimble and standardized Data Vault structure with dbt Cloud This is the story of how dbt Cloud enabled McDonald’s Nordics to implement Data Vault 2.0 ### Company details - Headquarters: Stockholm, Sweden - Solution: Food service retail chain operator - Data stack: Snowflake, Fivetran, dbt Cloud, Power BI ### Results - €1.3bn in sales tracked across restaurants, drive-throughs, and delivery - 4 markets with different data stacks centralized - 5x faster delivery times for historical data > “What we do with Data Vault wouldn’t be possible without dbt. We don’t need to think about the underlying technical stuff and can instead focus on modeling our business concepts.” > > — Cristian Ivanoff, Data Engineer ### Unifying data across McDonald’s in the Nordics McDonald’s Nordics operates in four markets (Sweden, Denmark, Norway, and Finland) with over 400 restaurants. When McDonald’s Nordics was purchased by the master franchise Food Folk, a new need arose: to report on all Nordic franchisees to both Food Folk and McDonald’s Global. However, each market had its own IT department and distinct data warehouse solution: “Denmark had a data warehouse alongside Tableau for reporting, Sweden was using an Oracle database with homemade applications, and so on. Everyone had their own solution,” said Cristian Ivanoff, Data Engineer at Food Folk. “We needed to create one platform, one data warehouse, for all countries.” #### Vast amounts of data—from sales to products McDonald’s collects granular sales data from all orders in their franchises, from in-person and online orders to orders from third-party sources such as Uber Eats. The data includes receipt-level detail, down to the amount of time required to serve each burger. They use this data to improve performance across the business—for example, by tracking and fine-tuning the amount of time to fulfill a drive-through order, starting from the customer's approach to the order window. In addition to order and transaction data, the company also collects: - financial data for accounting, reporting to the board, and optimizing each franchise location. - operational data such as product and menu information, which allows the company to account for regional differences (such as distinct product lines and ingredients by country) in analytics downstream. ### Creating a centralized data stack with Snowflake, Fivetran, and dbt Cloud on Data Vault 2.0 #### Assessing and selecting vendors Since each market had its own solution, it was up to the parent organization to define what the new centralized data stack would look like. Cristian and the data engineering team began by assessing the company's data use cases. “We started to look at our outputs and tracked backward to all our existing tools and vendors,” he said. After evaluating capabilities and making a short list, they selected Snowflake for storage and Fivetran for data loading. “And, because I believe in an SQL-first approach, I wanted to use [dbt](https://www.getdbt.com/product/dbt),” added Cristian. #### Why they chose dbt Cloud McDonald’s Nordics initially implemented [dbt Core](https://www.getdbt.com/product/dbt-core-vs-dbt-cloud), but upgraded to dbt Cloud to take advantage of its [built-in scheduler](https://www.getdbt.com/product/deploy) and [data lineage](https://www.getdbt.com/product/dbt-catalog). Besides the appeal of these dbt Cloud capabilities, they found it practical to keep data transformations decoupled from other parts of the stack. “We needed these features, but we also needed to be more independent from the other tools we picked for our new stack, such as our new data warehouse and data visualization tool,” said Cristian. “We don’t know if we’ll keep these products forever, so we wanted to separate our models with dbt Cloud.” #### Why they chose Data Vault 2.0 The data engineering team initially used a traditional Kimball structure to organize their data but ran into challenges soon after rollout. Every developer had a slightly different approach to formatting data or handling slowly-changing dimensions (SCDs). “The data started getting messy, with strange transformations,” said Cristian. The variability in their data development code led to different versions of the same dimensions. At that point, the team decided to explore Data Vault 2.0, a data modeling technique used to organize data from many source systems in a data warehouse. With Data Vault, data is separated into: - Hubs: business entities / keys - Satellites: descriptive context - Links: relationships between hubs ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/c11b423a5cdd393b03f2b4f41c9accf16d9f9eee-1093x339.png) This separation of concerns keeps data organized at the ingestion point, to make downstream data modeling easier. Like many teams using the Data Vault 2.0 approach, Cristian's team built data marts (dimensional models or star schemas) on top of their Data Vault models, to prepare the information for use in downstream applications and BI tools. “We needed to have rules to coordinate the slowly-changing dimensions, and to do it in such a way where we could still reuse all transformations easily. We required a standard and that’s what Data Vault gave us,” Cristian explained. ### Accelerating Data Vault Implementation in dbt Cloud To streamline their initial setup, the McDonald's Nordics team used [AutomateDV](https://hub.getdbt.com/datavault-uk/automate_dv/latest/) to generate pre-built, Data Vault-compliant models and macros in their warehouse. This sped up their implementation and allowed Cristian and his team to focus on refining business logic. “We didn’t need to do a lot of investigation or setting new rules. The AutomateDV package handles the technical parts, with macros and materialization macros, creating the SQL,” explained Cristian. “We don’t need to think too much about the technical stuff and can instead focus on modeling our business concepts.'' “What we do with Data Vault wouldn’t be possible without dbt,” he shared. “I’ve done Data Vault with Matillion and Dataform, but in both cases, it required significant amounts of work even just to get started. We had to create all the packages and macros ourselves from scratch.” ### Reaping the benefits of a new data structure #### Access to historical data even when not previously in-scope With the Data Vault structure, McDonald’s Nordics could now access all historical changes made to a dimension. When the platform team built the restaurants’ menus’ views, the original brief from stakeholders detailed they’d only need to report on the latest version of each menu. However, once the report was delivered, the requirements changed. Stakeholders requested they’d also like to access how menus had changed over time. By visualizing which items were added and removed, they could analyze how they affected operations and sales. “With the previous structure, we’d have to make many changes in the dimensions to get this data,” said Cristian. “But because we had all the historical changes and links in the Data Vault, we could quickly implement the changes. It was very satisfying how little time it took to deploy the new code.” #### Easier troubleshooting and higher trust from business users As well as delivering new data requests faster, there’s a secondary advantage to the Data Vault structure: data traceability. And, alongside [dbt’s data lineage](https://docs.getdbt.com/terms/data-lineage), McDonald’s Nordics can connect all the dots from raw data to reporting views. “I can see when I got the record, and I can trace the record back to the raw file or raw table created by Fivetran,” said Cristian. “With data lineage, I can also see which views were affected. This makes troubleshooting a lot easier and gives us a more robust data warehouse.” #### Power BI developers servicing business users Although traceability is not directly visible to business users, the robust data warehouse it creates leads to faster delivery times, better governance, and more trustworthy data. “People don’t care how they get the data, just that they have the correct, quality data,” said Cristian. “With dbt, we ensure this is the case.” Based on business logic and views delivered by the data platform team, their analytics team—including 5 Power BI developers—can create reports and dashboards. These developers have their own schema in Snowflake and can use it to access the dimensional and fact tables. Business users can go straight to the Power BI team whenever they have data questions or report requests. ### Moving forward: democratization and process Currently, Power BI developers aren’t exposed to the raw data that is in the “hubs,” “links,” or “satellites.” This leads [analysts](https://www.getdbt.com/product/analyst) to request guidance from the platform team. A new process, where BI developers can view and understand all the assets in the Data Vault, is coming soon. To increase the democratization of the new data stack, McDonald’s Nordics also has a project to train all business analysts and other business users on creating their own reports. By furthering who gets to access and model the data—while maintaining a governed standard—the team can further increase the velocity of data insights. --- --- title: "Zip unifies audience segmentation with Snowflake, Census, and dbt Cloud" description: "Discover how Zip modernized its data platform and built a scalable, fit-for-purpose modern data stack." url: "https://www.getdbt.com/case-studies/zip" date: "2023-03-31" industry: "Banking & Financial Services" --- # Zip unifies audience segmentation with Snowflake, Census, and dbt Cloud Discover how Zip modernized its data platform and built a scalable, fit-for-purpose modern data stack. ### Company details - Headquarters: Sydney, Australia - Data stack: Snowflake, dbt Cloud, Snowplow, Fivetran, Census ### Results - 1000+ models in production after 18 months > "We like to be Zip by name and Zip by nature. So we needed to update our old technology choices that were holding us back from moving as quickly as we wanted to." > > — Moss Pauly, Sr. Product Manager ### Challenge: Unifying customer data As a fast-growing company, unifying customer data was one of [Zip’](https://zip.co/)s biggest challenges. For example, Zip’s growth team was already segmenting audiences in Braze, their CRM tool, but they weren’t able to serve the same segmented offers in their own Zip app. Their product & marketing teams needed more reliable and self-service access to data to power the business. As a two-sided marketplace, Zip needed more data to not only personalize offers for their ‘Buy Now, Pay Later’ customers, but also to serve their merchants, who wanted to deliver cash-back offers to granular segments of customers. ### Solution: Building a Best-of-Breed Modern Data Stack Zip’s Sr. Product Manager Moss Pauly worked with the Data Engineering team to modernize their data platform and build a fit-for-purpose modern data stack. As Zip’s team was evaluating data solutions, their top priorities were seamless integration, cost scalability, and real business use cases. “Cost scalability is a key consideration for us. We’ve been burnt before with event volumes so we went into cost scalability with eyes wide open," said Moss Pauly, Sr. Product Manager at Zip. Zip built out a best-of-breed modern data stack with Snowflake, dbt, Snowplow, Fivetran, and Census. One of the biggest benefits for Zip was that each of these tools was the best-of-breed in its domain, yet they had tight integrations with the other components. With their new stack, the Zip team focused on: - **Self-service data access:** By implementing Census, Zip was able to supercharge the marketing team with full access to all their 360° customer data in Snowflake. - **Cost-scalability & performance:** As a company with an extremely high event volume, Zip needed scalable and performant, yet cost-effective tools for their stack. - **Single source of truth:** Ingesting first-party and third-party data into Snowflake provides a centralized repository that powers business operations. #### Snowflake Data Cloud As compute and storage are the core of the modern data stack, choosing a data warehouse was Zip’s most critical decision. They evaluated multiple solutions extensively and ultimately decided on the Snowflake Data Cloud. Over the past 18 months of using Snowflake, Zip’s data team has been very satisfied with its ease of use, performance, and seamless integration with dbt. Some of the questions Zip’s data team considered during their evaluation include: - If you wanted to run a quick query, what’s the time to result? (Opening the tooling, navigating it if required, waiting for a cluster spin up etc…) - What granularity of cost visibility can we have easily? - How well written is the documentation and how easy is it to find answers to questions? - What would the management impost be on the small team responsible for operating and maintaining the platform? - How easily can we manage PII redaction in this stack to protect our customers’ privacy? #### Data Transformation: dbt Cloud Zip needed to store business logic and transforms to [build data models](https://docs.getdbt.com/category/models) in a scalable, future-proof way. Dependency management and documentation were both significant pain points of their previous transformation stack. They chose dbt Cloud and haven’t looked back, with 1000+ models in production after 18 months. The cloud-based IDE has been a game changer, and they’re also diving deep into the power of macros and incremental models. #### Event Collection: Snowplow With millions of customers, Zip’s previous stack was unable to deal with its sheer volume of raw events. Snowplow appealed to the data team because it was open source, flexible, and didn’t have a SaaS cost tied to Events/Month. Zip’s data team was explicit that they did not want a solution where cost concerns would limit what they could track, and they wanted to retain first-party ownership of their events. #### Ingestion: Fivetran With their first-party event collection solution solved, Zip knew they needed a solution for third-party data ingestion. They didn’t want their engineers spending time wrangling third-party data APIs and wanted to capitalize on standard models in dbt for third-party data sets where possible. They evaluated a few options in this space, but Fivetran clearly came out ahead. They had coverage for all their third-party integrations, thorough documentation of data structures, and pre-packaged dbt transform availability. “Recently, our CIO wanted to query Twilio data and pinged me about the Twilio table structure while I was getting coffee. I was able to send back a link to the Twilio ERD in Fivetran about 5 seconds later that fully explained everything. Well-documented third-party integrations are really valuable,” explained Moss. #### Data Activation: Census Once they had the core elements of their stack built out, the data team realized they had several business use cases they couldn’t solve without connecting their data platform to their business tools. In particular, their product and marketing teams needed data in Braze, their customer engagement tool, to enable granular segmented offers. They evaluated a variety of options with Census coming out on top due to its performance, cost scalability of syncing high volume of data to Braze, and easy-to-use UI. After implementing Census, Zip was able to supercharge the marketing team with access to all their 360° customer data in Snowflake. Zip’s growth team uses [Census Segments](https://www.getcensus.com/segments), a visual audience builder, to segment customers then sync those audiences to all their marketing tools. “[Census’ Entity](https://www.getcensus.com/entities) models have been a game changer to enable non-technical users to create segments that would normally need complex joins across a large number of data sources,” said Moss. Census Entities combined with Census’ seamless integration with dbt and their visual audience builder has been a force multiplier to make unified, trusted data available to their business teams to take action. “This is the first time we've been able to have one unified audience that we sync from Census. This means that we're more confident in the messages and the offers we have live. We can actually get more granular with who we're targeting and what we're saying to them,” said Bianca El-Jalkh, Growth Product Manager of Shop & Rewards at Zip. ### Looking Ahead Zip’s modern, future-proof data stack helps empower marketing and data teams alike to do their best work. Moss views building this modern data platform as really “the entry ticket into being a modern data-led company”. They're currently underway with their next phase of activities: to further enable their data science teams, visualize model performance using Streamlit applications, and implement more monitoring and alerting for data SLA breaches. For more, read Moss Pauly’s blog on[ Building a fit-for-purpose modern data stack](https://medium.com/zip-technology/building-a-fit-for-purpose-modern-data-stack-3941d3562a67). --- --- title: "TIER Mobility manages 350,000+ e-scooters and e-bikes with dbt Cloud" description: "Discover how TIER uses dbt Cloud to support rapid growth and streamline data operations." url: "https://www.getdbt.com/case-studies/tier-mobility" date: "2023-03-23" industry: "Transportation & Logistics" --- # TIER Mobility manages 350,000+ e-scooters and e-bikes with dbt Cloud Discover how TIER uses dbt Cloud to support rapid growth and streamline data operations. ### Company details - Headquarters: Berlin, Germany - Solution: Micro-mobility - Data stack: AWS, Etleap, Snowflake, GitHub, dbt Cloud, Looker, Segment, Amplitude ### Results - 500x more rides in 4 years - 6 to 60 data team members in under 2 years, while increasing analysts' output - 1 hour to onboard new data team members to dbt > “dbt Cloud allowed our data platform team to strengthen our data infrastructure and our analysts to build scalable models on top—unlocking the velocity and insights we need to scale.” > > — Kumar Aman, Team Lead Data Engineer ### No data, no ride [TIER Mobility](https://www.tier.app/en/) is the world's leading shared micro-mobility provider, with a mission to c_hange mobility for good._ By providing people with a range of shared, light electric vehicles—from e-scooters to e-bikes—TIER helps cities reduce their dependence on cars. Founded in 2018, TIER currently operates in 560+ cities across 31 countries, including London, Paris, Berlin, San Francisco, and Dubai. The entire rental experience relies on fast, quality data. TIER needs to know how many of their vehicles, and at what time, should be allocated to each of their tens of thousands of pick-up locations. The data team is therefore tasked with two essential missions: 1) balancing supply/demand and 2) providing a positive user experience. Their operational data also has an additional use case: safety. Measuring vehicles’ placements enables TIER to stay compliant with municipal regulations and become the preferred micro-mobility partner for cities around the world. ### The challenge: exponential growth in operations and hiring Data has been at the core of TIER’s operations and business strategy since day 1. One of their first 10 hires was a Head of Business Intelligence. But, amid their growth stage, their data started to overwhelm the team. In a period of just four years, the company went from operating in a couple of cities to 560+. During that time, TIER completed more than 380 million rides with 350,000+ vehicles, and the data team grew from 6 to 60 people. As the data volume increased exponentially, so did the needs from the data team. The number of people who consumed TIER’s data products skyrocketed. On Mondays, internal dashboards had thousands of viewers. “The data team became a bottleneck,” said Kumar Aman, Team Lead Data Engineer at TIER Mobility. “We needed everyone to effectively contribute and develop.” In hypergrowth, the question became: how could TIER continue successfully leveraging its operational data while increasing efficiency and output from a fast-scaling team? ### Migrating to a modern data stack Thankfully, TIER optimized for growth and set up its data stack accordingly from the beginning. After a 1-month [dbt Core](https://www.getdbt.com/product/dbt-core-vs-dbt-cloud) trial, the team confirmed their intention to move to [dbt Cloud](https://www.getdbt.com/product/dbt). TIER’s leadership knew that provisioning an environment with built-in version control and software engineering best practices for the data team was the key to enabling velocity without exposing the business to risk. With dbt Cloud, TIER built the workflows and processes with 50 people that would support them now, at 5,000 people. Analysts with the appropriate business context could jump right into the data models they needed to go straight to business value: “dbt abstracts most of the underlying data, such as creating and dropping tables, defining data types, and version controlling code,” explained Kumar. “The data practitioner doesn’t need to know the intricacies of the underlying data to use a table. Our analytics team can go straight to building scalable models on top of the existing infrastructure.” ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/4085ff754f104cd8344213c3100b71132b9792dd-1600x880.png) dbt enabled TIER’s data team to build more with data—a good thing when it comes to delivering value but a challenge when scaling their platform with larger data volumes. “We started with Redshift, but two years later, Snowflake was a better fit for our scaling data needs,” said Kumar. With dbt handling the orchestration and compilation logic, the migration from Redshift to Snowflake was completed in six weeks—without downtime or stopping development. ### Leveraging fleet data to run and improve TIER #### Meeting supply and demand Fleet management is the backbone of TIER’s data team. They’re responsible for predicting when and how many vehicles should be allocated to each location in the 560+ cities they operate in. Without this prediction model, TIER would not be able to operate as efficiently—with users struggling to rent an e-scooter or e-bike during rush hour, and vehicles sitting unused in lower traffic areas. This data allows the operations teams to identify which areas need more e-scooters and e-bikes and when. Armed with this information, users are then incentivized to use idle vehicles and park them in areas with higher demand in exchange for free minutes. Maintaining and optimizing its fleet management strategy is essential for TIER’s profitability. Each improvement to its supply and demand prediction model has a direct impact on the bottom line: “Through demand prediction we operate in a much more efficient manner, unlocking profits of hundreds of thousands or even millions of dollars,” emphasized Kumar. #### Fleet data as a competitive advantage for tenders and compliance The same fleet data used for fleet management also serves safety and compliance purposes. “Fleet data backs our decisions on safety processes,” explained Kumar. “It informs us if users are parking their e-scooters and e-bikes in the designated, compliant spots.” This compliance and transportation dataset provides an additional benefit for TIER. With dbt, TIER’s city-focused analysts can build models on top of the existing data infrastructure to provide reliable data to cities—giving the company a competitive edge in tender processes. ### Coming up for TIER #### Decentralizing data dbt Cloud decreased the data team’s dependency on engineers, [empowering data analysts](https://www.getdbt.com/product/analyst) to take greater ownership and deliver more value. “With dbt, all members of our data team have the flexibility to create jobs on schedule. They can define the frequency of the data refresh and run their own models,“ said Kumar. Today, the data team at TIER is contributing to a mono repo within a single dbt project. But with the acquisition of Nextbike and Spin, TIER is planning on [implementing a multi repository multi project setup](https://discourse.getdbt.com/t/how-to-configure-your-dbt-repository-one-or-many/2121) to further their move toward a data mesh strategy to gain better visibility on costs by domain. #### Modern data stack expansion dbt Cloud was one of the first tools in TIER’s modern data stack. After their efficient migration from Redshift to Snowflake, they’re looking into expanding their data stack [around dbt](https://www.getdbt.com/product/integrations/). “For example, could we utilize a reverse ETL or a data catalog tool that communicates well with dbt?” asked Kumar. Reverse ETL could further improve TIER’s fleet management and profit margins; by synching product analytics to a CRM, marketing teams could send lifecycle communications, like push notifications, to users and incentivize certain behaviors. “dbt is the heart of our data stack. It all begins with centralizing and defining models and metrics in dbt. From there, our platform team can push the data to a variety of tools to facilitate their work.” #### Building python models TIER originally introduced Airflow to supplement dbt workflows; now that there’s inbuilt support for python models in dbt Cloud, it becomes easier to schedule everything within dbt. For use cases such as customer lifetime value estimation and demand/supply prediction models, data team members no longer need to go through the technicalities of Airflow and can manage everything in dbt Cloud. "TIER started out strong on the modern data stack, and we plan to stay at the forefront of these technologies to drive our competitive advantage in the market," concluded Kumar. --- --- title: "Deputy improves accuracy and speed with dbt Cloud" description: "This is the story of how Deputy used dbt Cloud to improve the accuracy and speed of their data products, building confidence and trust" url: "https://www.getdbt.com/case-studies/deputy" date: "2023-03-20" industry: "Industrial Automation" --- # Deputy improves accuracy and speed with dbt Cloud This is the story of how Deputy used dbt Cloud to improve the accuracy and speed of their data products, building confidence and trust ### Company details - Headquarters: Sydney, Australia - Solution: HR Software - Data stack: Snowflake, Stitch, Workato, GitHub, dbt Cloud, Tableau ### Results - +100 data team NPS — rated by internal stakeholders - 1 week to build new dashboards — down from months - 20% decrease — in computing costs > “The business now feels very comfortable asking us quirky ad hoc questions, and because we have a model in place, it's easy for us to either build it or answer that query. Productivity is skyrocketing.” > > — Huss Afzal, Data Director ### Data that informs internal and external stakeholders, as well as customers [Deputy](https://www.deputy.com/) makes it easy for employers and their staff to schedule and manage their shifts. Founded in 2008 in Sydney, it’s currently used by one in 10 Australian shift directors. Data is crucial for Deputy and it starts from the product itself. Customers use Deputy for insights on how employees are working and what they choose to do with their time. Deputy’s data is so comprehensive that during Covid, the big four consultancy firms reached out to access their raw workforce data for insights on how work patterns were changing with the pandemic. At the same time, data helps Deputy understand and improve their business. It gives them an overview of if they’re investing in the right place – from marketing to product to sales. ### Data becomes a bottleneck However, their large product usage dataset proved challenging for the data team to leverage internally. As Deputy grew, its data architecture became so convoluted that the team could no longer efficiently use it. It was a complicated setup where a series of Snowflake tasks would move raw data into an analytics layer. The team tried to resolve the complexity by using Airflow instead, but it didn’t make data access or modeling any easier. The data team still struggled to access and model the data they needed. #### Messy data + no documentation = no deliverables This complicated architecture was further harmed by a lack of documentation. With messy data and no clarity on how it worked, analysts couldn’t perform their jobs. Dashboards at Deputy would break, but the team couldn’t identify why. Debugging was extremely difficult. It was customary for a broken table to take half a day of work to be fixed. Since there was no documentation, analysts would have to manually trace everything back to the source in order to understand how tables were built. With such a complex, time-consuming workflow, the data team struggled to fulfill "simple" requests. “When I arrived at Deputy, I was all excited,” said Sara Young, Senior Manager of Data Analytics. “Stakeholders would make requests that didn’t seem complicated, but it would take me so long. I couldn't explain properly to the stakeholders that the data was such a mess. It was making me and the data team look bad.” #### Trust in the data team hit rock bottom All Deputy data team’s stakeholders—finance, customer success, marketing, product, sales—were unhappy with the team. The company’s low confidence in their abilities reached the point that when stakeholders needed data, they would circumvent the team. They neither trusted the data they’d deliver nor their ability to complete it in a timely manner. “All leadership had nothing but horror stories about working with the data team,” admitted Huss Afzal, Data Director. “The VP of Customer Success shared that she asked for a dashboard that the data team said they couldn’t deliver for nine months, because that’s how long it would take them to build it.” “Stakeholders knew the data team was not a good team. They weren’t surprised to be disappointed,” shared Sara. #### High employee churn The challenges the team was experiencing also harmed the team’s morale. They knew they had the skills to deliver the requested data products but, blocked by their tech stack, they faced their stakeholder’s disappointment every day. They were stressed, and team members were quitting because of the pressure. “The team got to a breaking point”, said Sara. “A lot of people started leaving. First it was the director of data left. And then just more things started breaking. Then we lost one of the lead data engineers and he left with little documentation behind… things just turned off.” ### Starting over from scratch, with dbt Cloud When Huss joined the team, he worried even more people were going to leave the data team. He suggested re-building Deputy’s data infrastructure from scratch to remove the technological blockers and allow the team to do the work they knew they could do for stakeholders, well. Both a new data engineer hire and Altis—the data consultancy Deputy worked with—suggested implementing [dbt](https://www.getdbt.com/product/dbt) at the center of their new data infrastructure. After ingestion via Stitch and Workato, all different sources—payment, CRM, Segment, Google Analytics—now go into a raw layer in Snowflake. dbt sits on top for transformation. After the models are created in dbt, they are fed into Tableau and exposed as reports. “Learning dbt was a small learning curve,” said Sara. “[Analysts](https://www.getdbt.com/product/analyst) proficient in SQL can pick it up once they see how it works. Once you get your head around, the rest falls into place.” ### Building confidence in the data #### One source of truth dbt unified Deputy’s data and created a single source of truth within their new data stack. Instead of different databases, different views, and different logic, they have [one model](https://www.getdbt.com/product/develop). “You have one fact table and everyone is feeding off that table using the same metrics and getting consistent numbers,” said Andy Kwier, Lead Data Engineer. “Having that single source of truth grew confidence in the data. Today, if I run a report and you run a report, we get the same metric.” #### Using data freshness to identify issues earlier Another step in building trust in the data was to spot issues before business users did. With [freshness testing](https://docs.getdbt.com/reference/resource-properties/freshness), whenever a field or table stops updating, the data team receives a notification. “Freshness testing has helped a great deal to pick up on things straight away and get them fixed,” said Andy. “We can now proactively flag and solve issues before anyone notices. This happened the other day with a table that was failing because of a null value that wasn’t allowed in the field.” #### Greater visibility: reduced costs and easier debugging dbt’s documentation features—such as [data lineage](https://www.getdbt.com/product/dbt-explorer)—finally enabled the Deputy data to see what was running under the hood of their reporting. “Before, we had over 400 tasks running: streams and pipes going everywhere. We didn't even know which ones were still being used for dashboards or reporting,” shared Andy. “We saved 20% computing costs on Snowflake by cutting what we didn’t need.” Their newfound ability to report back to their source tables and corresponding data sources also made debugging easier than ever. “Since we migrated to dbt, I haven’t had a problem that's taken more than a couple of hours to fix. Before, it’d take a minimum of half a day,” said Andy. “We had no lineage so you’d need to look into multiple pipes and streams. Now I can see what’s downstream or upstream.” ### Faster turnaround times for new dashboards and metrics The improvement in governance paired with a clear overview of their data lineage has enabled the Deputy data team to work better with other teams. “The pace of delivery has been phenomenal. Analysts are collaborating directly with the business,” said Huss. “When you can answer questions almost immediately, then you’re invited by business stakeholders to join meetings and provide valuable insights on the spot.” The deputy team has launched 26 data products in the last 10 months, most of them in the last 3-4 months after they got dbt up and running. “On average, I’d say dashboards that’d take months now take a week,” smiled Andy. Following the implementation of dbt, the [trust in the data team](https://www.getdbt.com/product/build-trust-in-data-and-data-teams) has done a full 180. In just a few months, the team has gone from “the usual disappointment” to a total delight. Last quarter, they polled their internal partners (CS, marketing, product, finance) on how they’d rate the data team’s work; they received a perfect NPS score of 100. “The business feels very comfortable asking us quirky ad hoc questions, and because we have a model in place, it's easy for us to either build it or answer that query,” shared Huss. “Productivity is skyrocketing.” ### Moving forward: scaling dbt and creating new revenue opportunities from data #### Data science-led predictive metrics With the data team moving faster than ever, Deputy is now looking toward predictive analytics. For example, calculating and predicting the likelihood to churn. “Outside of that we also want to introduce some data science to some of our metrics,” said Huss. “dbt’s integration with Python will come in handy there.” For example, by adding data science to churn prediction and net user expansion metrics, Deputy will be able to help their customer success team focus on the customers that have a higher likelihood of expanding or churning so that they can better serve them and their needs. #### Expanding into other dbt features Now fully onboarded to dbt, Deputy is exploring more dbt features, such as the [semantic layer ](https://www.getdbt.com/product/semantic-layer)and the [metadata API](https://docs.getdbt.com/docs/dbt-cloud-apis/discovery-api). Both features will continue Deputy’s journey into building trust and improving governance. “With the metadata API, we’ll be able to understand which sources and which tests generally fail,” said Andy. “This will give us more visibility into where our problems are.” #### Data as a service During the pandemic, Deputy saw the value of their aggregated data in spotting trends and telling a story. They’re doubling down on this, by creating a revenue stream for the internal data they capture. “This is something we’re building next year. We’re super excited about it,” said Huss. “And dbt will play a big role in modeling that data and serving our new product.” --- --- title: "SpotOn reduces time to actionable insights by 6x" description: "This is the story of how SpotOn built a scalable data platform with Snowflake, dbt Cloud, and Metaplane" url: "https://www.getdbt.com/case-studies/spoton" date: "2023-02-17" industry: "Industrial Automation" --- # SpotOn reduces time to actionable insights by 6x This is the story of how SpotOn built a scalable data platform with Snowflake, dbt Cloud, and Metaplane ### Company details - Headquarters: San Francisco, CA - Solution: POS, Payment Software - Data stack: Meltano, Fivetran, Heap, Snowflake, dbt Cloud, Tableau, Metabase, Metaplane ### Results - 600% decrease — in time to actionable data - 8x increase — in engineering contribution - $110,500 saved — in annual engineering costs > "We’ve grown incredibly in the last 2.5 years, and there’s no way that growth would have been possible without bringing in Snowflake, dbt, Metaplane, and the modern data stack" > > — Ben Cohen, Data Engineering Lead [SpotOn](https://www.spoton.com/) is a rapidly growing business that offers mobile payment processing and management software for restaurants and small businesses. As the company has scaled, data has increasingly become a differentiator to drive the business forward. More than 500 team members rely on data on a daily basis to make decisions and SpotOn’s customers rely on data for merchant reporting and a recommendation engine to power better online ordering experiences. With this widespread integration of data across the business came new challenges for the data team. Ben Cohen, the Data Engineering Team Lead at SpotOn, and his team were running into bottlenecks in the performance, accessibility, and engineering workflows for iterating on data. The performance of their Postgres database was routinely slow, causing delayed ETL jobs and degraded BI reporting experiences. The database was undersized and tuning was difficult; “ingestion jobs would fail because of whatever the new Postgres error was,” explained Ben. Data was either missing or delayed. These performance issues had a ripple effect—data was not accessible because it was often slow or broken. Users were unable to run queries because resources were being consumed by upstream ETL jobs for several hours every morning. New and advanced analytics use cases were impossible to create on top of the existing warehouse because Postgres couldn't handle reprocessing large-scale data aggregations. When the data team needed to address incoming requests or improve models, only a small subset of the team had the skills necessary to deploy changes in a timely manner. Ben and his team couldn’t keep up with the number of data requests, and they wanted to start using data for new use cases that could unlock more growth for the entire company. The data stack became a bottleneck, and the team needed to move quicker as the company scaled. With these challenges in mind, Ben decided to implement Snowflake, [dbt Cloud](https://www.getdbt.com/product/dbt-cloud), and Metaplane to scale the analytics capabilities and create a new team culture, all without adding undo complexity or cost. ### Solution: Scaling analytics capabilities with Snowflake, dbt Cloud, and Metaplane ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/12cd571a52bb1c3b6aa05b079f863666bce14ff9-1336x734.png) #### How Snowflake improved performance, made data more accessible, and improved engineering efficiency When Ben and his team migrated from Postgres to Snowflake, new advanced analytics use cases were immediately possible. For example, the recommendation engine for online ordering platforms is powered by large-scale pre-aggregated data augmented from multiple sources. Whereas Postgres could not support these types of aggregations, Snowflake handled them with ease thanks to the scaling capabilities provided by the separation of storage and compute. “With the scale of data and infrastructure we had, we couldn’t even begin to solve these problems using Postgres. We knew we had to upgrade to a data cloud like Snowflake” Other work that used to take up much of the team’s time, like tuning Postgres instance sizes, re-indexing data, or designing physical table structures, simply went away with the power of the Snowflake data cloud. Snowflake’s easy-to-use integrations with tools like Snowpipe made ETL’ing and operationalizing data fast and simple. It wasn’t necessary to build out complex pipelines that needed to be maintained and scaled. For example, the SpotOn data team used Snowpipe to ingest a weather data set from OpenWeather to augment order data. They were also able to seamlessly integrate with their product analytics platform, Heap, and use data sharing to repurpose data without needing to build new pipelines. “There are pretty expansive, geospatial cross-joins that we need to do on a regular basis that Postgres would never be able to handle,” explained Ben. Snowflake also became a key part of data integration after SpotOn acquired Appetize—an existing Snowflake customer. The data team could securely and instantaneously share data between SpotOn and Appetizes’ Snowflake instances. Migrating from Postgres to Snowflake was a game-changer for not only the data team, but the entire company. The performance improvements led to faster reporting load times and more advanced analytics use cases. Next, Ben turned to building on top of their new warehouse with dbt Cloud to make this data even more accessible and improve engineering workflows. #### How dbt Cloud made data accessible and improved engineering contribution by 8.5x When Ben joined SpotOn, there were only two engineers who could consistently contribute to their ETL project. Adding data sources took months, data models were built in siloes, and the testing and deployment process was painful. Their engineering workflows were centered around custom Python jobs in Airflow which required advanced engineering skills. If data wasn’t already in Postgres, the team needed to update Python scripts to ingest this new data. The Airflow instance required testing and deployment, and if historical data was needed, the team had to backfill Postgres—which could take days or a week depending on the size of the dataset. Only then would they start the modeling process…which meant another Airflow deployment. After all of this, there was no guarantee that the data was modeled in the way the team needed. If a model was delivered, and then an additional piece of data was needed, the entire process would start over again. The cumbersome workflow limited the team’s velocity and capacity. The only way to scale was to hire more people, which was not a reasonable path forward. “Before dbt, only two engineers could contribute to modeling. Bringing in dbt Cloud changed all of that for us. The number of contributors has grown by 750% and there is no limit to scale.” After Ben and his team adopted Snowflake to improve performance and scalability, they migrated all existing models to dbt Cloud—financial, operation, sales, and product analytic data were all transformed using dbt. With their core business logic in dbt, the data team’s workflow became much easier due to greater collaboration and accessibility. “dbt Cloud enables more people to build models, self-serve, and feel empowered. That translates to people being more engaged all around.” After SpotOn implemented dbt Cloud, creating data models went from taking days to hours. dbt quickly became the workhorse that powered the entire SpotOn warehouse, internal BI, and analytics. Ben and his team also operationalized the data to be used in the SpotOn product. By using dbt to pre-aggregate data, the product teams can power merchant-facing reporting as well as a recommendation engine. For example, they can augment ordering data with weather data to make better recommendations in the online ordering platform. With dbt Cloud’s analytics workflow in place, SpotOn went from two to 17 contributors to their ETL code. Product teams, engineers, and finance all contribute to their dbt project, including documentation. “We use existing [dbt packages](https://docs.getdbt.com/docs/build/packages) to push all of our definitions directly to Snowflake and Metabase in an effort to make our documentation available where our internal stakeholders work everyday. We even have some people on the finance team contributing to YAML files and building out documentation.” Not only are more people contributing, but dbt Cloud bakes in software engineering best practices, challenging the team to learn new skills and implement testing and version control as a default. “An analyst can work with our data pipeline and feel engaged and empowered. Stakeholders are happier and ask for help in creating new data use cases. It’s a virtuous cycle of improvement.” #### How Metaplane increased trust in data and reduced time to identify issues from days to seconds Ben and his team quickly noticed that as they improved performance and made the data more accessible, they created a data feedback loop: more capabilities allowed the organization to move more quickly, which resulted in more teammates asking for data and analytics. On the one hand, this feedback loop was driving SpotOn’s entire organization to leverage data and make more informed decisions. But with this came more attention and scrutiny of the data team’s work, and trust in data became top-of-mind for the data team. “As stakeholders use more data and have new capabilities, they ask more from your team, and you need to move quickly. You can’t test everything yourself. The other way to grow is to hire more people that add data quality checks, but that doesn’t scale well from a cost and efficiency standpoint.” Metaplane helped the SpotOn data team scale this feedback loop. By providing observability across their data stack, the team was able to build and retain trust; with Metaplane’s machine-learning-based testing approach and ability to automatically add hundreds of tests, they saved engineering time and always received context about potential root causes and downstream impact when data incidents arose. For example, the data team ingests transactional data from payment processor partners. This data is critical in helping the company determine and report on key KPIs. The source systems can be legacy databases and don’t always deliver data on a regular schedule. At times, the data was incomplete or contained values outside the expected data contract. All of this fed data into reports that executives used every morning. By proactively catching data incidents like these, Metaplane helped Ben’s team get in front of any issues that would impact downstream stakeholders, helping the data team maintain trust in the data. After receiving data incident alerts, Ben and his team could pull back scheduled reports until they verified that the data was fixed after an issue. “We were always behind the 8 ball in terms of communicating with the executive team when there was an issue. We were starting to lose trust and they weren’t going to use the reports. If they can’t use our data, that’s bad not only for our team but also for our business. Metaplane helped us get in front of those issues.” Data quality issues don’t just impact executive reporting. SpotOn is core to their customers’ businesses because they process all of their transactions. When SpotOn experiences data quality issues, it affects their customers and potentially impacts how they earn money and serve their own customers. Being the first to know about data quality issues allows the data team to quickly triage and fix issues, preventing them from ever impacting downstream customers. Ben’s team went from chasing down data bugs and data anomalies to proactively finding out about them and spending more time actually fixing the issues. Time-to-identify data quality issues went from hours or days to seconds. “Metaplane is key to preserving trust in our data. You can spend all your time moving to this great modern stack, but if you lose people’s trust and they won’t use it, that work is for nothing.” ### Results - After migrating from Postgres to Snowflake, Ben and his team don’t need to plan database changes like tuning and migration nor do they have to build complex ETL pipelines for product analytics and clickstream data, saving the team at least 10 hours every week. - With dbt Cloud, the SpotOn data team increased engineering contribution by 750% percent and building models went from taking days or weeks to hours. - After SpotOn adopted Metaplane, the time to identify data quality problems dropped from hours to seconds. The organization went from mistrusting their data, to asking for more data capabilities to power advanced analytics use cases. - With Snowflake, dbt Cloud, and Metaplane, SpotOn quickly integrated with Appetize’s data stack post-acquisition. Zero-second latency data sharing between databases, a shared skillset of dbt, built-in documentation, and guaranteed data quality helped the teams work together. ### What’s Next For SpotOn Over the next year, Ben and his team will continue to invest in Snowflake, dbt Cloud, and Metaplane. To create more consistency, the data team plans on migrating some legacy pipelines that ingest and cleanse important analytical and operational data. In addition to this migration, Ben and his team will be looking into ways to continue to use dbt to power modeling and decrease latency so the data can be provided closer to real-time. In an effort to reduce the number of data quality issues introduced by code changes, the SpotOn data team is adopting Metaplane’s CI/CD tooling to automate impact analysis and data test previews. Lastly, the company’s organic growth and the Appetize acquisition means a large number of new team members need to be incorporated into the larger data culture. Ben is focused on leveraging the [dbt semantic layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-semantic-layer) to help bridge this gap and make analytical data more accessible and consistent across the company. --- --- title: "Car & Classic boosts reporting speed 10x with Metaplane, Snowflake, and dbt Cloud" description: "See how Car & Classic built trust in their data and achieved 10x faster report load times with Metaplane, Snowflake, and dbt Cloud." url: "https://www.getdbt.com/case-studies/car-classic" date: "2023-02-17" industry: "eCommerce" --- # Car & Classic boosts reporting speed 10x with Metaplane, Snowflake, and dbt Cloud See how Car & Classic built trust in their data and achieved 10x faster report load times with Metaplane, Snowflake, and dbt Cloud. ### Company details - Headquarters: London, England, UK - Solution: Online Automotive Marketplace - Data stack: Meltano, Snowflake, dbt Cloud, Metaplane, Hightouch, Metabase ### Results - 10x report load time improvement with Snowflake - 8 hours/week saved identifying data incidents > "With Snowflake, dbt, and Metaplane, trust in data is a lot higher. As a result, our team is using data to drive the business forward." > > — James Sharwin, Head of Data [Car and Classic](https://www.carandclassic.com/) is Europe’s leading classic car marketplace trusted by millions of buyers and sellers monthly. Because cars are bought and sold on the platform, often quickly and at high prices, the infrastructure needs to be reliable, robust, and secure. James Sharwin, Head of Data, leads a nimble team of himself and one other data specialist. They’re responsible for collecting, organizing, and understanding Car and Classic’s influx of data, which is used by more than 100 stakeholders every day. This includes data in the core database, the CRM, marketing platforms, and even the finance platform. “Our chief responsibility is helping all teams across the company use data to drive better decision-making,” said James. “That means organizing the data, providing dashboards and visualizations to our teams, and working hand-in-hand with our engineers so we have the data they need to drive the business forward.” Even though widespread accessibility was a priority, it was difficult to make it a reality with their current tools. After thorough evaluation of different solutions, James leveraged Snowflake, [dbt](https://www.getdbt.com/product/dbt-cloud), and Metaplane to overcome the challenges the data team faced. ### Challenges: database performance, lack of centralization, and long iteration cycles James and his team faced a number of challenges when it came to making data accessible to the team. Because the team was using a MySQL database and often ran multiple complex queries at the same time, the data was either slow and took more than 30 seconds to load, or was unusable because it never loaded at all. The data was also not centralized nor easily accessible. Not having transformations defined in one place coupled with the slowness of analyzing data had several painful side effects like repetition and redundancy; teammates would often create different definitions of metrics without realizing it. Attribution in some areas, such as marketing spend, was virtually unanswerable. Performance and a lack of centralization not only impacted how frequently data was used — it also introduced delays in the iteration cycles of the data team. It’s difficult to empower teammates to analyze data and find insights when the underlying database is slow and they have no consistent metrics to rely on. “The amount of time it took to go from the inception of an idea to its implementation in our analytics was far too long.” All of these challenges led to the degradation of trust in the quality of the data. “Our data lacked integrity and our team didn’t trust it– they were even skeptical of our raw data,” said James. “You can have the best data setup and fanciest dashboards but if people don’t trust the data then it’s all for nothing.” James knew he needed to redesign how Car and Classic approached data so that it could be trusted throughout the organization and be a driver in pushing the business forward. With trustworthy and accessible data, all teams at Car and Classic could make better-informed decisions. ### Choosing Snowflake, dbt, and Metaplane To solve his challenges, James implemented Snowflake, [dbt Cloud](https://www.getdbt.com/product/dbt), and Metaplane together. Each tool offered benefits on its own, but the best results came with all three tools working in sync. ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/41736758e237fceb949655ba27e28a74d94cc1c2-1494x800.png) ### Snowflake: seamless set-up with unmatched ease-of-use James began looking for a data cloud platform and quickly landed on Snowflake. For James and his team, Snowflake was an easy decision because of the immediate performance gains, their relatively straightforward cost model, integrations with every tool they wanted to use, and their quick and easy setup process. “Snowflake was ridiculously easy for us to set up. It took 2 hours, maybe less, for it to be ready to go.” After adopting Snowflake, James and his team felt the benefits immediately. Given the data cloud’s ability to handle any analytics use case and the helpful built-in functions, the team improved load times of every report by 10x. “Snowflake was easy to use for any analytical use case and has every function under the sun. Load times of reports became near-instant, with at least 10x performance improvements.” ### dbt Cloud: engineering best practices, data quality, and velocity For James and his team, there was no better tool than dbt Cloud to increase the speed at which data products were built, in addition to helping teammates have visibility into the data and gain an understanding of how things are derived. Before dbt Cloud, the data team was creating massive nested queries using Metabase. After implementing dbt Cloud, James and his team created simpler modeling and orchestration all while introducing [engineering best practices](https://www.getdbt.com/resources/the-analytics-development-lifecycle) like opening pull requests, reviewing code, and adding tests. “In some ways, it’s going slow to move fast,” shared James. By defining models using code, the team avoided having to re-develop the same models 50 times—a problem that wasted time and resources, and caused bugs in production. Tools like the dbt Cloud IDE made it simple for the team to get started and scale. The straightforward API made it easy to run more complex orchestration outside of the dbt Cloud environment when required. For example, James added a job to push metadata from dbt Cloud to Metabase, which provided documentation to teammates where they viewed the data. “With the [dbt Cloud IDE](https://docs.getdbt.com/docs/cloud/dbt-cloud-ide/develop-in-the-cloud), we have a consistent environment where everything is the same. You don’t need to deal with Python package dependencies or VS Code extensions”—simplicity essential to onboarding new members of the team. James and the team felt the benefits of dbt Cloud by increasing the speed at which data products were built and reducing the number of data quality issues. “Before I joined and implemented [dbt Cloud and Snowflake](https://www.getdbt.com/data-platforms/snowflake), it used to take ages to build data products because the data team was re-inventing the wheel with code each time. If you wanted to build a dashboard about auction sale rate, you’d need to copy and paste the same 50 lines of SQL and adjust it,” shared James. Not needing to maintain hundreds of definitions scattered across different places also translated into fewer data quality issues. With dbt Cloud’s standardized model views of data, data engineers and analytics can adjust and re-use previous work in one place, reducing time and complexity in developing dashboards. The result was game-changing; analysts now build dashboards with three lines of SQL. The way dbt maps out complex transformations was also immediately helpful to the broader team. For example, engineers might have questions about a particular line of data in our dashboards, and the data team can point them to specific models, SQL, and lineage. “With dbt, I can give someone access to the project; they can look through and see everything like metric definitions and lineage to see where every piece of data comes from and which reports it feeds.” In terms of where dbt fits in the stack, James describes it as the “top of the Christmas tree.” Engineers can see everything from the tool. It provides a repeatable and consistent way to provide visibility to the broader organization. ### Metaplane: trust across the data stack For James, one of the worst recurring things he’s experienced is a downstream data consumer telling him that published data is incorrect. There have been numerous times in his career where someone mentions data that looks wrong in a dashboard, and the problem is obvious, but even with tests, it’s easy to miss given the scale of data he manages. "You will never be able to write a test suite that covers every eventuality, and there is nothing worse as a data professional than when bugs and erroneous source data slip through your controls and you're one of the last to notice. It erodes trust in both the technical architecture and you as a data practitioner." James wanted to make sure there were no issues slipping through the cracks. Metaplane offered a solution that integrated better into Car and Classic’s existing stack than any other tools. After a 15-minute setup, the platform has been automatically monitoring their data. Since implementation, the team has caught issues in production-level systems at least four times, saving hours of work necessary to identify the issue and weeks of work to fix affected data. In one case, Metaplane proactively alerted James’ team about an anomalous decrease in the volume of website log data. This data included important metadata about auctions like auction statuses and timelines. After receiving the alert, James was able to identify a cron script that was accidentally deleting data older than a year. If Metaplane had not identified this issue, data would have been lost and difficult to recover. Another interesting discovery was brought to James’ attention because of Metaplane’s ML-based monitoring strategy. Over time, the distribution of auction prices changed in a way that was not consistent with seasonal changes. This was an interesting insight that the team would not have noticed without active monitoring but were able to explain to the broader team. In addition to receiving proactive alerts, James’ team leverages Metaplane’s downstream impact analysis to make every alert actionable. For example, James has received alerts regarding important data that is operationalized by the marketing team; he can immediately notify the team and show the organization that his team is on top of any data incident, further preserving trust. James is now the one catching the issues as opposed to having team members bring them to his attention. “Before anyone in the organization brings an issue to my attention, I know about it,” he said. “Thanks to Metaplane, I’m the one telling people, which puts me in the driver’s seat and fosters more trust.” ### Results - A scalable, trusted data stack built on Snowflake, dbt Cloud, and Metaplane that can serve various teams and over 100 stakeholders across the organization. - Immediate performance improvements thanks to Snowflake; seamless integration with Car and Classic’s existing and future tooling. - Dependable centralization and visibility provided by dbt, particularly in helping the team understand data lineage. - A consistent development environment with near-zero engineering overhead via dbt Cloud. - Increased [data trust](https://www.getdbt.com/product/build-trust-in-data-and-data-teams) with Metaplane; the data team is now catching issues before stakeholders. - 8 hours/week saved by catching issues proactively with Metaplane. - Reduced time to identify data quality issues from weeks to hours. ### What’s Next James is planning on increasing the team size over the next six to twelve months to tackle new projects. He hopes to increase the use of event-level data to enable Car and Classic to better understand marketing attribution, provide recommendations, and more thoroughly test and measure changes to their platform. “I believe investing in the right tools to do the job is often better than hiring two full-time engineers to do the same thing,” concluded James. --- --- title: "Nasdaq empowers business users with dbt Cloud and a modern data stack" description: "Nasdaq modernized data workflows with dbt Cloud, boosting collaboration, data access, and business insights." url: "https://www.getdbt.com/case-studies/nasdaq" date: "2023-01-23" industry: "Banking & Financial Services" --- # Nasdaq empowers business users with dbt Cloud and a modern data stack Nasdaq modernized data workflows with dbt Cloud, boosting collaboration, data access, and business insights. ### Company details - Headquarters: New York, NY - Founded: 1971 ### Results - 100-125 billion transactions operated per day - 55 business users onboarded to modeling and reporting > “Before, 9 out 10 times the sales and executive team had to wait months to receive a data point they requested. By then, the data wasn’t relevant anymore or the new business was already lost.” > > — Michael Weiss, Senior Product Manager ### The global engine of financial markets ventures into the cloud [Nasdaq](https://www.nasdaq.com/) manages 30 stock exchanges across North America and the Nordics, and is the home of over 4,000 public companies. They also provide technology to 2,200 financial institutions in 130 markets. To say Nasdaq is the engine of the financial markets is not a hyperbole. In 2012, before most of the financial community made the jump, Nasdaq started their journey to the cloud. They quickly progressed—launching a data warehouse, new products, and new marketplaces. After encountering scalability issues, they migrated to a data lake and optimized their data infrastructure by 6x. The migration saved significant maintenance costs and opened the door for analytics, ETL, reporting, and data visualization use cases. #### An opportunity arises from a better data infrastructure: analytics #### Opportunity 1: Empowering business users Business users at Nasdaq previously struggled to access the data they needed. Data requests were a complicated process, requiring business stakeholders to submit tickets in order to receive the reports they needed. The reports could be provided by any of Nasdaq’s four distinct data teams—who owned which data product was unclear, which created confusion and overlap. “Our vision was to have business users participate in the transformation of our data logic,” said Michael Weiss, Senior Product Manager at Nasdaq. #### Opportunity 2: Provide new types of data access to clients As the cloud-native data stack developed, so did the requirements of Nasdaq’s business clients. Delivering static reports was no longer enough. With their unique datasets in hand, Nasdaq saw an opportunity to build data products for these new needs. “Clients were becoming increasingly more likely to ask for dynamic access data, whether it be APIs or visuals,” said Michael. “We then searched for tools that could solve this new demand.” But to get there, Nasdaq still needed to overcome several roadblocks. ### Data for analytics: not as simple as it sounds #### Data optimized for performance, not for querying Nasdaq’s unique trading data is both a valuable differentiator and their biggest challenge for analytics use cases. “Trading system data is really complex. It is optimized for performance, not for querying or analytics as most message data is,” said Michael. The large data volume also added an extra layer of complexity. In 2019, Nasdaq operated 25 to 35 billion transactions daily through their “data warehouse-lake.” But, after 2020, it increased to 100-125 billion per day. #### Legacy ETL tools locked business users out One of the main blockers in onboarding business users to data modeling was Nasdaq’s legacy data tools. They used a combination of Pentaho and “a lot of SQL scripts running around for various things.” “Even though business users wanted to engage with the data, our legacy data infrastructure didn’t allow it,” explained Michael. #### Teams working in silos Nasdaq’s teams therefore resorted to creating their own data stacks for their reports. Because new data requests took weeks or even months to address, teams would try to solve their own data requirements. “Business users couldn’t sit around and wait for the reporting team to add a column to a report, get insights into making a pricing change, or see what's actually happening in their markets,” said Michael. #### Monolithic SQLs script SQL scripts were an important part of Nasdaq’s data infrastructure. Different data teams were using SQL for analysis data ingestion. However, a lack of documentation and collaboration meant they were becoming monoliths. “SQL scripts were hard to maintain. When the person in charge of them would leave the company, we’d have no idea what was in there. And then we had to rewrite the code,” said Michael. ### Moving to a secure, modern data stack with the support of dbt Cloud ![data stack](https://cdn.sanity.io/images/wl0ndo6t/main/fdcb624d48d67ccbfe9d198936423d71e3b64dac-1600x800.png) #### Building for end users: sales and executives It was clear from the get-go that a move to a different data stack wasn’t just about the tooling, but about the end user. Nasdaq’s team had a vision to onboard business users into the data modeling process—their new tools needed to enable team members across the company with different technical skills to participate. “It’s about the people you pick tools for,” explained Michael. “And it’s hard to pick tools for sales people and executives.” Success meant unlocking a previously-unattainable velocity for Nasdaq’s business teams. They would no longer have to wait for reports and miss the window where an insight could change the outcome of a business decision. “To me, self-service is not just necessarily about the capability of someone to get access to data, but how quickly you can give them the answer to something that they don't have the answer to today,” shared Michael. #### dbt’s Enterprise support sets the new data stack up for success One of Nasdaq’s concerns about moving to a new stack was onboarding [analysts](https://www.getdbt.com/product/analyst). Setting up development environments was something most analysts had never done before. “I didn't want to have to sit with each team member and guide them through a complex install as we brought in new analysts,” said Michael. [dbt Labs’ services](https://www.getdbt.com/services) team stepped in, helping Nasdaq’s large teams onboard with ease and move into a productive implementation in a matter of weeks. The team at dbt Labs continued to assist with Nasdaq’s large-scale dbt roll-out to answer more challenging setup questions and optimize their models at scale. “dbt Labs laid out good practices and helped really get our options business, which was the first business we got online, moving quickly,” said Michael. “Most recently, we had performance issues with a model that does 15 to 20 billion messages per day. The dbt Labs services team helped us solve that. The model took 45 minutes to an hour to run in a day, but they got it down to 10 minutes.” ### Streamlined analytics for both data and business teams #### Better collaboration with SQL Following the set-up of [dbt](https://www.getdbt.com/product/dbt), SQL models provided new value for Nasdaq. The models were now understandable, maintainable, and repurposable. “I was very skeptical of improving collaboration through SQL, but it actually worked,” laughed Michael. The combination of SQL’s accessibility and dbt’s embedded software engineering best practices simplified and centralized Nasdaq’s analytics workflow across teams. #### Insights from previously inaccessible data points Nasdaq’s most competitive business, their options business, was the first to fully onboard to dbt Cloud. They can now count on several hundred models that lead into a couple key models they use day-to-day. “Because of these [dbt models](https://docs.getdbt.com/docs/build/models), our sales and executive teams, can now see data points they never had access to before,” said Michael. “For example, exchanges do pricing changes on a relatively monthly basis in the US. It has an impact on all types of different metrics you care about: your revenue, your volume, who's trading with who, things like that.” #### No more missed business opportunities from slow turnaround times The move to a modern data stack completely changed how business teams access data at Nasdaq. “Before, our sales and executive teams would put a request into the economic research team or the data team. If they were lucky, it would be in a report somewhere or several reports, and maybe they could piece that answer together. Nine times out of 10, that was not true,” shared Michael. The wait time for these reports could take months. “By the time business users would receive what they requested, it might not even be relevant anymore. They could have missed a new business opportunity or lost revenue. Data velocity is important,” explained Michael. “In today's world, because we have built a robust set of models and metrics, people can model or view the data themselves.” ### What’s next: Building on top of the modern data stack From the start, Nasdaq’s approach to building a modern data stack was agile. They built piece by piece to isolate successes, mitigate risk, and iterate without a complete infrastructure overhaul every time. #### Onboarding new teams and markets The options business is fully online and reaping the benefits, but there are still other businesses to go. For example, the equities team is next in line to implement dbt. Currently, Nasdaq is onboarding their economic research team to dbt. They have their own backlog of scripts and models ready to be exposed and repurposed as dbt models. Nasdaq is also looking at how they can extend their new data stack to international markets, like the Nordics. #### Adding new tooling and capabilities Nasdaq signed new contracts with Atlan and Monte Carlo, and are now working on integrating the tools. Atlan will unlock an improved view of their data lineage, and Monte Carlo will enable a more proactive approach to data pipeline breaks. Another tool—[a reverse ETL](https://www.getdbt.com/blog/reverse-etl-playbook)—is potentially on the horizon: “Our work on this is never done. We're always going to constantly evaluate and make sure we're picking the right tools for the right job,” emphasized Michael. --- --- title: "Blend scales the impact of reliable data with dbt Cloud and Monte Carlo" description: "Discover how Blend used dbt Cloud and data observability to deliver faster, more trustworthy insights for top financial institutions." url: "https://www.getdbt.com/case-studies/blend" date: "2023-01-12" industry: "Banking & Financial Services" --- # Blend scales the impact of reliable data with dbt Cloud and Monte Carlo Discover how Blend used dbt Cloud and data observability to deliver faster, more trustworthy insights for top financial institutions. ### Company details - Headquarters: San Francisco, CA - Data stack: Snowflake, dbt Cloud, Monte Carlo ### Results - 4 months reduced in time-to-value when compared to internal POC frameworks > “Blend is enterprise software handling data for large financial institutions—our data is a product. Without dbt, we wouldn't be able to live up to our agreements on providing customers visibility into their data" > > — DC Chohan, Senior Data Engineering Manager [Blend](https://blend.com/) powers billions of dollars in transactions each day in its mission to reimagine the future of online banking. Central to this mission? Fast, accessible, and most importantly, reliable data. DC Chohan, Senior Data Engineering Manager, is helping lead the charge toward data adoption at Blend. “As we’re growing, the number of sources is growing—the number of people relying on data is increasing… Today, we have around 32 sources of data, and 12 different teams utilize the data on top of our platform,” DC said. As Blend’s data needs began to grow, it became increasingly difficult to harness their data effectively on their existing platform. More data was critical to power their internal insights and customer-facing products—but they also needed a data platform that could operationalize that data in a way that was fast and accurate. ### The challenge: optimizing data modeling and improving reliability at scale As an enterprise cloud-banking platform supporting the likes of Wells Fargo, Blend’s data needs are as large as they are complex. With three distinct data teams and dotted lines to both growth and revenue operations, Blend’s data function provides operational data supporting everything from product development to standard measuring and performance reporting. In addition to internal data users, Blend also leverages data in a variety of external data products as well—including the integration of sensitive PII data from their customers’ end users. But manual SQL querying and a lack of sufficient data quality monitors left Blend with limited engineering resources to develop new products and anomalous data that was prone to cause incidents for users downstream. #### Manual transformation querying draining engineering resources In the early days, the data team was relying entirely on a series of manually scheduled SQL queries in Airflow to manage operational workflows and feed external data products. But as the company—and its data—began to scale, so too did the bottlenecks. While Blend's SQL queries were technically capable of delivering data to downstream products and BI tooling, the queries were becoming increasingly difficult to run and iterate, requiring constant redeploy fixes to keep pipelines moving. And because each query was manually created and executed in Airflow, teams like product analytics were fully dependent on data engineering to create and implement new workflows. If a new transformation request came in from analytics, either another priority project would need to be paused or analysts would be left waiting until engineering could action the request. #### Insufficient data quality monitoring When it came to managing data reliability, things weren’t any simpler. While more tables and sources meant more opportunities for data discovery, they also meant more opportunities for anomalous data to cause mischief downstream. “We had an incident where our revenue numbers were skewed…and it was because we weren't calculating data quality well,” DC said. Understanding the importance of [data quality](https://www.montecarlodata.com/blog-data-quality-issues/), the data team attempted to create a POC data quality framework in-house, using orchestrated validation queries. “It was supposed to run on some schedule to try to detect anomalies in data…but when we ran it, [the monitors] completely overwhelmed our warehouse… Not only were queries taking a long time, but it was taking a lot of CPU on the Redshift cluster,” said Albert Pan, software engineer at Blend. Recognizing the imminent need to reevaluate their platform’s infrastructure, the team began looking for tooling resources that would allow them to operationalize their data more effectively for modeling and discovery while also ensuring the data was reliable for end users. With limited time to stand up new integrations, the team needed accessible and easy-to-scale solutions that worked out of the box. ### The solution: [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) and Monte Carlo The Blend data team had three objectives: reduce bottlenecks, improve data reliability, and maintain velocity as their data needs grew. The solution to their scalability and reliability woes lay in the combination of [dbt Cloud](https://www.getdbt.com/product/what-is-dbt/) and [Monte Carlo](https://www.montecarlodata.com/)—where efficient transformation and modeling meets hyper-scalable end-to-end quality coverage. ### Improve speed-to-insights with modular SQL and automated anomaly detection #### Modular SQL development dbt Cloud was the natural evolution of Blend’s existing workflows. While it operated under a similar SQL framework to what the data team had created within Airflow, dbt Cloud was purpose-built for the challenges Blend was facing at scale. Unlike the manual queries the data engineering team had been laboriously creating for each use case, dbt introduced [software engineering best practices](https://www.getdbt.com/resources/the-analytics-development-lifecycle), allowing the team to run workflows and develop data products faster. And because the system was modular, it enabled data practitioners like analysts to self-serve development as well, operating as ad-hoc engineers to create queries and transform data as needed—optimizing engineering resources for priority projects. “Because a lot of the work in understanding data is within the product analytics team, we didn't want to have data engineering be a roadblock,” said Albert. “We wanted to give analysts a self-serve tool to be able to develop on, and dbt Cloud was exactly what we needed.” #### Automated end-to-end anomaly detection But faster insights are only as good as the data they’re built on. And up to that point, Blend had struggled to qualify the health of their data in a meaningful way. What the data team needed was an efficient way to [monitor](https://docs.getdbt.com/docs/deploy/monitor-jobs) for anomalies across the full breadth of Blend’s sources and production tables without the need for manual configuring. “A year ago, data discovery was the most important. We were asking, ‘what should we do with this data?’ Now we’ve reached a stage where data reliability is the most important,” DC said. “We want to calculate ROI for everything we invest in. The market has changed the strategy of how we reliably use the data. And that's where Monte Carlo comes into the picture providing that insight in a very reliable way.” Utilizing Monte Carlo’s no code implementation, Monte Carlo’s ML-powered monitors were instantly deployed across 100% of Blend’s production tables out-of-the-box— evaluating freshness, volume, and schema changes automatically without any configuring or thresholding. “It really started becoming a very automated workflow where we didn't have to tune it or babysit it a lot,” Albert continued. Monte Carlo’s automatic field-lineage also meant the team could root-cause incidents at a glance, resulting in faster resolutions and a deep understanding of the downstream impact of anomalies across pipelines. ### Meet external data SLAs with system stability and Slack alerts for breakages While Blend’s Airflow framework provided a means for the data engineering team to manually schedule queries, the system was bulky, tedious, and prone to frequent breakages, making it nearly impossible to use as a reliable solution for external data products. “Given that Blend is enterprise software handling data for large financial institutions, our data is a product,” said Noah Hall, product analyst at Blend. “Without dbt, we just wouldn't be able to live up to the kind of agreements we have in terms of providing customers visibility into their own data…or it would take many more months to actually get these products out the door.” And if anything within their customer-facing data products did break—or anywhere else for that matter—Monte Carlo would be there first with automated alerting through Slack to alert to breakages in real-time, while also giving Albert, DC, and the rest of the team the tools to fix them. “We had many cases where folks would say, ‘Hey, we're not seeing this data. We're not seeing these rows. Where are they?’” recalled Albert. “It's never great when a customer tells you that something is wrong or missing. So, we wanted a proactive solution that could tell us when something is wrong, and we can fix it before they even know. That's one of the main reasons why we use dbt and Monte Carlo.” ### Minimize compute-power with pre-processed operations and metadata monitoring #### Modular SQL operations to pre-process data Blend divides its codebase into a variety of microservices to support customer applications. Unfortunately, most of the processing required for these SQL operations are too compute-intensive to execute live as dashboards are opened by users. To reduce the strain on Blend’s compute resources, the product analytics team uses dbt’s modular SQL framework to pre-process complex SQL operations before creating the dashboards that will ultimately be consumed by internal and external users. “An application might use different microservices for different income verification integrations in the mortgage application process. So the foreign key out of one of those microservices that will allow you to join it back to the application ID might be stored in a JSON blob,” said Noah. “We use dbt to modularize that different microservices data and extract some of those key pieces of information so that we can build out the clean end-user data into a table that will be queryable efficiently by our dashboarding tools.” #### Comprehensive metadata monitoring for freshness, volume, and schema changes across production tables Unlike Blend’s in-house POC program which relied on time-intensive and CPU-heavy queries within their Redshift cluster, Monte Carlo’s end-to-end data observability minimized impact on warehouse performance. “One of the things Monte Carlo offers is a very light way to [monitor data] by collecting metadata, which doesn't impact warehouse performance on a significant level like our original POC project did,” said Albert. Monte Carlo’s monitors avoid querying data directly by using that metadata, pulled from native objects such as information schema, to automatically generate thresholds for freshness, volume, and schema changes. Those thresholds continue to be updated over time as well, to ensure that updated processes are accurately reflected in data quality rules. This solution saved on Blend’s compute costs while simultaneously giving the data team visibility into the health of its pipeline from ingestion to consumption. In addition to broad monitoring across all of its tables, Monte Carlo also allows the data team to opt-in to deeper custom data monitors on critical data sets to identify when the data fails a specific user-defined logic as well. By using Monte Carlo’s automated monitoring for common anomalies and limiting custom monitors to critical production data sets, the data team was able to provide powerful cost-effective quality coverage across Blend’s pipelines without sacrificing deep monitoring where it counts. ### Outcome: Immediate time to value, reduced compute costs, and improved data reliability dbt’s cloud-based transformation workflow [democratized data development](https://www.getdbt.com/blog/managing-data-democratization) at Blend by offering a modular approach to SQL queries that gave development power to analysts and data practitioners while improving the stability of the platform. Meanwhile, Monte Carlo provided scalability for data quality by offering a solution that went beyond the warehouse to provide convenient data quality coverage for Blend’s entire pipeline—from ingestion to consumption—right out of the box. By combining the operational efficiency of dbt Cloud’s data transform tools with the scalability of end-to-end data observability through Monte Carlo, Blend has: - Reduced time-to-value by 4 months for a new data quality solution when compared to internal POC frameworks - Improved data trust for internal and external users with automated data quality checks across tables - Dramatically improved speed-to-insights with modular SQL development for practitioners - Reduced warehouse and compute costs with pre-processed SQL operations and broad automated metadata monitoring ### What’s next? Data transformation and observability solutions in hand, the data team has already set its sights on building out new infrastructure and data products utilizing these new resources. “At the infrastructure level, we’re going to focus on building real-time pipelines and a tracking framework… And on the other side, we also want to keep enhancing our products, which is the STP that we service our external users,” said DC. Continuing to champion Monte Carlo to data users outside their team and integrating observability into existing platforms also remains a high priority on Blend’s data roadmap. Integrating dbt Cloud and Monte Carlo has not only enabled Blend to harness their data faster—it puts their data consumers at the top of the priority list. --- --- title: "Reforge uses dbt to 3x their team and save 18 hours a week" description: "Reforge is a learning and community platform that helps tech professionals do the best work of their careers. Reforge connects its members to industry experts and industry-leading content, and teaches these materials in cohort-based courses" url: "https://www.getdbt.com/case-studies/reforge" date: "2022-12-02" industry: "Education" --- # Reforge uses dbt to 3x their team and save 18 hours a week Reforge is a learning and community platform that helps tech professionals do the best work of their careers. Reforge connects its members to industry experts and industry-leading content, and teaches these materials in cohort-based courses ### Company details - Headquarters: San Francisco, CA - Data stack: Stitch, Segment, Snowflake, dbt, Metabase, Hightouch, Amplitude, Metaplane ### Challenge Growing a data team as the demand for insights increases ### Solution Building a scalable data stack on Snowflake with dbt and Metaplane ### Results - 350% team growth rate thanks to their newfound ease of onboarding - 18 hours saved per week > "Engineers can see how the data they model in dbt becomes data that other teams are acting on. They can understand why we make certain requests. It makes a world of difference in their motivation and satisfaction." > > — Daniel Wolchonok, Head of Data Data has always played a critical role in driving [Reforge](https://www.reforge.com/)’s business forward. Since day one, data was leveraged to understand important metrics like conversion rates and user engagement. As Reforge experienced rapid growth, their Head of Data, Daniel Wolchonok, wanted to make sure the data team kept up with this pace and made new metrics and insights available to the larger team. > _"Now, we have a larger number of users and all the responsibilities of a SaaS business. We’re focused on metrics like new revenue, retained revenue, expansion, contraction, and churn. Our data needs have exploded and I expect that will only accelerate."_ As the number of stakeholders relying on data skyrocketed, Dan and his team ran into scalability issues with their data warehouse performance and data engineering development practices. The first data warehouse that Dan and his team built used Postgres, a popular open source OLTP database. When more teams started using data, queries routinely timed out for different stakeholders, making it harder for them to use data. After continually increasing to more powerful instances, Dan’s team felt like they were spending more time figuring out how to optimize their Postgres resource utilization rather than on what really mattered - bringing insights to the business. Because Reforge didn’t previously have a dedicated data team, the data stack was developed by the engineering team using software engineering tools. The data models lived in the engineering team’s GitHub repo and were managed like a Ruby on Rails migration schema. Those systems worked well enough when the company was small, but they weren’t designed to scale with more complex data needs or a fast-growing team. Changes to data models in the core GitHub repo required submitting a pull request and waiting for engineering resources, so they’d take the quick and dirty route — making changes directly in the data warehouse. Of course, that meant there was no version control and sometimes those changes would break something downstream. The result was, as Wolchonok described it, _“a bad experience for everybody.”_ ### Building a scalable data stack on Snowflake with dbt and Metaplane To Daniel Wolchonok and others at Reforge, it was clear that they needed to migrate to a scalable data stack. As a newly-formed and growing data team, they wanted a data stack built specifically to fit their needs, not one cobbled together from tools designed for software engineering. > _"I want to empower each function with the best-of-breed tools, rather than one monolithic platform. There’s a wealth of tools out there that are really easy to stitch together._ ![Data Stack](https://cdn.sanity.io/images/wl0ndo6t/main/13bc226268f57fdab69f0d33a569e9f8e463beb2-2672x1236.png) Reforge uses the best-of-breed tools to build a cutting-edge data platform that has helped power Reforge's use of data and rapid team growth. #### Snowflake as an easy-to-use and scalable warehouse Searching for a warehouse that is easy to get started with but could scale with their needs over time led Dan and his team to adopt Snowflake. Snowflake made it simple to set up development, staging, and production databases so that the data engineers could easily move data from Postgres to Snowflake. After finishing the migration, it was immediately obvious how the separation of storage and compute allowed the data team to focus on helping their teammates answer questions about the business, rather than worrying about provisioning resources or storage costs. > _"It’s blown me away how quickly you go from getting by with some quick and dirty hacks to needing tools like Snowflake and dbt to break up and model datasets, manage testing, staging environments, deployment, all that stuff."_ #### [dbt Cloud](https://www.getdbt.com/product/dbt-cloud) for implementing data engineering best practices and showcasing the impact of data Introducing dbt has been a milestone for Reforge because it encouraged data engineering best practices and made it easier to migrate to Snowflake. When Reforge switched to Snowflake, all they needed to do was swap the source and the transformations would run in their new Snowflake environments. Using dbt Cloud made it easier to build and maintain models over time, rather than data engineers making changes directly to the warehouse. When dbt became a part of Reforge’s CI/CD process, the data team felt empowered to own data modeling changes end to end. With dbt, it also became easier for Reforge to showcase how data is used across the organization. Being able to see a visualization of the DAG and data lineage helped the team understand and share how data is being used. > _"Engineers can see how the data they model in dbt becomes data that other teams are acting on. They can understand why we make certain requests. It makes a world of difference in their motivation and satisfaction."_ #### Metaplane for Data Observability Even before his work at Reforge, Dan was aware of the challenge of data monitoring and had experienced first-hand the chaos that can happen when data issues go unnoticed. Because the data team is growing and wears many hats, they found themselves responsible for managing thousands of objects in Snowflake. Before Metaplane, it wasn’t uncommon for them to find out tables had been missing weeks of data or the distribution of data had considerably drifted from source systems. Dan and his team also found it hard to keep up with the fast pace of changes from upstream application databases. While they were at the center of data, they could be the last to find out about a data quality issue. Dan and his team were often digging around in the database looking for new tables and trying to understand changes just from column names and relationships. It was fine early in Reforge’s existence, but it was clearly not scalable. Since implementing Metaplane, Dan and his team have been the first to know about data quality issues and changes in their data. When an engineer accidentally loads contacts to the wrong place, the data team finds out immediately, allowing them to fix the problem before it compounds or goes unnoticed for weeks. With Metaplane monitoring the data pipelines and reporting back, Reforge’s Head of Data has built a different reputation with his team - one of trust. > _"Because of Metaplane, I spook people with how quickly I find data issues now. I know when events are created with crappy names, when there are new attributes, when someone puts a unique ID in a segment event. It helps me understand what's going on with our data and helps me catch issues that are painful to untangle."_ ### Results - Reforge now has a scalable data stack built on Snowflake, dbt, and Metaplane that can service the entire organization and scale with a rapidly growing team - Using [dbt and Snowflake](https://www.getdbt.com/data-platforms/snowflake), the data team was able to easily migrate from a poorly performing Postgres instance to a Snowflake environment where they didn’t need to constantly tune resources. - After migrating to Snowflake, BI queries and dbt transformation queries that used to take hours now take minutes. - Functions across the company have access to consistent and accurate data based on sophisticated data models built in dbt — [everyone is working from the same set of data](https://www.getdbt.com/product/build-trust-in-data-and-data-teams). - Using Metaplane, the Reforge data team is the first to know about data quality issues. When issues occur, the team has a starting point to find the root cause and an understanding of the downstream impact to models and BI dashboards. ### What's Next at Reforge **Focusing on what matters most for Reforge’s continued success: insights from data** As Reforge’s team grows, they can now focus on what is most important to the business - enabling the larger organization to ask questions and find insights in the data. They no longer need to worry about tuning their warehouse to handle increased loads because Snowflake can handle any workload size. The data team now manages data models in a scalable way and can be confident changes aren’t being made directly to the warehouse. And, as the company grows, Dan can easily show how the data is being used. Lastly, the data team has preserved trust with the growing organization because they are the first to know about data quality issues and can broadcast issues they are fixing to the larger team. --- --- title: "Ramp drives 25% increase in new business with dbt" description: "Discover how Ramp created a new sales lead source with Snowflake, dbt Cloud, and Hightouch." url: "https://www.getdbt.com/case-studies/ramp" date: "2022-11-28" industry: "Banking & Financial Services" --- # Ramp drives 25% increase in new business with dbt Discover how Ramp created a new sales lead source with Snowflake, dbt Cloud, and Hightouch. ### Company details - Headquarters: New York, NY - Data stack: Fivetran, Snowflake, Airflow, dbt, Hightouch, Looker, Retool, PostgreSQL ### Results - 20% reduction in data platform costs - 33% increase in data transformation speed - 25% of sales pipeline generated by their new personalization engine > "All of our models are born and bred in dbt. When people think of clean data, they think of dbt models. There's very little that isn't powered by dbt at Ramp." > > — Kevin Chao, Senior Analytics Engineer ### About Ramp [Ramp](http://www.ramp.com) is the first and only finance automation platform and corporate card designed to help businesses spend less time and money. With over 1,000 integrations with financial products like Netsuite and Quickbooks, Ramp makes it easy to issue corporate charge cards, make bill payments, automate accounting processes, and monitor expenses. Ramp powers $5B+ in payment volume across corporate and small business card and bill payment data, serving 10,000+ businesses and 200,000+ cardholders. Its customer mix is diverse—ranging from small businesses to venture-backed startups, mid-market, and public enterprise players. Businesses spend an average of 3.5% less and close their books 8x faster by switching to Ramp. Founded in 2019, Ramp powers America's fastest-growing corporate card and bill payment software. ### The challenge of exponential growth Ramp has seen exponential growth since its start, increasing revenue and cardholders by 10x and 15x, respectively, year-over-year in 2021. Ramp's recent fundraising round, backed by Goldman Sachs, Citi, Founders Fund, Stripe, Coatue, and Thrive, valued the company at $8.1B. Over the last 12 months, Ramp's data team has tripled in size to keep pace with the company's rapid growth and business demands. This influx of new users created concurrency issues—resulting in costly, hard-to-debug failures, forcing the data team to spend too much time managing workloads. Whenever too many users tried to access a specific database, Ramp's entire database would deadlock, and nobody could access the data. Scaling up to meet the demands of the various business teams and internal stakeholders took a lot of work. _"We started to see a lot of issues when it came to scaling our previous data warehouse. We were spending too much time fine-tuning our workloads just so we could keep everything up and running,"_ explained Kevin Chao, Senior Analytics Engineer at Ramp. Ramp's data team also wanted to give their business users access to the data they cared most about in the tools they use every day. This objective, mixed with the challenges of Ramp's previous data warehouse, led the company to adopt a new data stack that could not only scale with Ramp's data team but also provide a way for business teams to access the data sitting in the analytics layer. ### Ramp's new data stack In search of a data stack that could facilitate these needs, Ramp quickly turned to Snowflake, [dbt](https://www.getdbt.com/product/what-is-dbt), and Hightouch. ![Data Stack](https://cdn.sanity.io/images/wl0ndo6t/main/3c49f19d8416f843efc899e5a51fc0f58abae2ff-1600x459.jpg) ### Snowflake Looking for a scalable and flexible solution that could act as a single source of truth, Ramp landed on [Snowflake](https://www.snowflake.com/). As a fully managed data platform, the Data Cloud eliminates all of the bottlenecks that Ramp faced previously. Snowflake automatically manages all of the underlying maintenance in the background, allowing the data team to focus their time on transforming the data and [building models](https://docs.getdbt.com/category/models) that can positively impact Ramp's bottom line. Snowflake scales automatically as Ramp grows, and the data team no longer has to worry about key tables getting locked or dashboards freezing. With Snowflake, Ramp’s data can automate all of the manual tasks that the admins were forced to manage in the previous platform. _"Snowflake essentially handles everything out of the box without us having to think about anything so we can focus on modeling our data and extracting value,"_ said Kevin. Since adopting Snowflake, Ramp can handle more workloads faster. Transformation jobs in Snowflake now complete 33 percent more quickly, and Ramp has seen a 20 percent decrease in overall cost compared to its previous data platform. In addition to this, the team no longer has to worry about contentious deadlocks thanks to Snowflake's near-unlimited concurrency. ### dbt Ramp collects valuable customer data from an array of different data sources, including Postgres, Salesforce, Hubspot, Zendesk, Outreach, and various ad platforms. Spinning up additional software infrastructure to clean and transform this data in the analytics warehouse was difficult. The team built new data models manually without testing or version control, resulting in variations of core metrics. Ultimately, this led the data team to adopt dbt as their data transformation tool of choice. Rather than building and maintaining ad hoc transformation jobs, Ramp uses dbt to standardize around a core set of metrics and fully automate, schedule, and run every transformation job. _"We can use dbt as a testing ground before fully committing ourselves to a new data product. dbt lets us abstract all of the nitty-gritty details that come with building data models so we can focus on delivering value to our business teams."_ Meeting the right individuals with the right message at the right time is vital to Ramp's success. With [dbt jobs running in Snowflake](https://www.getdbt.com/data-platforms/snowflake), Ramp's data team can aggregate, transform, and enrich all of their customer data in Snowflake to build a complete 360-degree view of the customer. This data is then used to create risk profiles for specific users and improve personalization, whether it's in the app, on the website, or through automated marketing campaigns. _"All of our models are born and bred in dbt. When people think of clean data, they think of dbt models. There's very little that isn't powered by dbt at Ramp,"_ emphasized Kevin. ### Hightouch Building data models in dbt is one thing, but activating them in downstream sales and marketing channels is another. Ramp's data team set it sights on enriching Salesforce with all the valuable dbt models living in Snowflake. Anytime Ramp wanted to move data out of Snowflake, they were forced manually to build and replace various python scripts. This was a challenge because Ramp's business teams wanted access to this data in Salesforce and Hubspot to improve personalization. In search of a scalable solution that could solve this problem, Ramp turned to Hightouch. Since adopting Hightouch for Reverse ETL, Ramp has created an entirely new outbound automation team (OATs) which now drives 25 percent of all sales pipeline. Staffed by data engineers, this team collaborates closely with Ramp’s marketers to identify target prospects and deliver customized emails at scale to the right person at the right time. OATs has become the single lowest-costing customer acquisition channel for Ramp. _"Having enriched data available in Salesforce means the sales team has one view with everything they need to understand what's going on in their pipeline,"_ said Kevin. Ramp also uses Hightouch to enrich Hubspot and Outreach with relevant customer metadata, key events, product usage data, and other sales-related information. Before Hightouch, A/B testing various campaigns was a nightmare. Using Hightouch, Ramp can sync custom audiences directly to various ad platforms helping the marketing team optimize return on ad spend (ROAS) and increase conversions from paid ads. Ramp also leverages Hightouch to automate the company’s entire underwriting and application process by syncing data directly to Postgres. _"Thanks to Hightouch, we're able to use Snowflake for all the heavy computations, allowing our production databases to focus on operations,"_ said Kevin. Syncing data to Slack is another major use case for Ramp. With Hightouch, Ramp's data team is notified about potential problems in their data stack before they escalate. ![Data freshness](https://cdn.sanity.io/images/wl0ndo6t/main/e064a38cebc0b7692761ae1480adc8775e5446f7-650x51.png) _"We have very aggressive SLAs for data freshness, and we want to know when something goes wrong in our warehouse. With Hightouch, we get notified immediately."_ ### What’s Next Since adopting a modern data stack, Ramp can go from ideation, to validation, to iteration, and set up fully functional operational workflows and marketing campaigns in less than a day. _"With Hightouch, Snowflake, and dbt, we can go from zero to one as fast as possible,"_ affirmed Kevin. For Ramp, the future is continuing to find ways to help businesses become better, more profitable versions of themselves through finance automation that maximizes the output of every dollar and hour. They're looking to build out their automation platform to reach customers across every expense, payment, purchase, application, and insight (from reporting to forecasting). --- --- title: "Sunrun enables last mile modeling with dbt Cloud" description: "This is the story of how Sunrun got 40 analysts from across the business to develop together" url: "https://www.getdbt.com/case-studies/sunrun" date: "2022-11-17" industry: "Renewable Energy" --- # Sunrun enables last mile modeling with dbt Cloud This is the story of how Sunrun got 40 analysts from across the business to develop together ### Company details - Headquarters: San Francisco, CA - Data stack: Github, Snowflake, dbt Cloud, Tableau, Informatica ### Results - 100% participation with 10 engineers and 40 analysts now sharing one common development framework - 50% reduction in engineering tickets to diagnose and resolve complex data issues - 75% acceleration in time to deployment > “With data teams spanning several business functions, we didn't just need a way to standardize development. We needed a way for those processes to be easily understood by everyone—seasoned engineers, new analysts, the CFO... everyone.” > > — Jared Stout, Head of Data Management ### The hard part of becoming the largest solar provider in the nation [Sunrun](https://www.sunrun.com/) is the largest residential solar company in the United States, servicing more than 750,000 customers nationwide and growing fast. “Data drives efficient growth. In order to grow correctly, maintain all of our inventory, retain employees, execute a very aggressive sales strategy, and report and forecast effectively, we need to have our data in order,” said Jared Stout, Head of Data Management at Sunrun. After acquiring Vivint Solar in 2020, the team was faced with the expected but no less daunting challenge of merging vastly different data structures. ### Unifying data development post-acquisition One of the main difficulties was various people developing and merging code at the same time. Vivint Solar was using homemade python scripts to manage data transformations within their data warehouse. Sunrun was using the composer feature on Google Cloud and Airflow to transform and schedule their data. Across both companies, “people didn't know how we were building things,” explained James Sorensen, Senior Data Engineer. “They didn't know the code, they didn't know the lineage, and they didn't know how things were moving and transforming through our different stages of development.” Each approach was less than efficient in their own right, both were inaccessible to anyone without python skills, and in any event, the two ways of work were incompatible with one another. They needed to standardize, with a more accessible and reliable way of work. With code conflict management top-of-mind, the team searched for a solution that would allow them to: - See model dependency - Troubleshoot and run up and downstream models in the DAG - Provide an isolated sandbox for individuals to make changes, separate from the whole team. #### Choosing dbt Cloud The Sunrun team had heard about dbt, and hoped that it could solve the issues they faced. They started a self-managed trial of the open source offering, dbt Core, just to ensure the workflow felt right for their business. However, Jed, the data team lead, quickly realized that managing an open source project that required command line proficiency, might still be limiting to members of the data team they hoped to activate for transformation work. So, he turned to dbt Cloud. With dbt Cloud, analysts could use the same development pattern as the central BI team and create their own data products. “The number one reason why we switched from the open source offering to dbt Cloud was the ability to more meaningfully collaborate across data teams and with business stakeholders,” said Jared. “Sure, we could have paired dbt Core with Airflow to add things like testing and version control—for the small number of folks that have python and command line skills. Quality would be up, but we wouldn't be moving fast enough for anyone to care. Instead, we have 40 analysts self-serving via the IDE, and business stakeholders answering their own questions with human-readable documentation.” ### Setting up the foundation #### Migrating existing models into dbt With dbt Cloud in place, the question now was, how much work would be involved? And how long would the migration process take? Turns out, way less than expected. “Us building out two complete data warehouses in two different platforms within a year and a half speaks volumes about the kind of velocity that dbt allowed us to achieve. Once the models were built out, moving them to a different platform was, for the most part, a pretty simple process…just porting over the code and making a few tweaks and syntactical changes that were necessary,” said Jared. #### Automating code deployment for SOX compliance With their data models in dbt, the BI team’s first priority was to eliminate manual complexity to reduce opportunity for error and speed time to production—a challenge when it comes to deploying code at a publicly traded company under the strict controls of SOX compliance. SOX compliance requires a rigorous checklist: testing, code review and approval, independent sign-off, system lockdown, and a history of all included checklist items—all things that can block automation efforts. “dbt Cloud enables all of the above by having a place where we can easily target code executions for only files edited. It keeps a history of the success or failure of jobs and more importantly, tests,” explained Jared. James and Jared created a GitHub Actions workflow that utilized Sunrun’s CI/CD tools: dbt for transformation and model generation layer, Jira for ticketing, and Github for code repository. “We're able to essentially create Git workflows and actions that check all of the Jira tickets to make sure they're in compliance with requirements and run the dbt models in the database, making sure they succeed before we allow the tickets to be deployed and go out into production.” ![Workflow](https://cdn.sanity.io/images/wl0ndo6t/main/034001e0ce4f55ba1dfbe055a8853dabc59e68ef-2500x1781.jpg) The combination of dbt’s version control and Jira and Github’s peer review process allows Sunrun to automate deployment while remaining SOX compliant. “Our deployment has sped up significantly—up to 75%,” said Jared. “We've become so dependent on the automation we've built and the ease of use that dbt Cloud offers that it’s hard to imagine ever going back to manual deployment.” ### Speeding up data development With dbt Cloud in place, the data team increased their velocity and data quality simultaneously. #### Leveraging the power of macros To keep up with the growing data scale, the team needed to develop and iterate quickly. "For me, the aha moment was when I realized the power of macros and how you could reproduce code automatically and dynamically generate SQL based on a macro that you defined,” recounted James. Using reusable chunks of code the team had already written and tested unlocked a level of velocity that was previously inaccessible. #### Speeding Q&A with automated data lineage and documentation Before dbt, troubleshooting was a tedious process requiring cascades of data engineering tickets. James recounted a scenario he faced often: “I’d spend an entire afternoon hunting down a broken link in the chain somewhere and trying to figure out what was going on that caused a report error. It was up to me to iterate through the code, going from file to file, from view to table, and table to view back and forth to see the dependencies and lineage.” The lack of visibility led to constant questions like "what's the source data for this dashboard?" "Why did this number change from last month?" "How are you calculating this?" James recalled, “I remember working with the collections department on this one specific Salesforce field…” They had several fields that did the same thing, with slight, poorly-defined differences that the collections team used separately and therefore didn't know where their data was coming from within Salesforce. “Those types of questions come up all the time and they're legitimate questions that should be answered,” said James. “So having dbt’s auto-documentation, having the ease of use of everything being in one spot, helps business units self-serve answers. If somebody has a question about whether something's up to date, dbt makes it really quick and easy to just check the refresh jobs and make sure that everything is running as expected.” At Sunrun, self-service not only involves building reports in a BI tool but also understanding where the data is coming from, the lineage of the source of that data… even how it’s calculated. “The documentation and dbt status is really worth its weight in gold and saves everybody time,” emphasized Jared. #### Upleveling data quality with testing To further reduce time spent troubleshooting, the Sunrun team leveraged dbt's four out-of-the-box [generic tests](https://docs.getdbt.com/docs/building-a-dbt-project/tests): not null, unique, accepted values, and relationships. “So far, we applied a standard set of tests across all of our models so nothing is built or delivered to stakeholders without having at least one level of testing on the table’s primary keys,” said Jared. The BI team is working towards expanding test coverage for alerts and notifications throughout the company when data quality is not up to par. “dbt tests add a ton of value with the clarity and quality they enable. As we talk to different business units and deploy dbt Cloud in their areas, they definitely see the vision there of ‘wow with one or two lines of code, I can make sure my data is quality before it even gets to the stakeholder’ which is huge,” emphasized Jared. ### Uniting data across the business Documentation as a hub for collaboration As dbt adoption spreads across the business, the BI team foresees documentation becoming a centralized place where anyone can gather information about the data contents and relationships within their warehouse. “Defining our data in a centralized place will be a central line of discussion in our governance committee meetings where we discuss what data means, how it's defined, how it's calculated… all of these things are an opportunity to use dbt documentation as a discussion hub in finalizing those decisions,” said Jared. ### The last mile: empowering analysts at Sunrun Those definitions start with the central BI team, but their “customers”—analysts embedded across business units—complete the data modeling process. “Our vision is to be the centralized place to get data, and then other business units would be the users who take it the last mile. Using the same system, they can give their stakeholders the same type of data dictionaries and the same documentation from a dbt standpoint so everybody's working on the same platform,” said Jared. #### The workflow 1. The business intelligence team first ingests everything. 2. They standardize the data, the field names, and the data types—flattening or otherwise preparing the data from an engineering standpoint. 3. Finally, they add the definitions and pieces relevant across the whole business. From there, the data is handed off to embedded analysts across business units who prepare the data for useful insights within their department’s context. 1. Senior analysts across electric operations, customer service, sales, install operations, commissions, accounting & finance, performance, etc., customize mappings only relevant to their departments, import spreadsheets to the data warehouse, and add their “special sauce” 2. With their completely modeled data, analysts then use their data to create Tableau reports for their stakeholders. For example, the performance team ingests IOT data that is created by Sunrun’s solar panels—the data about how much electricity they’re generating—and reports it back to the business. Additionally, they port individual installation performance to customers via their website portal. “It’s a huge data set that’s critical to ensure our panels are efficient and provide value to the customer. And it’s up to the analysts on the team to work with the dataset’s unique complexities and model appropriately to generate useful insights,” explained Jared. ### Ramping analysts of every skill set quickly The BI team is measuring initial success through adoption metrics, and the results are impressive. “We set up half a dozen GitHub repositories by departments and we check on their activity every week: we’ll pull up Github and show them the trends on adoption and collaboration, and most business units are on an upwards trajectory,” said Jared. More people committing their code to GitHub means greater dbt adoption and momentum towards quality, collaborative, governed data. “We have 50 dbt Cloud user licenses right now, at 100% utilization, which speaks to how easy it’s been for analysts to ramp. dbt’s inline web interface is very good, and the IDE enables exactly what we need the business to do. They’re excited to get into dbt, and there’s a queue of non-senior analysts wanting access too,” said Jared. ### Coming up: the BI team’s roadmap Sunrun’s BI team is continuing to make progress on the foundations of their dbt implementation: locking down and applying additional tests, onboarding new data sources and business units, building out documentation, and ensuring that documentation is widely available to anybody in the business. For example: ![Roadmap](https://cdn.sanity.io/images/wl0ndo6t/main/c073456d3f867294f35f1384a2a27755fce61425-2500x1932.jpg) With their central BI team to embedded analysts to business stakeholders model, Sunrun is also keeping collaboration top of mind. “We're in the process of building a dbt development community with regular meetings to share code, talk about techniques, and really build a cohesive place where people can share information and prevent developing in siloes,” said James. “We're on the edge of our seats to see where dbt is going and what further additions are coming up that we can take advantage of” added Jared. “There’s so much opportunity to improve, and we’re just getting started.” --- --- title: "Red Ventures reaches the right customers with data and AI" description: "This is the story of Red Ventures optimizes marketing campaigns to boost value for its clients with Databricks, Fivetran, and dbt" url: "https://www.getdbt.com/case-studies/red-ventures" date: "2022-10-14" industry: "Media" --- # Red Ventures reaches the right customers with data and AI This is the story of Red Ventures optimizes marketing campaigns to boost value for its clients with Databricks, Fivetran, and dbt ### Company details - Headquarters: Fort Mill, SC - Founded: 2000 ### Results - 100 hours saved on a typical data integration - 80% less time spent on data processing jobs - 30% more clients supported without increasing IT headcount > “With Databricks, Fivetran, and dbt, we can use data and AI in ways that help us reach the right customers with our clients’ marketing campaigns. That means we can deliver better results for their investment.” > > — Brandon Beidel, Director of Product Management ### Data ingestion across clouds causes a bottleneck When companies want to optimize their marketing strategy, they turn to RV. The [company’s Red Digital division](https://www.redventures.com/) provides end-to-end performance marketing services that help business-to-consumer (B2C) services providers attract new customers. RV uses a modern, and scalable, data solution to optimize search campaigns, social campaigns, display and landing pages in ways that will turn top-of-funnel marketing leads into repeat customers for their clients. “For some of our clients, we manage the entire customer journey,” said Brandon Beidel, Director of Product Management at Red Ventures. “Many clients don’t have the in-house marketing expertise to run sophisticated campaigns. It’s our job to help them drive customer growth over time, and we use data and AI to give them the best results possible.” For RV, delivering greater value to clients — and consumers — is about using timely insights from data to get in front of the right consumers. But the company doesn’t just think of the bottom line. Red Ventures wants customers and prospects to have an outstanding user experience as they explore products and decide whether to buy. “Today’s consumer expects to be able to complete the entire purchase process for most products online,” said Beidel. “But when they’re buying complex services, it’s challenging for businesses to get their message across effectively. Our goal is to create for our clients an optimal ad experience or many iterations of a website and then connect each of their customers with the user experience that will best help them understand the product. For us to tailor marketing to that level, we need to effectively use data and AI.” To ensure data integrity, RV maintains each client’s data in a separate cloud environment. From there, RV must integrate each client’s data and perform intensive transformations to generate machine learning predictions. Until recently, data engineers wrote custom scripts to ingest data for each client — a tedious and time-consuming task. ### A modern data stack helps simplify and automate ETL workloads Seeking to process data more efficiently, RV implemented Databricks to scale data engineering pipelines and speed up insights. The company also uses [Fivetran](https://www.fivetran.com/) to perform data ingestion and [dbt](http://www.getdbt.com) to apply data transformations. Databricks is now the engine that performs RV’s heaviest computing tasks. For each client environment, RV has set up a dedicated Databricks workspace, Fivetran connectors and dbt projects that only designated employees can access. All three solutions feed data into a machine learning pipeline that drives functions such as budgeting for clients’ advertising spend. “Due to the nature of our business, we must maintain a separate data stack for each client,” Beidel explained. “But Databricks, Fivetran and dbt help us avoid reinventing the wheel as we build out the processing, computation and data management pieces for each client. We’ve been able to take work our engineers have done and reuse it across multiple clients. We’ve also put Fivetran in the hands of our marketers to let them set up integrations for ingesting our clients’ data, and we’ve given them access to dbt to define data transformations. Because these tools are so simple, we often don’t need to involve our engineers.” Most of the onsite events from clients’ websites stream into RV. The company receives click data, event views, page views, scrolling depth and more into its warehouse. From there, RV has simplified and automated its ETL workloads. “Databricks has allowed us not to have to think about managing infrastructure,” said Beidel. “For many workloads, we simply select the compute size, define our script and let it run. Our data engineers don’t have to be experts in managing Spark clusters. With Databricks, they can focus on the shape of the data and the business logic.” Red Ventures has decreased its data processing and troubleshooting time — and continues to review data to improve the results they deliver for clients. ### Delivering better results for clients with less data work RV now processes client data more efficiently than ever. The company used to spend 100 to 150 hours per data integration. Today, that figure is 10 hours — which frees up data engineers to work on higher-value tasks. Processing jobs that ran overnight and took up to 20 hours now finish in four or five. And because Fivetran provides the same data set for every client and dbt gives everyone a clear view of data transformations, RV’s engineering team has reduced troubleshooting from 50% of its time to less than 20%. Automated cluster management capabilities in the Databricks Lakehouse Platform have yielded further time savings. “Not having to manage clusters is a huge time-saver for us,” Beidel remarked. “Our biggest grievance with other warehousing technologies is the lack of effective autoscaling. It’s a must-have capability when you’re working with a large volume of analytical workloads and have fluctuations in demand throughout the day. Having that variability in usage taken care of for us by Databricks probably saves one engineer out of 10.” This greater efficiency has enabled RV to support 30% more clients. “One of the things that drew me to dbt is that everything is in SQL,” Beidel explained. “It serves as a common language for our team. Not everyone can write SQL, but everyone can read it and have a better understanding of what’s going on underneath.” Because dbt and Fivetran are so easy to use, more RV employees are involved in data transformations. Engineers will often use dbt to show marketing staff how specific values were calculated — leading to deeper discussions about the data and to knowledge transfer from engineering to business users. This data democratization is contributing to better results for RV’s clients. “Using the highly intuitive machine learning pipeline we’ve built, we’ve helped our clients increase cost efficiency in some channels by 20 to 30%,” Beidel concluded. --- --- title: "Yummy & dbt create a new standard for startups in Latin America" description: "Learn how Yummy used dbt Cloud to unify data across four business verticals and streamline operations." url: "https://www.getdbt.com/case-studies/yummy" date: "2022-10-13" industry: "Industrial Automation" --- # Yummy & dbt create a new standard for startups in Latin America Learn how Yummy used dbt Cloud to unify data across four business verticals and streamline operations. ### Company details - Headquarters: Caracas, Venezuela - Solution: Super-app - Founded: 2020 ### Results - 5 analysts trained with no previous data engineering experience - 5x more models developed with dbt scheduler & staging environments - 50% reduction in model build time by enabling more people on the data team to safely develop > “dbt opens up the visibility and capabilities of our analysts so they can do more than just reporting. Because they need to do more. There’s no other way to be efficient enough.” > > — Luis Noguera, Head of Analytics ### Scaling fast [Yummy](https://www.ycombinator.com/companies/yummy-inc) was founded in April 2020 as a food delivery business in Caracas, Venezuela. Since then, it has become a full super-app with ridesharing, groceries, entertainment, retail shopping, and payments—becoming Venezuela’s largest tech startup and the youngest company in Y Combinator’s top company list. With less than two years in business, Yummy is now operating in 6 countries: Venezuela, Chile, Peru, Panama, Ecuador, and Bolivia. **The aim:** grow the business in Latin America as a full superapp for underserved markets. **The challenge:** 800 employees, 4 verticals, 6 countries, and 1 company acquisition led to many disparate data systems. “Our rapid growth and diversification are why it's so important yet uniquely difficult for us to create a single source of truth for our data,” said Luis Noguera, Head of Analytics. “As we scale the data team and the organization, our data infrastructure has to be scalable, self-service, and trusted by the entire organization.” ### Fragmented data, fragmented team To meet the growing business’ needs, Yummy’s small data team needed to be as full-stack as possible. Before 2022, the team was building everything as a scheduled query in BigQuery. “There was only one person modeling the data for four verticals, and it was a bottleneck that really slowed time to insights” described Luis. “Analysts were often repeating models that had already been built for different use cases,” resulting in tangled data made of single-use tables. There was no reusability or collaboration, and the bottleneck backlog kept increasing,” said Luis. Analysts struggled to create their reports without the context of how tables were built or modeled. Luis joined Yummy with [dbt](https://www.getdbt.com/product/what-is-dbt) experience and saw the team’s challenges as an opportunity. “The fact that we could implement dbt as a new data team was exciting. We had a lot of work to do, but there was so much potential for building our foundations with dbt. I think every new team should start out using it. The earlier the better,” said Luis. ### Ramping the team on dbt: from zero to one Analysts brought deep business context and critical thinking, but didn't have experience with transformation tooling previously only accessible to data engineers. #### Learning core concepts The team went through the [dbt Fundamentals course](https://courses.getdbt.com/courses/fundamentals) and leaned on the [community](https://www.getdbt.com/community/) to get started. “And just like that, they jumped into developing. It was hard work but very fast to ramp from zero to one” said Luis. “The hardest part was understanding data modeling. But seeing the DAG for the first time was when things clicked. We didn’t have that visibility with our scheduled queries,” said Luis. #### Standardizing a workflow Internally, the team established SQL conventions so everything moving forward was traceable and understandable. They implemented collaborative code reviews; when someone submitted their first pull requests, the team workshopped improvements and how to better adhere to the new SQL conventions. “It was the first time we had worked collaboratively as a team. And it really accelerated our work” said Luis. With dbt’s SQL best practices and automatic documentation, the team has been able to keep up with the company’s rapid growth. “ Before, we were building for now, or for the next hour. But by taking a little bit longer in the moment, we’ll reap the rewards of velocity in the future,” said Luis. #### Getting hands-on without risk As the team was getting comfortable with data modeling and the new workflow, the staging environment helped them learn by doing without risk. “[Analysts](https://www.getdbt.com/product/analyst) could do demo runs before pushing anything into production, which they couldn’t do before. dbt lets them get hands-on without messing anything up, and I think that’s what made our 3-month implementation so fast. Having a safe playground to develop and test was essential,” said Luis. ### Uniting data across verticals Once comfortable with [dbt](https://www.getdbt.com/product/dbt-cloud), the team focused on a major problem area: catching errors before production. “We were growing so fast that breaking things seemed unavoidable. As a result, we had business users answering business questions with in-app events-based data instead of data coming from documented and well-maintained tables,” said Luis. Without a single source of truth, timeliness was another issue. Yummy's early data architecture was distributed across multiple models built for the same or similar purposes, making analytics work slow and repetitive. This lack of speed also impacted customers. “Our app has notifications for when there aren’t enough drivers in an area. If our data models aren’t centralized, our ops team lacks visibility, and we’ll fail to meet the demand of the specific time and place” said Luis. The operational notification wouldn't be triggered, the problem wouldn’t escalate, and drivers wouldn’t get deployed to the area to meet the demand. By unifying their data models and establishing clear and well-defined data definitions with dbt docs and tests, the team reduced model build time by 50% and addressed their operational needs across business areas. ### Training future analysts with dbt With the success of Yummy’s first [dbt-trained analyst](https://www.getdbt.com/product/analyst) cohort, Luis is set on repeating the process with future local team members. “We want to hire analysts in Venezuela and the rest of Latin America. And the last three months have proved all you need is a willingness to learn. I’d prefer to train someone local with strong analytical capabilities than hiring someone who already has dbt experience” explained Luis. Yummy hopes to set an example for other Latin American startups and analysts. “Efficient, skilled data teams are essential to a tech startup scene, and we need more of them in Latin America. Leaning into dbt and analytics engineering is the best way we know how to build the future of tech here” said Luis. ### Coming up for dbt <> Yummy With Yummy’s core business data in dbt, the team is working on adding more data sources to their project to start to address business questions that were previously modeled through a tangled network of disconnected data sources, like CAC and demand flux. As they’re expanding their scope, the team is testing a self-service model to transition repetitive reporting away from the data team, and maturing their testing with more complex custom tests. “dbt has been a huge unlock for us. The possibilities with a strong foundation are endless, and it’ll sit at the center of our data strategy as we continue to scale” said Luis. --- --- title: "Whatnot leverages data to pioneer social commerce" description: "This is the story of how Whatnot uses dbt and Hex to unite data teams and speed data development" url: "https://www.getdbt.com/case-studies/whatnot" date: "2022-10-12" industry: "eCommerce" --- # Whatnot leverages data to pioneer social commerce This is the story of how Whatnot uses dbt and Hex to unite data teams and speed data development ### Company details - Headquarters: Los Angeles, CA - Data stack: Airflow, Python, Amazon S3, Snowflake, dbt Cloud, Sigma, Hex, Retool, Spell ### Results - 4-8x increase in speed from idea to production by enabling data analysts and data engineers to use common dbt models - 1 week for new analytics engineer hires to start shipping by speeding up onboarding with dbt data catalogs and DAGs - 10x decrease in maintenance costs by increasing data quality and accessibility > "dbt and Hex make the data development environment so much easier to work with than any other combination of tools...I can instead focus on scaling my team, and building the best live shopping platform." > > — Emmanuel Fuentes, Head of Machine Learning & Data Platforms [Whatnot](https://www.whatnot.com/) was founded in 2019 as a live stream platform for collectors to auction off coveted items like Funko Pops! and Pokemon cards. Last year alone, sales grew by 20x. Today, it’s [valued at nearly $4 billion](https://www.forbes.com/sites/laurendebter/2022/07/21/livestream-shopping-stays-hot-as-whatnot-takes-on-ebay-valuation-more-than-doubles-to-37-billion/?sh=318a7c1cc4ed) and offers shopping sessions in more than 100 product categories. **The aim:** Built the best live stream shopping experience by increasing the speed at which data is analyzed and pushed to the product. **The challenge:** A complex engineering-heavy data structure didn’t incentivize cross-team collaboration and made hiring in a tight market even harder. “Our previous stack of AWS Glue, EMR, Kinesis, Athena, and Airflow required knowledge of complex frameworks & languages,” said Emmanuel Fuentes, Head of Machine Learning & Data Platforms. “In order to scale, we needed a consistent end-to-end approach.” ### Fast growth and complex data Data was a foundation at Whatnot from day one. And it had to be. The product is a combination of Twitch and eBay. It’s a modern, mobile-first, influencer-friendly QVC. Data is necessary for the operation of the product and, being both a marketplace and a live video platform at the same time, Whatnot has a lot of data. A lot of very diverse data. ![Whatnot App](https://cdn.sanity.io/images/wl0ndo6t/main/d2aed26d483d931be79e9dc7b62443ac537d5d5a-1390x781.png) Data is also essential for platform safety (ensuring a healthy experience for everyone during live shows), business metric reporting, and machine learning (serving the right products, to the right users, at the right time). In China, live stream shopping has exploded in popularity and [is expected to reach $423 billion in sales by 2022](https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/its-showtime-how-live-commerce-is-transforming-the-shopping-experience). In the US, it’s gaining steam, fast. Last year, Whatnot's revenue grew by 20x, and it expanded into more product categories and ways of buying. The product was changing fast, data volume was growing at an exponential rate, and they needed more and more qualified data professionals to keep up. Their data stack, however—built wholly on top of AWS services like Glue, EMR, Kinesis, Athena, and Airflow—wasn’t providing the flexibility and speed they needed. It was an engineering-heavy workflow that required knowledge of complex frameworks and languages. This had multiple consequences. One of them was that data engineers were spending too much time on non-scalable activities like building custom crawlers and maintaining brittle CI/CD pipelines for Airflow DAGs. And, because this was a complex workflow, the engineers were responsible for the bulk of the work and analysts weren’t delivering value up to their full potential. For a data-led startup that was scaling, moving, changing and hiring quickly, its data workflow risked hampering its growth. In order to build the leading live shopping platform in the US, Whatnot needed to build their data-led foundations on a system that enabled them to scale. They required new data tools and processes to: increase headcount effectively, maintain the quality of the live shows, expand offerings, optimize algorithms, inform the growing number of stakeholders, and remain agile in the face of fast-paced pivots. ### The vision: SQL as the lingua franca Emmanuel Fuentes, Head of Machine Learning & Data Platforms, was tasked with improving the efficiency of data teams at Whatnot. He had a clear vision: to employ the analyst-friendly SQL as the lingua franca across the whole data organization. This change would: 1. Increase the output from analysts and rely less on custom data products built by engineering. 2. Enable greater collaboration between the different teams within the data organization. 3. Facilitate the growth of the data team, since SQL is a cornerstone language known by millions of people. Emmanuel searched the market for SQL-first tools that would allow analysts, machine learning and data engineers to all work together. “I looked for tools that would allow us to act as if we were 10x size,” said Emmanuel. “I wanted us to all use common features to be aligned on how calculations are made.” In June 2021, they started to migrate their data warehouse to Snowflake. At the same time, they signed with dbt, Hex and Sigma. In a quarter and a half, they had shifted everything over to this modern data stack. Today at Whatnot, dbt is used for modeling and documenting data. Sigma is used for dashboarding, where leadership and operations visualize and consume data. Hex sits in the middle, for the exploratory stage of analysis. ![Data Stack](https://cdn.sanity.io/images/wl0ndo6t/main/f2fbfb2ae0a67f8121216211158e4f593eddbb8c-980x561.png) “This workflow is really useful for business reporting, performance marketing, and machine learning”, said Emmanuel. “We could have the same setup with Airflow DAGs. But I’d need a bigger, more specialized team and our maintenance costs would be 10 or 20 times higher,” he emphasized. ### The benefits of collaborating with SQL With analysts, machine learning, and data engineers all using SQL and the same tools, collaboration between data teams improved, as did efficiency. #### Re-utilizing models At Whatnot, SQL is used and repurposed across the whole journey from research to presentation (BI) or production. Analysts and data engineers use Hex as an exploratory internal development environment. They explore models, prepare forecasts, and share findings with each other. Once they confirm the model is solid, they push the logic to dbt. Since dbt and Hex share a native integration, dbt docs and metrics can be accessed on Hex so users don’t need to switch between the two tools. In dbt, models are cataloged and exposed to their whole data organization. This exposure allows data and machine learning engineers to incorporate existing models further into other data products, algorithms or business reporting—increasing speed from insight to action by up to 8x. “Analysts, engineers, and machine learning are all working on similar problems. When teams use common data models, that enables a lot of efficiency,” said Emmanuel. “Someone might have built something for an analytics report that gave an idea to a machine learning person. Instead of having to rebuild it from scratch, they can reference their dbt model and then move on.” “This new workflow decreased the delta between an idea and publishing to production by a fourth or eighth,” said Emmanuel. “We’ve gained a lot of efficiencies because we use SQL across our models, transformation layer, analysis, and dashboards.” #### A bigger pool for hiring Another benefit of SQL as a common language is that it has allowed Whatnot to hire from a bigger, more diverse pool of candidates, like bootcamp and non-traditional college graduates. “When you're growing in a tight hiring market, you need really smart people that can scale very quickly,” said Emmanuel. “A lot more people know SQL than languages like Scala. A SQL-first approach to data pipelines, transformations, QA, and analysis is essential.” The community aspect of dbt and Hex further boosts the collaborative nature of the SQL-first strategy—leveraging community expertise to answer questions, encouraging continued learning, and even sourcing two new hires from the [dbt Community Slack](https://www.getdbt.com/community/). ### Shortening the path to business value By using dbt models and documentation as the base of their data workflow, Whatnot can now bring new external data sources online and make them usable within a few days—as compared to weeks or months before. “We can skip redoing this layer that people get stuck in—like dimensional modeling, Kimball models or Snowflake schemas—and go straight to business value,” said Emmanuel. Whatnot’s upfront investment in dbt and Hex to bring engineering workflows to the data team returns compound value to the organization. As they scale and hire faster and faster, the speed at which new team members onboard has a big impact. “With our new stack, it takes new analytics engineering hires just one week to learn our stack. They read the docs, open up the dbt DAG, explore our data in Hex, and are ready to push code to production.” ### Looking ahead: tackling the data complexities of category expansion In the near future, Whatnot is planning to launch new product options, formats, and ways of buying. However, each different product category has different data requirements, which adds complexity. “Because each business is so different—sneakers, vintage clothing, Pokemon cards, food and beverage—they all have different data assets being generated,” said Emmanuel. “That results in thousands of fields that we need to be constantly computing and articulating to stakeholders. As you go wide, that data diversity only grows.” Hex and dbt will enable Whatnot’s team to tackle that complexity, from bringing new data sources online in less than a week to expanding their library of reusable data assets for future work. “dbt and Hex make the data development environment so much easier to work with than any other combination of tools. Since it’s all native, we don’t need to wait for or build a custom adapter,” said Emmanuel. “I can instead focus on scaling my team, and building the best live shopping platform.” --- --- title: "Aktify democratizes data access with Databricks Lakehouse Platform and dbt" description: "T his is the story of how Aktify uses Databricks and dbt to eliminate manual tasks and errors from its data transformations." url: "https://www.getdbt.com/case-studies/aktify" date: "2022-09-16" industry: "Industrial Automation" --- # Aktify democratizes data access with Databricks Lakehouse Platform and dbt This is the story of how Aktify uses Databricks and dbt to eliminate manual tasks and errors from its data transformations. ### Company details - Headquarters: Menlo Park, CA - Solution: Customer Retention - Data stack: Databricks Lakehouse, dbt, Tableau ### Results - 80% reduction in data engineering hours - 95% reduction in time to onboard new employees - 6 figure annual savings in IT headcount costs > “Databricks Lakehouse Platform and dbt have eliminated the manual tasks and errors from our data transformations. We’ve been able to stand up new solutions for internal clients in half a day compared to 3-5 days previously.” > > — Brandon Smith, Director of Data and Analytics [Aktify](https://aktify.com/) aims to help its clients convert their customers through conversational AI. Clients can use a conversational AI agent to conduct thousands of SMS conversations with sales prospects at the same time. With Aktify’s solutions, the average customer generates $227,000 in revenue per year and achieves a 19x ROI. To fine-tune its AI agents to client needs, Aktify must be able to drill into massive volumes of data and find overlooked insights. Using [Databricks Lakehouse Platform and dbt](https://www.getdbt.com/data-platforms/databricks), Aktify has taken the manual effort and risk out of its data transformations, democratizing data throughout its organization. Stakeholders are making better-informed decisions—and Aktify is saving six figures per year on its IT operational costs. ### Complex data dependencies prevent data democratization Aktify’s customers use conversational AI agents to do the work of thousands of live customer service agents. Behind the scenes, Aktify works to make these AI agents as effective as possible. The insights Aktify and its customers need are hidden in the company’s massive data volumes. “A typical use case for our customers is to look at the data points in our software and wonder when and why their customers and prospects are dropping off their chat funnel,” said Brandon Smith, Director of Data and Analytics, Aktify. “My team and I help them figure out how they can best move leads from initial engagement to positive engagement, which is when they’re responding to texts and want to schedule a call with their customer's call center. ” To deliver this level of customer service, Aktify seeks to help data teams interact with data how and where they need it. A data scientist may need raw data, whereas an executive team would want all their data to be pre-aggregated so they can quickly find answers to their questions. Self-service is great, but in practice, you need the right tools to prevent bottlenecks. “I knew if we used SQL Server Integration Services (SSIS), everyone would be afraid to work with data in production because it would probably break something downstream,” Smith explained. “We would probably have people complaining that we broke their favorite dashboard—and fixing these kinds of problems can be difficult and time-consuming. We didn’t want complex dependencies at Aktify. We wanted to ‘democratize’ our data.” ### Taking the risk out of data transformations Seeking to minimize complexity around its data, Aktify implemented Databricks Lakehouse to make its data management more simple, flexible, and cost-effective. The company also began using [dbt](https://www.getdbt.com/product/dbt) to make data transformations easier and more reliable. Right away, Smith and his team felt more comfortable letting employees with minimal SQL skills work directly with data. Databricks makes data in production accessible to anyone on Aktify’s staff who might need to tap into it, while dbt provides [testing](https://docs.getdbt.com/docs/build/data-tests) and [documentation](https://docs.getdbt.com/docs/build/documentation) that prevent downstream disasters. “Data transformations are no longer a scary thing for Aktify,” Smith said. “From day one, we can let new employees loose on our data without wrapping a lot of red tape around their hands. With Databricks and dbt, we’ve been able to onboard employees into working with our data systems in half a day, compared to two weeks previously, and if they have some SQL experience, they’re usually comfortable with the systems in about three days. This wasn’t possible with any other system I’ve used in my career.” With Delta Lake on Databricks, Aktify ensures it will get reliable performance. The time travel functionality helps Aktify streamline troubleshooting so it can roll out new functionality releases more quickly. “With everything I have going on in Databricks and dbt, the time travel features in Delta Lake enable me to stop after making changes and see what actually changed,” Smith explained. “If anomalies surface, I can go back and spot when I introduced the issue and why. That makes fixing the problem so much faster and easier.” ### Building data pipelines in half a day With Databricks and dbt, Aktify now finds the insights its customers need in a fraction of the time previously required. Instead of having its data science team build complex scripts to query databases—a process that could take three to five days—the company uses Databricks Utilities to find answers in less than a day. When a manager in Aktify’s customer success organization needed to know how much revenue the company was generating from the first 30 days of an account, Smith found an answer in hours. “When colleagues come to me and ask for new tables, pipelines or extractions, I can crank them out in half a day,” reported Smith. “It’s incredible how quickly I can engineer data pipelines with Databricks and then use dbt to model the data so it will surface the right metrics. Our stakeholders see new information they couldn’t access before and they’re making better decisions, whereas before they always complained that they were flying blind.” Databricks and dbt help Aktify meet its data needs without bursting its budget. At one point, Smith was the only member of the company’s data team, but he kept up with the demand for new insights. “dbt cuts out so many of the little tasks I used to have to do,” he said. “Without dbt, I would need at least two more people on my team. It’s saving us six figures per year by scaling the impact of the people we already have. If you’re still using SSIS, you’ll never meet your data needs because you can’t pivot fast enough and it’s too error-prone. dbt is a must-have tool.” Smith compares [Databricks](https://www.databricks.com/) favorably to competing solutions. “Other platforms may give you similar query speeds, but that’s not the deciding factor,” Smith said. “With Databricks, you can stand up new solutions much more quickly because the open-source tooling removes barriers. That’s the kind of speed that’s most important to us.” Using the combination of Databricks Lakehouse Platform and dbt—as well as a one-click Tableau connector in Partner Connect that allows all users to spin up their own Tableau dashboards—Aktify is getting better data processing performance from a simpler technology footprint. “We no longer stage data in Snowflake because all our data, including about 85 gigabytes of operational data, is instantly available in the Databricks Lakehouse,” Smith concluded. “Partner Connect also helps our team discover new data and AI solutions, bringing us closer to our vision of organization-wide data literacy.” --- --- title: "Pepperstone creates data decision makers can rely on" description: "This is the story of how Pepperstone uses dbt to create data decision-makers can rely on" url: "https://www.getdbt.com/case-studies/pepperstone" date: "2022-09-08" industry: "Banking & Financial Services" --- # Pepperstone creates data decision makers can rely on This is the story of how Pepperstone uses dbt to create data decision-makers can rely on ### Company details - Headquarters: Melbourne, Australia - Solution: Forex Broker - Data stack: Amazon DMS, Redshift, dbt Cloud, Airflow, Tableau, Sagemaker ### Results - 30% increased speed to delivery through lineage, data democracy, and a new streamlined workflow - 80% decrease in inconsistent reports thanks to easier troubleshooting and tested sources and models - 10 inconsistent data sources identified and resolved per month, before any downstream issues > “Trust is so important because we are the experts in data analysis. And if you have a good level of trust, your insights are more likely to be robustly discussed.” > > — Sam Ellett, Lead Data Scientist ### A Major Forex Player Founded in 2010, [Pepperstone](https://pepperstone.com/en-au/) has rapidly grown to become one of the world’s leading foreign exchange (forex) brokers. The business now operates from offices across the globe, has more than 300,000 traders on its books, and handles more than US$12.5 billion worth of exchanges on an average day. However, with the impressive growth of the business came an increase in the amount of data Pepperstone’s data team needed to process. As the number of incoming sources began to stack up, the team found that their existing systems weren’t able to keep up with the growth. ### Growing pains At the time, Pepperstone’s data team was operating with an analytics schema in Redshift, created from a collection of sources in their data warehouse. The BI-reporting schema read from all the other source schema in Redshift, and the team would build out tables and views to then ingest into Tableau. “It was created when the team was a bit smaller and the company was a bit smaller,” explained Sam Ellett, Lead Data Scientist at Pepperstone. “At that point in time, it suited our needs, but as the business grew, we quickly hit the ceiling as new sources and new team members came in. As a result, we couldn’t scale our data sets effectively, and we started to lose a sense of lineage as the number of datasets blew out.” As the growing business added more and more data sources, it became difficult for the data team to have complete confidence in its output. “We would start to lose track of how an upstream change would impact a downstream data set that would be exposed in Tableau,” said Sam. “We would often have the same metric across various dashboards being different.” ### Consistent inconsistency These inconsistencies presented a significant issue for the Pepperstone team. Of course, every industry relies on accurate and trustworthy data, but there are few where reliability is more important than forex. Sam explained: “Pepperstone works in an industry where price matters and price can change every millisecond. Small inaccuracies can blow up to very large reporting issues.” These small inaccuracies showed up as the same metrics, such as revenues and retention numbers, appearing as different numbers across dashboards. These discrepancies began to undermine the data team’s ability to deliver accurate, clear reporting to the business. “Our business stakeholders lost confidence in our numbers,” said Sam, explaining that as a result, other teams within the organization tried to perform their own analytics without involving the data team—pulling numbers from older spreadsheets and potentially making decisions that weren’t backed by reliable data. “That confidence is really hard to build…and it’s really easy to lose,“ he added. This lack of confidence had consequences across the team. It became harder, for example, to interrogate unexpected data when it appeared on a dashboard. Was an unusual data point a sign of something interesting buried in the numbers? Or was it the result of an inconsistency in the system? “It’s good to be surprised if your confidence in your model is high because that’s an insight,” explained Sam. “It’s bad to be surprised if you have no confidence in your model…that’s just a search for the mistake.” ### A strong lineage Sam began working with dbt Cloud after a recommendation from a former colleague. After experimenting with a small project, he quickly recognized the potential benefits. “I realized dbt was good for us as soon as I saw the lineage,” he explained. “It was easy to educate others because it was visual. You could just read it; you could actually see it. ![Data Stack](https://cdn.sanity.io/images/wl0ndo6t/main/ee796a76bfe5f00aa897a9115e949b98cf7547e6-6637x3312.png) One of the team’s main goals in switching to dbt Cloud was to slash the number of inconsistencies in their reports, with an aim of limiting them to only a handful of cases each year. In addition, Sam and his team wanted to set up comprehensive test coverage of their most important reporting. Any mistakes needed to be caught before the reports are sent out - “it's better for us to delay a report rather than send out an inaccurate one,” he noted. The DAG and documentation that clarified Pepperstone’s data lineage have not only reduced errors by 80% but also allowed the team to scale its data sources. “We get new data sources a lot of the time,” said Sam. “We’ve been able to keep pace with the changes now because we have a really easy way to add or update documentation.” ### Realizing the benefits of dbt Cloud Since introducing dbt Cloud, the data team at Pepperstone has been able to keep up with the business’s rapid growth. #### Onboarding new team members with ease One of the issues the Pepperstone team encountered was that their legacy systems made it difficult to bring on new talent. New hires had to absorb fragmented context, which was time-intensive to teach and learn. “We’ve got a growing team worldwide and needed a better way to onboard people,” said Sam. “Before, we were very much at capacity.” Onboarding new data team members the old way slowed analytical capacity and hindered the team’s ability to scale with the company. After implementing dbt Cloud, easy access to information simplified the onboarding of new team members. Documentation became widespread and detailed, with systems in place for users to request any missing information from the team. “We’ve already managed to release company-wide documentation across our data sources, which we never had before,” said Sam. This knowledge hub, accessible to all, allowed analysts of all skill levels to plug into the right models and sources without the previously-required extensive context. #### Enabling analysts to do more The move to dbt Cloud has also allowed many existing team members to expand their analytics skills with Jinja templating. “A lot of people had reached a bit of a ceiling in terms of what they were learning in regards to SQL,” Sam explained. “It allows analysts to build macros and think a bit more programmatically in SQL, which has raised that ceiling.” #### Boosting velocity One of the most apparent improvements to Pepperstone’s workflow came from a 30% increase in speed to delivery. Beyond improving the team’s data modeling efficiency, “the increased speed gives analysts more time for insights and more time for discussion with stakeholders about the report that they’ve built, which should be the priority,” noted Sam. “We aren’t just building a data set; we should be talking about it.” #### Building a trusted source of truth With more people contributing to data modeling, faster, quality remained top-of-mind. The clear lineage and out-of-the-box test allowed the team to quickly identify any potential errors or inconsistencies, making it much easier to track any issues back to their source. “It makes our life easier if we’re able to identify root data quality issues in sources and raise that visibility to the engineering team who can fix the root cause,” said Sam. The ability to test for quality and easily find and resolve issues allows Pepperstone’s analysts to be more confident in their data sets and deliver insights without worrying about unreliable data. “Trust is so important because we are the experts in data analysis,” Sam explained. “And if you have a good level of trust, your insights are more likely to be robustly discussed.” #### Confident compliance Pepperstone holds trading licenses across the globe. Its international market presence means the organization is held accountable to a multitude of regulatory environments. From anti-money-laundering transaction monitoring programs to frequent standardized reports, failing to supply reliable data could cause regulators to impose financial penalties or withdraw their licenses. On top of strict requirements, earning and maintaining each license is further complicated by the fact that no two licenses have the same requirements. “So you have to be nimble, you have to be able to scale compliance, and you have to be able to handle many different requirements. “Working with dbt helps us achieve that,” explained Sam. ### What’s next for Pepperstone? Looking ahead, the Pepperstone data team is looking at improving its workflow and efficiency. As part of this, the team plans on performing meta-studies on its existing projects. “We’re going to run analytics on our dbt analytics, getting across more of the metadata reporting,” said Sam. “I want to show all our tests to the rest of the business.” Beyond looking to ensure that his team is the most trusted source of analytics, Sam also plans to enable other teams throughout Pepperstone to do their digging and research using dbt. “My job is to make it as easy as possible for the team to get their work done,” said Sam. “And dbt helps me do just that.” --- --- title: "SafetyCulture gets serious about company OKRs with dbt Cloud" description: "This is the story of how SafetyCulture united its data workflow to drive the company’s yearly objectives and provided scalable and trustworthy data with dbt" url: "https://www.getdbt.com/case-studies/safetyculture" date: "2022-04-22" industry: "Industrial Automation" --- # SafetyCulture gets serious about company OKRs with dbt Cloud This is the story of how SafetyCulture united its data workflow to drive the company’s yearly objectives and provided scalable and trustworthy data with dbt ### Company details - Headquarters: Sydney, Australia - Solution: Safety & operations management platform - Founded: 2004 ### Results - -20 to +69 Data team eNPS, an increase of 89 points after implementing dbt - 2x increase in new customer retention - 80% of data rebuilt using dbt > “The data team knew there was a better way. We needed to invest in building the foundations to be able to operate at the right level.” > > — Agnieszka Hatton, VP of Data & Analytics ### Safety lies in the data [SafetyCulture](https://safetyculture.com/) is a global software company, headquartered in Australia, that creates solutions for safety and operations management in the workplace. Their flagship product, iAuditor, captures data from sensors and reports from frontline workers so companies with distributed workforces can identify common issues, implement processes to prevent them, and make operational improvements. _“We are a data organisation at the core,”_ said Agnieszka Hatton, VP of Data & Analytics at SafetyCulture. _“Good data is essential for our customers to make well informed decisions about their operational and safety processes."_ SafetyCulture experienced rapid growth over the last year, 3xing the amount of data processed daily, and they needed to build a high performing team and scalable data architecture to keep up. ### A workflow problem _“We were using LookML for all of our transformations, which just wasn't scalable. There's also a limit to what you can do—no testing, no DAG... it required tons of supervision to ensure alignment with existing architecture,"_ said Oscar Lukersmith, Lead Analytics Engineer at SafetyCulture. At first the team tried using Airflow to schedule and orchestrate LookML models, but quickly ran into accessibility and architecture challenges. _“We had to manually build all the dependency graphs, and we ended up with our transformation layer sitting across two different tools. We were just adding more complexity to a system only a few folks had the skills to operate.”_ said Oscar. As a result, business stakeholders often didn't know where the source of truth for the data was. _“Where to go for the right data was confusing. It resulted in mistrust and a lot of wasted time”_ added Agnieszka. _“From an internal data team perspective, it also meant we had to redo things all the time. We lacked reusability, which dragged out time to insights and affected morale.”_ Oscar knew that part of the solution was technology, but there was more to it. _“Half of dbt is the tool but the other half is the workflow it enables,”_ said Oscar. While onboarding [dbt](http://getdbt.com/product/dbt), the team invested in new capability around data modeling. _“We brought in an experienced Data Modeler to work closely with Oscar and the rest of the team to design the future state data model. And then we used that data model to rebuild our data,"_ explained Agnieszka. **With the foundations in place, they set 3 objectives for the year:** 1. Transform SafetyCulture’s internal data analytics capabilities 2. Apply advanced analytics to strategic business use cases 3. Empower stakeholders to make informed decisions ### Transforming SafetyCulture’s internal data capabilities _“Internal data capabilities is the stuff under the waterline, where the work in dbt is setting us up for success. It’s often the things that the business doesn't see because it's below the waterline but they feel the impact,”_ said Agnieszka. ##### Redshift infrastructure The team redesigned their data architecture in Redshift to increase compute by 50% and disk by 100x at the same cost. The re-platforming reduced their average pipeline query time from 17 to 4 minutes, laying the foundation for the speed that’s essential to the data work. ##### Data model and transformation layer Consolidating all of their transformation work in one place, the team designed future state conceptual, logical, and physical data models in dbt. _“The work we've done using dbt to improve our data model and data transformation approach has meant that people can actually trust the data. We’ve now got these component parts that we can reuse, and as a result, speed has really improved over time,”_ said Agnieszka. ##### BI/Visualization Most of the internal capability work happened within SafetyCulture’s existing data stack, with the exception of migrating from Looker to Tableau. _“People love Tableau, and we’ve gotten some great feedback from our stakeholders in the business,”_ said Agnieszka—in particular calling out multi-use dashboards, an intuitive interface, and faster load times. ##### Increasing investment in the data team SafetyCulture’s data team grew from 6 to 15 people over the last year. _"Existing team members at the beginning were frustrated because they could see there were better ways to work. We needed to invest in the system, process and people capabilities”_ said Agnieszka. The work involved changing the operating structure, introducing Analytics Engineer and Data Management roles, and defining how to work together. _“Shifting into the dbt workflow was a huge shift. But 3-4 weeks in, you could see the moment when [Analysts](http://getdbt.com/product/analyst) went from learning each component to seeing the big picture,” said Oscar. “Part of the benefit is from dbt itself. But then the other half is the workflow it enables.”_ From January to December, the team’s eNPS went from -20 to +69—an 89 point increase. _“The team was part of creating the solution and it was really the Data Analysts and Data Engineers who drove the change. They created the human process together and are really advocates of dbt now,”_ said Oscar. ### Applying advanced analytics to strategic use cases One of SafetyCulture’s strategic objectives is to broaden its offering to deliver an operations management platform. _“We’re branching beyond our initial safety focus and building a platform that empowers our customers to improve their operations management. To do that, we need to understand how customers manage safety and operations today and how our software enables them,”_ said Agnieszka. ##### Mapping current customer trends To understand their customers, SafetyCulture needed to analyse a massive amount of existing customer data. _“We're using AI to better understand our customers, looking at about 150 different customer variables from demographics to product usage and behavior,”_ said Agnieszka. Answering questions like “who are our customers” and “which customers are most likely to expand or churn,” Analysts who previously only focused on one business area collaborated in dbt and identified 7 customer segments. These segments were applied in Salesforce and used by AEs and CSMs for customer expansion and retention. ##### Understanding obstacles for first time customers _“One of our biggest challenges was that we had a lot of new users coming into our platform but 97% of our first time users dropped off after 28 days. And we didn't really know why”_ said Agnieszka. Combining data analysis and customer interview insights, the team worked across the business to identify the key customer pain points and defined five initiatives to improve the customer experience. Six months later, the number of customers staying on beyond 28 days has more than doubled—from 3 to 6.5%. _“The increase in new customer retention has compounding effects on our active user numbers, with a projected 40% increase in MAU,”_ said Agnieszka. ##### Exploring commercial models The team also used dbt to explore new pricing models based on how customers derive value from the product and expand usage. _“We built a scenario modeling tool that enabled us to model different pricing constructs across our whole customer base. What if the pricing model was like this? What would that mean for customer expansion? How would it impact our revenue? We combined that with research, working with Qualtrics to add voice of customer data into the model,”_ said Agnieszka. _“Underneath that, the whole model was built on [dbt tables](https://docs.getdbt.com/docs/build/creating-models),”_ added Oscar. ### Empowering stakeholders to make informed decisions To ensure that data was embedded in daily decision making the team focused on empowering stakeholders to make decisions with that data. ##### Goal tracking and GTM reporting In the last 12 months the team developed clear goal tracking across the company. _“We worked with GTM and Product teams to develop goals and built the Tableau dashboards that they use on a weekly basis to track performance and to take corrective action,”_ said Agnieszka. _“Building the data underneath in dbt has enabled having a global unified approach to the way that we look at our data. For the first time, there’s now one way of approaching targets and drilling into global sales pipeline across geographies.”_ The dashboards inform pipeline, onboarding, and customer success tracking, supporting MAU and ARR growth. ### What’s next for SafetyCulture This year, the team’s strategic focus will be exploring 5-star rating and benchmarking so customers can look at how they perform compared to peers in their industry. Over time, the initiative will involve combining product usage and customer data to provide next-best-action recommendations in-app to enable customers to improve their safety and operations management. Internally, the team is finalizing its migration from Looker to Tableau, implementing a data literacy program across the business, and determining a ‘fit for purpose’ data governance program. As the team continues to build new data assets, dbt is being used to provide a standardised approach to transforming data to ensure quality, speed and reusability. --- --- title: "Fortune 500 oil & gas company embraces self-service analytics with agile data management" description: "Discover how a Fortune 500 oil & gas company ditched legacy data warehouses for agile, self-service analytics." url: "https://www.getdbt.com/case-studies/oil-and-gas" date: "2021-05-10" industry: "Oil & Gas" --- # Fortune 500 oil & gas company embraces self-service analytics with agile data management Discover how a Fortune 500 oil & gas company ditched legacy data warehouses for agile, self-service analytics. ### Company details - Headquarters: Oklahoma City, OK - Founded: 1989 ### Results - $10M reallocated back into the business - 3 weeks of work eliminated from regulatory reporting - 2x increase in the number of people collaborating on data modeling > "I remember reading the dbt viewpoint and thinking this is fantastic — a simple way to focus on SQL as a way to manage data objects and at the same time solve our scheduling and dependency problems. dbt was an off-the-shelf solution that took our ideas to the next level. It was revelatory." > > — Ryan Goltz, Principal Data Architect For Ryan Goltz, Principal Data Architect at a Fortune 500 oil & gas company, the demise of the enterprise data warehouse (EDW) began between 2012 to 2017. During that time, two technology trends were developing in parallel–big data platforms and broad adoption for [version control](https://docs.getdbt.com/docs/cloud/git/version-control-basics) and [continuous integration](https://docs.getdbt.com/docs/deploy/continuous-integration). _“CI/CD stems from open-source development as a means to ensure governance in a distributed development environment. This change complimented the technical capabilities and performance of the big data platforms,”_ Ryan said. _“The idea was to use these technologies to empower information consumers to become information producers.”_ Ryan was convinced that the value of big data platforms wasn’t the size of the data: _“Most of these people weren’t writing Spark. They were writing Impala views. This wasn’t big data, it was just big queries.”_ Instead, he saw the value of big data platforms as being in the new workflows. If you could apply software engineering best practices like code control and reusability to the data warehouse in a simple language like SQL, you could empower a whole new set of users to build and deploy their own data sets. ### Why the enterprise data warehouse couldn’t support self-service Like many in the Fortune 500, the company has a long history with the centrally-managed EDW. Over 10 years ago the team started to build a system for consolidating, cleaning, and organizing information from more than 70 operational systems into a single repository. _“It served its purpose,”_ Ryan said. However, the EDW would never allow for self-service. It came from a time that assumed the fewer people touching your data sets, the better. _“We understood that for the company to get value from our information we need to change how we work with that information,”_ Ryan said. He saw two barriers blocking people from ever contributing to the enterprise data warehouse: 1. **The tools to manage the EDW were inaccessible.** At the oil and gas company, an organization of nearly 2300 employees, there were 15 people on the team with the knowledge needed to modify or update the EDW. _“Provisioning information models in an Oracle data warehouse requires a lot of up-front knowledge about the data and how that data will be accessed,”_ Ryan said. _“Moreover, to implement that model, we relied on legacy tools like Informatica PowerCenter to perform incremental loads. There was a lot of complexity involved in building performance into the system just to accomplish our data loads and as the calculations became more complex so did the data loads.”_ 2. **Managing model dependencies required deep knowledge of the EDW.** _“We have hundreds of models in our EDW,”_ Ryan said. _“Managing these dependencies adds an additional layer of complexity. There is too much risk in asking a business user to understand these dependencies.”_ Ryan had a new vision for how data teams should work. The question he found himself asking was, “How do we provide a mechanism that allows more people to contribute to the value chain of the data warehouse?” The approach he imagined would have four things: _“First, you need to co-locate the data so that users can assemble models that span operational systems. Second, you need a really fast database with certain characteristics that make the system easy to use. Third, you need some tooling which allows models to be built in a language that the users understand. Finally, you need integration with workflow-based release processes.”_ This setup would unlock self-service for the company. ![Before and After](https://cdn.sanity.io/images/wl0ndo6t/main/e7a237a4b6c4212a22987f936e2e026565cb4605-1872x1500.jpg) ### How dbt + Snowflake enabled self-service For Ryan, the choice to go with Snowflake was an easy one. The team evaluated other warehouses, but he was sold on the ease of use: _“Snowflake makes performance easy for us. We rarely think about performance and when we do the solution is incredibly simple.”_ For an organization trying to get non-IT people to work with the data warehouse, this ease of use was a must-have. He was still looking for a way to solve transformation when a colleague sent him the dbt viewpoint. _“We had spent the last 10 years building systems which templatized the EDW development process and we firmly believe in the advantages of standardization,”_ Ryan said. dbt provided a standardized process, but packaged in a new workflow that was accessible to non-IT people. _“I remember reading the dbt viewpoint and thinking this is fantastic–a simple way to focus on SQL as a way to manage data objects and at the same time solve our scheduling and dependency problems. dbt was an off-the-shelf solution that took our ideas to the next level. It was revelatory. We took this idea to the EDW team warehouse team that had been doing PowerCenter work and we were like, guys, this is it. The spaceship had landed.”_ Together, [Snowflake and dbt](https://www.getdbt.com/data-platforms/snowflake) hit Ryan’s four most important requirements: 1. Co-located data that spanned operational systems 2. Data warehouse that didn’t require sophisticated tuning or complex maintenance. 3. Transformation process that required limited proprietary knowledge and where business logic could be expressed in SQL. 4. Workflow that allowed non-IT practitioners to apply best practices from the software development lifecycle to the data warehouse. The final piece of the puzzle was figuring out how to put an architecture in place that would preserve trust in the data. The primary reason that everyone at the company used the EDW was simple–they trusted it. The EDW was carefully governed by a central team. It had a set of common definitions and data quality was high. If Ryan was going to make self-service a reality, the team would need an architecture that retained the high quality of the EDW even as more people were building and deploying data models. _“Cool tools are just cool tools,” Ryan said. “But they wouldn’t do anything for us if we didn’t have the right architecture in place.”_ The company needed a logical architecture that would allow for _“innovation at the business domain level without propagating risk to everyone else.”_ Key data that every business domain needed would be managed at the enterprise level and owned by the enterprise team. Domain level knowledge would be managed by that domain. _“Snowflake allowed us to add segmentation in the logical architecture. This helps us manage the company’s risk while providing a clear path to expand usage as needed,”_ Ryan said. _“If someone needs data from another business domain, the enterprise team can apply the necessary data governance to move the appropriate dbt models to the enterprise level.”_ This architecture follows an important principle of software engineering–domain-driven design. This is the concept that code should match the business domain it supports. Ryan was able to implement domain-driven design using Snowflake and dbt by giving each business domain its own dbt project, Snowflake database, and a separate warehouse for processing. These domains are built around the business process—planning, operations, production—as well as corporate functions like finance. Within a given domain, business users fully own their dbt project. Users can also “reach into” the enterprise domain to reuse or extend enterprise business objects. ### The business impact of increased self-service Ryan is candid about the work that still needs to be done. Introducing the ability to self-serve your own data models comes with some challenges–there’s still an education gap on how to do this work really well, and he imagines implementing education sessions and quarterly business reviews for department groups to meet and review their dbt projects. But the self-service model has already proven to be incredibly valuable to company. Ryan points to a few key results: **Reallocated $10 million back into the business:** Energy companies have high operational costs. With more agile analytics, business users can better analyze these operational costs and forecasts. Mike Green, Manager of the company's Accounting Center of Excellence, says, "It can be challenging to articulate the value of empowering the business with an operating model and technology to analyze and innovate against data at scale. The capital accrual process is a great example using significant data volumes from multiple systems. With data collocated in a platform and an operating model that allows the business to analyze it, we've freed up millions of dollars that were tied up in the capital accrual process." **Eliminated three weeks of work on regulatory reporting:** The company routinely participates in audits for each state where the company operates. The states typically bring in one of the major accounting firms to assist in the audit which involves analyzing General Ledger data. _“The auditors would coordinate with IT to run and extract data in batches then load the data in their databases to rebuild the dataset. This process is tedious for all involved, and normally takes several weeks at best,”_ Ryan said. With Snowflake and dbt, this process has become not only better and faster, but also more secure. _“We make a model of the data needed for the audit, publish to a Snowflake Share and when the audit is complete, we turn off their access to the data. We are able to keep the data resident in our systems under our control and we’ve saved weeks of pain extracting data.”_ **2x the number of people collaborating on data modeling:** The EDW was previously managed by a team of 15 people. Today? _“There are currently three people building the enterprise structures inside Snowflake.”_ Instead, domain experts are now collaborating on data modeling. The team has nearly 30 people who contribute to this work–a massive expansion in raw data engineering capability. _“With Snowflake and dbt, the people who have the business problem now have the tools to go and solve their business problem. It creates a fundamentally different relationship with IT. We’re no longer the bottleneck or order taker,”_ Ryan said. With self-service analytics, the entire organization is empowered to get value out of the company’s data. --- --- title: "HubSpot empowers analysts to own their tools with dbt" description: "This is the story of HubSpot’s journey toward creating a more productive, happier, and scalable analytics team." url: "https://www.getdbt.com/case-studies/hubspot" date: "2021-05-10" industry: "Industrial Automation" --- # HubSpot empowers analysts to own their tools with dbt This is the story of HubSpot’s journey toward creating a more productive, happier, and scalable analytics team. ### Company details - Headquarters: Cambridge, Massachusetts - Data stack: Github, Snowflake, dbt Cloud > "We think empowering analysts to own their tools is the only way to build a productive analytics team at scale. dbt makes it easier to do data modeling the right way, and harder to do it the wrong way." > > — James Densmore, Director of Data Infrastructure [James Densmore](https://www.linkedin.com/in/jamesdensmore/) has spent the last 10 years of his career working in data. In his current role at [HubSpot](https://www.hubspot.com/), a customer experience platform used by over 78,000 companies, he’s the Director of Data Infrastructure. The biggest change he’s seen over the course of his career is how modern cloud warehouses have effectively solved some of the most challenging problems in data management. HubSpot was an early adopter of [Snowflake](https://www.snowflake.com/), a modern cloud warehouse platform. _“I can’t emphasize this enough,”_ James said. _“We don’t have a dedicated database admin or anything like that on the team because with Snowflake, we just don’t need one.”_ As an example, James points to dynamic scaling, _“We often have times where we have a critical backfilling job. With Snowflake you can scale up for that one job, run the backfill, then scale right back down. Dynamic scaling with Snowflake is something that analysts can do on their own -- no waiting for a data engineer.”_ Snowflake cloning is another previously sticky data management problem. _“Even five years ago, data teams would have maybe one dev instance and one staging instance of their warehouse,”_ James said. _“Snowflake is more like an ecosystem. You don’t even think of it as an instance because people can use cloning to spin up so many different databases.”_ This functionality means that anyone who knows SQL can quickly—and more importantly, safely!—work with data in the warehouse without breaking anything for someone else. Snowflake’s data lake serves as the central storage for all of HubSpot’s data, _“We’re pulling in data from our in-house databases, APIs, a lot of flat files sitting in S3, and Kafka streams. It all lands in our Snowflake data lake,”_ said James. Snowflake makes managing this data easy, but with that set of problems out of the way, HubSpot’s data team started to feel pressure in another part of the ELT stack–transformation. ### Attempting to solve data transformation with Airflow _“At any data organization in any company, you typically have a lot of analysts and fewer technical resources. This always creates a blocker to productivity. Whenever an analyst needs a new column or data grain, they have to go to a data engineer to get it,”_ James said. _“We’re an organization of over 3,500 people. If we need to hire a data engineer for every 2-3 analysts, that’s just not going to be cost effective. It doesn’t scale.”_ The early attempt at solving this blocker was using Apache Airflow to build and deploy SQL models. The combination of Snowflake and Airflow meant that highly technical analysts could build data sets without going through data engineers, but the process opened up new problems: 1. **A messy code base:** There was a lot of copy/pasting between models and the result was a code base that was getting harder to maintain. _“Even a little bit of copy/pasting makes maintenance difficult. If you need to change one line of code in a model, you need to find all the places where that code was copy/pasted."_ 2. **Difficulty determining model dependencies:** Analysts needed to manually define the dependencies of every query meaning that deploying new models was time consuming and stressful. _“In some cases the data engineering team needed to write custom tooling to determine dependencies.”_ As the number of SQL queries deployed in Airflow grew, so did the cognitive load for analysts who needed to make sense of all that code. 3. **Challenging to troubleshoot:** When deploying new models is difficult, analysts make an obvious choice–fewer models. _“Analysts were incentivized to have fewer models because the more complex the DAG is in Airflow, the more times you have to copy/paste the code,” James said. When something inevitably broke, finding the issue was painful. “An analyst would look back at a query after a week or month later and have no idea what it was all about.”_ Even with best-in-class data warehousing, analytics velocity remained slow. HubSpot still had a huge gap between where they were and where they wanted to be in their journey toward [empowering analysts](https://www.getdbt.com/product/analyst). ![old vs new](https://cdn.sanity.io/images/wl0ndo6t/main/689995d6c544f0f5016503e5a085438761a1af41-1560x786.png) ### Introducing dbt _“As an organization, we use the term ‘empowering’ a lot,”_ James said. _“We have a highly autonomous company culture. We hire smart, trustworthy people and we want to make sure they have the tools they need to get their work done. That’s what we’re trying to do on our data team as well. We think empowering analysts to own their tools is the only way to build a productive analytics team at scale.”_ Ultimately, it was the analyst community at HubSpot who discovered the secret to owning their own tooling–[dbt](https://www.getdbt.com/product/dbt). _“We had some early adopters who advocated for dbt. They came to us and said, ‘Hey, I’m trying this out. I’d love to use it.’ And we said, ‘Great. Let’s do it.’”_ Two teams in particular were eager to start using dbt so the infrastructure team got things up and running. These early adopters found that dbt empowered them in three important ways: **1. Empowered to do data modeling the right way:** _“We talk about guardrails vs. gates. We don't want to put up gates for people, but we do want to provide guardrails. Guardrails are best practices that allow people to move quickly and confidently.”_ dbt provided the technical infrastructure for analysts to own data transformation along with a set of best practices. _“dbt makes it easier to do [data modeling](https://www.getdbt.com/blog/data-modeling-techniques) the right way, and harder to do it the wrong way.”_ One example James points to is [incremental models](https://docs.getdbt.com/docs/build/incremental-models-overview). _“In the past, most analysts on our team weren’t comfortable with incremental models. So they would either work with a data engineer or just do a full refresh every day.”_ In other words, analysts had a choice–wait for an engineer until you could do it the right way, or do it the faster but more expensive way. With dbt, an analyst can configure an incremental model using the is_incremental() macro. Because dbt provides clear guardrails around how to build an incremental model, analysts are more likely to use them. **2. Empowered to define model dependencies:** With Airflow, defining model dependencies was limited to data engineers or the most technical analysts. In dbt, [model dependencies](https://docs.getdbt.com/faqs/Models/create-dependencies) are defined using a “ref” function to indicate when a model is “referencing” another model. dbt sees these references and automatically builds the DAG in the correct order. This intuitive approach to “layering” data models means any analyst who knows SQL is able to model data. ![dbt ref](https://cdn.sanity.io/images/wl0ndo6t/main/947d927ee1f9317ef22a1b26ab1f98b9c0411fa8-1400x979.gif) **3. Empowered to update and troubleshoot models:** The ease of referencing models ends up, _“totally changing the way people write SQL,”_ James said. When dependencies are easy, analysts start to break their queries into smaller pieces. This follows one of the oldest best practices in software engineering–modularity. Modularity makes it easier to update and collaborate on models as well as troubleshoot issues. Well-named models are organized into schemas. So, if your sales pipeline numbers aren’t updating one morning, instead of sorting through 200 lines of SQL lumped into one query, you can jump into the sales schema, find the troublesome model, and make the fix. ### Rolling out dbt to the analyst community at HubSpot The data team at HubSpot is a hybrid model. James’ team acts as the centralized resource owning infrastructure, tooling, and some core data models. Analysts are largely decentralized, sitting within a given business function. Today, two of HubSpot’s analyst teams have fully migrated to dbt. The first was HubSpot’s partner team. _“HubSpot works with a lot of sales and marketing service providers. The classic example is a digital marketing agency who works with HubSpot to both sell the HubSpot product to a new small business and help them really get set up on HubSpot,"_ said Ashley Sherwood, the Senior Data Analyst on HubSpot’s partner team. _“Our partner channel is an important part of our business, which leads to us asking a lot of questions about the program.”_ The question: “How many months since this partner’s last sale?” is straightforward. Things get complicated when you want to track changes to this answer over time. Using dbt, Ashley was able to fan out sales records by month, then group by partner. The resulting table showed one record per partner per month. With this [data grain](https://www.getdbt.com/blog/guide-to-data-grain), the business users on the partner team were empowered to use Looker to do their own iterative exploration independent of analyst support. This exploration resulted in a far more nuanced understanding of partner states such as who was active, dormant, or reactivated, and who was at risk of becoming dormant. _“We'd been able to report on how many new sellers that we had, but never the attrition of sellers falling off. This led to us bringing a lot of pieces together for 2020 planning in a really exciting and new way.”_ In addition to getting more analyst teams using dbt, James’ team is also migrating HubSpot’s testing infrastructure to dbt. “We have a legacy testing framework that works quite well for us. But being separate from dbt, it doesn't make a lot of sense,” James said. “We’re currently adopting the [dbt testing framework](https://www.getdbt.com/product/test-and-observe) so analysts will be able to write their own data quality tests as well.” Finally, they want to help all of the analysts at HubSpot to flex their SQL muscles. _“Most analysts learn SQL on a very small, simple SQL database. When you switch to a highly scalable, columnar database, it opens up new possibilities,”_ James said. _“This is about education. We want to continue building an internal community around Snowflake and dbt to empower our analysts to get the most out of what these tools can do together."_ --- --- title: "JetBlue eliminates data engineering bottlenecks with dbt" description: "Learn how JetBlue modernized workflows and infrastructure to democratize data access company-wide." url: "https://www.getdbt.com/case-studies/jetblue" date: "2021-05-10" industry: "Transportation & Logistics" --- # JetBlue eliminates data engineering bottlenecks with dbt Learn how JetBlue modernized workflows and infrastructure to democratize data access company-wide. ### Company details - Headquarters: Long Island City, NY - Data stack: Fivetran, Snowflake, Azure Blob Storage, Azure Data Factory, dbt Cloud ### Results - 99.9% uptime of the data warehouse and pipelines - 3 month migration of 26 data sources & 1200 models to dbt - $0 increase in total cost of ownership > "The new workflow with dbt and Snowflake isn't a small improvement. It's a complete redesign of our entire approach to data that will establish a new strategic foundation for analysts at JetBlue to build on." > > — Benjamin Singleton, Director of Data Science & Analytics [JetBlue’s](https://www.jetblue.com/) data infrastructure is managed by a central data team. This team maintains the data warehouse, owns data transformation, and ensures that data used for reporting meets compliance standards. Over the past few years, as data volumes have grown exponentially, the centralized data team structure has begun to reach its limits–there is too much data for a single team to own. _“Centralized data teams are often asked to be a catch-all for data related issues. We deal with infrastructure issues, business logic issues, data quality issues, and we have to be honest with ourselves that we create bottlenecks,”_ said Ben Singleton, Director of Data Science & Analytics at JetBlue. _“If we can let people take greater ownership of the data that they themselves are experts in, we’ll all be better off.”_ ### From central ownership to shared collaboration In Ben’s first week at JetBlue, he joined a meeting with a team that is one of the primary producers of analytic insights for the airline to discuss a number of data concerns. _“My welcome to JetBlue involved a group of senior leaders making it clear that they were frustrated with the current state of data,”_ Ben said. _“It made it real to me on day one that there was a lot of work to be done. That was my call to action.”_ In Ben’s mind, the legacy approach had to go completely: _“I knew that we couldn't just work on making our Microsoft SSIS pipelines incrementally better, or making configuration changes to our Microsoft parallel data warehouse to reduce our nightly eight hour maintenance window to six and a half hours. We needed to rebuild from the ground up.”_ Rebuilding included both data infrastructure and data workflows, with the goal of adopting a setup that eliminated the current data engineering bottleneck and gave data analysts and consumers at JetBlue a bigger role to play in owning their own data sets. _“Data engineers spend all of their time in data, but they’re not necessarily experts on how it’s used by different business functions to generate insights,”_ Ben said. He offers a seemingly innocuous example: the number of passengers on a plane. Simple, right? Not so fast. Do you mean the number of paying customers or number of seats occupied, which includes commuting crewmembers? Do you include infants in lap? Do you mean the total souls on board including the flight crew? What about pets?! The people with this nuanced understanding of the data frequently aren’t the data engineers–they are the analysts and decision-makers in specific business units. _“The only way we can support JetBlue in becoming a data-driven organization is if more people can participate in the data transformation process. My data team can’t scale fast enough to meet the needs of the company. We need to enable a distributed model of data management,”_ Ben said. The ability for an analyst to create the data assets they need to do their job is the secret to removing the data engineering bottleneck. If more people could contribute to data workflows, they would increase their own productivity while also reducing the burden on Ben’s team. A democratized approach to data access increases productivity, but more people involved in the data development process also raises concerns about risk: _“There are security considerations, Sarbanes–Oxley Act (“SOX”) compliance requirements, and risks to internal systems throughout the development process.”_ Ben said. The answer? Do both. The status quo just wasn’t an option. _“We don’t have much of a choice. Given the challenges we're facing as an airline during an unprecedented global pandemic, we have to do more and be more efficient with less. Not only do we support regulatory reporting requirements, we need to enable critical decision-making that will ensure our recovery from this crisis. We can’t give up on or deprioritize data. We have to modernize, and we’re doing it quickly.”_ ![From central ownership to shared collaboration](https://cdn.sanity.io/images/wl0ndo6t/main/dda1414efe330a660c49232f1b888e1c859b95c5-767x720.jpg) ### Choosing Snowflake + dbt _“I came to JetBlue knowing that introducing [Snowflake and dbt](https://www.getdbt.com/data-platforms/snowflake) would be the right game-changing move. I had bought into the vision that Tristan Handy and other folks in the dbt community articulated about modern data warehousing and the analytics engineering workflow. I still had to validate the technology and ensure it would meet JetBlue’s needs, but I trusted the approach.”_ Ben said. Ben saw dbt and Snowflake as the perfect fit for helping JetBlue adopt a more modern, collaborative data workflow that would eliminate the current data engineering bottleneck. Together, these tools enabled his team to overcome obstacles familiar to any data team as well as some particularly sticky problems unique to a publicly-traded, Fortune 500 company working in a highly regulated industry: #### 1. [Ensuring compliance and security](https://www.getdbt.com/security) All publicly traded companies need to comply with the Sarbanes-Oxley Act. An important part of SOX compliance, particularly for data teams, is to ensure that code is reviewed and data is validated when changes are promoted from development environments to production. Because dbt is built to leverage the git workflow, this process becomes much simpler. _“Version controlled code ensures that changes can be implemented following audit-approved processes. This is a new workflow to most analysts, but they are often very willing to adopt new processes if it means gaining greater control over their data.”_ Ben said. #### 2. Delivering real-time data JetBlue has two classes of data: batch data delivered at various intervals and real-time operational data. _“Airlines can succeed or fail based on how they respond in a snowstorm,”_ Ben said. _“JetBlue’s data infrastructure not only needs to help business leaders make long-term decisions, it needs to help crewmembers in system operations decide whether or not to keep the doors open an extra few minutes to accommodate a delayed connecting flight. These are high-pressure situations that require real-time data.”_ _“People might think of dbt as a tool that only supports batch workflows which update a few times a day or once daily, but dbt can be used for any kind of data transformation in your warehouse, including real-time use cases,”_ Ben said. To pull this off, the team used dbt to implement [lambda views](https://discourse.getdbt.com/t/how-to-create-near-real-time-models-with-just-dbt-sql/1457), which union historical data with the most current view of the data in Snowflake. Lambda views provide real-time access to flight information, bookings, ticketing, and check-in’s, among several other data sources–the data that ultimately makes or breaks the customer experience. _“Delays in operational data result in suboptimal decisions, and one suboptimal decision across a network of a thousand flights a day can disproportionally affect the customer experience and be expensive,”_ Ben said. _“Traveling can be stressful, especially in the time of coronavirus. We want to be really smart about using data to improve that experience.”_ #### 3. Catching data quality issues before they impact end users _“Previously, we monitored our data transformation jobs, but we didn’t actually monitor our data,”_ Ben said. They could tell if a job failed, but they couldn’t tell if the correct data was coming through. _“Missing data is inevitable, but the data team should be the first to catch it. Incorrect data is not only embarrassing, but can be costly. We frequently weren’t aware that something broke until our users told us.”_ [dbt makes testing easy](https://docs.getdbt.com/docs/build/data-tests), and as a result, people write more tests. _“dbt makes testing part of the development process. It’s just part of the work you do every day, so it doesn’t feel like an extra burden,”_ Ben said. And this daily work adds up. More testing improves data quality, [data quality improves trust](https://www.getdbt.com/product/build-trust-in-data-and-data-teams), and trust creates a happier and more productive data team. _“Catching data quality issues before your users helps the culture of your team. No one likes to have fingers pointed at them. It’s an overlooked benefit of dbt–what it can do for morale,”_ Ben said. #### 4. Implementing a workflow that was accessible for analysts Data availability was an urgent problem for JetBlue. Data pipeline jobs took eight hours to run, and during that time, the data warehouse was unavailable to analysts because the system couldn’t simultaneously handle the load of maintenance and user queries. _“Our availability hovered around 65%,”_ Ben said. _“In an age where people expect 99.9% uptime, 65% doesn’t remotely cut it anymore.”_ The legacy infrastructure required frequent repairs as the team worked to maintain pipeline jobs that were written 5, 10, or 15 years ago: _“We had to watch our Microsoft SSIS transformation pipelines and our APS warehouse jobs because they frequently and unexpectedly failed throughout the night.”_ All of this was a result of the explosion of data that occurred in the last five years. With Snowflake, a data platform built for modern data volumes, these maintenance windows are virtually non-existent, meaning data is available to analysts anytime they need it. _“My hope is that by making our data really reliable and solid, we can get more people interested in asking questions of the data and getting their hands dirty,”_ Ben said. _“We want our Excel power users to be motivated to learn SQL, and our SQL users to get inside the warehouse and build data models. We want our decision-makers to feel comfortable building their own dashboards and getting answers for themselves without always having to rely on analysts, even to answer basic questions.”_ Self-service success is dependent on accessible tooling. The first part of the Snowflake and dbt migration involved the data engineering team migrating existing data transformations. With a solid foundation in place, the team is ready to roll-out the new workflow to analysts with dbt Cloud. Ben is betting on dbt Cloud, which enables analysts to orchestrate the analytics engineering workflow from their browser, making data modeling accessible for analysts. _“Our first group of 20 analysts have been trained on Snowflake and exposed to the data warehousing concepts guiding our approach, and the next step in the plan is to have interested data analysts trained on dbt Cloud.”_ #### 5. Improving data transparency and documentation Airlines rely on a large number of specialized systems, most of which generate data that are analyzed to improve efficiencies. When Ben first began to explore JetBlue’s data warehouse, he found it difficult to answer even basic questions because he didn’t know what a column name represented, what source system created a data set, or how a metric was calculated. _“We need to approach data warehousing like any other product IT might develop. It needs to be intuitive, well-organized, and simple,”_ Ben said. Creating transparency in the warehouse starts with naming conventions–_“No more abbreviations or acronyms.”_–but also encompasses documentation. [dbt documentation includes information about a project including model code and tests](https://docs.getdbt.com/docs/build/documentation), information about the warehouse like column data types and table size, as well as descriptions that have been added to models, columns, and sources. _“Documentation is important for newcomers, and also important for getting advanced analysts involved in the data transformation process.”_ ### The journey toward data democratization In the first stage of the migration, the data engineering team migrated 26 data sources to Snowflake and dbt. These data sets represented JetBlue’s most critical operational and customer data. In the next three months, Ben anticipates that this number will double. _“We’ll be focused on building out our customer data models to provide enhanced ‘Customer 360’ insights. We have a fantastic loyalty program that we’re aggressively working to expand. My team is very excited to support those efforts through data,”_ Ben said. Ben’s hope is that with Snowflake and dbt his team will be able to solve the frustration he heard in his first week at JetBlue and achieve a 100% improvement in analyst NPS. _“I’ve been the angry analyst frustrated with the data engineering team in previous jobs. Now I’m the person who analysts will be pointing fingers at if they’re not happy!”_ Ben said. _“So my team is working extremely hard to provide a world-class data stack and regain the trust of our analyst community. I’m confident it’s going to be a radical transformation. The new workflow with dbt and Snowflake isn’t a small improvement. It’s a complete redesign of our entire approach to data that will establish a new strategic foundation for analysts at JetBlue to build on.”_