Accelerating and Scaling dbt for the Enterprise
Guide for large scale dbt projects. .
A curated list of awesome dbt resources
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
Guide for large scale dbt projects. .
A comprehensive canvas guide to dbt, from the basics to advanced topics.
A beginner tutorial to understand dbt with a real world example.
Paid Udemy course that covers theory, building a dbt project from scratch, and deploying to dbt Cloud.
Complete course covering both theory & practice through real-world Airbnb use-case.
Data engineering course on cutting edge tools including dbt.
Paid course offered by Uplimit covering the basics of dbt.
Another paid course by Uplimit covering the advanced dbt topics.
Official free course offered by dbt. Excellent for learning the basics of dbt Cloud.
Another dbt labs offered free course on dbt refactoring and CTE supercharging.
Guides you through a setup paired with Snowflake (decorated with extras).
An overivew of best practices on how to integrate dbt with Redshift, that includes information about performance tunning and dbt code optimizations.
Jinja Functions Cheat Sheet that covers the Jinja additions in dbt Core.
Use-cases and examples for the dbt Cloud Discovery API.
How to deploy dbt docs as a static website with App Engine and GitHub Actions.
How to get started with the team dbt workflow.
Combining ingestion-time partitioning and partition copy is a great way to achieve better performance for your models.
How to use grouped checkes in dbt-utils to keep our data "on track".
Use dbt & BigQuery dry run jobs to validate our 1000+ models in under 30 seconds.
Automatically generate ERDs and display in your docs site.
Best practices in Business Intelligence standards for integrating with dbt.
Jinja cheatsheet for dbt development.
Leverage Snowflake Zero-copy-clones to run slim ci checks.
How the dbt team structures its dbt projects.
Primer on how you should properly set up and configure your dbt workflow.
Yet another tutorial for using dbt Cloud.
Configuring Bigquery with your dbt project.
A dbt & Snowflake workshop on financial data.
Setting up teh integraton with Snowflake.
Leverage GitHub Actions to set up CI/CD with dbt Core.
Mono-repo or not mono-repo?
Pocket guide on optimization best practices with Snowflake.
An alternative guide to set up your dbt-external-tables workflow.
Standards for well organized base layer with Airbyte ingestion.
Tips from community members.
Standards for well organized base layer with Airbyte ingestion.
An overview about the package dbt-unit-testing.
A step by step guide on how to build Kimball dimensional models with dbt.
CLI tool to copy data from any source to any destination with a single command. Load data into your warehouse before dbt transforms it. Supports 50+ sources including Postgres, MongoDB, Salesforce, Shopify.
MCP for dbt CLI.
MCP tools to interact with dbt.
A curated collection of agent skills, built and maintained by dbt Labs, to help AI coding agents work more effectively with dbt.
Governed, multi-tenant MCP access to customer data. Turn your dbt models into a secure, per-customer MCP for AI agents.
Stream Deck plugin enables you to view the status of models and jobs as actions in your Stream Deck.
Automate and streamline the alerting/ notification process for dbt test results using this versatile CLI companion tool. Receive detailed alerts & test metadata seamlessly on various platforms, promoting improved collaboration on dbt project issues 🐞🚀.
Tabula is an end-to-end automation platform for data management tasks.
This repo gives some code to run dbt jobs/actions using modal which is a serverless application framework.
Expose warehouse dbt tests in CI to upstream data consumers so production changes never break the warehouse.
Gives a quick print out summary of changes so you can move fast and (not) break stuff!
Queries the dbt Cloud API to return some useful information about your models (number of tests, time they took to run etc…).
APIs, Caching, and Access Control on top of dbt Metrics.
Business Intelligence platform with deep dbt Cloud and CLI integration.
Raycast integration to monitor dbt Cloud Jobs.
Data Observaibility layer on top of your dbt + BI project.
Landscape of ML utilities around dbt.
Integration of Soda's data observability platform and dbt.
Offically supported database adapters.
Open source Looker alternative with deep dbt integration.
Open source visualization layer for your Modern Data Stack.
Getting started with the dagster-dbt library.
Add multi-language support (Python) to your dbt project.
Anonymize data in your dbt project.
Collection of Prefect integrations for working with dbt with your Prefect flows.
A unified data catalog, governance, and observability platform. Use it to view models, docs, test results, and column-level lineage across your dbt projects and downstream dashboards.
A rewrite of dbt docs using Next.js with React Server Components and SSG.
A dbt<>Superset connector that leverages Superset's API capabilities and dbt's manifest.
Eliminate the need for a data modeling semantic layer in BI.
How to leverage dbt as a data catalog and semantic layer (joins, synonyms, etc.) that BI tools can just plug into.
How Monzo built an extension framework for dbt.
Save by replacing dbt Cloud with GitHub Actions.
An experiment in productionizing LLMs.
A dbt<>Superset connector that leverages Superset's API capabilities and dbt's manifest.
Eliminate the need for a data modeling semantic layer in BI.
Reflection on one-year usage of dbt.
How to own your real-time transformation workflows like batch-based alternatives.
Behind the community of analytics engineers.
Behind the upheaval of the analytics engineer profession.
Using project artifacts to identify anomalies and room for refactoring.
Demonstration of integration with Azure.
Things you should be aware of when using external tables with dbt.
Best practises from the Vimeo Data team.
Generate known-answer seed and test data for dbt models: declare the expected aggregates (revenue curves, rates, rollups) and assert your models return them exactly.
Make simple storing test results and visualisation of these in a BI dashboard leveraging 6 Data Quality KPIs.
Stale data detection with dbt and BigQuery dataset metadata.
A dbt package that provides data anomaly detection as dbt tests.
Guide on how to run unit tests in dbt dynamically.
Port between dbt and great_expectations to extend out-of-the-box tests.
A dbt package for montioring metrics and detect anomalies.
Zero-config data quality CLI that complements dbt test with auto-detected anomalies (volume, schema drift, freshness, distribution, cardinality) on every materialized model after dbt run.
Fails the pull request when a model stops producing what a downstream consumer expects. Reads columns from catalog.json/manifest.json (falling back to a sqlglot parse of the SQL) and compares them against reverse-ETL destinations such as HubSpot CRM and against FastAPI/Pydantic services, in CI or…
Suggestions on testing your data powered by the community.
Unit test using source defer and generic custom tests.
Data breaks. Servers break. dbt and other tools break. Observability and alerting across and down your data estate. Save time with simple, fast data quality test generation and execution.
Agentic data quality framework that runs YAML rules against your warehouse (DuckDB, BigQuery, Snowflake, Databricks, Athena, Postgres), then uses LLMs to diagnose failures, trace root causes, and propose SQL fixes. Pairs naturally with dbt models as the validation layer after dbt run.
This setup is designed to demonstrate and implement best practices for testing and deploying dbt models.
This post dives into the use of CI/CD for dbt Core, providing insights on dbt Cloud's Slim CI CICD job pattern and how to implement this using dbt Core.
How to setup slim CI on Bitbucket.
A GitHub action for exporting dbt model docs to a Notion database.
Checkpoints to consider when reviewing an analytics engineer PR.
Great and detailed blogpost on setting up Slim CI in dbt Cloud.
A very tidy and fail-safe way to run dbt in production by using two parallel production enviromnents.
Limit data in long-running CI checks to improve developing experience.
A demo CI/CD implementation using GitHub Actions, including slim CI and unit tests.
A guide to implementing a complete CI/CD process for a dbt project. The defer feature (for slim CI) and unit tests are used for continous integration (CI). For continuous delivery (CD), automatic deployment is advised in lower environments, and the Write-Audit-Publish pattern using the dbt clone…
Showcase of advanced options when running CI for dbt.
A GitHub action for downloading dbt artifacts from dbt Cloud CI jobs.
terraform plan for dbt: diffs compiled SQL between two revisions to predict the DDL dbt run would emit. Names the columns an on_schema_change: sync_all_columns model will drop, the downstream models and data tests that drop breaks, and enforced-contract violations. Reads files only -- no warehouse…
Run and schedule dbt-style SQL transformations without Airflow. Adds data ingestion (50+ sources) and built-in data quality to the transformation layer. Open-source CLI or managed Bruin Cloud for teams who want dbt Cloud-like experience with ingestion included.
Run your dbt Core projects as Apache Airflow DAGs and Task Groups.
Leveraging the dbt manifest in Airflow.
Yet another article on extracting value from the manifest file.
Demonstration of a data orchestration project with Airflow.
Primer about dbt on Azure Data Stack.
Hosted dbt-core scheduler. Connect your repo, pick dbt-core and adapter version, set a cron, run dbt-core in isolated Docker containers. No Airflow to manage.
Open-source platform that runs dbt-core transformations alongside dlt-powered extract and load, with 36 connectors, cron scheduling, and run history in one web UI. AGPL-3.0, self-hostable with Docker Compose or Helm.
Terminal footer that pays you while dbt runs - one sponsored line on the bottom row during long commands (dbt Core and dbt Cloud CLI), 50% ad revenue share. Open-source Go CLI; never reads your code, queries, or output.
TypeScript SQL parser and static analyzer: type inference, schema diagnostics, and column-level lineage. Reads Jinja-templated SQL (dbt models) natively, without rendering.
CLI health check and linter for dbt projects (SQL, YAML, Jinja): 190+ custom rules, optional SQLFluff, 0–100 score, GitHub Actions CI, and coding-agent skills.
Gives coding agents instant dbt lineage: no dbt compile, no Python, no grep. Fast model-level lineage CLI parsing SQL files directly (or a manifest.json), with experimental column-level lineage too.
Open-source data engineering harness with 100+ deterministic tools for building, validating, optimizing, and shipping data products — usable from any LLM, across your warehouses. Ranked #1 on ADE-Bench (78%).
VS Code extension that puts a visual ERD designer inside your dbt repo. Two-stage (logical/physical) canvas, drift detection against manifest.json, auto-generated selectors.yml, and AI-readable semantic models (YAML/JSON) with a built-in harness for Claude, Copilot, Gemini, and Codex.
Rust-powered 'detective'/linter for dbt project/metadata best practices
Documentation Build Tool - Generate YAML documentation for dbt models with optional AI assistance. Built with Streamlit for an intuitive and familiar web interface.
Tool to configure and enforce conventions for your dbt project.
Tool to visualize the column level lineage of dbt models.
Various tools for working with data stacks.
Extract column level linage from dbt projects.
A sweet and speedy code generator for dbt.
Turn your diff into docs with the help of GPT-4o.
A local web application that provides a user-friendly interface to monitor and manage dbt runs.
Linter for dbt metadata.
RAG based LLM chatbot for dbt projects.
An AI-powered CLI tool for converting dbt SQL files to YAML using OpenAI.
AI teammate for engineers to ensure best practices in their SQL.
Automate the creation of dbt exposures from different sources.
The open-source Python library for data loading.
A handy docs composer and column-level lineage.
A dbt-core plugin to weave together multi-project dbt-core deployments.
A dbt-core plugin that automates the management and creation of dbt groups, contracts, access, and versions.
An LLM-powered chatbot with the added context of the dbt knowledge base.
A Column Level Lineage Graph for dbt.
Low-code application framework that turns your dbt projects into web apps.
A tool to help you stay in flow state while developing dbt models.
A self-updating dbt library that will maintain a list of current IANA/ICANN recognized top level domains.
A Streamlit web app to find currently running dbt models.
A Streamlit web app to explore the dbt Cloud API.
Feature Flags in dbt models.
A Neovim plugin for dbt model editing.
Cookiecutter template for dbt projects.
TurboVault4dbt is an open source tool that automatically generates dbt models according to datavault4dbt-templates.
Generate DBT Vault files from yml metadata (supporting dbtvault package).
All the basics to get a nice containerized dbt development environment.
DAG auditing tool that audits the DBT DAG and generates a summary report.
Makes your sql less bad.
CLI to generate DBML file from dbt manifest.json.
Generate dbt yml files using the CUE language.
It enables us to deal with catalog.json, manifest.json, run-results.json and sources.json as python objects.
This allows to always have the newest code commit running in the CI job without having to wait for the stale job runs to finish.
Unaffiliated python interface to various dbt Cloud API endpoints.
Enhance the developer experience significantly with workbench, output diffs, and YAML management.
Pytest dbt core is a pytest plugin for testing your dbt projects.
Generate lookml from dbt.
A version manager for dbt.
This tool formats your dbt SQL code so you don't have to.
SQL linter that supports dbt and Jinja templating.
Multi-dialect SQL linter and formatter in Rust. Postgres, MySQL, SQLite, BigQuery, Snowflake. Single binary, ~448x faster than SQLFluff on plain SQL, with SARIF output for GitHub code scanning and a baseline mode for legacy adoption.
Package to build GraphQL API on top of your dbt project.
Handy bash function to run changed models since last commit.
Things I wish I would have known when started working with dbt. Tools and hacks to improve developing experience.
Search dbt models interactively from terminal.
VSCode extension to give more clarity on model dependencies.
Jetbrains IDE plugin for dbt lineage and more.
Checklist on items necessary for a successful dbt project.
Developing styleguide often referred in PR templates.
Clean out warehouse models which are not existent in the project.
Excellent companion to your dbt practice with rich collection of tips.
Understanding the scopes of dbt tags.
Pre-commit hooks for checking data integity before schema change commit.
An open UI for dbt providing model browsing, lineage visualization, run orchestration, documentation, environment management — without vendor lock-in. Designed for local, on‑prem, and air‑gapped deployments.
Python library for generating models from a declarative YAML configuration. It supports multiple data modeling approaches including Data Vault 2.0, Anchor Modeling, and Dimensional Modeling, with template packages for popular dbt libraries.
A modern web-based user interface for dbt-core projects
Open-source notebook for running, exploring, and debugging dbt projects with interactive data analysis.
Currency exchange rate macros and models for dbt, powered by the UniRate API.
This package lets you materialize Semantic Views via dbt and reference them from downstream models.
Feature Engineering in dbt.
A dbt package for Snowflake Dynamic Data Masking.
A dbt package for Snowflake Streams.
A dbt package for monitoring airflow DAGs and tasks.
Data-diff solution for dbt-ers with Snowflake.
Tag-based masking policies management in Snowflake.
Generate dbt tests based on sample data.
Takes dbt runs and turns them into OpenTelemetry traces.
Package to assert rows in-line with dbt macros.
Write your dbt models using Ibis, the portable Python dataframe library.
The TimescaleDB adapter plugin for dbt.
Daily updated fake data for dbt learning and projects.
Package to calculate dbt Cloud usage-based cost.
A dbt package containing reconfigured macros.
A collection of dbt macros for working with Census data.
A dbt adapter for working with Microsoft Fabric Data Warehouses.
A code package from Microsoft for enabling dbt to work with Synapse Spark in Microsoft Fabric.
Dbt Fabric Spark Notebook Generator by Insight based on/forked form dbt-fabricspark by Microsoft.
Data Quality Observation of Data Vault layer.
Translate numbers into words.
A dbt adapter for working with Excel.
Linear regression in SQL using dbt.
Automatically tag dbt-issued queries with informative metadata.
Yet another package to monitor Snowflake usage.
Provides insights on the database/table level usage informations from Snowflake.
Package for dbt that allows users to train, audit and use BigQuery ML models.
This repo represents my attempt to build a fast version of DBT which gets very slow on large projects (3000+ data models). This project attempts to be a direct drop in replacement for DBT at the command line.
A dbt package to help you monitor Snowflake performance and costs.
Macros for staging and creation of all DataVault-Entities you need, to build your own DataVault2.0 solution.
Perform DataOps & administrative CI/CD on your data warehouse.
Checks that columns defined in YAML also exist in SQL.
A command-line tool and Python library to efficiently diff rows across two different databases.
This package highlights areas of a dbt project that are misaligned with dbt Labs' best practices.
Generate database constraints based on the tests in a dbt project.
Date logic and calendar functionality.
Macros to make it easier to protect your customers' data.
General macros and helpers.
Macros to support secondary calculations and generate business metrics.
Model synchronization from dbt to Metabase.
CLI tool for generating a scaffold for your dbt project.
Data profiling and doc block generator.
General macros library. A must have.
Macros for data audits that compare columns values and schemas between tables.
A SQL port of python's scikit-learn preprocessing module, provided as cross-database dbt macros.
Macros to stage your external sources.
Macros to build a feature store right within your dbt project.
Macros that generate dbt code, and log it to the command line.
Create a project and populate as much of the dbt project as possible.
This package builds a mart of tables from dbt artifacts loaded into a table.
This packages generate ERD diagrams from a dbt project.
IAC in dbt Cloud via Terraform.
Generate Looker views for dbt models.
Checks dbt docs and tests coverage.
Yet another coverage testing.
Push and pull metadata between dbt to Superset.
Package for generating and executing ETL for Data Vault 2.0.
CLI for creating, updating, and deleting dbt property files.
Package which contains macros to support unit testing.
A free to use dbt package with helper macros
Macro to join dbt snapshots.
Generate an ERD via dbdiagram.io from a dbt project.
A conference for data teams.
A survey of pains, gains, and areas of investment for global data teams.
Official TikTok channel of dbt Labs.
A Slack community of aspiring analytics leaders discussing and sharing lessons learned and challenges from their experiences in using data.
Global online community of data enthusiasts. Podcasts and blogs, etc. are distributed with high frequency.
Weekly substack about metadata, the metrics layer and MDS.
Great curated list of upcoming data analytics conferences.
Worldwide community driven analytics conference with a handful of talks fitting to the dbt stack.
Automatically generate ERDs and display in your docs site.
Recordings of Coalesce conferenfes from 2022 and after.
Second iteration of the analytics engineer conference.
Annual dbt conference full of fascinating use-cases.
List of community led dbt meetups.
Official dbt Labs newsletter on topics of the MDS.
Tought-provoking reads from founder of Mode.
Weekly newsletter of recent trends in Data Engineering.
One of the most popular data engineering podcasts covering great concepts and new products.
Official podcast of dbt Labs.
Energy-filled hub of analytics engineers (Highly recommended).
Subreddit of data engineering topics.
Special guests discussing big data, business intelligence, modern data stack.
Newsletter about dbt, with tips and tutorials on various topics.
Production-ready template deploying dbt project to Snowflake using GH Actions.
Add 4 more integrations to your dbt CI pipeline: Slim CI, pre-commit hooks, Data Diffs, and Slack notifications.
Using the official NBA API, this repo explains how to ingest, store, transform, and serve insights.
A curated list of awesome public dbt projects.
MDS in a box deployed anywhere.
A demonstration of production gatekeepers in Snowflake and BigQuery.
Demo project for open source MDS.
F1 Data Pipeline.
E2E dbt project for scraping and transforming football data from Transfermarkt.
A dbt project for GitHub Archive data on BigQuery.
A workspace template for dbt demos.
A dbt project to monitor cloud costs.
Repo containing data and dbt template of the survey.
Explore the topic of fake GitHub stars.
Dagster's ability to create a global dependency graph between different dbt projects.
GitLab's open source dbt project.
A worked example to demonstrate how to model customer attribution.
A worked example to demonstrate how to model subscription revenue.
Set up your dbt environment with pre-installed extensions.
Dbt-Airflow-GreatExpectations Stack.
A self-contained dbt project for testing purposes.
Sample dbt project with Spotify user data.
Deploy BigQuery + Airflow.
Demonstration of Airflow integration.
How to build a small and modern data infrastructure.
A production grade ELT with tests, documentation and CI/CD (GHA) about french open data (housing, demography, geography, etc). Can be used to learn with voluminous and ambiguous data. Contributions are welcome.
VoltAgent/awesome-openclaw-skills
The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
awesome-dsh-plugin/awesome-dsh-plugin
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
Kristories/awesome-guidelines
Programming style, best practices, and coding conventions.
sindresorhus/awesome
😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]
ai-boost/awesome-prompts
Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.
matiassingers/awesome-readme
A curated list of awesome READMEs