FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Tools

dbt: when client-site SQL needs git, tests and history too

For a forward deployed engineer, the tool's value is not the SELECT statement. It is what you leave behind once you have walked away from the client's warehouse.

In brief

  • dbt turns plain SELECT statements into modular models, with version control, CI/CD and documentation
  • A test passes when its query returns no bad rows; incremental models and snapshots handle large and frequently changing tables
  • The v2 generation is written in Rust, catches SQL errors before they reach the warehouse, and runs on an Apache 2.0 open-source runtime
ShareLinkedInFacebookX
Four cards linked by arrows: 01 SELECT model (stg_orders.sql contains only a SELECT from the raw table), 02 four built-in tests highlighted in orange (unique, not_null, accepted_values, relationships; zero rows returned means the test passes), 03 Incremental (processes only new or changed rows, cutting run time), 04 Snapshot (Type-2 SCD for the customers table, keeping old versions). Bottom bar: wrapped in git, review, CI/CD and docs; tests run before production.
A dbt test passes when it returns no rows. Four steps wrapped in git, review and CI/CD are what let a client's analysts fix the pipeline themselves.

A dbt test passes when its SQL returns no rows. It sounds backwards, but it captures the tool’s philosophy: you write queries that hunt for bad data, and silence is good news.

For a forward deployed engineer, that small detail says a lot. Picture yourself at a client site, facing a warehouse full of raw tables from which your product has to read clean numbers. What keeps the project alive after you leave is not clever SQL. It is how that SQL is managed.

dbt Labs, the team behind dbt, describes it as a tool that turns raw data in the warehouse into trusted data products. The mechanism is simple: you write plain SELECT statements, and dbt assembles them into modular, maintainable models.

What separates dbt from a folder of scripts is everything wrapped around the SELECT: version control, CI/CD and documentation applied to analytics code, so that anyone on the team who knows SQL can contribute to the production pipeline.

Why do FDEs need it more than in-house data engineers?

Consider a familiar situation. The client has a raw orders table and a customers table that changes often, and wants your product to read clean figures every morning. The quickest route is to write a few SQL scripts, put them on cron and fly home. Three months later someone edits a column by hand, the pipeline breaks and nobody knows why.

dbt exists to prevent that. When logic lives in versioned models, every change has a history, a reviewer and tests that run before it reaches production. That is why it suits FDEs: you need to hand over something the client’s analysts can fix themselves, not something only you understand.

SQL scripts left on site

  • Logic scattered, nobody knows who changed what
  • Errors surface only when the client sees wrong numbers
  • Only the author dares to touch it

A dbt project

  • Each model is a versioned, reviewed SELECT
  • Tests hunting for bad rows run before production
  • Anyone on the client team who knows SQL can change it

A minimal example, from model to snapshot

Start with the first model: a file, stg_orders.sql, containing only a SELECT that pulls id, customer_id, amount and status from the raw table and renames columns for consistency. No CREATE TABLE, no DDL. dbt handles that heavy lifting.

Next come tests. dbt ships with four generic tests: unique, not_null, accepted_values and relationships. You declare that id must be unique and not_null, that status may take only a few predefined values, and that customer_id must exist in the customers table. Each test is a query that looks for violating rows; if none come back, the pipeline moves on.

When the client’s orders table grows to hundreds of millions of rows, you switch the model to incremental. After the first run, dbt transforms only new or changed rows. The dbt documentation states plainly that this limits the amount of data to process and sharply cuts run time, which is exactly what the client asks about first when the compute bill rises.

For the customers table, where addresses and service plans change over time, snapshots are the answer. A snapshot implements a type-2 slowly changing dimension over a mutable source table: every time a row changes, the old version is kept. When the client asks “which plan was this account on last month?”, you have data rather than a shrug.

The new generation catches errors before they reach the warehouse

The current generation of dbt, known as v2, is written in Rust and understands SQL across the dialects of different engines. The practical consequence: it catches SQL errors before a statement reaches the warehouse. At a client site, where every failed run costs both compute and trust, that is the most valuable feature.

v2 is the default experience when you install dbt, built on an open-source CLI runtime under the Apache 2.0 licence. dbt is free; some features unlock when you sign in to a dbt platform account.

Limits to be honest about with the client

dbt starts from raw data already sitting in the warehouse; that is the starting point dbt Labs itself describes. So before writing the first model, ask the client whether the source data is already in the warehouse and how it got there, because that question lies outside the SELECT statements dbt manages.

Tests catch only what you declare. The four built-in tests check structure, not meaning: negative revenue passes not_null without complaint. You have to write additional business tests, ideally together with whoever understands the data best on the client side.

Incremental models process only new or changed rows, so agree with the client from the start on what marks a row as “new” in their system. Treat it as a design decision to discuss with the client, not a line of configuration.

What to learn first

Solid SQL is a prerequisite, not an advantage. What needs practice is the engineering discipline around SQL: git, review, automated tests. On a CV, a line such as “built a dbt project with tests and snapshots for customer tables” says more than “proficient in SQL”.

When reading job descriptions, treat phrases such as “analytics engineering”, “data model” and “warehouse” appearing alongside FDE as a signal to have a story ready about a dbt project you have built.

Imagine an interviewer hands you a raw orders table and asks what you would do in your first week on site. The credible answer is not “write SQL” but to walk through the cycle: a SELECT model, the four tests, a snapshot for the table that changes, then an explanation of why the client’s analysts must be able to fix it themselves.

Practise telling that story in five minutes, with a real project on your machine.

Clients will not remember the SQL you wrote. They will remember that the pipeline still runs, and the tests are still silent, long after you have left.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
6 sources
Read next on the roadmap · Stage 2: Broad engineeringAWS Cloud Practitioner Essentials: 13 modules that teach you to talk to a client's IT teamAWS's beginner cloud course will not turn you into an infrastructure architect. It will help you keep up in your first meeting on a client's systems.