Fundamentals of Data Engineering, four years on: the map to take on customer sites
Data tools have changed several times since 2022. The six "undercurrents" that Joe Reis calls the most useful part of the book still give a forward deployed engineer the right questions for a first customer meeting.

In brief
- The book came out on 26 July 2022. Four years later, Joe Reis says he would write it largely the same way again, with a few caveats.
- Most of its value is in the six undercurrents, not in any particular tool. Some of the storage sections go into a lot of detail and can be skimmed.
- For FDEs, the book gives a way to ask the right questions and to trace bugs step by step, which matters most when a customer's environment spans several clouds.
“Four years ago yesterday.” Joe Reis opens his look back at Fundamentals of Data Engineering with exactly that line. The book came out on 26 July 2022, and his post went up on 27 July 2026.
He asks himself whether he would write it differently today. His short answer is that he would mostly keep it as it is, with a few caveats.
Data tools change quickly. When an author still stands behind his framework four years later, that is worth noting.
It matters even more if you want to work as an FDE. You will not get to choose the stack on a customer site. What you can bring is a way of looking at systems, and that is what the book by Reis and Matt Housley teaches.
Five stages, one way to debug
The core of the book is a five-stage data lifecycle: Generate, Store, Ingest, Transform, Serve. It sounds like a slide diagram, but it proves its worth when something breaks.
Imagine you are an FDE at a retail chain. The head of operations complains that this morning’s revenue report is lower than the figures the cashiers sent in. Many developers would open the dashboard, which is the Serve stage, and start checking the SQL.
The lifecycle tells you to work back from the start instead. When does the stores’ POS system (Generate) finish syncing? When does the job that pulls the data in (Ingest) run?
Say the POS finishes syncing at 2am and the Ingest job runs at 1am. Every transaction since the previous sync will be missing from the report, even though your SQL is entirely correct.
The bug is not in Transform or Serve, so changing the SQL fixes nothing. The sensible fix is to enforce ordering: Ingest runs only after the POS reports that its sync is complete, not at a fixed time.
The five stages give you an order in which to rule things out. On a customer site, that order can save hours of guesswork.
The six undercurrents are where the value is
The example above is really an Orchestration problem, and Orchestration is one of the six “undercurrents” the book describes. They are not stages. They run through every stage, and Reis himself calls them the most useful part of the book.
For an FDE, the undercurrents are valuable because they turn into questions you can ask on a customer site. Applied to the retail chain above, each one becomes a specific discovery question:
| Undercurrent | Suggested question for a retail customer |
|---|---|
| Security | Does the POS data include customers’ phone numbers, and who can read it? |
| Data management | Does “revenue” mean the same thing in every department? |
| DataOps | When a pipeline breaks, who is told, and how quickly? |
| Data architecture | Which systems does store data pass through before it reaches a report? |
| Orchestration | Which jobs must finish before which others, and who checks that order? |
| Software engineering | Is pipeline code reviewed, tested and version-controlled? |
A tip for your first discovery session: do not arrive with a feature list. Arrive with six questions like these and note where the customer hesitates. That is usually where the project will break.
When the bug spans two clouds
Now take the example one step further. Suppose the retail chain stores its POS data on one cloud and runs its dashboards on another. Google Cloud calls this kind of setup multicloud, meaning a combination of at least two public cloud providers, and one reason it gives is to avoid being tied to a single vendor.
The timing bug above is now much harder to spot, because Ingest and Serve live in two places with two sets of logs.
Start by checking the point where data leaves one cloud for the other, with three questions. When does the transfer job run relative to when the source finishes writing? Do both sides record times in the same time zone? If a transfer fails, who gets the alert, and on which cloud?
If someone suggests moving that transfer job to serverless to tidy things up, remember that Corey Quinn has written that the promise of serverless has, unfortunately, still not been fulfilled. Before you agree, ask whether it actually solves the problem of run order and alerting at this boundary.
On AI/ML, Reis is still undecided about adding it as a seventh undercurrent, but he is clear that AI/ML will feature heavily in data engineering either way. For FDEs deploying agents or models for customers, this is a reminder that output quality still depends on the five stages upstream.
Who should read it, and in what order
Johnny Winter, writing on Greyskull Analytics, recommends that anyone working in data, not just data engineers, should own a copy. For backend or full-stack developers aiming for FDE roles, it is the book to read before any tool-specific documentation.
Winter also points to a weakness: some of the storage sections go into a lot of detail. Read the lifecycle and undercurrents sections closely first, apply them straight away to a pipeline you are working on, and come back to the storage sections when a project actually needs them.
When writing your CV, borrow the book’s language. Instead of “optimised ETL”, write that you found a bug in the Ingest stage that was skewing reports in the Serve stage, and explain how you fixed it. When reading job descriptions, look for phrases such as “end-to-end data pipeline” or “work with customer data infrastructure”. Those roles need exactly this way of thinking.
Tools will change many more times. A framework its own author still wants to keep after four years is worth bringing to every data project you take on next.
4 sources
- The Data Engineering Lifecycle and Undercurrents, 4 Years Later · 2026-07-27
- Book Review: Fundamentals of Data Engineering (Greyskull Analytics, Johnny Winter) · 2025-03-02
- Build hybrid and multicloud architectures using Google Cloud · 2024-10-24
- The Unfulfilled Promise of Serverless (Corey Quinn) · 2021-11-10