FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Books & courses

Reading "Designing Machine Learning Systems" as an FDE: start with chapters 3, 4, 9 and 10

Chip Huyen's book is long, but the problems an FDE meets every day at a customer site fit into four chapters: data, training data, model updates and infrastructure.

Cover of Designing Machine Learning Systems
Designing Machine Learning Systems · Chip Huyen · Cover: Open Library

In brief

  • Chip Huyen's book (O'Reilly, 2022) grew out of the CS 329S course at Stanford and treats an ML system as a whole.
  • For FDEs, chapters 3, 4, 9 and 10 are the core: data, training data, continual learning and MLOps infrastructure.
  • Three ideas worth keeping: no algorithm rescues bad data, class imbalance is normal, and build versus buy is a question that keeps coming back.
ShareLinkedInFacebookX

This is a book about machine learning, yet its most memorable line says the algorithm will not save you: if your training data is bad, your algorithm cannot perform well. For a Forward Deployed Engineer, that sentence is close to a job description.

The book is Designing Machine Learning Systems by Chip Huyen, published by O’Reilly in 2022. It grew out of CS 329S, the course Chip Huyen taught at Stanford, which set out to provide an iterative framework for developing ML systems that run in the real world.

You do not need to start on page one. An FDE at a customer site usually wrestles with four things: connecting the customer’s data, getting training data that is good enough, keeping the model from going stale, and answering whether to build or buy the infrastructure. Chapters 3, 4, 9 and 10 cover exactly those four.

Why does a “holistic” book suit FDEs?

The author’s book page describes it as a holistic approach to designing ML systems. Every decision, from how to process and create training data to which features to use, is weighed against the goals of the whole system rather than considered in isolation.

The book’s official GitHub repo opens with the observation that ML systems are both complex and unique, and aims at systems that are reliable, scalable and maintainable. FDE work revolves around that uniqueness: every customer has its own kind of data and its own constraints, so systems thinking is worth more than any tuning trick.

A data format is a commitment to the customer

Chapter 3, Data Engineering Fundamentals, contains a line worth copying down: with structured data, the code that writes the data has to assume its structure. When you integrate a customer’s data, choosing a storage format is therefore not a minor detail. It is an agreement between your system and theirs.

The chapter also separates two paths. Data processed in batches yields static features; data moving through real-time transport is usually processed by a stream computation engine and yields dynamic features.

Picture a delivery company that wants to predict which orders will be cancelled. “A buyer’s average number of orders per month” changes slowly, so recomputing it nightly from an export file is enough: that is a batch feature. “The number of times the buyer changed their address in the last 10 minutes” is meaningless if you wait until tomorrow morning: that is a stream feature.

Get into the habit of classifying features this way in the first week, because the answer determines what architecture you propose. If every important feature is batch, you may not need to touch streaming infrastructure yet.

95% accuracy can catch nothing at all

Chapter 4, Training Data, argues that the quality of training data determines performance, and no clever algorithm can make up for it. It also states an uncomfortable fact: class imbalance is the norm in real-world data, not the exception.

Do a small calculation. Suppose a customer hands over 1,000 transactions, of which 950 are normal and 50 are fraudulent. A lazy model that predicts “normal” for everything will be right 950 times out of 1,000, an accuracy of 95%, while catching 0 of the 50 frauds.

That 95% looks fine on a demo slide, but the 50 frauds are the reason the customer hired you. So at the first working session, ask for the label distribution before discussing which model to use, and report results per class rather than as a single number.

The real work starts after the demo

Chapter 9 covers continual learning and test in production, arguing that this is a problem specific to ML but one that mostly requires infrastructure solutions. The author also pushes back on the fear that streaming is hard and expensive: streaming technology has matured considerably.

Chapter 10, on infrastructure and tooling for MLOps, calls ML platform an emerging team. When building infrastructure, the build-or-buy question haunts engineering managers and CTOs alike, and the book admits there is still no consensus on what an ML platform should contain.

That gap is where FDEs have a voice, because you are the one who sees first-hand what the customer is missing. A one-page build-or-buy analysis written from the customer’s own reality is usually more persuasive than any generic architecture diagram.

Which chapter should you open for which situation?

One reading is not enough; the value of these four chapters lies in knowing which page to return to when a project stumbles. The table below is a suggested checklist to pin beside your desk.

When the project hits this situation Open chapter What to do right away
The customer sends data exported from a legacy system 3 Write down the structure the data-writing code assumes, and settle the format before building the pipeline
Some features need updating by the minute 3 Route static features through batch and separate the dynamic features that need stream processing
The demo model looks good but the customer sees it getting things wrong 4 Count the label distribution and report results per class
The model needs frequent updates after go-live 9 Treat it as an infrastructure problem, and do not dismiss streaming just because it seems hard and expensive
The CTO asks whether to build or buy 10 Write a one-page build-or-buy analysis grounded in the customer’s reality

Which reading order gets you into FDE work fastest?

A sensible order is 3, 4, 9, then 10, returning to the other chapters when a project calls for them. Before opening the book, skim the chapter summaries at github.com/chiphuyen/dmls-book to get a map.

The author’s page is at huyenchip.com/books, and the course materials are at stanford-cs329s.github.io. If you work in backend or data and already know pipelines, these four chapters will follow naturally from your existing experience.

When reading an FDE job description, try sorting each requirement into the four chapters: requirements about data and training data map to chapters 3 and 4; those about operating models and infrastructure map to chapters 9 and 10.

On a CV, a small project that includes both batch and stream features, with a note on how you spotted and handled class imbalance, says more than a line reading “knows machine learning”.

Finishing these four chapters will not make you much better at algorithms. But you will ask customers the right questions sooner, and for an FDE, that is the real advantage.

4 sources
Read next on the roadmap · Stage 3: Applied AIChip Huyen's AI Engineering: a book with little code that teaches FDEs how to decideChip Huyen says plainly that this is not a tutorial book. That admission is the reason a forward deployed engineer should read it before any step-by-step coding guide.