FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Books & courses

Neural Networks: Zero to Hero: Andrej Karpathy's eight lectures, from 150 lines of code to GPT

When a client asks why a model gave the wrong answer, an FDE who has written backprop and a tokenizer by hand can explain it. Someone who only knows how to call the API cannot.

In brief

  • The course has 8 lectures: it opens with micrograd (2h25m), runs through 5 makemore lectures, then GPT and the GPT Tokenizer (2h13m). The syllabus is still marked 'ongoing'.
  • Lecture 1 needs only basic Python and some high-school calculus, but the full course demands solid Python programming.
  • For an FDE, the course's main value is being able to explain model behaviour to clients, not training models yourself.
ShareLinkedInFacebookX
GraphicThe FDE skill each lecture builds
What you buildFDE skill
Lecture 1: microgradAn autograd engine of about 100 lines, to understand backpropExplaining training without treating it as magic
Lectures 2–6: makemoreCharacter-level model, MLP, BatchNorm, manual backprop, WaveNetMeasuring each layer to diagnose rather than guess
Lecture 7: Let's build GPTBuilding GPT from scratch, in codeUnderstanding what sits beneath the API the client uses
Lecture 8: GPT TokenizerA tokenizer, with the code released as the minbpe repoShowing clients how their text is split into tokens

In each lecture you write one part of the model yourself, and each part maps to a job you will do on a client site.

Graphic: FDE Times

The first lecture runs 2 hours 25 minutes, and by the end you have an autograd engine of about 100 lines plus a neural network library of about 50. That is micrograd, the starting point of “Neural Networks: Zero to Hero”. Andrej Karpathy describes it as a course on building neural networks from scratch, in code.

The course does not teach you to call an existing model. It teaches you to write one. The official GitHub repo describes it as a course that starts from the most basic foundations.

For a developer moving into FDE work, this approach is worth the time. Picture a client asking why a model just did something strange. Even if you never train a model yourself, a correct explanation requires understanding what sits beneath the API.

Eight lectures, taken in order

The official page on karpathy.ai lists the syllabus in a clear sequence. Lecture 1 builds micrograd to explain backpropagation. Lectures 2 to 6 form the makemore series: from a character-level language model to an MLP, then activations, gradients and BatchNorm, then writing backprop by hand, and finally WaveNet.

Lecture 7, “Let’s build GPT”, is the best known. Lecture 8 builds the GPT Tokenizer and runs 2 hours 13 minutes. The syllabus ends with the word “ongoing”, so more lectures may follow.

The advice here is not to jump straight to lecture 7. The syllabus builds from the foundations up, and skipping ahead easily turns you into someone watching another person type code. Even the nanoGPT repo, which accompanies the GPT lectures, recommends watching Zero To Hero for context on GPT and language models.

Who should take it, and do you need strong maths?

The two official sources state slightly different prerequisites, and the difference is useful. The repo says lecture 1 needs only basic Python and a vague memory of high-school calculus. The official page sets a higher bar for the whole course: solid Python programming and introductory maths such as derivatives and Gaussian distributions.

If you have written Python professionally for two or three years, you qualify. The maths you can brush up on as you go. The harder part is patience: these are long videos that expect you to type the code along, step by step.

Every lecture comes with exercises, listed in the YouTube video description. Do not skip them. Watching the videos without doing the exercises is like reading deployment documentation without ever deploying.

Idea one: gradients are just the chain rule

The first idea worth keeping comes from micrograd. Try a small example yourself: take f = a·b + c, with a = 2, b = -3, c = 10, so f = 4. By the chain rule, the derivative of f with respect to a is b, or -3; with respect to b it is a, or 2; and with respect to c it is 1.

All of backprop is that calculation run backwards through the graph, one node at a time. Once you have worked out those three numbers by hand, you stop treating training as magic.

Idea two: learn to read the health of a network

makemore part 3, on activations, gradients and BatchNorm, teaches a habit very close to FDE work: look inside the system to diagnose, rather than guess. Before concluding that a model “learns poorly”, check whether the signal is actually flowing through the layers.

Try a small calculation with tanh. If a neuron’s input is 5, its output is tanh(5) ≈ 0.9999, and the local derivative 1 − tanh² is only about 0.0002. The gradient flowing back through that neuron is multiplied by a number close to zero, so the weights before it barely move even while the loss stays high.

A neuron in this state is called saturated. If it is saturated for every example in the data, it almost never receives a meaningful gradient and is effectively “dead”. Detecting this is simple: plot a histogram of each layer’s outputs and count the share of values with an absolute value above 0.99.

If most of a layer’s values pile up at −1 and 1, suspect the weight initialisation or the input scale before you suspect the architecture. BatchNorm is one way to pull the activation distribution back into the range where gradients can still flow.

Beginners tend to make three mistakes: watching only the loss curve rather than the activations and gradients of each layer; initialising weights too large, so the network saturates from the first step; and rushing to change the architecture or raise the learning rate before measuring anything.

On a client site, this habit carries over unchanged. When a pipeline produces poor results, the first job is to measure each layer in turn, before you start rewriting prompts or swapping models.

Idea three: the tokenizer is a separate layer

Lecture 8 gives the tokenizer a lecture of its own, and its code is released as the minbpe repo, with a written version of the lecture in lecture.md. The lesson is that the model does not see text, only tokens, and how text is split into tokens is decided in a step that sits outside the neural network.

Imagine a client in Vietnam asking why a Vietnamese prompt costs more tokens than an English prompt with the same content. After lecture 8, you can open the tokenizer, show the client which pieces their text is split into, and explain it using their own data. That kind of answer is how an FDE earns a client’s trust.

Turning the course into evidence on your CV

“Watched Zero to Hero” on a CV is worth almost nothing. What has value is a repo where you rewrote micrograd yourself, did the makemore exercises, and wrote a README explaining each part in your own words.

When reading FDE job descriptions, look for requirements about understanding how models work or debugging model behaviour: that is where to point to this repo.

After the eight lectures, the logical next step is to open nanoGPT and minbpe, the two repos that accompany the course, and read the code yourself rather than just watching videos. When a client asks “why”, you will have an answer, because you have written the very thing that is giving the wrong answer.

5 sources
Read next on the roadmap · Stage 3: Applied AIPractical Deep Learning for Coders: nine lessons from code to a working modelJeremy Howard's free course has you running a model on real data in the first lesson. The theory comes later.