Hamel Husain and Shreya Shankar's AI Evals course: the most valuable lesson is reading the output
Maven's highest-grossing course has taught more than 5,000 engineers and PMs, and its central argument is that error analysis matters more than any eval tool.
In brief
- The 'AI Evals For Engineers & PMs' course on Maven teaches one loop: build a real agent, find where it fails, fix it with evals you can trust
- Core view: error analysis is the most important activity; eval criteria emerge when humans grade output, they are not set in advance
- Not for beginners: read the free Field Guide and FAQ on hamel.dev first, and pay only once you have a product in production
The highest-grossing course on Maven does not sell a tool. After teaching more than 5,000 engineers and PMs, Hamel Husain and Shreya Shankar sum it up in one short line in their official FAQ: error analysis is the most important activity in evals. Not dashboards, not LLM-as-judge, but sitting down and reading the model’s outputs one by one.
For anyone who wants to work as an FDE, this is the detail that matters most. Hamel writes in his Field Guide that evaluation criteria cannot be fully defined before humans have graded the LLM’s output. If that holds, an engineer deploying an agent for a customer cannot wait for the customer to hand over the criteria. They have to read the output and find the criteria themselves.
Finding where an agent fails, then turning that finding into an eval suite the customer trusts, is exactly the skill this course trains.
What the course actually teaches
Its full name is “AI Evals For Engineers & PMs”, and it is taught live in cohorts on Maven. Hamel Husain describes himself as an ML engineer with more than 20 years of experience; Shreya Shankar holds a PhD in EECS from UC Berkeley and will join CMU as faculty in 2027. The course has had more than 5,000 students from more than 500 companies, with a rating of 4.7 from 1,011 reviews.
Its promise comes down to a loop: build a real AI agent, find where it fails, then improve it with evals you can trust. Students get more than 200 pages of course reading, 4 homework assignments with solutions and walkthroughs, more than 10 office hours, and lifetime Discord access.
Recordings are kept permanently, and you can retake any future cohort at no extra cost.
The listed price is $4,200, which is not a small sum for many engineers. The more useful question, then, is whether it is worth it for you, and the answer depends on whether you already have a product running.
Why you would pay to learn to “look at the data”
It sounds obvious. Yet common practice runs the other way: pick an eval framework first, write criteria first, then run. In his Field Guide on rapidly improving AI products, Hamel says plainly that nothing replaces the insight gained from looking at real examples.
Combine that with the argument that criteria only emerge when humans grade output, and you get a way of thinking that inverts many engineers’ habits. Criteria are not an input to evals; they are an output of error analysis.
Imagine you are an FDE deploying a ticket-answering agent for a logistics company. Before reading the traces, you guess the main failure is hallucination. After reading them, you may find the agent answers correctly but in the wrong tone, or ignores a field in the customer’s internal system. None of these failures appears in any benchmark; they only surface when someone sits down and reads.
Who should take it, and in what order
The course page is explicit: this is a deep dive, not an introduction to LLMs or prompt engineering, so it is best if you are already building something. If you do not yet have an agent running with real users, save your money. The assignments will have nothing to apply to, and office hours will turn into listening to other people’s questions.
The sensible order is to read for free first and pay later. The Field Guide and the AI Evals FAQ on hamel.dev set out the core ideas, enough for you to decide. The FAQ in particular, according to the two authors, collects the questions they were asked most often while teaching more than 5,000 engineers and PMs.
If, after reading them, you still want someone to guide you through each assignment and a community to ask, that is when the course is worth the money.
For FDEs, the extra value lies in the word “PMs” in the course title. You learn alongside the people who make product decisions, which means practising how to talk about evals in language they understand, something you will have to do every week in front of customers.
What to take away, and what to put on your CV
The other point worth keeping is that the loop has to close. The course’s promise does not stop at “find where it fails”; it continues to “improve with evals”, then goes back to finding new failures. Evals, in this view, are not a gate at the end of a project but a daily working rhythm.
To show this skill on a CV, do not write “experienced with evals”. Write how many traces you read, how many failure categories you identified, and which category the eval you built caught. A line like that says more than any framework name.
The course fee can wait. The habit of reading output should start this week, because that is what customers are actually paying for.