Fix retrieval first, prompts later: lessons from Jason Liu's RAG course
Jason Liu's course on Maven teaches a discipline every FDE needs once the demo is over and the customer wants to know why the chatbot still gets answers wrong.
In brief
- Jason Liu's course on Maven teaches how to improve RAG with measurements, not by whether the demo looks fine.
- According to Liu, the most common mistake is optimising answer generation while search is still not working properly.
- Read the free RAG Playbook article and listen to TWIML episode 709 before deciding whether to enrol.
“Stop building RAG systems that impress in demos but disappoint in production.” That is the opening line of the landing page for Jason Liu’s Systematically Improving RAG Applications on Maven, part of the Applied LLMs product line. It is a marketing line, but it names a moment every FDE has lived through.
You finish a Q&A chatbot over a customer’s internal documents. The demo goes smoothly. Two weeks later, real users start complaining that answers are wrong, incomplete or cite the wrong document.
At that point the question is no longer “how do we make it better” but “where is it failing, and what do we fix first”. This course teaches how to answer the second question with numbers.
Who teaches it, and why does that matter?
Liu describes himself as a staff ML engineer and AI consultant who spent 8 years building search and recommendation systems, and as the creator of the Instructor library. That background explains the course’s point of view: RAG is a search problem first and an LLM problem second.
In The RAG Playbook, a free post on his blog jxnl.co, Liu identifies the mistake the course targets: teams get absorbed in optimising answer generation while search is still not working properly.
They tweak prompts, swap models and add instructions, while the passage containing the answer never makes it into the context. No prompt can rescue a retriever that returns the wrong documents.
What does the course teach?
According to the official page, the material covers three areas. The first is measuring retrieval quality with standard IR metrics (precision, recall, MRR) to find weak points.
The second is building a synthetic data pipeline and evaluation set: using an LLM to generate realistic question-answer pairs, then establishing a baseline with tools such as LanceDB.
The third is architecture: hybrid search combining BM25, embeddings and metadata, along with query routing.
The course page also cites specific gains, such as a 20% accuracy improvement from re-ranking and 14% from cross-encoders. Treat these as figures published by the seller, not independently verified.
The claim of more than 400 engineers trained is also self-reported. The cohort listed on the page runs from 17 to 30 November; the page does not say which year.
Three ideas worth keeping
One: measure retrieval separately from generation. Imagine a question with 3 relevant document passages. The retriever returns a top-5 that includes 2 correct passages, at positions 2 and 4.
Recall@5 is 2/3, about 67%. Precision@5 is 2/5, or 40%. The reciprocal rank is 1/2 because the first correct passage is at position 2, and MRR is the average of this value across the whole question set.
Those three numbers tell three different stories. Low recall means the answer is not reaching the context, so the job is to fix search. If recall is high but the answer is still wrong, only then does the fault lie in generation. Once you have measurements, you know which file to open first.
Two: do not wait for real users. Liu calls synthetic data a “secret weapon”. Have an LLM read each document passage, generate a question that passage can answer, then check whether the retriever finds that same passage again. In a single afternoon you have a precision and recall baseline before the customer even opens the system to staff.
For an FDE, this is a real advantage. You usually have to show progress in the first few weeks, when there are no query logs yet. A baseline table with results after each change is more persuasive to a customer than any “it looks better now”.
Three: measurement replaces guesswork, and gradually becomes intuition. Liu argues that clear metrics help you build intuition about what actually makes a difference. Once you have seen hybrid search lift recall on questions containing product codes, the next time you meet a customer with similar data, it is the first thing you will try.
To get a firm grip on all three ideas, work through a small exercise. Suppose you have 5 questions, each with exactly one correct passage; the retriever places the correct passage at positions 1, 3, 2 and 5 for the first four questions, and fails to find it in the top-5 for the fifth.
Recall@5 is 4/5, or 80%. The reciprocal ranks are 1, 1/3, 1/2, 1/5 and 0; they sum to about 2.03, and dividing by 5 gives an MRR of about 0.41.
Read the two numbers together: the retriever almost always finds the answer, but often ranks it too low. That is the signal to try re-ranking before touching the prompt.
Who should take it, and where to start?
The course suits engineers who have already shipped at least one RAG system to users and are stuck at the “sometimes right, sometimes wrong” stage.
If you have no system yet, you can still start, because synthetic data lets you experiment without waiting for users. Even so, the advice here is to build a small RAG project of your own first, so the metrics have something to attach to.
Before enrolling, read The RAG Playbook on jxnl.co, then listen to episode 709 of the TWIML AI Podcast (November 2024), in which Liu discusses test datasets, data-driven experimentation and metrics for each type of use case. If these two free resources hit the problem you are facing, the Maven course is the next step.
For your career, this skill is easier to demonstrate than you might think. On a CV, instead of writing “built a RAG chatbot”, state how many questions your evaluation set contained and how recall changed after you switched search strategy.
When an FDE job description mentions RAG or LLMs, ask in the interview how the team measures retrieval quality. The answer tells you whether you would be joining a project with measurement discipline or one still running on gut feel, and the question shows the interviewer that you think like someone who has done the work.
A demo only shows whether the system runs. Whether the customer keeps using it after the demo depends on whether you can measure how well you are doing.
Was this article useful?
Thanks for the feedback!