Data Engineering Zoomcamp: a free course for aspiring FDEs who need to build data pipelines from scratch
The most valuable part of the course is not the six modules. It is the last three weeks, when you build a complete pipeline yourself.
In brief
- A free course from DataTalks.Club that requires no prior data engineering experience.
- The GitHub README says 9 weeks; the 2026 documentation says 10 weeks (7 weeks of modules plus 3 weeks of project).
- Treat the GitHub repo as the source of truth and put most of your effort into the end-to-end final project.
In 2025, Data Engineering Zoomcamp dropped Mage and switched to Kestra as its pipeline orchestration tool. For a free course, that sounds like a backstage detail. But it points to the most important thing about how to approach the course: the tools will keep changing, while the thinking behind building pipelines stays.
The course is built by DataTalks.Club. Its official GitHub repo describes it as a free nine-week course on building production-ready data pipelines. All videos, materials and homework are free and open source, and no prior data engineering experience is required.
For developers looking to move into forward deployed engineering, that is reason enough to pay attention. Picture your first project on the customer side: there is a good chance their data has to flow reliably through a processing pipeline before any model or agent can do anything useful. Zoomcamp is not a course about FDE work, but it teaches exactly that foundation.
Nine weeks or ten?
The course’s two official sources disagree. The GitHub README says 9 weeks, and a KDnuggets roundup from September 2026 also says nine. The 2026 documentation site from DataTalks.Club, however, says the course runs for 10 weeks: 7 weeks of modules plus 3 weeks for the final project.
The curriculum page explains the structure more clearly: six modules spread over seven weeks, followed by a final project that takes up the last three weeks of the cohort. The same page advises treating the GitHub repo as the source of truth for videos, code and the exact homework questions. The practical advice: plan around 10 weeks to be safe, and when the documentation contradicts itself, open the repo and check.
That habit is an FDE skill too.
Who should take it, and in what order?
The course suits backend developers with two to five years of experience who are comfortable writing services but have never owned a data flow from end to end. Because it assumes no data engineering background, you can start straight away. If you have already worked as a data engineer for a few years, the value lies mainly in the final project and in catching up on newer tooling.
The sensible order is to follow the cohort’s pace: work through the six modules in turn, do each module’s homework in the same week, then give the final three weeks entirely to the project. Do not binge the videos and leave the project for later. Homework is where you hit real errors, and real errors are what you can talk about in an interview.
Three ideas worth keeping
The first lies in the words “from scratch”. The documentation says the goal is to build production-ready pipelines from the ground up using industry-standard tools.
The second comes from the switch from Mage to Kestra. If even a course has to change tools between cohorts, your future customers are quite likely to be using a different stack from the one you learned.
So when you study the orchestration module, ask what problem the tool is solving: scheduling, retrying on failure, tracking dependencies between steps. Once you understand those questions, you will recognise the answers in whatever tool a customer happens to use.
The last idea, and the most important, is the final project. Three weeks spent building an end-to-end pipeline produces the clearest evidence of what you can do. Pick a problem close to the industry where you want to work as an FDE, such as logistics, retail or fintech, rather than a sample dataset everyone else uses.
What does a sample project look like?
Take a retail example: a chain of shops wants to see revenue per store every morning. Your pipeline picks up the nightly export from the point-of-sale system, loads it into storage, cleans and standardises it, then aggregates it into a table for a dashboard. The orchestration tool handles the nightly schedule and reruns any step that fails.
This sketch is deliberately small. What you need to prove is not that you can handle big data, but that every step has a reason and that when one step breaks, the whole flow does not collapse with it.
How should it go on your CV?
Do not list the course as a certificate. Take the sample project above as an example. The CV line before editing: “Completed Data Engineering Zoomcamp (DataTalks.Club)”.
The line after editing: “Built a nightly pipeline that ingests sales data, cleans it and aggregates revenue by store; configured scheduling and automatic retries when a step fails; README explains design decisions (repo linked).”
The second line tells a specific story: where the data comes from, which steps it passes through, how failures are handled and who uses the output. Those are the questions you should have answers ready for before an interview. When reading job descriptions, look out for phrases such as “data pipeline”, “ETL” or “customer data integration”.
When you see them, your project becomes a story to tell, not just a line on a CV.
A free course will not make you an FDE. But it gives you three weeks to prove that, handed someone else’s messy data, you know where to start.