FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Books & courses

Berkeley RDI's Agentic AI course: what to study before you take agents into the enterprise

For a Forward Deployed Engineer, the first lecture to watch in this course is about running agents in production, not about models.

In brief

  • The Fall 2025 edition (CS294/194-196) is taught by Dawn Song. Xinyun Chen co-taught only the Fall 2024 and Spring 2025 editions.
  • The material closest to FDE work is Clay Bavor's (Sierra) lecture on deploying agents in the real world, along with the lectures on evaluation and safety.
  • Three ideas worth keeping: agents are products, reliability matters more than occasional brilliance, and you should build agents to grade agents.
ShareLinkedInFacebookX

The Fall 2025 edition of UC Berkeley’s CS294/194-196 Agentic AI course includes a session on 10 November that Forward Deployed Engineers should watch before any other.

In it, Clay Bavor of Sierra talks about practical lessons from deploying AI agents in the real world. One learner wrote up their main takeaway on dev.to: agents are shifting from a technology to a product.

Turning an agent from a demo into a product that customers use every day is also the core of an FDE’s job at a client company.

The course belongs to a MOOC series that Berkeley RDI treats as its flagship programme. The institute’s education page says more than 40,000 people have enrolled.

If you are a developer with a few years of experience and want to move into FDE work, do not treat this course as a theory module on agents. You will get more from it if you treat it as a handbook on getting agents to run reliably inside an enterprise.

Who teaches which edition?

The MOOC series has three editions: Large Language Model Agents (Fall 2024), CS294/194-280 Advanced Large Language Model Agents (Spring 2025) and Agentic AI (Fall 2025). Dawn Song of UC Berkeley teaches the Fall 2025 edition and also taught Fall 2024.

Xinyun Chen, a researcher at Google DeepMind, was a guest co-instructor for the Fall 2024 and Spring 2025 editions. Her name does not appear on the official Fall 2025 page, so if you want to hear Chen lecture, look to the two earlier editions.

Beyond the Sierra talk, Fall 2025 includes lectures on agent evaluation and agent training. There is a lecture on multi-agent systems by Noam Brown (OpenAI) and one on agent safety and security taught by Dawn Song. Fall 2024 has more material on agents in the enterprise, including AI Agents for Enterprise Workflows by Nicolas Chapados (ServiceNow).

In what order should you study?

Berkeley recommends that its own students have a background in machine learning and deep learning before taking the course. The course page sets no separate requirements for MOOC learners. Even so, if you have never trained a model, the lectures on agent training may be hard going.

If you build agents for clients, start with Clay Bavor’s lecture. Then move on to the Fall 2025 lectures on agent evaluation and on safety and security.

Next, go back to Fall 2024 for the lectures on enterprise workflows. Leave the lectures on agent training and multi-agent systems until last, when you have a real agent that needs improving.

The Fall 2024 edition has now ended. It used to award certificates in several tiers: every tier required 12 quizzes and a passing written article, the Mastery tier added 3 labs, and the Ninja and Legendary tiers added a hackathon. The structure still works well for self-study: watch the lectures, write an article, do the labs, then build a real project.

Idea one: agents are now products

The dev.to learner’s observation sounds obvious. But it changes the questions you ask in your first meeting with a client.

If you think of an agent as a technology, you ask which model is strongest. If you think of it as a product, you ask who uses it, who is accountable when it gives a wrong answer, and whether there is a way to hand over to a human. At a client, write down the answers to those three questions before you write the first line of a prompt.

Idea two: reliability matters more than brilliance

The same learner wrote that when an agent handles millions of conversations, reliability matters more than a few moments of excellence. This is the learner’s personal view, not that of the course organisers. A quick calculation shows why it holds up.

Suppose a customer-service agent answers 99% of questions correctly. In a 20-question demo, it may well get none wrong. But across 1,000,000 conversations, a 1% error rate means 10,000 customers get a wrong answer, and each one could become a complaint ticket.

So when you demo to a client, do not just pick the questions the agent answers best. Show the error rate on a sufficiently large test set and explain what the agent does when it is unsure of the answer.

Idea three: use agents to grade agents

The Fall 2025 edition is tied to the AgentX–AgentBeats competition. In it, a benchmark is turned into an evaluating agent, called a green agent, which scores the competing agents. You can apply this idea at a client straight away.

Say you are building an agent that handles refund requests for an e-commerce marketplace. Instead of having the client’s QA team read logs line by line, you build an evaluating agent that plays the buyer. Sometimes it tries to claim a refund on an order past the deadline; sometimes it gives the wrong order number.

Every time you change a prompt or swap a tool, you rerun the full set of scenarios and get comparable numbers immediately. An evaluation suite like this is also worth more on a CV than a certificate.

A line such as “built an automated evaluation suite for an agent, catching failures before release” shows recruiters that you know how to measure agent quality. When you read an FDE job description and see the words evaluation, deployment and safety, watch the course lectures on those topics first.

The course certificate will lose its value at some point. The habit of grading an agent with numbers before handing it to a client will serve you for a long time.

6 sources
Read next on the roadmap · Stage 3: Applied AIHugging Face's LLM Course: how the first three chapters take an FDE from pipeline() to fine-tuningThe course is free and available in many languages. Its first three chapters can be read as the order of work at a client site: set a baseline first, inspect the data next, fine-tune last.