Measuring ROI for customers: count work that meets the bar, subtract checking and rework
When the customer's CFO asks whether the project was worth the money, "users say it feels faster" will convince nobody.

In brief
- Self-reported time is unreliable: in METR's trial, developers took 19% longer yet still believed they were 20% faster.
- Count only the work AI completes to the quality bar, then subtract checking time (always paid) and rework time (probabilistic).
- Measure the baseline before deployment and tell the customer plainly that payback usually takes years, not months.
- 1Baseline: 480 secondsProcessing one invoice by hand took 8 minutes before the system
- 2After deployment: 45 secondsThe system processes one invoice in 45 seconds
- 3Gross savings: 435 seconds480 − 45, the figure most easily misreported as ROI
- 4Minus checking: −60 secondsEach invoice needs 1 minute of human review, a cost that is always paid
- 5Minus rework: −24 seconds5% of invoices wrong × 8 minutes of manual work, weighted by probability
- 6Net savings: 351 secondsMultiplied by 10,000 invoices a month gives 975 hours, before system costs
In this hypothetical example, checking and rework eat 84 of the 435 seconds of gross savings.
Graphic: FDE Times
Three months after the system went into production, the quarterly review had a new face in the room: the customer’s chief financial officer. She did not ask which model was used or what the latency was. She asked one question: “Was this project worth the money we put into it?”
If all you have at that moment is a dashboard screenshot and a few kind words from users, the project has already lost half the battle. FDEs are not paid to ship features.
Palantir’s job description for a Forward Deployed Enablement Engineer says plainly that the role exists to maximise the outcomes of deployed products and workflows, tied to the customer’s most important business results.
That is why measuring ROI is not a job for sales or finance alone. It is an engineering skill: it takes data, a well-designed measurement and honesty. The method below follows an invoice-processing pipeline from baseline to the final figure placed in front of the CFO.
Why can’t you trust the feeling of being “faster”?
The first trap is asking the users. In 2025 METR ran a randomised controlled trial with experienced open-source developers. The result: when using AI tools, they took 19% longer than when working without them.
The striking detail comes next. Before the trial, these developers expected to be faster. Afterwards, even though the measurements said otherwise, they still believed AI had made them 20% faster.
The two numbers are not on the same scale: one is measured extra time, the other is a speed-up users believe they got. But they point in opposite directions, and that is what should worry you.
The lesson for FDEs is clear. A survey asking “how much time do you feel you saved?” helps you understand user experience; it cannot serve as financial evidence. What you hand the CFO must be measured time per task, with a before figure and an after figure.
Start from goals and a baseline, not from the model
Paul Parks of AICPA & CIMA, writing in the Journal of Accountancy, proposes an ROI process that begins by defining objectives: what the use case is and what leadership wants to achieve with the investment. His framework also accounts for benefits that cannot be expressed in money.
For an FDE, this is customer discovery: before writing a line of code, you need to know which number the customer will use to judge success.
The next step is the baseline. Delos recommends spending the first 30 days of a 90-day framework establishing baseline metrics before reporting any ROI figure. The reason is simple: without measuring the “before”, there is nothing to compare the “after” against.
Waiting until the system is running to look for a baseline is usually too late, because the old process no longer exists.
Parks also points out that the benefit side is inherently hard to measure, because most organisations run several initiatives in parallel. A rise in invoices processed might be thanks to your system, or because the accounts team just hired more people.
If you can, keep a control group, such as a branch or a document type not yet on the system, to isolate the project’s effect.
Example: an invoice-processing pipeline
Delos offers an easy-to-picture before-and-after example: the time to process one invoice drops from 8 minutes to 45 seconds. This is the right way to express it, because it is per task rather than a vague claim of “higher productivity”. But if you stop there, the figure is still inflated.
GSPANN calls the missing part the repair burden: the time people spend checking and fixing AI output. They sum it up this way: checking is always paid, while rework depends on probability.
GSPANN gives a net-savings formula to account for this. In words: take gross hours saved, subtract repair hours, then subtract the system’s other costs.
Apply it to a hypothetical case. The customer processes 10,000 invoices a month, each invoice needs 1 minute for a person to check the result, and 5% of invoices are wrong and must be redone by hand, taking 8 minutes.
| Component (per invoice) | Calculation | Seconds |
|---|---|---|
| Gross savings | 480 − 45 | 435 |
| Checking (always required) | 1 minute × 100% | −60 |
| Rework (probabilistic) | 8 minutes × 5% | −24 |
| Net savings | 351 |
Across 10,000 invoices, net savings come to 3,510,000 seconds, or 975 hours a month. Delos advises converting hours to money by multiplying time saved by the fully loaded hourly cost of the person who used to do the work, meaning insurance, benefits and overhead as well as salary.
Then you subtract the cost of the model, the infrastructure and the staff who run the system.
The second figure to give the finance side is cost per successful task. According to GSPANN, you count only the work AI completes that clears the agreed quality threshold, add up the full cost of producing it, and divide.
In the example above, if you assume that invoices redone by hand do not count as work completed by AI (even though, once fixed, they may have met the threshold), the denominator is 9,500, not 10,000.
Doing it for your own project
Start with a session with the budget holder, not the end users. Ask which business metric they want to improve, and what quality threshold counts as “done”. Write it down and get them to confirm it.
Then measure the baseline on the old process: time each task, count volumes, record the current error rate. Once the system is live, log three things for every task: processing time, human review time, and whether the task had to be redone. Without that third column, you will never be able to calculate the repair burden.
Finally, talk about payback time from the start. A Deloitte survey, cited by ACCA, found that a typical AI use case needs two to four years to achieve ROI, and only 6% of respondents saw payback in under a year. Promising payback within a quarter sets you up to lose.
Mistakes that get the numbers thrown out
The most common mistake is reporting gross savings: taking 8 minutes minus 45 seconds and multiplying it out, ignoring checking time. Someone in finance will immediately ask who is reviewing the output, and at that point your figure collapses.
The second is counting failed tasks in the denominator, which makes cost per task look cheaper than it is. The third is attributing every improvement to your project when the customer is running three other initiatives at the same time.
The fourth is relying on perception surveys. METR showed that perception can be wrong in direction, not just in magnitude.
Where does this skill belong on your CV?
For developers looking to move into FDE work, this is a skill that makes a CV stand out. When reading job descriptions, look for phrases such as “business outcomes”, “value realization” or “customer success”. They signal that the company needs people who can measure results, not just people who can build.
On your CV, replace a line like “built an invoice OCR pipeline” with one that includes the baseline, post-deployment figures and the pass rate.
This week’s exercise
Take a feature you have shipped, at work or in a personal project, and build a table like the invoice table above: gross savings, checking time, rework time, net savings. Any cell left empty for lack of a real measurement is exactly the data you must collect from day one of your next project.
When the CFO turns up at the review, the winner is usually not the person with the slickest demo but the one who measured the baseline on the first day.
6 sources
- Palantir Technologies - Forward Deployed Enablement Engineer - Customer Success
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity · 2025-07-10
- Why Nobody Can Total the AI Bill (GSPANN Insights) · 2026-09-03
- AI Agent ROI: How to Measure the Business Impact of Autonomous AI Workers · 2026-07-22
- AI investment returns elusive · 2025-12
- Generative AI's toughest question: What's it worth? · 2025-02-01