Design error responses with RFC 9457 so customer ops teams can fix incidents themselves
If your API returns only a 500 with "Something went wrong", every incident on the customer's side becomes a ticket for you. A few JSON fields in the right places can change that.
In brief
- RFC 9457 has replaced RFC 7807 and is the current standard for application/problem+json error responses.
- Machines classify errors by type, people read detail, and operations teams use instance to search logs. Never make clients parse detail.
- Add extension members such as runbook or retryable so customers can resolve issues themselves, and keep stack dumps out of responses.
Same incident, but a structured response lets the customer resolve it themselves instead of opening a ticket for you.
Graphic: FDE Times
Picture the first week after you hand an integration over to a customer. At midnight, their operations dashboard turns red. The log holds a single line, 500 Internal Server Error, and the body reads {"error": "Something went wrong"}.
The customer’s on-call engineer cannot tell whether the fault is in their system or yours, whether to retry, or where to look in the logs. So they open a ticket. By morning you have ten tickets, and every one begins with “Your API is broken”.
For an FDE, error responses are part of the handover, not just a technical detail. Design them well and the customer can handle most incidents on their own. Design them badly and you become the permanent on-call engineer for someone else’s system.
A status code tells only half the story
RFC 9457 starts from a practical observation: an HTTP status code on its own is often not enough for the recipient to understand what happened. A 422 says the request had a problem. It does not say what the problem was, which account it affected, or what the recipient should do next.
To fill that gap, the RFC defines a JSON format with its own media type, application/problem+json. You may have heard of RFC 7807 from 2016. RFC 9457 formally replaces it, so cite RFC 9457 when writing documentation or talking to the customer’s architects.
The strength of the standard is that each field is written for a specific reader. The table below maps the fields to the questions an on-call engineer typically asks:
| Field | Who reads it | Question it answers |
|---|---|---|
type |
The customer’s machines | What kind of error is this, so it can be handled automatically? |
title |
People | In short, what is this kind of error? |
detail |
People | What specifically happened this time? |
instance |
Operations team | What ID does this occurrence carry, for searching and cross-checking logs? |
| Extension members | Both | Which account is affected, and which guide should be followed? |
Example: an API that pushes orders to a warehouse
Suppose you are integrating a retail chain’s ordering system with your company’s warehouse API. One account has used up its daily order quota. A poor implementation returns this:
{ "error": "Bad request", "code": 4021 }
The on-call engineer has no idea what 4021 means, so another ticket gets opened. Here is the RFC 9457 version (the example uses Vietnamese strings; the title means “Daily order limit exceeded”):
{
"type": "https://docs.example.vn/problems/vuot-han-muc-don",
"title": "Vượt hạn mức đơn trong ngày",
"detail": "Tài khoản KH-0192 đã dùng hết hạn mức đơn hôm nay. Hạn mức được đặt lại lúc 00:00.",
"instance": "/incidents/7f3a2c",
"account_id": "KH-0192",
"runbook": "https://docs.example.vn/runbook/han-muc-don",
"retryable": false
}
Reading this, the on-call engineer knows at once that the fault is on their side and when it will clear: the detail says account KH-0192 has used up today’s quota, which resets at 00:00. They have a link to the runbook, and if they still need to call you, they have the ID 7f3a2c so you can find the exact log line. account_id, runbook and retryable are extension members.
The RFC allows fields like these and requires clients to ignore any they do not recognise, so you can add them gradually without breaking older clients.
The detail field here states only the problem and how to fix it. Postman likewise advises that error messages should contain exactly those two things. Database table names and Java class names are of no use to the customer’s on-call engineer.
Letting the customer’s machines handle errors too
The on-call engineer is only half the picture. The other half is the customer’s code calling your API, and that code needs to know what to branch on.
RFC 7807 was already explicit: clients must use type as the primary identifier for the kind of error and should not parse the detail string for information. The reason is visible in the example above. The detail string is meant for people, so next week you might reword it, switch it to English or add figures.
If the customer’s code is reading that string with a regex, it will break silently. The illustrative code below shows how the client side should branch on the response from the warehouse example:
PROBLEMS = "https://docs.example.vn/problems/"
def xu_ly_loi(resp, method):
if not resp.headers.get("Content-Type", "").startswith("application/problem+json"):
return mo_ticket(None)
problem = resp.json()
if problem.get("type") == PROBLEMS + "vuot-han-muc-don":
return doi_den_ngay_mai(problem.get("account_id"))
if problem.get("retryable") is True and method != "POST":
return thu_lai_sau()
return mo_ticket(problem.get("instance"))
Not one line of this function reads detail. However often you reword it, the function keeps working. And when it does have to open a ticket, it attaches instance so both sides can find the same incident.
The retryable flag answers the question every customer asks: should I try again? To be clear, retryable is not in the RFC. It is an extension member you define yourself, so you must document its meaning, on the very page that type points to.
A Hackernoon article on retries advises retrying only on certain transient error codes, and generally skipping POST to avoid creating duplicate records. In the warehouse example, if the customer’s code automatically resends an order-creating POST after a timeout, the warehouse may receive two identical orders, because a timeout does not reveal whether the first order was recorded.
That advice is only complete with an idempotency key. If the customer genuinely needs to retry POSTs, design it so the client generates a unique key for each order and sends that same key on every attempt. When the server sees a key it has already processed, it returns the original result instead of creating a second order.
Until that mechanism exists, retryable should be false for every error on order-creating POSTs.
An afternoon’s work with Spring
If your API is written in Java/Spring, most of the work is already done. Spring Framework supports RFC 9457 through the ProblemDetail class. Spring Boot has a spring.mvc.problemdetails.enabled property that makes its built-in exceptions return problem details automatically:
spring.mvc.problemdetails.enabled=true
For your own business errors, build a ProblemDetail in an exception handler:
@ExceptionHandler(QuotaExceededException.class)
ProblemDetail handleQuota(QuotaExceededException ex) {
ProblemDetail pd = ProblemDetail.forStatusAndDetail(
HttpStatus.UNPROCESSABLE_ENTITY, ex.getMessage());
pd.setType(URI.create("https://docs.example.vn/problems/vuot-han-muc-don"));
pd.setTitle("Vượt hạn mức đơn trong ngày");
pd.setProperty("account_id", ex.getAccountId());
pd.setProperty("runbook", "https://docs.example.vn/runbook/han-muc-don");
pd.setProperty("retryable", false);
return pd;
}
The hard part is not the code. It is sitting down with the customer’s operations team to list the errors they actually encounter, naming a type for each one and writing a runbook for every error. That is closer to customer discovery than to programming.
Common traps
The most dangerous trap is leaking implementation details. RFC 9457 specifically warns against exposing things such as stack dumps through the HTTP interface. Write the stack trace to internal logs tied to instance, and return only that ID to the customer.
The second trap is every endpoint returning errors in its own shape: one uses error, another uses message, a third returns HTML. Postman advises that error responses be clearly and consistently structured. A single exception forces the customer’s code to add a special-case branch.
The third trap is a type that points to an empty URI. The on-call engineer at midnight will click that link. If it leads to a 404, you have lost the chance for them to fix the problem themselves.
An answer to “How did you reduce tickets after handover?”
In an FDE interview, a question such as “What did you do to reduce tickets after handover?” is a good opening to talk about this skill. Do not say “I handle errors carefully”; describe the actual sequence: standardising errors on RFC 9457, writing a runbook for each type, and the customer’s operations team closing tickets on their own that previously had to be escalated to you.
The next time a midnight error is resolved by the customer without anyone calling you, that is the sign you designed your error responses correctly.