APIs and webhooks for FDEs: build every integration as if everything arrives twice
The costliest bug in an enterprise integration rarely takes the system down. It quietly charges a customer twice, or drops an event that nobody notices is gone.

In brief
- After a network error or a 500, you do not know what the server did. Retry with the same idempotency key and the same parameters, then reconcile through webhooks.
- Webhooks can arrive more than once and out of order. Deduplicate by event ID and never rely on ordering.
- Return 2xx fast, push heavy work onto a queue, and know which providers will not resend webhooks for you.
At nine on a Monday morning, the customer’s accountant sends you a screenshot: the same order has been charged twice. The logs show that the network was unstable for a while over the weekend, the payment-creation request timed out, and your job dutifully retried, exactly as it had been told to.
The code has no syntax errors and it did not crash. What was wrong was the assumption that a request arrives exactly once, in the right order, and always gets a clear answer. For a Forward Deployed Engineer, that assumption breaks almost every day.
Palantir describes the FDE role as embedding engineers alongside customers to solve their most urgent problems. Working inside a customer’s systems also means that sooner or later you will have to touch the pipes connecting those systems to others. Build those pipes solidly and you earn trust.
Build them carelessly and you will spend a week explaining why the numbers do not match.
When you get an error, you do not know what happened
Start with an uncomfortable fact. Stripe’s error-handling documentation says it plainly: on a network error, the client does not know whether the server received the request. A 500 response to a write operation must also be treated as an indeterminate result: the operation may have run, or it may not.
So blind retries are dangerous, while not retrying risks losing orders. The way out is an idempotency key: a string generated by the client, up to 255 characters long, for which Stripe suggests a UUID V4. You attach it to the request so the server can recognise resends.
Stripe’s mechanism is simple. It stores the status code and body of the first request made with each key, even if that request failed, and returns exactly that result for every retry. If you reuse a key with different parameters, the idempotency layer returns an error to catch the mistake.
Three details often escape newcomers. Keys may be removed after 24 hours, so do not treat them as a permanent record. Every POST accepts a key, while sending one with GET or DELETE has no effect, because those methods are already idempotent.
And Stripe’s advice after a network error is to retry with the same key and the same parameters until you get a result from the server.
A payment job, rewritten correctly
Back to Monday’s incident. The most important change is that the key must be generated once, when the order is created, and stored with the order, not regenerated on each retry. If it is regenerated, every retry looks like a new transaction and idempotency does nothing.
def create_payment(order):
key = order.idempotency_key # UUID v4, generated at order creation, stored in DB
payload = {
"amount": order.amount,
"currency": order.currency,
"metadata[order_id]": order.id, # for reconciliation later
}
return send_with_retry(payload, key, order)
The retry loop is separate, and every attempt uses that same key:
def send_with_retry(payload, key, order):
for attempt in range(6):
try:
r = requests.post(PAYMENTS_URL, data=payload,
headers={"Idempotency-Key": key}, timeout=10)
except (requests.ConnectionError, requests.Timeout):
time.sleep(2 ** attempt) # retry with the same key
continue
if r.status_code == 429:
time.sleep(retry_after(r, attempt))
continue
if r.status_code >= 500:
mark_indeterminate(order) # wait for webhook/reconciliation
return None
return r
mark_indeterminate(order)
return None
The 429 branch deserves attention. Rate limits are routine in integration work, and according to MDN a 429 response may include a Retry-After header saying how long to wait before sending a new request. If the server gives you a number, wait exactly that long rather than guessing:
def retry_after(r, attempt):
wait = r.headers.get("Retry-After")
return int(wait) if wait and wait.isdigit() else 2 ** attempt
The 500 branch does not retry blindly. The order is marked “indeterminate”, and the truth arrives from the other side: the webhook, plus the order_id metadata to match the transaction to the order. That is also what Stripe recommends: reconcile through webhooks and metadata.
Webhooks: late, twice, out of order
The other direction is harder, because you do not control when someone else calls you. Stripe states that an endpoint may occasionally receive the same event more than once, and that events may arrive out of order. Crossmint’s documentation says its webhooks guarantee “at least once” delivery, which means the receiver has to be idempotent.
A good handler does four things, in this order.
@app.post("/webhooks/payments")
def receive():
raw = request.get_data() # raw body, not parsed
sig = request.headers.get("Stripe-Signature")
try:
event = stripe.Webhook.construct_event(raw, sig, WEBHOOK_SECRET)
except Exception:
return "", 400 # bad or stale signature
if not processed_events.insert_if_absent(event["id"]):
return "", 200 # already seen: skip
queue.enqueue(handle_event, event["id"])
return "", 200 # return 2xx immediately
The first step is to read the raw body. Stripe signs every webhook with HMAC-SHA256 via the Stripe-Signature header, and verification needs the original body exactly. If your framework has parsed and re-serialised it, a single changed whitespace character is enough to break the signature, and you will lose half a day suspecting the secret.
The signature also includes a timestamp to prevent replay attacks, and Stripe’s library by default tolerates a gap of up to 5 minutes between that timestamp and the current time. In practice, if the clock on the customer’s server drifts significantly, every webhook will be rejected even though the secret is perfectly correct.
Next comes deduplication by event ID, then returning 2xx before any heavy work. Stripe recommends processing events through an asynchronous queue to absorb sudden spikes. In the worker, do not assume a “paid” event always arrives after a “created” event. It is safer to read the object’s current state before updating the order.
Every provider promises something different, so read the promise carefully
This is where FDEs earn trust: reading each provider’s documentation closely enough to know who is responsible for resending when your system goes down.
| Stripe | GitHub | |
|---|---|---|
| Response deadline | Return 2xx before running heavy logic | Must return 2XX within 10 seconds |
| On failed delivery | Live mode resends automatically for up to three days, with exponential backoff | No automatic redelivery |
| Your job after an outage | Cope with a backlog of old events arriving at once | Redeliver missed webhooks yourself once the server is back |
This difference shapes the architecture. With Stripe, a two-hour outage usually heals itself, as long as the handler can absorb the backlog. With GitHub, without a redelivery script or procedure, missed data is gone for good, and the runbook handed over to the customer must spell out that step.
The mistakes that keep recurring
The most common mistake is generating a new idempotency key inside the retry loop, which makes the protection useless. Close behind is reusing an old key with different parameters, such as a changed amount, and then being surprised by the error. Changing parameters makes it a new operation, which needs a new key.
The second mistake is treating a 500 as a “definite failure” and recreating the transaction with a different key. The third is doing all the business logic inside the webhook handler: writing to the database, calling the ERP, sending emails. One slow ERP call pushes you past the timeout, the provider treats the delivery as failed, and the resend cycle begins.
The last mistake is rarely discussed: having no reconciliation path. However good the code, you still need a scheduled job that compares state on both sides via metadata, because one day an event will be missed.
Putting this skill on your CV
When reading job descriptions for FDE or solutions engineer roles, look for phrases such as “integrations”, “webhooks”, “customer systems” and “data pipelines”. When you see them, have a story about idempotency and webhooks ready. Show that you have dealt with real failures, not just called APIs.
So instead of writing “Integrated Stripe”, be specific: added idempotency keys and webhook-based reconciliation to eliminate duplicate transactions; moved the webhook handler to queue-based processing with deduplication by event ID. If asked in an interview, walk through the incident, its cause and how you proved it would not happen again.
This week’s exercise: take an integration you are running, pull the network cable in the middle of a POST, then send the same webhook three times. If the data is still correct after both tests, you have done the part of the job customers will never see, and that is exactly the part that decides whether they trust you.
Was this article useful?
Thanks for the feedback!
7 sources
- Idempotent requests (Stripe API Reference)
- Advanced error handling (Stripe Docs)
- Receive Stripe events in your webhook endpoint (Stripe Docs)
- Best practices for using webhooks (GitHub Docs)
- Best Practices (Crossmint Docs, Webhooks)
- Retry-After header - HTTP | MDN · 2025-11-21
- Palantir Technologies - Forward Deployed Software Engineer