FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

APIs and webhooks for FDEs: build every integration as if everything arrives twice

The costliest bug in an enterprise integration rarely takes the system down. It quietly charges a customer twice, or drops an event that nobody notices is gone.

APIs and webhooks for FDEs: build every integration as if everything arrives twice
Photo: Negative Space / CC0

In brief

  • After a network error or a 500, you do not know what the server did. Retry with the same idempotency key and the same parameters, then reconcile through webhooks.
  • Webhooks can arrive more than once and out of order. Deduplicate by event ID and never rely on ordering.
  • Return 2xx fast, push heavy work onto a queue, and know which providers will not resend webhooks for you.
ShareLinkedInFacebookX
Infographic in two parts. Sending: an order has an idempotency key generated once. Attempts 1 and 2 time out and attempt 3 returns a result, all sent with the same key. The server remembers the result by key, so the customer is charged only once. On a 500 error, the order is marked “unknown” and the sender waits for the webhook. Receiving: a webhook passes four steps: get the raw body, verify the HMAC-SHA256 signature with a clock skew of at most 5 minutes, deduplicate by event ID, then return 2xx at once and push to a queue. A worker then reads the current state to do the heavy processing. Notes: Stripe live mode retries webhooks for up to three days, while GitHub does not retry automatically.
When sending, every retry must use the same key and parameters. When receiving, webhooks must be deduplicated before they are processed.

At nine on a Monday morning, the customer’s accountant sends you a screenshot: the same order has been charged twice. The logs show that the network was unstable for a while over the weekend, the payment-creation request timed out, and your job dutifully retried, exactly as it had been told to.

The code has no syntax errors and it did not crash. What was wrong was the assumption that a request arrives exactly once, in the right order, and always gets a clear answer. For a Forward Deployed Engineer, that assumption breaks almost every day.

Palantir describes the FDE role as embedding engineers alongside customers to solve their most urgent problems. Working inside a customer’s systems also means that sooner or later you will have to touch the pipes connecting those systems to others. Build those pipes solidly and you earn trust.

Build them carelessly and you will spend a week explaining why the numbers do not match.

When you get an error, you do not know what happened

Start with an uncomfortable fact. Stripe’s error-handling documentation says it plainly: on a network error, the client does not know whether the server received the request. A 500 response to a write operation must also be treated as an indeterminate result: the operation may have run, or it may not.

So blind retries are dangerous, while not retrying risks losing orders. The way out is an idempotency key: a string generated by the client, up to 255 characters long, for which Stripe suggests a UUID V4. You attach it to the request so the server can recognise resends.

Stripe’s mechanism is simple. It stores the status code and body of the first request made with each key, even if that request failed, and returns exactly that result for every retry. If you reuse a key with different parameters, the idempotency layer returns an error to catch the mistake.

Three details often escape newcomers. Keys may be removed after 24 hours, so do not treat them as a permanent record. Every POST accepts a key, while sending one with GET or DELETE has no effect, because those methods are already idempotent.

And Stripe’s advice after a network error is to retry with the same key and the same parameters until you get a result from the server.

A payment job, rewritten correctly

Back to Monday’s incident. The most important change is that the key must be generated once, when the order is created, and stored with the order, not regenerated on each retry. If it is regenerated, every retry looks like a new transaction and idempotency does nothing.

def create_payment(order):
    key = order.idempotency_key          # UUID v4, generated at order creation, stored in DB
    payload = {
        "amount": order.amount,
        "currency": order.currency,
        "metadata[order_id]": order.id,  # for reconciliation later
    }
    return send_with_retry(payload, key, order)

The retry loop is separate, and every attempt uses that same key:

def send_with_retry(payload, key, order):
    for attempt in range(6):
        try:
            r = requests.post(PAYMENTS_URL, data=payload,
                              headers={"Idempotency-Key": key}, timeout=10)
        except (requests.ConnectionError, requests.Timeout):
            time.sleep(2 ** attempt)     # retry with the same key
            continue
        if r.status_code == 429:
            time.sleep(retry_after(r, attempt))
            continue
        if r.status_code >= 500:
            mark_indeterminate(order)    # wait for webhook/reconciliation
            return None
        return r
    mark_indeterminate(order)
    return None

The 429 branch deserves attention. Rate limits are routine in integration work, and according to MDN a 429 response may include a Retry-After header saying how long to wait before sending a new request. If the server gives you a number, wait exactly that long rather than guessing:

def retry_after(r, attempt):
    wait = r.headers.get("Retry-After")
    return int(wait) if wait and wait.isdigit() else 2 ** attempt

The 500 branch does not retry blindly. The order is marked “indeterminate”, and the truth arrives from the other side: the webhook, plus the order_id metadata to match the transaction to the order. That is also what Stripe recommends: reconcile through webhooks and metadata.

Webhooks: late, twice, out of order

The other direction is harder, because you do not control when someone else calls you. Stripe states that an endpoint may occasionally receive the same event more than once, and that events may arrive out of order. Crossmint’s documentation says its webhooks guarantee “at least once” delivery, which means the receiver has to be idempotent.

A good handler does four things, in this order.

@app.post("/webhooks/payments")
def receive():
    raw = request.get_data()                       # raw body, not parsed
    sig = request.headers.get("Stripe-Signature")
    try:
        event = stripe.Webhook.construct_event(raw, sig, WEBHOOK_SECRET)
    except Exception:
        return "", 400                             # bad or stale signature
    if not processed_events.insert_if_absent(event["id"]):
        return "", 200                             # already seen: skip
    queue.enqueue(handle_event, event["id"])
    return "", 200                                 # return 2xx immediately

The first step is to read the raw body. Stripe signs every webhook with HMAC-SHA256 via the Stripe-Signature header, and verification needs the original body exactly. If your framework has parsed and re-serialised it, a single changed whitespace character is enough to break the signature, and you will lose half a day suspecting the secret.

The signature also includes a timestamp to prevent replay attacks, and Stripe’s library by default tolerates a gap of up to 5 minutes between that timestamp and the current time. In practice, if the clock on the customer’s server drifts significantly, every webhook will be rejected even though the secret is perfectly correct.

Next comes deduplication by event ID, then returning 2xx before any heavy work. Stripe recommends processing events through an asynchronous queue to absorb sudden spikes. In the worker, do not assume a “paid” event always arrives after a “created” event. It is safer to read the object’s current state before updating the order.

Every provider promises something different, so read the promise carefully

This is where FDEs earn trust: reading each provider’s documentation closely enough to know who is responsible for resending when your system goes down.

Stripe GitHub
Response deadline Return 2xx before running heavy logic Must return 2XX within 10 seconds
On failed delivery Live mode resends automatically for up to three days, with exponential backoff No automatic redelivery
Your job after an outage Cope with a backlog of old events arriving at once Redeliver missed webhooks yourself once the server is back

This difference shapes the architecture. With Stripe, a two-hour outage usually heals itself, as long as the handler can absorb the backlog. With GitHub, without a redelivery script or procedure, missed data is gone for good, and the runbook handed over to the customer must spell out that step.

The mistakes that keep recurring

The most common mistake is generating a new idempotency key inside the retry loop, which makes the protection useless. Close behind is reusing an old key with different parameters, such as a changed amount, and then being surprised by the error. Changing parameters makes it a new operation, which needs a new key.

The second mistake is treating a 500 as a “definite failure” and recreating the transaction with a different key. The third is doing all the business logic inside the webhook handler: writing to the database, calling the ERP, sending emails. One slow ERP call pushes you past the timeout, the provider treats the delivery as failed, and the resend cycle begins.

The last mistake is rarely discussed: having no reconciliation path. However good the code, you still need a scheduled job that compares state on both sides via metadata, because one day an event will be missed.

Putting this skill on your CV

When reading job descriptions for FDE or solutions engineer roles, look for phrases such as “integrations”, “webhooks”, “customer systems” and “data pipelines”. When you see them, have a story about idempotency and webhooks ready. Show that you have dealt with real failures, not just called APIs.

So instead of writing “Integrated Stripe”, be specific: added idempotency keys and webhook-based reconciliation to eliminate duplicate transactions; moved the webhook handler to queue-based processing with deduplication by event ID. If asked in an interview, walk through the incident, its cause and how you proved it would not happen again.

This week’s exercise: take an integration you are running, pull the network cable in the middle of a POST, then send the same webhook three times. If the data is still correct after both tests, you have done the part of the job customers will never see, and that is exactly the part that decides whether they trust you.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
7 sources
Read next on the roadmap · Stage 2: Broad engineeringWhat FDE recruiters look for: code gets you in, customer communication sets the rankingYou cannot work as an FDE without coding. Once candidates have shown they can code, what decides who gets hired is how well they communicate with customers and win them over.