FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

L4 and L7 load balancers and reverse proxies: running your service behind a customer's network

When your service has to run behind a customer's load balancer, three faults tend to surface together: users' real IPs disappear, sessions jump between servers and health checks report the wrong thing. All three can be fixed once you know what each network layer can see.

In brief

  • L4 sees only addresses and ports; L7 can read HTTP headers. Find out which one the customer uses before writing your first line of configuration.
  • Behind a proxy, your server sees only the IP of the last proxy. Trust X-Forwarded-For only from proxies you have identified as trustworthy.
  • A service is easy to put behind a load balancer when it keeps no sessions in RAM, has a real health check endpoint and lets the load balancer handle TLS if the customer wants that.
ShareLinkedInFacebookX

Imagine it is your third day at a bank. Your internal document Q&A service has been running smoothly on the staging machine. Then the infrastructure team sends a message: all traffic must go through the bank’s load balancer, no exceptions.

An hour after the switch, the logs record the same IP address for every request. Your per-IP rate limit blocks an entire department, and users are occasionally logged out mid-conversation. There is nothing wrong with your code. The problem is that you do not yet understand how the customer’s network sees your requests.

This is everyday work for an FDE. You rarely get to build infrastructure from scratch; you have to fit your service into an existing system, run by other people, with its own rules. To do that well, you need to tell three things apart clearly: the layer 4 load balancer, the layer 7 load balancer and the reverse proxy.

What does the customer’s load balancer see?

According to F5’s glossary, a layer 4 load balancer routes using transport-layer information, meaning addresses and ports, and does not read packet contents. It simply passes a TCP stream to a server, with no idea which URL or cookie is inside.

A layer 7 load balancer makes decisions based on characteristics of the HTTP header, such as the URL or a cookie. Because it reads HTTP, it can send /api to one cluster and /static to another. It can also insert extra headers before passing the request inside.

This matters because each layer determines what your service receives, for example whether there is an X-Forwarded-For header inserted by a proxy. TLS is a separate question: do not guess from the words L4 or L7 alone. Whether decryption happens at the load balancer or at your service depends on how the customer has configured it, so you have to ask.

A reverse proxy is not a load balancer

The two terms are often used interchangeably. A reverse proxy takes a request from a client and forwards it to a server that can handle it, while a load balancer distributes requests across a group of servers. That is why a reverse proxy is still useful even with only one server behind it.

Even if the customer already has a load balancer, you should still run a reverse proxy of your own, such as NGINX, directly in front of your app. It is the place you control: recovering the real IP, setting timeouts and answering health checks. You do not have to ask the infrastructure team to change their equipment every time you need a configuration tweak.

NGINX suits this role because of its architecture. The NGINX engineering blog explains that the common way to design network applications is to assign a thread or process to each connection, and that this incurs context-switching costs.

NGINX opts for an event-driven architecture: a master process handles privileged work such as reading configuration and binding ports, while worker processes do the main work.

A worked example: recovering the user’s real IP

Back to the bank. A request now travels: employee’s browser → the bank’s L7 load balancer → your NGINX → app. When a connection passes through any proxy, the server sees only the IP of the last proxy, so your logs are full of the load balancer’s IP.

The fix is the X-Forwarded-For header: the proxy inserts the client’s original IP into this header before forwarding the request inside. The app reads that header to learn who is really calling.

But MDN warns that if you use this header for security purposes, such as rate limiting or IP-based access control, you must use only the IPs added by trusted proxies. The reason is that a client can send a forged header of its own.

Suppose the infrastructure team tells you their load balancer sits in the 10.20.0.0/24 range. Your NGINX configuration would look like this:

# Chỉ tin X-Forwarded-For khi request đến từ load balancer của khách
set_real_ip_from 10.20.0.0/24;
real_ip_header   X-Forwarded-For;
real_ip_recursive on;

upstream app {
    server 127.0.0.1:8000;
}

server {
    listen 80;

    location /healthz {
        proxy_pass http://app;
    }

    location / {
        proxy_set_header X-Forwarded-For   $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $http_x_forwarded_proto;
        proxy_set_header Host              $host;
        proxy_pass http://app;
    }
}

(The comment on the first line reads: “Only trust X-Forwarded-For when the request comes from the customer’s load balancer.”)

The set_real_ip_from line is the most important in the whole configuration. Without it, anyone who sends X-Forwarded-For: 1.2.3.4 can impersonate someone else and get past your rate limit.

With it, NGINX trusts the value only when it comes from the load balancer’s IP range. The exercise at the end shows how to verify this with curl.

Why sessions jump, and why sticky sessions will not save you

Next, the logouts. Suppose you run 3 instances and configure balancing by source IP. HAProxy’s architecture documentation describes this method as follows: the same IP always reaches the same server, as long as the number of servers does not change.

Look at how both conditions break. If you hash on the IP your layer sees, every request carries the load balancer’s IP, so your 3 instances collapse into 1 overloaded instance. If next week you scale to 4 instances, the hash redistributes, many users are moved to a different server and lose the session held in RAM.

The more durable approach is to keep no sessions in the app. In a horizontal scaling model, the servers sit behind the load balancer and sessions live in a shared data store that every app server can reach.

At a customer site, that store is usually a Redis instance or a database they already run. In the same bank scenario, the two approaches play out like this:

Scenario Sticky by IP, sessions in RAM Sessions in a shared store
Every request carries the load balancer’s IP All 3 instances collapse into 1 overloaded instance Any instance can read the user’s session
Scaling from 3 to 4 instances The hash redistributes and many users are logged out Users stay logged in

Five steps before deployment day

  1. Ask the infrastructure team three questions. Is the load balancer L4 or L7, where does TLS terminate, and what is the proxies’ IP range? These three answers determine most of your configuration.
  2. Decide who handles TLS. The load balancer can take on encryption and decryption so the servers can focus on their main work. If the customer already does this, read X-Forwarded-Proto to tell whether the original request was HTTPS, rather than configuring certificates yourself.
  3. Write a /healthz that checks for real. According to AWS documentation, Elastic Load Balancing checks the health of its targets and sends requests only to healthy ones. An endpoint that always returns 200 while the app has lost its database connection will keep sending users to a broken instance.
  4. Move sessions and temporary data to a shared store. Then the customer can add or remove instances without anyone being logged out.
  5. Agree on timeouts. Ask what the equipment’s timeout is. Requests that call an LLM can run longer than the default, and being cut off midway is a very hard bug to track down.

Common mistakes

The most common mistake is trusting every X-Forwarded-For without restricting its source, then using it for rate limiting. The second is enabling IP-based sticky sessions without realising that every request appears to come from the same address.

The third is assuming that a load balancer removes the need for a reverse proxy, so every configuration change means waiting on another team’s ticket. The fourth is harder to spot: a health check that only confirms the process is running, not whether the app can actually serve requests.

For developers moving into FDE roles, this is a skill worth spelling out on a CV. Do not just write “knows NGINX”. Describe a time you put a service behind an existing proxy, kept the real client IP and ran the service on multiple instances without losing sessions.

When a job description mentions “customer environment” or “on-prem deployment”, you can expect to be asked about exactly these things.

Exercise: try to fool your own proxy

Set up NGINX with the configuration above in front of a small app that does just one thing: log the client IP. Set set_real_ip_from to the IP range of a machine you treat as the load balancer.

From a machine outside that range, run curl -H "X-Forwarded-For: 1.2.3.4" http://<proxy>/. If the log shows 1.2.3.4, your proxy is being fooled. If it shows the sending machine’s real IP, the configuration is correct.

Then remove the set_real_ip_from line, run the command again and compare the two log lines. The customer’s network is something you cannot change, but a ten-minute test like this tells you whether your service will hold up inside it.

7 sources
Read next on the roadmap · Stage 5: DeploymentCutting LLM latency and cost: four levers and the order to pull themWhen a system is slow and expensive, the usual first move is a cheaper model. That lever belongs near the end, and only once an eval is in place.