Writing an MCP client for a customer's remote server: from a 401 to the first tool call
A customer has just sent you an MCP URL, and you have one afternoon to get the first tool call working. JSON-RPC is rarely the problem. A few missing headers usually are.
In brief
- Streamable HTTP is the standard transport for remote MCP servers. There is one endpoint, every message is a new POST, and the client must accept both JSON and SSE.
- The token goes in the Authorization header on every request and never in the query string. Use only a token issued by that MCP server's own authorization server.
- After initialize, every request must send back the session ID and the negotiated protocol version. On a 404, run the handshake again. Every request needs a timeout.
- 1POST initializeAccept lists both application/json and text/event-stream; set a timeout
- 2Receive 401Read the WWW-Authenticate header to find the auth metadata
- 3Fetch metadataFrom /.well-known/oauth-protected-resource, find the token-issuing auth server
- 4OAuth with PKCESend the resource parameter with both the authorize and token exchange requests
- 5Re-initializeBearer token, store Mcp-Session-Id, read the version from JSON or SSE
- 6Initialized, call a tooltools/call carries the token, session ID and MCP-Protocol-Version
Most failures when connecting to a customer's MCP server are in the headers, not the agent logic.
Graphic: FDE Times
On Monday a customer drops you a one-line Slack message. It contains the URL of the internal MCP server they have just stood up and says “just point your agent at this”. You try a quick POST and get a 401. The customer’s team is not in the room, the documentation is thin, and the demo is on Thursday morning.
For a forward deployed engineer, this is an ordinary week. You do not own the server or the authentication infrastructure. You do own the client that calls them, and it has to work reliably. This guide goes from that 401 to the first tool call, one request at a time.
Start with this: at this stage most failures come from HTTP headers, not from agent logic. If you know a handful of rules from the spec, you can read the customer’s server errors instead of guessing at them.
What protocol does a remote server speak?
The MCP spec defines two standard transports. Streamable HTTP is the one for remote servers. It replaces the HTTP+SSE mechanism from the 2024-11-05 version. Sample code that opens separate connections for SSE and for POST is almost certainly out of date.
The Streamable HTTP model is simple. The server exposes exactly one path, called the MCP endpoint, which accepts both POST and GET. Each JSON-RPC message from the client is a new POST. The Accept header must list both application/json and text/event-stream, because the server may answer with plain JSON or with an SSE stream.
The lifecycle also runs in a fixed order. Initialize must be the first interaction. Once the server responds, the client sends an initialized notification, and only then do normal operations such as tool calls begin. The order is mandatory: initialize, then initialized, then tools.
Example: probing the customer’s server over raw HTTP
Before reaching for an SDK, send the requests yourself with httpx. You see exactly what the server returns, so when the SDK later reports a vague error you know where to look. Assume the customer’s endpoint is https://mcp.khach.vn/mcp.
import httpx
URL = "https://mcp.khach.vn/mcp"
H = {"Accept": "application/json, text/event-stream",
"Content-Type": "application/json"}
init = {"jsonrpc": "2.0", "id": 1, "method": "initialize",
"params": {"protocolVersion": "2025-06-18", "capabilities": {},
"clientInfo": {"name": "fde-probe", "version": "0.1"}}}
Send it, and set a timeout from the very first line. The spec recommends a timeout on every request, with the request cancelled when it expires, so connections do not hang. That matters even more when the server sits behind the customer’s VPN.
r = httpx.post(URL, json=init, headers=H, timeout=10)
print(r.status_code)
print(r.headers.get("www-authenticate"))
If the server is protected, the response will be a 401 with a WWW-Authenticate header. The spec requires clients to parse this header and respond to it correctly, because it points to the authentication metadata.
From a 401 to an access token
Protected resource metadata lives at a default well-known address: /.well-known/oauth-protected-resource. From there the client finds which authorization server issues tokens for this MCP server.
meta_url = "https://mcp.khach.vn/.well-known/oauth-protected-resource"
meta = httpx.get(meta_url, timeout=10).json()
as_url = meta["authorization_servers"][0]
The next step needs two addresses on the authorization server: one that accepts the authorize request and one that issues tokens. How you look them up depends on the customer’s infrastructure. In the example below they appear as two placeholder variables, AUTHORIZE and TOKEN_ENDPOINT, along with CLIENT_ID and REDIRECT.
Then comes the OAuth flow, where the spec sets two hard requirements. The client must use PKCE, and it must send the resource parameter so the token is bound to the intended MCP server. The setup looks like this.
import base64, hashlib, secrets
from urllib.parse import urlencode
verifier = secrets.token_urlsafe(64)
digest = hashlib.sha256(verifier.encode()).digest()
challenge = base64.urlsafe_b64encode(digest).rstrip(b"=").decode()
auth_params = {"response_type": "code", "client_id": CLIENT_ID,
"redirect_uri": REDIRECT, "code_challenge": challenge,
"code_challenge_method": "S256", "resource": URL}
login_url = f"{AUTHORIZE}?{urlencode(auth_params)}"
The user opens login_url in a browser and signs in, and the authorization server sends a code to the redirect. The client exchanges the code for a token, sending the verifier and repeating the same resource value.
tok = httpx.post(TOKEN_ENDPOINT, timeout=10, data={
"grant_type": "authorization_code", "code": code,
"redirect_uri": REDIRECT, "client_id": CLIENT_ID,
"code_verifier": verifier, "resource": URL}).json()
A practical tip: with a customer’s internal server, you will probably have to ask their team for CLIENT_ID, REDIRECT and the two endpoints. Put them on the question list for the first call, before you write any code.
The second handshake, done properly
With a token in hand, send initialize again. The token goes in the Authorization header. The spec requires this header on every request, even when the requests belong to the same logical session.
H["Authorization"] = f"Bearer {tok['access_token']}"
r = httpx.post(URL, json=init, headers=H, timeout=10)
sid = r.headers.get("mcp-session-id")
if sid:
H["Mcp-Session-Id"] = sid
If the server assigns an Mcp-Session-Id during initialization, the client must send it back on every later request. The client must also send an MCP-Protocol-Version header with the negotiated version. Do not hard-code the version you proposed. Read the version the server returns in the initialize result.
Because Accept allows SSE, the initialize response may arrive as an event stream instead of JSON. The snippet below handles both cases. For SSE, it extracts the data line before parsing.
import json
ctype = r.headers.get("content-type", "")
if ctype.startswith("text/event-stream"):
line = next(l for l in r.text.splitlines() if l.startswith("data:"))
body = json.loads(line[5:])
else:
body = r.json()
H["MCP-Protocol-Version"] = body["result"]["protocolVersion"]
Only now send the initialized notification, with the full set of headers.
note = {"jsonrpc": "2.0", "method": "notifications/initialized"}
httpx.post(URL, json=note, headers=H, timeout=10)
If you call .json() directly without checking Content-Type, you can lose half an hour to a meaningless parse error.
The first tool call
H now holds Accept, Authorization, Mcp-Session-Id and MCP-Protocol-Version. Pick a harmless tool to call, such as a read-only lookup. The tool name and arguments below are hypothetical (an order lookup with an order ID). Replace them with a real tool that the customer’s server exposes.
call = {"jsonrpc": "2.0", "id": 2, "method": "tools/call",
"params": {"name": "tra_cuu_don_hang",
"arguments": {"ma_don": "DH001"}}}
r = httpx.post(URL, json=call, headers=H, timeout=10)
print(r.status_code, r.headers.get("content-type"))
A 200 with a result means you have completed the whole path, from a 401 through OAuth to a real tool call. Keep the logs from this run. They will be your best reference in the days ahead.
Once you understand the headers, hand the heavy lifting to the SDK
The raw HTTP probe is there for understanding, not for production. The official Python SDK’s README says that passing a URL to Client selects Streamable HTTP, and that tools are invoked with call_tool. The official client-building tutorial requires SDK version 2.0.0 or later, so check your installed version before copying any sample code from the web.
The safest order of work is this. Probe over raw HTTP until tools/call succeeds, and record every header at each step. Then move to the SDK and add the OAuth flow following the guide for your exact version. Next, use call_tool to repeat the same read-only call you tested. Only then wire it into the agent.
When the SDK throws an error, go back to the header log from the probe. Usually it shows straight away which header is missing.
The mistakes that sink demos
The most common mistake is putting the token in the URL, such as appending ?access_token= as a quick test. The spec explicitly forbids this, and URLs are often written to proxy and gateway logs on the customer’s side. The token belongs only in the Authorization header.
The next mistake is subtler: using the wrong token. The customer already has an API key for another internal system, and someone suggests “just use that for now”. The spec forbids clients from sending an MCP server any token not issued by that server’s own authorization server. The resource parameter exists for much the same reason.
Another is treating a 404 as fatal. When the server returns a 404 to a request carrying a session ID, the client must start a new session from initialize. Your code needs a re-handshake path with a retry limit, instead of throwing an exception up to the agent.
Three smaller mistakes still waste hours: an Accept header that lists only one content type, forgetting to send MCP-Protocol-Version after initialize, and requests without timeouts that leave the agent hanging when the customer’s VPN is flaky. Add copying old HTTP+SSE code or SDK snippets from before 2.0, and you have a checklist to run before every demo.
How to put this skill on your CV
When reading job descriptions for FDE roles, look for phrases such as “integrate with customer systems”, “OAuth”, “MCP” or “agent tooling”. They tell you the job will involve exactly the situation at the start of this article.
On a CV, one specific line carries more weight than a list of technologies. For example: “Built an MCP client over Streamable HTTP with OAuth (PKCE, resource), handling session expiry and timeouts when calling servers behind customer VPNs.” Interviewers will follow up on 401s and 404s, and your own probe script will have given you the answers.
The next time a customer sends a URL with “just point it at this”, you will not start with the SDK. You will start with a short script that prints the headers at every step.
Was this article useful?
Thanks for the feedback!