FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

Web-reading agents for clients: choosing a crawler, keeping sources honest, and why robots.txt is not a licence

A demo web-search agent can be built in an afternoon. The hard part arrives the following week, when a site returns 403, a cited link does not say what the agent claims, and the legal team wants to know who authorised the crawl.

Web-reading agents for clients: choosing a crawler, keeping sources honest, and why robots.txt is not a licence
Photo: Negative Space / CC0

In brief

  • Each tool suits a different job: Crawl4AI for a pipeline you control, Jina Reader for quick experiments, Firecrawl when you want a ready-made platform, though its credit-based pricing climbs quickly at scale.
  • An agent that cites sources is not enough: the cited source must support the specific claim, not just share its topic.
  • On 15 December 2025 the S.D.N.Y. court treated robots.txt as a mere request under DMCA §1201, but the ruling does not rule out contract or copyright claims.
ShareLinkedInFacebookX
Diagram of a web-reading agent pipeline. Four linked steps: a source list with robots.txt and terms signed off by legal; a crawler chosen by real weekly page volume; a CSS selector that keeps only the review block and is a tested config; and a source record for each claim, with a verbatim quote. A 403 branch leaves the crawler step: treat it as a signal, log the domain and pass it to the client for review. At the end is a second check that assigns one of three support labels: direct goes to the client, while topical and none are dropped or flagged. Alongside is a quote: a topical link that does not truly support the claim is more dangerous than no link.
A web-reading agent is only trustworthy when every claim has a verbatim passage that truly supports it (the direct level) and every source has been approved by the client's legal team.

A client wants an agent that reads competitors’ product reviews on G2 every week and summarises them for the product team. The demo works on the first afternoon.

By the second week the crawler is getting 403 errors, one summary sentence links to a page that says nothing about the point it makes, and the head of legal asks: who authorised taking data from this site?

Those three questions map to the three jobs an FDE has when putting a web-reading agent into production: choosing how to fetch data, controlling sources, and making the legal risk clear to the client.

The sections below follow that G2 scenario, and each ends with the first thing to do on the client’s side.

Choose a crawler for the job, not for its popularity

Fetching a page’s content is only the first step. The real questions are how much customisation you need, what scale you will run at, and who pays when the page count grows tenfold. The three common tools sit in quite different places.

Crawl4AI Jina Reader Firecrawl
What it is Open-source crawler built to plug into LLMs, agents and data pipelines; outputs Markdown optimised for RAG Page-reading service you can start using without an account Popular scraping platform for AI
Cost Open source, runs inside your own pipeline Free up to one million tokens via the API Credit-based, split into plans
Limits Cannot get past sites with strong bot detection Little customisation, not suited to large-scale crawling Costs rise sharply at scale

The table suggests a sensible order. Use Jina Reader to validate the idea on day one, while the free million tokens go a long way. Move to Crawl4AI when the pipeline needs to run long term and needs more customisation than Jina allows.

Choose Firecrawl only after estimating credit costs against the client’s real page count, not the demo’s.

First thing to do on the client’s side: ask how many pages need reading each week, across how many domains. That number decides the tool more than any comparison table.

A 403 is a signal, not just a bug

In a tutorial combining Crawl4AI with DeepSeek, Bright Data records that the first attempt on G2 returned 403. Crawl4AI could not get past such heavily protected pages, and they had to use a remote browser (Scraping Browser) to retrieve the content.

Technically, that is a fix. But for an FDE, a 403 first of all says the site owner does not want bots reading it. Whether to get past it is the client’s business and legal decision; it should not be a default buried in your code.

When you hit a 403, record the domain and add it to a list for the client to review, rather than quietly switching to a proxy.

Web pages are often longer than the context window

In the same G2 example, the review page was too long to pass whole into the DeepSeek model running on Groq, because it exceeded the token limit. The fix was to use a CSS selector to extract only the relevant part before passing it to the model.

Picture a review page with a navigation bar, ads, a footer, dozens of recommended products and a single block holding the actual reviews. Feed in the whole page and you pay for the clutter, risk overflowing the context, and give the model more chances to quote the wrong passage.

Keep only the review block and each call becomes cheaper and, more importantly, every claim can be traced back to a specific passage.

Selectors should therefore be treated as per-domain configuration, stored with the pipeline and covered by tests. When a site changes its layout, the selector breaks, and you want to find that out through a failing test, not through an empty summary sent to the client’s boss.

Citations are not the same as evidence

PraisonAI’s documentation for its Web Search Agent gives simple advice: tell the agent to list the URLs it used. That is the minimum level of source control, and many teams stop there.

Stopping there is not enough. A guide to choosing AI search tools on Useful AI warns that having a source does not mean the source says what the answer says. It proposes an evaluation criterion: the cited source must support the specific claim, not merely share its topic.

Apply this to the G2 example. The agent writes “users complain a lot about onboarding time” and cites a review page for that very product. The page is real and on topic. But if nothing in the review block mentions onboarding, the claim has no basis, even though at a glance it looks fully sourced.

The fix is to make the agent return a source record for each claim, not just a list of URLs. A minimal record should contain:

{
  "claim": "Users complain about onboarding time",
  "url": "https://...",
  "selector": "div.review-body",
  "quote": "<verbatim passage taken from the crawled block>",
  "support": "direct | topical | none"
}

(Here claim holds the claim, “users complain about onboarding time”, and quote holds the verbatim passage taken from the crawled block.)

The quote field forces the agent to point to the exact passage. A second check, either a separate LLM call or a human grading a sample, assigns the support label. Any claim rated only topical or none is dropped or flagged before it reaches the client.

PraisonAI also recommends enabling memory=True so the agent draws on earlier queries instead of searching for the same content on every run. In a weekly pipeline, this cuts the number of search calls.

First thing to do on the client’s side: agree that “correct” means direct, then measure that rate on 20 sample answers before discussing any expansion.

robots.txt: what the court said, and why you should still respect it

On 15 December 2025, in Ziff Davis v. OpenAI, the S.D.N.Y. court held that robots.txt directives are merely requests, not a mechanism controlling access to copyrighted works under DMCA §1201. According to the court, ignoring robots.txt is not “circumvention” under the DMCA.

Read closely, the ruling is narrow. It answers one question about one statute. It does not rule out other claims, such as breach of contract or copyright infringement. So “the court said robots.txt isn’t binding” cannot become a reason to crawl everything.

The safe approach for an FDE is to treat each domain’s robots.txt and terms of use as inputs to the scoping session. List the sources, record what each site permits, and have the client’s legal team sign off. You are not a lawyer, and you should not play one.

Your job is to make the risk visible so that the people with authority can decide.

Five steps for your next web-reading agent project

  1. Draw up the source list: which domains, how many pages, how often, and what each site’s robots.txt and terms of use say.
  2. From the scale figures, choose the tool and estimate costs against the real page count.
  3. Write a selector for each domain, with tests that catch layout changes.
  4. Design per-claim source records, enable memory to reduce repeated searches, and add a support check before results go out.
  5. Hand the client an approved source table along with a procedure for handling 403s.

The most common mistakes

The most common mistake is estimating costs from the demo and then being surprised when the system runs at real scale. The second is feeding whole pages into the model, which wastes tokens and makes citations less precise. The third is checking only whether an answer has links, not whether those links support the exact claim.

The fourth is harder to spot: treating bot-detection bypass as a purely technical problem, then citing the robots.txt ruling as if it were a licence.

These mistakes usually surface only once the agent runs in production, so having dealt with them is valuable material for an FDE CV or interview.

Instead of writing “built a RAG agent on web data”, describe concrete work such as “designed a per-claim citation verification layer” or “built a legally approved data source catalogue”, and be ready to tell the story of a time you hit a 403 and how you handled it.

A web-reading agent is not judged by how many pages it finds. It is judged by whether every sentence it writes can be defended before the product team and, when necessary, before the legal department.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
5 sources
Read next on the roadmap · Stage 3: Applied AIWriting tool definitions for a customer's API: six steps to get the model to pick the right tool and fill in the right parametersThe model never reads the customer's code or Swagger file. Everything it knows about the API is in the few lines of description you write, so a wrong tool call usually starts with those lines.