FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Tools

DataHub and Atlan: using a data catalog to understand a client's data in week one

Atlan's documentation says you can connect your first source in minutes. For an FDE, the hard part comes after that: finding out who owns each dataset and which data touches personal information.

In brief

  • DataHub has an open-source core under Apache 2.0, with ingestion through the UI or the CLI. Atlan is proprietary SaaS and does not publish its prices.
  • Each has more than 80 connectors, so the choice usually depends on whether the client wants to run the tool itself or buy a managed product.
  • Connecting sources is the easy step. The value comes from assigning owners, tagging PII and classifying sensitivity.
ShareLinkedInFacebookX
Infographic on the first week with a data catalog. Five linked steps: fix scope, ingest, search, assign owner (highlighted), flag PII. Example: searching "order" returns 40 assets; 12 have an owner and 28, or 70%, do not. Assets are then classified into four levels: public, internal use, restricted, confidential. The product catalog sits at internal use, while the customer table holding phone numbers sits at restricted or confidential.
At a client site, connecting sources is only the start. The hard part is showing who is accountable for each data asset and which data touches personal information.

Atlan’s documentation says you can set up the platform, connect your first data sources and learn the core concepts “in minutes”. That is the vendor’s claim, and it may be true of the technical connection step.

Any FDE who has started with a new client knows those minutes do not answer the real questions of the first week: which tables can be trusted, who is responsible for them, and which columns hold personal information.

That is why data catalogs are worth learning. When you are asked to build an agent or a pipeline on a client’s data, the costliest mistake is rarely the model. More often it is using the wrong table, or sending sensitive data somewhere it is not allowed to go.

DataHub and Atlan represent two different approaches: one is built on an open-source core, the other is proprietary SaaS. If you understand what each does, and above all the order in which to use them, you can map a client’s data within the first week.

Two tools, two routes to the same map

DataHub describes itself as a context management platform for AI and data that brings discovery, governance and observability together across the whole data stack. Atlan calls itself a “context layer” for AI that links data to business definitions and access policies.

On connectors, Atlan’s documentation advertises more than 80 sources, and the independent comparison site Modern DataTools rates the two as equal, with more than 80 each. The real difference is the deployment model.

DataHub

  • Open-source core, Apache 2.0 licence
  • DataHub Core is free; DataHub Cloud is for enterprises
  • Self-hosted or managed
  • Ingestion through the UI or the CLI

Atlan

  • Proprietary SaaS
  • Prices not published
  • Core capabilities: catalog, governance, lineage, quality, glossary

So the first question at a client site is not which tool is better but what the client already has. If the client has an infrastructure team and wants everything kept on its internal network, a self-hosted DataHub is the obvious choice. If the client has already bought Atlan, your job is to use it well, not to talk them into switching.

What does week one with a catalog look like?

Picture a retail company that wants an agent to answer questions about orders. It has a data warehouse, a few dashboards and a folder of raw files that nobody remembers creating. On day one, the first job is not writing prompts but connecting these sources to the catalog.

With DataHub you can load metadata through the interface for speed, or use the CLI so the configuration is saved, reviewed and rerun. At a client site the second approach is worth more, because it leaves a record for whoever comes next.

Once the metadata is in, the most useful tool is the search box. DataHub lets you search across the whole ecosystem, including dashboards, datasets, ML models and raw files. Type “order” and, for the first time, you see every place the concept of an order appears, rather than only the table your guide mentioned.

A simple calculation shows the real work

Suppose that search returns 40 assets related to orders. You ask around and can identify owners for only 12 of them. That leaves 28 assets, or 70%, with nobody taking responsibility for them, and that is the figure you take to the end-of-week meeting.

DataHub lets you define ownership and track PII inside the catalog. The next step, then, is to assign owners to the 12 known assets and tag the columns holding personal information such as names, phone numbers and delivery addresses.

Then comes classification. Palo Alto Networks defines data classification as organising data by sensitivity, importance and predefined criteria. The common levels are public, internal use, restricted and confidential.

Applied to the example above, a product catalogue table might be internal use, while a customer table with phone numbers clearly belongs in restricted or confidential.

Two mistakes that derail the first week

The most common mistake is connecting every source before agreeing the scope with the client. With more than 80 connectors this is very easy to do. But a catalog full of metadata unrelated to the orders problem makes search results noisier, and you may end up touching client systems you have not been cleared to see.

Agree the list of sources in writing before the first ingestion.

The second mistake is tagging PII without assigning a named owner. If a column is labelled as a phone number but nobody is responsible for it, then when the agent needs read access, no one has the authority to approve or refuse. That is why, in the example above, assigning owners comes before tagging PII.

A catalog will not do the hard part for you

The first limit is the promise of speed. Atlan’s “minutes” is a vendor claim about the connection step. Assigning owners needs people to answer questions, and no tool can work out on its own who is responsible for an anonymous raw file.

The second limit concerns how to think about context for AI. In a blog post in April 2026, DataHub itself argued that a semantic layer is necessary but not sufficient: agents also need lineage, ownership, data freshness and access permissions. The post also argued that governance is not a feature to add later.

This is the view of a company that sells a catalog, so it has a motive, but the order of work it suggests matches what happens in real deployments.

The lesson for FDEs is clear. If you build the agent first and ask about permissions afterwards, you will have to redo exactly the part the client cares about most.

What to learn first, and how to show it

Start with DataHub Core, because it is open source and you can run it yourself without a contract. Learn it in the order of real work: ingest through the UI and then the CLI, search, assign ownership, tag PII. Atlan is harder to get access to unless a client already uses it, so reading its documentation to get familiar with glossary and lineage concepts is enough at the start.

When interviewing for an FDE role, ask one question: “At your most recent client, who decided which tables the agent could read, and where was that decision recorded?” The answer tells you whether the company treats governance as the FDE’s job or someone else’s.

On your CV, do not write “experience with data catalogs”. Write what you achieved, with numbers: how many sources you connected, the share of assets with an owner before and after, and how many PII columns you tagged.

Almost anyone a client hires can build the model. What earns their trust is producing, by the end of the first week, a list that states clearly which data is safe for the agent to read.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
7 sources
Read next on the roadmap · Stage 2: Broad engineeringAirbyte: pulling customer data into one place without writing scriptsIn the first week on site, a Forward Deployed Engineer usually spends the most time collecting data from five different systems. The model often goes untouched. Airbyte exists to take some of that work away.