FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Tools

AWS Glue, Azure Data Factory and Google Dataflow: how to read the ETL pipeline a client already runs

On a client site, an FDE rarely gets to choose the ETL tool. Usually the job is to take over a pipeline that is already running, so the first task is to understand it before changing anything.

In brief

  • Glue and ADF data flows both run on Spark managed by the service. Dataflow is built on Apache Beam, which uses one model for both batch and streaming.
  • At an Azure client, ask early whether they use ADF or Data Factory in Microsoft Fabric, because Microsoft is steering new users towards Fabric.
  • What to open first when you take over a pipeline: the Data Catalog on AWS, the integration runtime on Azure, the Beam code on Google Cloud.
ShareLinkedInFacebookX

“Are you on Azure Data Factory or Data Factory in Microsoft Fabric?” At a client running Azure, this should be one of the first questions an FDE asks.

The reason is on Microsoft’s own overview page. Microsoft still describes ADF as a managed cloud service for complex ETL, ELT and hybrid data integration projects. On the same page, however, it calls Data Factory in Fabric the “next generation” of ADF and points new users to Fabric.

The question reflects a wider reality: when FDEs work with data, they rarely start from scratch. By the time you arrive, the client is usually already running AWS Glue, Azure Data Factory or Google Dataflow. The data your agent or model needs to read flows through those pipelines.

You do not need to be an expert in all three. You do need to be able to read the client’s pipeline within your first week.

Three services, three vocabularies

AWS describes Glue as a serverless data integration service for discovering, preparing, moving and integrating data from multiple sources. Crawlers sit at the centre of Glue: a crawler infers the schema and writes it to the Glue Data Catalog. Glue Studio lets users build transformation workflows by drag and drop and run them on a serverless ETL engine based on Apache Spark.

Those who prefer to write code can use Spark, Python or Scala.

ADF has its own set of concepts: pipelines, activities, datasets, linked services, data flows and integration runtimes. Mapping data flows run on a Spark cluster that starts when needed and shuts down afterwards, so users do not manage the cluster. The integration runtime is the bridge between activities and linked services. In other words, it decides where the pipeline actually runs.

Google defines Dataflow as a unified batch and stream processing service that runs at scale. Dataflow was announced in June 2014, entered public beta in April 2015 and is now built on the open-source Apache Beam project. As load changes, the service adds or shuts down worker VMs automatically.

One problem, three places to look first

Picture a retail chain that wants an agent to answer questions about inventory. Every night, order files land in a data lake, while the inventory table sits in a database in the office. Your job is to bring both sources into one clean place the agent can query. The table below suggests where to look on each platform.

Step AWS Glue Azure Data Factory Google Dataflow
Know what the data contains Check the schema the crawler wrote to the Data Catalog Read the datasets and linked services Read the source-reading step in the Beam code
Reach the on-premises database Open the job, note the connection to the office database, then ask the network team about its route and access permissions Check which integration runtime acts as the bridge Find the database-reading step in the Beam code, then ask the network team whether the worker VMs can reach the office database
Transform Glue Studio or PySpark jobs Mapping data flows running on Spark Beam transforms
When it runs On a schedule, on demand or on an event Open the pipeline, look at the order of activities and ask the client what triggers it Ask the client how the job is launched, and whether it is a batch or streaming job

Read across the rows and the hardest question in this example sits in the second one. The integration runtime is the bridge between activities and linked services, so on Azure it is the first place to check whether the pipeline can connect to the office database.

Before writing any transform, sit down with the client’s network and security teams to pin down that bridge.

Where are the limits?

All three services hide much of the infrastructure, and that is exactly what makes FDEs complacent. Spark starting and stopping on its own, or workers scaling automatically, does not mean the transformation logic is correct. A schema inferred by a crawler from the data is still a guess, so check it with whoever owns that data source.

Azure adds a strategic question. Microsoft says existing ADF workloads can be upgraded to Fabric. Any proposal to build something new on ADF should therefore come with a direct question to the client: do they plan to move to Fabric? Skip it, and you may find yourself planning an upgrade just after handover.

With Dataflow, the difficulty lies in the programming model rather than the infrastructure. Beam uses one model for both batch and streaming. The model is powerful, but people used to writing sequential scripts will need time to adjust.

What to learn first

If you have only a month, learn Spark first. On Glue, you write ETL jobs directly in Spark. In ADF, Spark sits underneath mapping data flows and is managed by the service, so Spark knowledge helps you understand how data flows run rather than being something you operate yourself.

Then learn enough Beam to read a Dataflow pipeline. Finally, practise bringing pipelines into a release process: ADF fully supports CI/CD through Azure DevOps and GitHub, and enterprise clients will ask about it.

When reading FDE job descriptions, note which ETL services are mentioned to decide which platform to brush up on first. On your CV, do not write “knows AWS Glue”. Write that you read an existing pipeline, found where the schema had drifted and shipped the fix through CI/CD.

That sentence shows a hiring manager exactly the work an FDE does on a client site.

The client’s ETL tool was almost certainly chosen before you arrived. What you control is how quickly you understand it.

5 sources
Read next on the roadmap · Stage 5: DeploymentMLflow: answering “which model is running?” with runs, versions and aliasesAt a customer site, a good model is not enough. You also have to show which version is running in production, where it came from and how to rebuild it.