Inheriting a legacy Hadoop cluster at a bank or telco: read HDFS, YARN, MapReduce and HBase before you touch it
The system holding a client's transaction data doesn't need you to like it. You do need to know where it breaks before you run your first job.
In brief
- HDFS separates metadata (the NameNode) from data (the DataNodes). In Hadoop 1.x the NameNode is a single point of failure.
- YARN decides who gets to run; MapReduce decides how the job runs: map, shuffle by key, then reduce.
- HDFS is write-once, read-many. If you need real-time random updates, use HBase.
Picture your second day at a bank. You have SSH access to an edge node and a single instruction: the entire transaction history lives on the Hadoop cluster, so don’t bring it down.
Nobody left on the project team remembers how the cluster was built. You, meanwhile, have two weeks to pull data out of it and train a fraud-detection model.
If you plan to work as an FDE for banks or telecoms operators, expect a scene like this. You don’t need to become a Hadoop expert. You need to read four components quickly enough to know what is safe to touch and what will snap.
Four components, two questions: where does the data live, and who gets to run?
The easiest way to remember the cluster is as a stack of layers. At the bottom is HDFS, the distributed file system. Above it is YARN, which allocates CPU and memory. On top of YARN sits MapReduce, the batch-processing model. HBase stands beside the YARN/MapReduce branch and uses HDFS directly as its storage.
For each component you only need to answer two questions: where does the data actually live, and who decides what gets to run? Most incidents you meet on site come back to one of the two.
How many pieces is a 1 GB file on HDFS, really?
HDFS uses a master/slave architecture. The NameNode holds the namespace and tracks every file, its permissions and the location of each block. The DataNodes do one thing: store blocks.
Now do the arithmetic. The Hadoop 1.x documentation gives 64 MB as the typical block size. A 1 GB transaction file, or 1024 MB, is cut into 16 blocks. With a replication factor of 3, the cluster keeps 48 block replicas spread across the DataNodes, and the NameNode has to remember where all 48 are.
The number tells you two things. On load: millions of small files mean millions of metadata entries piling onto a single NameNode. On durability: HDFS always keeps at least one replica on a different rack, so losing an entire rack does not yet mean losing data.
Replication is set per file and can be changed after the file is created. So don’t assume every directory has the same level of protection. Newer clusters also tend to use blocks larger than 64 MB, so check the real configuration rather than trusting the number in the book.
One more ground rule: HDFS is write-once, read-many. You cannot edit a line in the middle of a file. If your model needs to update a risk score for each customer, HDFS is not the place to do it.
Why is the NameNode the first thing to ask about?
In Hadoop 1.x, the machine running the NameNode is a single point of failure for the whole cluster. If the NameNode dies, the DataNodes still hold the data, but nobody knows which block belongs to which file.
There is another behaviour that often causes panic: on startup, the NameNode enters Safemode and does not yet replicate blocks. An old cluster that has just restarted can sit in Safemode for quite a while, and your job will fail in baffling ways. That is not a bug in your code.
So the first question for the operations team should be: is NameNode HA enabled on this cluster, and how long did the last restart take? The answer tells you the real risk of everything you are about to do.
YARN and MapReduce: what is your job waiting on?
YARN splits resource management from job scheduling and monitoring, putting them in separate daemons: a ResourceManager, a NodeManager on each machine, and an ApplicationMaster for each application. The ResourceManager has the final say on how resources are shared among all applications.
When your job sits in a waiting state, don’t rush to change the code. Look at who the ResourceManager is giving resources to. At a bank, there is a good chance an end-of-day reconciliation job has priority over yours.
MapReduce decides how the work runs. A concrete example: counting transactions per branch. The snippet below is pseudocode in the style of Hadoop Streaming for readability; native MapReduce jobs are usually written in Java.
# Mã giả kiểu Hadoop Streaming (MapReduce gốc viết bằng Java)
# map: mỗi dòng giao dịch -> (ma_chi_nhanh, 1)
def map_fn(line):
cols = line.split(",")
yield cols[2], 1
# shuffle: hệ thống gom mọi cặp cùng key về một reducer
# reduce: cộng dồn theo chi nhánh
def reduce_fn(branch, counts):
yield branch, sum(counts)
(The comments read: the map step turns each transaction line into a (branch_code, 1) pair; the system shuffles all pairs with the same key to one reducer; the reduce step sums per branch.)
The map phase turns the data into key/value pairs. The shuffle phase sends every pair with the same key to the same reducer. The reduce phase then receives one branch along with its full list of ones, for example (CN_HN01, [1, 1, 1, …]), and adds them up to get that branch’s total transaction count.
If one large branch accounts for most transactions, a single reducer has to carry that load, and you will see the whole job waiting on one final task.
The limitation most worth remembering is that MapReduce generally does not keep data in memory between steps. Iterative jobs, such as model training, are therefore much slower than on Spark. Use MapReduce for extraction and aggregation, not for training loops.
Hadoop also has a philosophy: move the computation to the data rather than moving the data. The reflex to pull everything down to a laptop or push it to the cloud is usually both slow and at odds with security rules. Run the filtering and aggregation on the cluster, and take only compact results out.
HBase: when you need to look up a customer instantly
HDFS doesn’t allow in-place edits, so reading or updating individual records is HBase’s job. It is a non-relational wide-column store, modelled on Google’s Bigtable, running on HDFS or Alluxio. It provides real-time random reads and writes, automatic failover and automatic sharding.
The clearest division of labour for a fraud project is this. Years of transaction history stay on HDFS for batch training. The latest risk profile for each customer, which has to be looked up by key whenever a transaction arrives, goes in HBase.
Five things to do in your first week on site
First, read the real configuration. A few read-only commands are enough to get started without breaking anything:
hdfs dfsadmin -safemode get
hdfs getconf -confKey dfs.blocksize
hdfs fsck /data/giaodich -files -blocks -locations
yarn application -list
Next, ask about NameNode HA and the restart schedule. Then find out which YARN queue belongs to your team and when the cluster is quiet. After that, map the data: what lives on HDFS, what lives in HBase, and the replication factor for each directory. Finally, run a small job on a subset of the data before touching the full table.
If you are aiming for an FDE role, watch for job descriptions at banks or telecoms companies that mention Hadoop, HDFS or HBase. On a CV, a line such as “read fsck output and recovered a cluster stuck in Safemode” is worth far more than “knows Hadoop”.
The usual stumbles
The most common is writing thousands of small output files to HDFS. Every file adds metadata to the NameNode, on a cluster where the NameNode may be a single point of failure.
Next comes treating HDFS as a database and trying to update individual rows. Then believing the block size is still 64 MB, or that replication is the same everywhere. And blaming the code when the job is really waiting for resources, or the cluster is still in Safemode.
The most expensive mistake isn’t technical at all. It is running a large job during reconciliation hours without asking anyone. On a cluster the whole bank depends on, the operations team’s trust is harder to recover than any block.
A good FDE on a legacy system is not the one who rewrites it, but the one who understands it well enough to deliver new value without breaking anything that is already running.