FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

Classical ML or deep learning: choose by the shape of the client's data, not by fashion

On a ten-thousand-row table, tree-based models are still hard to beat. On images and text the balance flips, and a good FDE should be able to explain why in the first meeting.

In brief

  • On mid-sized tabular data (around 10,000 samples), tree-based models remain state-of-the-art, according to the benchmark by Grinsztajn et al.
  • Images and text are not naturally numeric; deep learning learns features from raw data, which is where it excels.
  • Four questions to ask the client before choosing: data type, data volume, need for explanation, and GPU.
ShareLinkedInFacebookX
GraphicTree models or deep learning for the client's data
Classical ML (tree models)Deep learning
Best-suited dataMid-sized tables, around 10,000 samplesImages, text and other data that is not naturally numeric
Feature engineeringFeatures designed by peopleLearns most features from raw data
Data requiredWorks well with mid-sized dataNeeds very large amounts of data
ExplainabilityEasier to interpretEasier to scale but harder to interpret
InfrastructureRuns on a laptopTraining should use GPUs

Mid-sized tables favour tree models; images and text favour deep learning, provided the client has enough data and GPUs.

Graphic: FDE Times

Picture your first week at a bank. You receive three things at once: a table of loan history, a folder of scanned documents, and an archive of customer complaint emails. Someone in the room asks the familiar question: “Shall we just use deep learning for all three?”

The right answer is neither “yes” nor “no”. The right answer is: it depends on the shape of the data, and you need to be able to give the reasons in two minutes. This is one of the first technical decisions an FDE makes, and it drives infrastructure cost, time to results, and whether the client trusts the model at all.

This article teaches a simple way of thinking built on four questions: what form is the data in, how many labels are there, who needs an explanation, and is there a GPU. In most cases those four questions already give you the answer.

Mid-sized tables are still home ground for tree models

The loan table is a textbook supervised learning problem. Microsoft Learn describes supervised learning as training data that contains both feature values and known labels, and uses exactly this example: predicting whether a bank customer will default based on income, credit history, age and other factors.

That is binary classification on tabular data.

For this kind of data, the evidence is fairly clear. The benchmark by Grinsztajn et al. on tabular data concludes that tree-based models remain state-of-the-art on mid-sized datasets of around 10,000 samples. If the client’s table is about that size, that is the first signal in favour of tree models.

The reason is not just received wisdom. The authors show that neural networks are biased towards overly smooth solutions, which makes irregular functions hard for them to learn. Tabular data is full of exactly those step changes: once income crosses a threshold, the risk shifts sharply; two or more late payments and the picture is entirely different.

Decision trees split on thresholds naturally.

Worked example: a baseline for the loan table

Suppose the table has about 10,000 rows, with columns such as income, age, num_late_payments, loan_type and the label defaulted. The first job is not designing a network architecture but building a gradient boosting baseline in a few dozen lines of code.

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import roc_auc_score

df = pd.read_csv("loans.csv")
X = df.drop(columns=["defaulted"])
y = df["defaulted"]
X["loan_type"] = X["loan_type"].astype("category")

X_tr, X_te, y_tr, y_te = train_test_split(
    X, y, test_size=0.2, stratify=y, random_state=42)

model = HistGradientBoostingClassifier(categorical_features="from_dtype")
model.fit(X_tr, y_tr)
print("AUC:", roc_auc_score(y_te, model.predict_proba(X_te)[:, 1]))

This baseline runs on a laptop, needs no GPU, and gives you a number to negotiate with. If anyone then wants to try a neural network, they have to beat that number on the same test set. That turns a debate about fashion into a measurement.

Neural networks for tables are not off limits. fast.ai’s Practical Deep Learning course has a full section on tabular data with categorical, continuous and mixed variables. But treat them as the second candidate, worth trying only once the tree baseline exists and the data is large enough.

With images and text, the balance flips

The document folder and the email archive are a different matter. IBM explains that text, images, social network graphs and user behaviour are not naturally numeric data, so they need far more feature processing.

With classical ML, you would have to invent a way to turn an invoice image into numeric columns yourself, which is laborious and error-prone.

Deep learning works on raw data and automates most of that feature engineering. Microsoft Learn also notes that deep neural networks can be used for regression, classification and specialised models for natural language processing and computer vision. For scanned documents and complaint emails, deep learning should be the starting point.

Four questions to ask the client before deciding

The data type gives you a direction, but the client’s circumstances decide whether you can follow it. IBM sets out the core trade-off: deep learning scales more easily but is harder to interpret than traditional ML, and it requires very large amounts of data. Microsoft Learn adds that neural network training should run on machines with GPUs, which are optimised for matrix operations.

Question for the client If the answer leans towards… Implication for the choice
What form is the data in? Tables, or images and text Tables: start with tree models; images and text: start with deep learning
How many labelled samples are there? A few thousand to a few tens of thousands Favour tree models, especially for tables
Who needs to understand why the model decided? Credit, audit, legal Tree models are easier to explain; record this requirement explicitly
Is there GPU infrastructure? No, or budget must be requested Deep learning brings extra cost and lead time

For the hypothetical bank, the loan table answers all four questions in favour of tree models. The images and emails justify investment in deep learning, but you need to ask in the first week: do they have enough labelled images, and where are the GPUs?

Five steps to apply at a client

Step one: classify each data source as tabular, image, text or mixed. Step two: count the samples with real labels, not the raw rows, because labels are what supervised learning needs.

Step three: ask who will use the results and whom they must explain them to. Step four: check the infrastructure: GPUs, where the data is allowed to live, and acceptable training time.

Step five: build the cheapest baseline suited to the data type first, then try heavier options on the same test set. Write up the results on a single short page the client can read.

Common mistakes

A subtler mistake is forgetting to ask about explainability. A model that is accurate but that the credit team cannot explain to borrowers or auditors will stay a demo forever.

If you are preparing to apply for FDE roles, put exactly this kind of decision on your CV: not just “used XGBoost” but “chose a tree model over a neural network because of mid-sized data and an explainability requirement, measured on the same test set”.

At the next meeting, when someone asks “shall we use deep learning for all three?”, you will have a shorter answer: trees for the table, neural networks for images and text, and a baseline to prove it.

5 sources
Read next on the roadmap · Stage 3: Applied AIRead the model card before choosing Qwen, Llama, Gemma or Mistral for a Vietnamese projectThe Llama 3.1 model card lists Thai but not Vietnamese. You need to catch details like that before you put an open model into a client's system.