How to pick a classification threshold by the cost of each error, not the 0.5 default
In a credit-scoring example from scikit-learn, the model stays the same and only the decision threshold moves. That one change makes the business outcome nearly twice as good.

In brief
- Choosing between precision and recall really means choosing which error the model is optimised to avoid. The answer depends on what each error costs the client.
- The default 0.5 threshold is probably not the best one for a real problem. In the scikit-learn example, tuning the threshold to cost moved the business gain from -209 to -143.
- AUC is for comparing models because it covers every threshold. A decision needs one threshold, and you should tune it on data that was not used for training.
In one credit-scoring example, the scikit-learn documentation keeps the model fixed and changes a single number: the decision threshold. The business gain moves from -209 to -143, which is nearly twice as good. There are no new features and no retraining.
If you want to work as an FDE, practise this before you meet clients. Data scientists tend to report AUC or accuracy. Clients want to know how accurate is accurate enough, and how much money or how many staff hours each mistake costs them. This guide walks through each step so you can answer that question in code, on your own laptop.
What you will do: write a cost function, sweep thresholds on a validation set, pick the best threshold and check it on a test set. What you need: Python 3 and any binary classifier that outputs probability scores. The code below uses plain Python with no libraries, so you can see every calculation. It is simplified for learning.
Are you choosing a metric, or choosing which error to avoid?
Two definitions to keep in mind. Precision is the share of positive predictions that are actually positive. Recall is the share of real positives the model catches. Domino’s data science dictionary makes a point that is easy to miss: deciding to prioritise precision or recall changes which type of error the model is optimised to avoid.
Google’s ML Crash Course gives a practical rule. Prioritise recall when false negatives cost more than false positives. Prioritise precision when positive predictions must be correct. Do not use accuracy on imbalanced data.
Why not accuracy? Picture 1,000 loan applications, 100 of them from bad borrowers. A “lazy” model that labels everyone a good borrower scores 90% accuracy and catches no bad borrowers at all. So the first step is a conversation with the client. Code comes later.
Step 1: get the cost of each error from the client
In the scikit-learn example, “positive” means a bad borrower. Lending to a bad borrower by mistake (a false negative) costs on average five times as much as wrongly rejecting a good one (a false positive). The gain matrix therefore gives -1 for each FP and -5 for each FN.
COST_FP = 1 # good borrower wrongly rejected
COST_FN = 5 # bad borrower wrongly approved
def business_cost(tp, fp, tn, fn):
return COST_FP * fp + COST_FN * fn
Check: does the client agree with the 1:5 ratio? The number does not have to be exact, but it has to come from the people who know the business. Engineers should not guess it.
Step 2: turn scores into labels
The model returns probabilities. To get labels you need a cut-off probability, called the classification threshold. Google notes that different thresholds usually produce different numbers of TP, FP, TN and FN.
def confusion_at(y_true, y_score, threshold):
tp = fp = tn = fn = 0
for y, s in zip(y_true, y_score):
pred = 1 if s >= threshold else 0
if pred == 1 and y == 1: tp += 1
elif pred == 1 and y == 0: fp += 1
elif pred == 0 and y == 0: tn += 1
else: fn += 1
return tp, fp, tn, fn
Check: tp + fp + tn + fn must equal the number of samples. If it doesn’t, your labels contain values other than 0 and 1.
Step 3: sweep thresholds instead of trusting 0.5
The scikit-learn documentation says plainly that the default strategy (a 0.5 threshold) is probably not optimal for the problem at hand. The best threshold is the one that gives the best value of the metric you chose. Here, that metric is business cost.
def sweep(y_true, y_score):
rows = []
for i in range(5, 96, 5):
t = i / 100
tp, fp, tn, fn = confusion_at(y_true, y_score, t)
precision = tp / (tp + fp) if tp + fp else 0.0
recall = tp / (tp + fn) if tp + fn else 0.0
rows.append((t, business_cost(tp, fp, tn, fn), precision, recall))
return sorted(rows, key=lambda r: r[1])
To see what the output looks like, go back to the 1,000 hypothetical applications above. The figures in the table are illustrative, not real measurements.
| Classification | TP / FP / FN | Precision | Recall | Accuracy | Cost |
|---|---|---|---|---|---|
| Label everyone a good borrower | 0 / 0 / 100 | — | 0% | 90% | 500 |
| Threshold 0.5 | 40 / 10 / 60 | 80% | 40% | 93% | 310 |
| Threshold 0.2 | 75 / 60 / 25 | 55.6% | 75% | 91.5% | 185 |
Look at what happens between 0.5 and 0.2: accuracy falls, but cost falls sharply. Recall goes up and precision goes down. scikit-learn reports the same kind of trade-off when it tunes the threshold in its credit example. If you report accuracy, you will pick the wrong threshold.
Step 4: never tune the threshold on training data
The scikit-learn guide says you should never use the same data to train the classifier and to tune the threshold, because it leads to overfitting. On the training set the model is more confident than it really is, so any threshold found there will be too optimistic.
The safe approach is a three-way split: train to fit the model, validation to run sweep, and test to measure the cost one last time at the chosen threshold. If the test cost is clearly higher than the validation cost, your threshold is fitting noise.
Check: run confusion_at on the test set with the chosen threshold, then compute business_cost. Put this number in your report, not the validation number.
So what is AUC for?
AUC and the ROC curve show how well the model separates the two classes across every possible threshold. That is exactly why AUC cannot tell you where to set the threshold. It is useful for comparing two models before you choose a threshold, but it does not replace the cost figure at the threshold you will actually run.
What this looks like on a client site
Imagine you have deployed a model that flags suspicious transactions for a bank. The operations team complains about too many false alarms. The risk team worries about missed cases. Both are right, and the argument only stops when everyone works from one shared figure: the cost of each type of error.
Start by sitting down with both teams and asking: how many times more does one missed case cost than one false alarm? Then run sweep, show them a table like the one above and let them pick the row that suits them. The threshold becomes a business decision with evidence behind it, rather than a technical parameter hidden in the code.
On your CV and in interviews, don’t just write “achieved 0.9 AUC”. Write that you turned the client’s requirements into a cost function, chose a threshold on a separate validation set, and cut the cost of errors by a given percentage compared with the default threshold.
When a job description asks for someone who can translate business needs into metrics, this example is evidence that you can.
The best model in your notebook can still be the most expensive one in production if the threshold is wrong. The engineers who get invited back are often the ones who ask the client “which mistake costs more?” before they run any code.
5 sources
- What is Model Evaluation (Domino Data Science Dictionary)
- Thresholds and the confusion matrix (Google ML Crash Course) · 2026-01-12
- Accuracy, precision, and recall (Google ML Crash Course) · 2026-01-12
- Post-tuning the decision threshold for cost-sensitive learning (scikit-learn)
- Tuning the decision threshold for class prediction (scikit-learn User Guide)