Five clients. Five briefs. There is no default technique — pick the one the
question and its constraints demand, then defend it. About five minutes.
scenario 1 of 5
Problem type
A major bank wants AI to set loan amounts for approved applicants.
They hold years of customer profiles and the actual amounts granted.
— client brief, national bank
What kind of machine-learning problem have they got?
B. Regression.
The output is a dollar amount — continuous. Banking context does not
make it classification; "approve/decline" would be, "how much?" is not.
Consultant's rule: when the business question is "how much?",
think regression. "Which category?" — classification.
Explainability is the requirement
A hospital network wants to flag patients at high risk of readmission.
"Any recommendations must be explainable to our medical staff and auditable for
regulatory compliance."
— client brief, Chief Medical Officer
Which technique do you recommend?
B. Decision tree.
When decisions must be explained to clinicians, patients, and regulators,
interpretability outranks small accuracy gains. A tree's if-then rules can be read,
challenged, and audited.
Consultant's rule: in regulated, high-stakes settings, a model nobody
can interrogate is a liability, whatever its accuracy.
Many features
An e-commerce platform wants to classify thousands of product reviews as
positive or negative. The dataset is large and wide: word frequencies, sentiment scores,
review length, and more.
What matters most in the technique you pick?
B. SVM.
Hundreds of text-derived features make this a classic high-dimensional
problem, and SVMs are built to find separating boundaries in exactly that geometry.
Consultant's rule: the shape of the data (wide, sparse, text-like)
narrows the shortlist before accuracy is even discussed.
Accuracy and insight, both
A retail chain wants to predict customer lifetime value to allocate its
marketing budget. It needs accurate numbers, and the marketing team also wants to know
which factors are driving the prediction.
How do you balance the competing requirements?
B. Random forest with feature importance.
The ensemble carries the accuracy; feature importance hands the marketing
team the drivers. Both requirements, one model.
Consultant's rule: when the client wants reliable predictions
and to understand them, ensembles with importance analysis are the usual
compromise — not the most accurate model, and not the most transparent one.
Check the data first
A mining company wants to predict equipment failure. The data: sensor
readings updated hourly, two years of complete maintenance records — but only from
machines that have not failed. Failed machines were replaced and their records
left the fleet.
— client brief, Operations Director: "lives depend on catching real failures,
but false alarms shut us down unnecessarily."
What is the biggest data concern to raise with the client?
C. Sampling bias — the textbook survivorship trap.
The model is asked to predict failure from data that only contains
non-failures. There are no examples of what failure looks like. The dataset's biggest
gap is precisely the outcome you are being paid to predict.
Consultant's rule: before choosing a technique, check the data
represents the full range of outcomes — especially the negatives. A technique
chosen over missing data inherits the gap.
What just happened
Five clients, five different right answers — same toolkit. There is no default
technique. The business question (how much? vs which category?), the constraints
(explainability, fairness, real-time), the shape of the data (wide? biased? complete?),
and the stakeholder requirements picked the model every time.
The five rules you just used
1. Question type: "how much?" → regression; "which category?"
→ classification; no label → clustering, if at all.
2. Stakes: regulated or high-stakes → pay accuracy for
interpretability (trees), deliberately.
3. Data shape: wide/sparse/text → high-dimensional methods
(SVM) before exotic ones.
4. Both-and: accuracy and insight → ensembles with
feature importance.
5. Data before model: does the data contain the outcomes you are
predicting — including the failures?
Score: 0/5
— the score is trivia; the five rules are the point.
Want the mechanics under the predictions? Play perceptron. Want to see why the
mining data can't be trusted? That answer was survivorship bias — the same
reason overfit models flatter you in overfit.