Pick your client
Each team selects one scenario. Multiple teams on the same scenario makes a better comparison panel at pitch time.
🏥 Melbourne Health Network — patient readmission prediction
Available data (2 years, 50,000 records)
- Demographics: age, gender, location, insurance
- Clinical: diagnosis, comorbidities, severity scores
- Treatment: length of stay, procedures, medications
- History: previous admissions, compliance
- Outcome: readmitted within 30 days (yes/no) — 12% base rate
Team tasks (45 min)
- Problem analysis (10): classification or regression? Which business requirement dominates?
- Technique selection (15): compare decision trees against alternatives
- Build & test (15): implement the recommended approach
- Business case (5): recommendation plus justification
Workflow starters
[File] → [Data Sampler] → [Tree] → [Tree Viewer] + [Test & Score]
[File] → [Data Sampler] → [Tree, SVM, kNN] → [Test & Score]
Widgets worth exploring: Tree Viewer (interpretable rules), Test & Score, Confusion Matrix, Feature Statistics.
Pitch deliverables
- Technique recommendation — which approach, and why
- Explainability strategy — how doctors read the recommendations
- Performance metrics — accuracy, precision, recall
- Implementation plan — integration into clinical workflows
- Risk assessment — what could go wrong, and mitigation
Debrief: the lesson is the tension between accuracy and interpretability under regulation — not the winning model. Cross-link: technique-picker, scenario 2.
💳 Aussie Fintech Startup — credit risk assessment
Available data (18 months, 25,000 applications)
- Traditional: credit score, income, employment history
- Alternative: utility payments, rent history, education
- Behavioural: application completion patterns, response times
- Social: references, guarantor information
- Outcome: default within 12 months (yes/no) — 8% base rate
Team tasks (45 min)
- Problem analysis (10): classification — but what are the fairness implications?
- Technique comparison (15): accuracy and bias across approaches
- Build & analyse (15): feature importance and demographic correlations
- Ethical assessment (5): fairness and regulatory compliance
Workflow starters
[File] → [Data Sampler] → [Tree, SVM, kNN] → [Test & Score] + [Confusion Matrix]
[File] → [Rank] → [Data Table]
[File] → [Select Columns] → [Scatter Plot]
Use the scatter plot to check score behaviour across demographic groups — the bias check matters more than the accuracy number.
Pitch deliverables
- Technique recommendation for accuracy and fairness
- Bias analysis — how protected groups are protected
- Performance comparison across techniques
- Regulatory compliance story
- Competitive advantage argument
Debrief: an 8% default base rate means "always approve" scores 92% — the accuracy headline is a trap. Connect to threshold-dial and k-anonymity.
🛒 National Supermarket Chain — dynamic pricing optimisation
Available data (3 years, millions of transactions)
- Sales: product, store, time, price point
- Inventory: stock levels, shelf life, supplier costs
- Competitor: rival pricing, market share
- Customer: loyalty data, purchase patterns
- External: weather, events, economic indicators
- Target: optimal price (continuous) or price tier (low/med/high)
Team tasks (45 min)
- Problem framing (10): regression (price) vs classification (price tier) — which serves the business?
- Technique testing (15): compare with real-time requirements in mind
- Build & evaluate (15): alternative pricing models
- Deployment strategy (5): rollout across 1,000+ stores
Workflow starters
[File] → [Data Sampler] → [Linear Regression] → [Predictions] + [Test & Score]
[File] → [Data Sampler] → [Tree] → [Tree Viewer] + [Test & Score]
[File] → [Data Sampler] → [Random Forest] → [Test & Score]
Pitch deliverables
- Problem-type decision: regression vs classification
- Technique recommendation for real-time pricing
- Performance analysis with speed considered
- Scaling strategy for a national chain
- Customer-trust impact
Debrief: the framing decision (continuous price vs price tier) changes everything downstream — most teams never notice they chose it. Connect to technique-picker, scenario 1.
⛏ Australian Mining Corp — equipment failure prediction
Available data (5 years, 200+ machines)
- Sensors: temperature, vibration, pressure, electrical
- Maintenance: service history, part replacements, downtime
- Environment: weather, operational intensity
- Fleet: age, manufacturer, model, usage patterns
- Outcome: failure within next 7 days (yes/no) — 3% base rate
Team tasks (45 min)
- Problem analysis (10): classification with extreme cost asymmetry — how do you weigh missed failures vs false alarms?
- Technique evaluation (15): reliability and explainability first
- Test (15): build models, read the confusion matrix carefully
- Safety strategy (5): AI augmenting human judgement, not replacing it
Workflow starters
[File] → [Data Sampler] → [Tree] → [Tree Viewer] + [Confusion Matrix]
[File] → [Data Sampler] → [Random Forest] → [Test & Score] + [Confusion Matrix]
The confusion matrix is the whole assessment: false negatives are dangerous, false positives are expensive. A 97% accuracy headline can hide every failure.
Pitch deliverables
- Technique selection: accuracy vs explainability balance
- Safety analysis: minimise false negatives, control false positives
- Confusion matrix analysis with error costs
- How maintenance teams use the predictions
- Human–AI interaction: augmentation, not replacement
Debrief: a 3% failure rate makes this threshold-dial's argument with lives attached — 97% accuracy by predicting "never fail" hides every real failure. Also ask: what's missing from five years of "complete" records? (Compare technique-picker, scenario 5.)
Pitches
Each team: 3 minutes presentation + 1 minute questions.
Pitch structure
- 30s — the business problem
- 60s — the technique decision and why
- 60s — evidence: workflows and results
- 30s — business impact
What good looks like
- Clear business logic — why this technique fits this problem
- Trade-off analysis — accuracy vs interpretability vs implementation
- Evidence — workflows and results that support the call
- Professional communication — technical concepts for business stakeholders
Debrief points
- Different business contexts demand different techniques — there is no default
- Rarely one correct answer; the defence is the deliverable
- Connect forward: k-means-stepper (clustering), perceptron (mechanics), threshold-dial (error costs)
- The durable skill: requirement analysis and technology selection, explained to non-technical stakeholders