06 / Predictive Analytics & Decision Support
Predictive Lead Ranking
Predictive ranking under a fixed contact capacity
Top-100 ranking performance
70 of the top 100 ranked historical records were positives.
Fixed 100-record capacity · Historical holdout evaluation
Source & implementationThe ranking question
I explored what a model could prioritise when only 100 records could be reviewed. This uses the same UCI bank-marketing dataset as Project 02, but keeps the contact capacity fixed rather than comparing several list sizes.
A fixed contact limit makes it necessary to look at both the quality of the shortlist and how much of the total positive population it misses.
My contribution
I compared ranking models, selected Logistic Regression and evaluated its top 100 records against observed outcomes and an expected-random baseline. I built the presentation around both the share of positives inside the shortlist and the share it leaves outside.
Results
Precision and capture at 100
Precision@100 is 70%: 70 of the 100 selected records were positives. Capture@100 is 6.62%: those same 70 records out of all 1,058 holdout positives.
At the holdout’s 11.7% prevalence, random selection would contain 11.7 positives per 100 records on average. The model’s top 100 therefore contains 5.98 times that expected-random concentration.
What I learned
I learned to explain model results with their denominators. A shortlist can contain a high proportion of positives while still leaving most of the total positive population outside it; whether that is useful depends on the contact capacity and objective.
Scope and limits
This is a historical ranking evaluation, not live scoring, calibrated conversion probabilities or evidence of causal uplift.
Data and method
The data
The same historical UCI dataset as the broader bank case: 70 positives in the top 100, representing 70% precision and 6.62% of all holdout positives.
Method
- Keep preprocessing inside training folds and omit fields unavailable at the assumed pre-contact decision point.
- Use three-fold training-only model selection; select Logistic Regression for the ranking task.
- Sort holdout scores and evaluate the fixed top 100 against observed outcomes and an expected-random baseline.
- Report precision and captured positives together, while inspecting leakage risks and repeated feature profiles.
Three-fold training-only model selection chose Logistic Regression. On the 9,043-record holdout, average precision was 0.3472 and ROC-AUC was 0.7204. Accuracy alone is less useful for this imbalanced, capacity-limited question.
Project background
The application displays stored evaluation results. It is not a production deployment or an autonomous sales agent.
AI helped with code, documentation and visual refinement.
Potential use
An analyst evaluating a model under a fixed 100-record capacity.
The decision
How many positives are in a top-100 shortlist, and how many remain outside it?
Original project outputs
Detailed limitations
The random holdout had been inspected previously. Scores are uncalibrated, and 375 holdout records share an included-feature profile with training records; matching profiles do not prove they are the same customers. Row identifiers are not unique customer IDs. Fairness, prospective performance and business impact remain unvalidated.
Prospective or time-separated validation, customer identity and eligibility checks, and group-level evaluation before use in a live contact process.