06 / Predictive Analytics & Decision Support

Predictive Lead Ranking

Predictive ranking under a fixed contact capacity

Top-100 ranking performance

70 of the top 100 ranked historical records were positives.

Fixed 100-record capacity · Historical holdout evaluation

Source & implementation

The ranking question

I explored what a model could prioritise when only 100 records could be reviewed. This uses the same UCI bank-marketing dataset as Project 02, but keeps the contact capacity fixed rather than comparing several list sizes.

A fixed contact limit makes it necessary to look at both the quality of the shortlist and how much of the total positive population it misses.

My contribution

I compared ranking models, selected Logistic Regression and evaluated its top 100 records against observed outcomes and an expected-random baseline. I built the presentation around both the share of positives inside the shortlist and the share it leaves outside.

The same 70 observed positives produce 70% precision@100 and 6.62% coverage of all holdout positives. Logistic Regression; historical evaluation only.

Results

Precision and capture at 100

Precision@100 is 70%: 70 of the 100 selected records were positives. Capture@100 is 6.62%: those same 70 records out of all 1,058 holdout positives.

At the holdout’s 11.7% prevalence, random selection would contain 11.7 positives per 100 records on average. The model’s top 100 therefore contains 5.98 times that expected-random concentration.

What I learned

I learned to explain model results with their denominators. A shortlist can contain a high proportion of positives while still leaving most of the total positive population outside it; whether that is useful depends on the contact capacity and objective.

Scope and limits

This is a historical ranking evaluation, not live scoring, calibrated conversion probabilities or evidence of causal uplift.

Data and method

The data

The same historical UCI dataset as the broader bank case: 70 positives in the top 100, representing 70% precision and 6.62% of all holdout positives.

Method

  1. Keep preprocessing inside training folds and omit fields unavailable at the assumed pre-contact decision point.
  2. Use three-fold training-only model selection; select Logistic Regression for the ranking task.
  3. Sort holdout scores and evaluate the fixed top 100 against observed outcomes and an expected-random baseline.
  4. Report precision and captured positives together, while inspecting leakage risks and repeated feature profiles.

Three-fold training-only model selection chose Logistic Regression. On the 9,043-record holdout, average precision was 0.3472 and ROC-AUC was 0.7204. Accuracy alone is less useful for this imbalanced, capacity-limited question.

Project background

The application displays stored evaluation results. It is not a production deployment or an autonomous sales agent.

AI helped with code, documentation and visual refinement.

Potential use

An analyst evaluating a model under a fixed 100-record capacity.

The decision

How many positives are in a top-100 shortlist, and how many remain outside it?

Original project outputs
A static evaluation dashboard for a retrospective top-100 scenario, not live scoring or campaign results.
Authentic saved notebook output: observed positives in the model’s top 100 versus the expected-random baseline. The shortlist captures 6.62% of all holdout positives.
Detailed limitations

The random holdout had been inspected previously. Scores are uncalibrated, and 375 holdout records share an included-feature profile with training records; matching profiles do not prove they are the same customers. Row identifiers are not unique customer IDs. Fairness, prospective performance and business impact remain unvalidated.

Prospective or time-separated validation, customer identity and eligibility checks, and group-level evaluation before use in a live contact process.

What I would do next

Validate on later or prospective data with clear customer identity and eligibility rules. Check group outcomes and calibrate scores if future decisions require probabilities.

Explore the repository