02 / Business Intelligence & Customer Analytics
Bank Marketing Customer Prioritization
Customer prioritisation under limited contact capacity
The business question
If a marketing team cannot contact everyone, which customers should it prioritise first?
Original university team submission: 9/10 · My role: Statistical Analysis Lead
Source & implementationWhere the project started
This began as a VU university group assignment on bank-marketing recontact strategy. The original submission received 9/10. I later revisited the analysis for this portfolio using the UCI bank-marketing dataset.
The comparison shows how many observed subscribers fall inside each contact limit: 10%, 20% or 30% of the list.
My contribution
As Statistical Analysis Lead in the original team, I contributed to data cleaning, the cleaned dataset and data dictionary, statistical feature-combination analysis, and interpretation of capacity, lift and leakage. I also helped coordinate the work and final presentation.
Machine learning, visual storytelling and the strategy memo had separate team leads. The models and visuals shown here belong to the later portfolio adaptation; the original team grade does not apply to this version.
Results
51.5% captured
The top-ranked 20% contained 545 of the 1,058 observed subscribers: 51.5% of subscribers in the historical test set.
The observed subscription rate was 30.1% in that selected group, compared with 11.7% overall.
Selecting 10% of records captured 35.6% of observed subscribers; selecting 30% captured 61.3%. A larger contact list reaches more of them, but includes a higher proportion of non-subscribers.
What I learned
I learned to compare contact limits by counting observed subscribers inside and outside each shortlist. I also had to check which information would be available before contact.
Scope and limits
These retrospective results do not show that contacting the selected records would cause additional subscriptions.
Data and method
The data
Historical UCI bank-marketing records; 9,043 holdout records with 1,058 observed subscribers; three predefined capacity scenarios.
| Capacity | Selected | Positives | Captured | Shortlist rate |
|---|---|---|---|---|
| 10% | 905 | 377 | 35.6% | 41.7% |
| 20% | 1,809 | 545 | 51.5% | 30.1% |
| 30% | 2,713 | 649 | 61.3% | 23.9% |
Method
- Exclude current-campaign contact information, call duration and other fields unavailable at the assumed pre-contact decision point.
- Compare Logistic Regression and Random Forest using five-fold, training-only cross-validation and average precision.
- Evaluate the selected Random Forest on the historical holdout; translate ranking quality into 10%, 20% and 30% capacity scenarios.
- Check both shortlist concentration and subscriber coverage, and document the timing assumptions for prior-contact history.
The source contains 45,211 records, with 9,043 records and 1,058 observed subscribers in the holdout. Random Forest was selected after five-fold training-only cross-validation; holdout average precision was 0.383 and ROC AUC was 0.734.
Project background
The displayed results come from a later adaptation of the original university team submission. The team grade and instructor feedback apply only to that original submission.
AI assistance supported the later portfolio adaptation, including analytical code and documentation.
Potential use
A marketing analyst discussing contact capacity with a campaign team.
The decision
How many observed subscribers fall inside the first 10%, 20% or 30% of a ranked contact list?
Original project outputs
Detailed limitations
The random holdout had been inspected in earlier work and is not fresh external validation. Scores are uncalibrated. Historical contact timing is assumed; the model ranks records rather than deciding when to call again. There is no measured campaign ROI or causal uplift.
Contemporary validation, contact eligibility, operating costs and a controlled pilot. Historical enrichment does not establish additional subscriptions caused by contact.