02 / Business Intelligence & Customer Analytics

Bank Marketing Customer Prioritization

Customer prioritisation under limited contact capacity

The business question

If a marketing team cannot contact everyone, which customers should it prioritise first?

Original university team submission: 9/10 · My role: Statistical Analysis Lead

Source & implementation

Where the project started

This began as a VU university group assignment on bank-marketing recontact strategy. The original submission received 9/10. I later revisited the analysis for this portfolio using the UCI bank-marketing dataset.

The comparison shows how many observed subscribers fall inside each contact limit: 10%, 20% or 30% of the list.

My contribution

As Statistical Analysis Lead in the original team, I contributed to data cleaning, the cleaned dataset and data dictionary, statistical feature-combination analysis, and interpretation of capacity, lift and leakage. I also helped coordinate the work and final presentation.

Machine learning, visual storytelling and the strategy memo had separate team leads. The models and visuals shown here belong to the later portfolio adaptation; the original team grade does not apply to this version.

At 20% capacity: 545 of 1,058 observed subscribers captured; a 30.1% shortlist subscription rate versus 11.7% overall. Retrospective enrichment, not causal uplift.

Results

51.5% captured

The top-ranked 20% contained 545 of the 1,058 observed subscribers: 51.5% of subscribers in the historical test set.

The observed subscription rate was 30.1% in that selected group, compared with 11.7% overall.

Selecting 10% of records captured 35.6% of observed subscribers; selecting 30% captured 61.3%. A larger contact list reaches more of them, but includes a higher proportion of non-subscribers.

What I learned

I learned to compare contact limits by counting observed subscribers inside and outside each shortlist. I also had to check which information would be available before contact.

Scope and limits

These retrospective results do not show that contacting the selected records would cause additional subscriptions.

Data and method

The data

Historical UCI bank-marketing records; 9,043 holdout records with 1,058 observed subscribers; three predefined capacity scenarios.

Historical holdout · 9,043 records / 1,058 positives
CapacitySelectedPositivesCapturedShortlist rate
10%90537735.6%41.7%
20%1,80954551.5%30.1%
30%2,71364961.3%23.9%

Method

  1. Exclude current-campaign contact information, call duration and other fields unavailable at the assumed pre-contact decision point.
  2. Compare Logistic Regression and Random Forest using five-fold, training-only cross-validation and average precision.
  3. Evaluate the selected Random Forest on the historical holdout; translate ranking quality into 10%, 20% and 30% capacity scenarios.
  4. Check both shortlist concentration and subscriber coverage, and document the timing assumptions for prior-contact history.

The source contains 45,211 records, with 9,043 records and 1,058 observed subscribers in the holdout. Random Forest was selected after five-fold training-only cross-validation; holdout average precision was 0.383 and ROC AUC was 0.734.

Project background

The displayed results come from a later adaptation of the original university team submission. The team grade and instructor feedback apply only to that original submission.

AI assistance supported the later portfolio adaptation, including analytical code and documentation.

Potential use

A marketing analyst discussing contact capacity with a campaign team.

The decision

How many observed subscribers fall inside the first 10%, 20% or 30% of a ranked contact list?

Original project outputs
At 20% selection capacity, the model captures 545 of 1,058 observed subscribers in the historical holdout.
Prior campaign outcomes provide descriptive context. Association in historical records is not a causal effect.
Detailed limitations

The random holdout had been inspected in earlier work and is not fresh external validation. Scores are uncalibrated. Historical contact timing is assumed; the model ranks records rather than deciding when to call again. There is no measured campaign ROI or causal uplift.

Contemporary validation, contact eligibility, operating costs and a controlled pilot. Historical enrichment does not establish additional subscriptions caused by contact.

What I would do next

Test the approach on a contemporary, time-separated sample. A controlled pilot would then need clear contact eligibility, costs and group-level outcome checks.

Explore the repository

Original university context

Original Bank Marketing
assignment

This project originated from a Business Intelligence & Analytics team assignment at Vrije Universiteit Amsterdam on re-contact strategy.

My documented role was Statistical Analysis Lead, contributing to data cleaning, statistical analysis and interpretation. The original team submission received 9/10.

The version shown in this portfolio is a later adaptation in which I reframed the analysis around contact capacity and customer prioritisation.

“your group achieved the highest grade among the group assignments.”

Edona Elshan · Course instructor · Vrije Universiteit Amsterdam