05 / Descriptive Data Analysis

Cyclistic NYC Bike Usage Analysis

Descriptive analysis of historical trip records

The business question

How do recorded bike trips differ between Customer and Subscriber user types?

Course-inspired portfolio case · Historical NYC aggregate data

Source & implementation

The learning project

I revisited a course-inspired bike-share case to practise analysing historical trip data. The analysis uses an aggregate export of NYC Citi Bike activity, with 126,816 recorded trips in the primary 2015 period.

Comparing trip duration and weekday patterns can suggest questions for customer research or a later operational analysis.

My contribution

I analysed the preserved export in Python, checked date coverage and used trip counts to weight duration averages. I compared recorded user types by weekday, month and geography, and checked how unusually long-duration groups affected the results.

Trip-weighted averages in the preserved NYC export: 126,816 recorded trips in 2015. Trip statistics, not unique customer characteristics or current demand.

Results

Trip duration and usage patterns

Customer trips averaged 29.40 minutes; Subscriber trips averaged 13.33 minutes in the 2015 data.

Subscriber activity was stronger on weekdays, while Customer activity was stronger on weekends. Calendar-normalised averages make the weekdays comparable.

Customer trips averaged 66.7 per Saturday versus 32.5 per Friday. Subscriber trips averaged 238.0 per Saturday versus 360.2 per Friday.

What I learned

I learned to check what each row represents before calculating an average. Here, each row combines trips, so I weighted duration by trip count and checked date coverage before comparing the user types.

Scope and limits

These are historical trip statistics, not unique-customer characteristics or a measure of current demand.

Data and method

The data

A preserved NYC aggregate export, with 126,816 recorded trips in the primary 2015 period; trip statistics describe neither unique people nor current demand.

Method

  1. Distinguish the wider 2013–May 2018 export from the complete 2015 analysis window.
  2. Weight duration summaries by recorded trip counts rather than averaging aggregate rows equally.
  3. Compare user types by weekday, month and geography, using calendar denominators and source-coverage checks.
  4. Run a sensitivity check excluding groups with mean durations above 24 hours.

The wider export spans 2013–May 2018; 2015 covers all 365 dates. Subscribers account for 88.05% of its trips and Customers for 11.95%. Excluding groups with mean durations above 24 hours changes average duration to 12.41 and 28.15 minutes respectively.

Project background

The saved learning material includes the fictional Chicago brief from Google Data Analytics and separate NYC business-intelligence planning documents. This portfolio analysis uses the preserved NYC export; the files do not establish original project completion or grading in either programme. The original SQL extraction, raw rides and Tableau workbook are unavailable, and I do not claim to have reproduced them.

AI assistance supported the current Python analysis, documentation and portfolio refinement.

Potential use

A bike-share analyst planning exploratory customer or operational research.

The decision

Which differences between recorded user types and days of the week warrant further investigation?

Original project outputs
Weekday averages use the number of calendar occurrences in 2015, making like-for-like day comparisons possible.
Detailed limitations

The preserved export does not establish complete NYC ridership coverage or current demand. Missing dates elsewhere in the export are not zero-demand observations. Commuting, tourism, weather effects and conversion intent are not demonstrated. Source CSV publication is withheld because rights remain unresolved.

Confirmed extraction coverage and source rights, plus contemporary ride, station-capacity and availability data before operational recommendations.

What I would do next

Confirm the extraction coverage and rights to the underlying records, then compare the patterns with current ride and station data before considering operational recommendations.

Explore the repository