The Wandering Fork Dataset
Data Description | Reference
This page documents the dataset used throughout the course. All data are entirely invented — no real personal or business data are included.
Downloads
| File | Format | Description |
|---|---|---|
wandering-fork.csv |
CSV | Full long-format table (240 rows × 16 columns) |
wandering-fork.xlsx |
Excel | Same data as an Excel Table (tblFork) plus subset sheets |
payroll-skew.csv |
CSV | Small skewed example (11 rows) for Modules 4 & 7 |
lakeside-bakery.csv |
CSV | Fresh 2-outlet dataset (120 rows) for Module 10 walkthrough |
The Excel workbook contains four sheets:
- data — the full 240-row table as an Excel Table named
tblFork - harbour_30d — the first 30 days of the Harbour franchise (hand-computation subset)
- campus_30d — the first 30 days of the Campus franchise (hand-computation subset)
- dictionary — the data dictionary below
Data dictionary
| Name | Statistical type | Unit | Meaning |
|---|---|---|---|
franchise |
Qualitative nominal | — | Franchise name: Harbour or Campus |
day |
Quantitative discrete | day number | Day index within franchise (1–120) |
date |
Date | YYYY-MM-DD | Calendar date |
day_of_week |
Qualitative nominal | — | Monday … Sunday |
month |
Qualitative ordinal | — | April … August (calendar order) |
weather |
Qualitative nominal | — | Sunny / Cloudy / Rainy |
temperature_c |
Quantitative continuous | °C | Daily temperature |
event_nearby |
Qualitative nominal (binary) | — | Yes / No — event near the truck |
promo_active |
Qualitative nominal (binary) | — | Yes / No — promotion running |
staff_on_duty |
Quantitative discrete | people | Staff working that day |
foot_traffic |
Quantitative discrete | people | Passers-by count |
customer_count |
Quantitative discrete | customers | Number of customers served |
items_sold |
Quantitative discrete | items | Total items sold |
daily_revenue_eur |
Quantitative continuous | € | Total revenue for the day |
waste_kg |
Quantitative continuous | kg | Food waste (right-skewed) |
satisfaction_rating |
Qualitative ordinal | 1–5 | Customer satisfaction rating |
Designed properties (verified by verify_dataset.py)
- Harbour days 1–30 match the legacy CSV exactly on the eight original columns.
- No missing values; no duplicate
(franchise, day)pairs. items_sold ≥ customer_count;customer_count ≤ foot_traffic; revenue > 0 everywhere.- Campus revenue is right-skewed (skewness > +0.5); Harbour revenue is near-symmetric (|skew| ≤ 0.5).
- Strong customer–revenue correlation in both franchises (r ≥ 0.80); weak temperature–revenue correlation (|r| ≤ 0.35).
- Weather × promo dependence visible in conditional profiles (≥ 15 pp row-profile gap in at least one franchise).
- At least one 1.5 × IQR outlier per franchise.
- Every category appears at least 5 times (no empty contingency cells).
Secondary datasets
payroll-skew.csv (Modules 4 & 7)
11 rows: one “CEO” salary (€180,000) plus ten staff salaries (€25k–€45k). Strong right skew for illustrating mean vs median vs mode and Pearson’s skewness coefficient.
lakeside-bakery.csv (Module 10)
Two outlets (Riverside, Hilltop) × 60 days = 120 rows. Variables: outlet, day, date, day_of_week, month, weather, temperature_c, foot_traffic, customer_count, items_sold, daily_revenue_eur, waste_kg, satisfaction_rating. Used for the Module 10 synthesis walkthrough on a fresh dataset.
How the data were invented
All values were generated by a seeded Python script (generate_dataset.py) using numpy/pandas. The script:
- Copies the original 30-day Harbour data verbatim for the eight legacy columns.
- Generates the remaining 90 days for Harbour and all 120 days for Campus with realistic seasonal, weekday, weather, and promotional effects.
- Adds the new columns (month, event_nearby, promo_active, staff_on_duty, foot_traffic, waste_kg, satisfaction_rating) for all 240 rows, including the first 30 Harbour days.
- Enforces logical constraints (e.g.,
customer_count ≤ foot_traffic). - Calibrates two Harbour outliers (days 45 and 80) and Campus event-driven skew to satisfy the invariants above.
The seed is fixed (42) so the dataset is fully reproducible.