The Wandering Fork Dataset

Data Description | Reference

Modified

October 6, 2026

This page documents the dataset used throughout the course. All data are entirely invented — no real personal or business data are included.

Downloads

File Format Description
wandering-fork.csv CSV Full long-format table (240 rows × 16 columns)
wandering-fork.xlsx Excel Same data as an Excel Table (tblFork) plus subset sheets
payroll-skew.csv CSV Small skewed example (11 rows) for Modules 4 & 7
lakeside-bakery.csv CSV Fresh 2-outlet dataset (120 rows) for Module 10 walkthrough

The Excel workbook contains four sheets:

  • data — the full 240-row table as an Excel Table named tblFork
  • harbour_30d — the first 30 days of the Harbour franchise (hand-computation subset)
  • campus_30d — the first 30 days of the Campus franchise (hand-computation subset)
  • dictionary — the data dictionary below

Data dictionary

Name Statistical type Unit Meaning
franchise Qualitative nominal — Franchise name: Harbour or Campus
day Quantitative discrete day number Day index within franchise (1–120)
date Date YYYY-MM-DD Calendar date
day_of_week Qualitative nominal — Monday … Sunday
month Qualitative ordinal — April … August (calendar order)
weather Qualitative nominal — Sunny / Cloudy / Rainy
temperature_c Quantitative continuous °C Daily temperature
event_nearby Qualitative nominal (binary) — Yes / No — event near the truck
promo_active Qualitative nominal (binary) — Yes / No — promotion running
staff_on_duty Quantitative discrete people Staff working that day
foot_traffic Quantitative discrete people Passers-by count
customer_count Quantitative discrete customers Number of customers served
items_sold Quantitative discrete items Total items sold
daily_revenue_eur Quantitative continuous € Total revenue for the day
waste_kg Quantitative continuous kg Food waste (right-skewed)
satisfaction_rating Qualitative ordinal 1–5 Customer satisfaction rating

Designed properties (verified by verify_dataset.py)

  1. Harbour days 1–30 match the legacy CSV exactly on the eight original columns.
  2. No missing values; no duplicate (franchise, day) pairs.
  3. items_sold ≥ customer_count; customer_count ≤ foot_traffic; revenue > 0 everywhere.
  4. Campus revenue is right-skewed (skewness > +0.5); Harbour revenue is near-symmetric (|skew| ≤ 0.5).
  5. Strong customer–revenue correlation in both franchises (r ≥ 0.80); weak temperature–revenue correlation (|r| ≤ 0.35).
  6. Weather × promo dependence visible in conditional profiles (≥ 15 pp row-profile gap in at least one franchise).
  7. At least one 1.5 × IQR outlier per franchise.
  8. Every category appears at least 5 times (no empty contingency cells).

Secondary datasets

payroll-skew.csv (Modules 4 & 7)

11 rows: one “CEO” salary (€180,000) plus ten staff salaries (€25k–€45k). Strong right skew for illustrating mean vs median vs mode and Pearson’s skewness coefficient.

lakeside-bakery.csv (Module 10)

Two outlets (Riverside, Hilltop) × 60 days = 120 rows. Variables: outlet, day, date, day_of_week, month, weather, temperature_c, foot_traffic, customer_count, items_sold, daily_revenue_eur, waste_kg, satisfaction_rating. Used for the Module 10 synthesis walkthrough on a fresh dataset.

How the data were invented

All values were generated by a seeded Python script (generate_dataset.py) using numpy/pandas. The script:

  • Copies the original 30-day Harbour data verbatim for the eight legacy columns.
  • Generates the remaining 90 days for Harbour and all 120 days for Campus with realistic seasonal, weekday, weather, and promotional effects.
  • Adds the new columns (month, event_nearby, promo_active, staff_on_duty, foot_traffic, waste_kg, satisfaction_rating) for all 240 rows, including the first 30 Harbour days.
  • Enforces logical constraints (e.g., customer_count ≤ foot_traffic).
  • Calibrates two Harbour outliers (days 45 and 80) and Campus event-driven skew to satisfy the invariants above.

The seed is fixed (42) so the dataset is fully reproducible.