Module 4 — Measures of central tendency
Data Description | Part II — Summarising one variable
The business question. “The Wandering Fork” has 120 days of daily revenue data per franchise. The owner wants “one typical number” to put in a report — but which one? The mean, the median, and the mode can give different answers, especially when the data is skewed. This module teaches you to compute all three, understand when they diverge, and choose the right one for the business question.
Learning objectives
By the end of this module, you will be able to:
- Compute the arithmetic mean, median, and mode of raw data.
- Compute the grouped-data mean (using class midpoints) and the weighted mean (for portfolio returns, GPA, etc.).
- Explain why the mean, median, and mode can differ, and choose the most appropriate one for a given business question.
- Prove that the arithmetic mean is the value that minimizes the sum of squared deviations — the property that makes it the natural partner of the variance (Module 6).
Time plan
| Block | Minutes | What you do |
|---|---|---|
| Theory | 70 | Mean, median, mode, grouped mean, weighted mean, sensitivity to outliers, which measure to choose |
| Excel lab | 50 | AVERAGE, MEDIAN, MODE.SNGL, SUMPRODUCT for weighted/grouped mean |
| Exercises | 60 | Three by-hand and three Excel exercises |
| Total | 180 |
Theory
Arithmetic mean
The arithmetic mean (or average) \(\bar{x}\) of \(n\) values \(x_1,\dots,x_n\) is their sum divided by their count. It uses every single observation, which makes it efficient but also sensitive to extreme values: one very large or very small \(x_i\) pulls \(\bar{x}\) toward it.
\[ \bar{x} = \frac{1}{n}\sum_{i=1}^{n} x_i. \] In plain terms: add up all the observations, then split the total evenly across them — the “if everyone got the same amount” value.
Median
The median \(Me\) is the value that splits an ordered series into two equal halves: at least 50% of observations are \(\leq Me\) and at least 50% are \(\geq Me\). Unlike the mean, the median only cares about the rank of each value, not its magnitude, so a single extreme observation cannot drag it far.
For \(n\) ordered values, the median sits at position \((n+1)/2\): if \(n\) is odd this is a single observation; if \(n\) is even, it is the average of the two middle observations. This is exactly the quantile formula with \(p=0.5\) (see Module 5).
Mode
The mode \(Mo\) is the most frequently occurring value (for discrete or categorical data) or the center of the most frequent class (for grouped continuous data — the “peak” of the histogram from Module 3). A distribution can have one mode (unimodal), two (bimodal), or more.
For continuous data grouped into classes of equal width \(h\), the mode is estimated by linear interpolation inside the modal class (the class with the highest frequency \(f_1\)): \[ Mo = L + h \cdot \frac{f_1-f_0}{(f_1-f_0)+(f_1-f_2)}, \] where \(L\) is the modal class’s lower bound, \(f_0\) is the frequency of the preceding class, and \(f_2\) is the frequency of the following class. Intuitively, the mode leans toward whichever neighboring class is “competing” harder with the modal class for observations: a much lower \(f_0\) than \(f_2\) pulls the estimated mode toward the left edge of the class, and vice versa.
Grouped-data mean
When raw data is unavailable and only a class frequency table exists, the mean is approximated using class midpoints \(m_i\): \[ \bar{x}_{\text{grouped}} = \frac{\sum_{i=1}^{k} f_i m_i}{\sum_{i=1}^{k} f_i} = \frac{\sum_{i=1}^{k} n_i m_i}{n}, \] where \(n_i\) is the count in class \(i\), \(m_i = (L_i + L_i + h)/2\) is the class midpoint, and \(k\) is the number of classes. This is a weighted mean where the weights are the class frequencies.
In plain words: treat every observation in a class as if it were exactly at the class midpoint, then compute the usual mean. The approximation error is exactly the information lost when data is grouped.
Weighted mean
When observations carry different importance (weights \(w_i\)), the weighted mean is \[ \bar{x}_w = \frac{\sum_{i=1}^{n} w_i x_i}{\sum_{i=1}^{n} w_i}. \] Common business uses: portfolio return (weights = amounts invested), GPA (weights = credit hours), price indices (weights = quantities).
In plain words: multiply each value by its weight, sum the products, and divide by the total weight. The ordinary mean is the special case where all weights are equal.
Choosing the right measure
- Mean: symmetric distributions, no outliers, when you need the value that feeds into variance/standard deviation (Modules 6–7).
- Median: skewed distributions, ordinal data, or when outliers would distort the mean.
- Mode: nominal data, or to identify the most common category (e.g., the most popular product).
- Weighted mean: when observations have different importance (portfolio, GPA, indices).
- Grouped mean: only when raw data is unavailable; always state it is an approximation.
Worked example
We use the Harbour 30-day subset (sheet harbour_30d) and the payroll-skew.csv dataset to illustrate the concepts of this module.
Harbour 30-day revenue (symmetric case)
| Statistic | Value |
|---|---|
| Mean \(\bar{x}\) | 628.34 |
| Median \(Me\) | 643.63 |
| Grouped mode \(Mo\) | 666.67 |
Class frequency table (class width 100, from Module 3):
| Class (€) | Frequency |
|---|---|
| [400, 500[ | 6 |
| [500, 600[ | 6 |
| [600, 700[ | 8 |
| [700, 800[ | 7 |
| [800, 900[ | 3 |
Step 1 — mean. Summing all 30 daily revenues and dividing by 30 gives \(\bar{x} = 628.34\): on average, “The Wandering Fork” brings in about €628 per day.
Step 2 — median. As found in Module 5, \(Me = 643.63\): half the trading days earned less than this, half earned more.
Step 3 — mode. The modal class is \([600,700[\) with \(f_1=8\), flanked by \([500,600[\) (\(f_0=6\)) and \([700,800[\) (\(f_2=7\)). With \(L=600\) and \(h=100\): \[ Mo = 600 + 100 \times \frac{8-6}{(8-6)+(8-7)} = 600 + 100 \times \frac{2}{3} \approx 666.67. \] The single most “typical” revenue band, estimated from the histogram shape, centers around €666.67.
Putting it together. The three indicators — 628.34, 643.63, and 666.67 — are all close to one another, which is itself informative: it signals a roughly symmetric, well-behaved distribution rather than one dominated by a skewed tail.
Payroll skew (right-skewed case)
The payroll-skew.csv dataset contains 11 salaries: one CEO at €180,000 and ten staff members between €25,000–€45,000.
| Statistic | Value |
|---|---|
| Mean \(\bar{x}\) | 47,231 |
| Median \(Me\) | 38,050 |
| Mode \(Mo\) | 25,166 (no unique mode; smallest value shown) |
Step 1 — mean. The CEO’s €180,000 pulls the mean up to €47,231, well above what 10 out of 11 employees earn.
Step 2 — median. The median (€38,050) is unaffected by the extreme value and better represents a “typical” salary.
Step 3 — mode. No value repeats exactly; the distribution is strongly right-skewed (skewness = 3.12), so the mean > median > mode pattern holds.
Using Excel
The complete list of function names, with their French equivalents, is on the Excel functions page.
| Concept | Excel function / steps |
|---|---|
| Arithmetic mean | =AVERAGE(range) |
| Median | =MEDIAN(range) |
| Mode, raw data (single mode) | =MODE.SNGL(range) |
| Mode, raw data (all modes, array formula) | =MODE.MULT(range) |
| Weighted mean | =SUMPRODUCT(values, weights)/SUM(weights) |
| Grouped mean | No direct one-cell function — use SUMPRODUCT(midpoints, frequencies)/SUM(frequencies) |
| Mode, grouped continuous data | No direct one-cell function — use FREQUENCY(range, bins) to get \(f_0, f_1, f_2\), then apply the interpolation formula manually |
For the full table (120 days per franchise) use the Excel Table tblFork and add the franchise as one more condition, for example =AVERAGEIFS(tblFork[daily_revenue_eur], tblFork[franchise], "Harbour").
Proof / derivation
Consider the function that measures the total squared “distance” between every observation and some reference point \(a\): \[ f(a) = \sum_{i=1}^{n} (x_i - a)^2. \] To find the value of \(a\) that minimizes \(f\), differentiate with respect to \(a\) and set the result to zero: \[ f'(a) = -2\sum_{i=1}^{n} (x_i - a) = 0 \;\;\Longrightarrow\;\; \sum_{i=1}^{n} x_i - na = 0 \;\;\Longrightarrow\;\; a = \frac{1}{n}\sum_{i=1}^{n} x_i = \bar{x}. \] The second derivative, \(f''(a) = 2n > 0\), confirms this is a minimum, not a maximum. In words: no other single number sits closer, on average in squared terms, to every observation than the arithmetic mean does. This is exactly why the variance (Module 6) is defined as the average squared deviation from the mean specifically — any other reference point would only make the total squared deviation larger.
Visual intuition
Picture each day’s revenue as a small weight placed on a number line. The mean is the physical balance point (center of mass): if the number line were a rigid rod, it would balance perfectly on a pivot placed at \(\bar{x}=628.34\). The median is a purely positional idea: the point that has exactly as many weights to its left as to its right, regardless of how heavy any of them are. The mode is simply the tallest bar on the histogram — the single most crowded spot on the number line. When the histogram is close to symmetric, all three landmarks cluster near the same spot, which is exactly the pattern seen above (628.34, 643.63, 666.67). When a single extreme weight is added (the CEO’s salary), the balance point (mean) shifts dramatically, while the positional median barely moves.
Interactive demo
TODO: an interactive mean/median/mode explorer (editable sample) is a candidate for a future Shinylive demo.
Exercises
Five employee monthly salaries (in €k) are: 32, 35, 33, 34, 210. Compute the mean and the median. Which one better represents a “typical” salary here, and why?
Solution. Mean \(= (32+35+33+34+210)/5 = 344/5 = 68.8\). Ordering the values gives \(32, 33, 34, 35, 210\), so the median (the middle value) is \(34\). The median is the better representative here: the single very high salary (210) pulls the mean far above what four out of five employees actually earn, while the median, using only rank, is unaffected by how extreme that one value is.
The owner of “The Wandering Fork” asks for “one typical number” for daily revenue. Given \(\bar{x}=628.34\) and \(Me=643.63\), which would you recommend, and why?
Solution. Since Module 5 showed zero outliers in the daily revenue data, the mean is not being distorted by extreme values, so either indicator is defensible here. The mean is generally the better recommendation in this specific case: it uses every one of the 30 observations (not just their ranks), and it is the indicator that feeds directly into the variance, standard deviation, and coefficient of variation computed in Module 6 — keeping the reporting consistent across the analysis.
A rival street-food vendor’s daily revenue has a modal class \([500,600[\) (\(L=500\), \(h=100\)) with frequency \(f_1=18\), preceded by a class with frequency \(f_0=10\) and followed by a class with frequency \(f_2=15\). Estimate the grouped mode.
Solution. \[ Mo = 500 + 100 \times \frac{18-10}{(18-10)+(18-15)} = 500 + 100 \times \frac{8}{11} \approx 572.73. \] The estimated most-typical revenue band for the rival vendor centers around €572.73 — lower than “The Wandering Fork”’s €666.67, and leaning toward the lower-frequency neighboring class (\(f_0=10\) vs. \(f_2=15\)) rather than the middle of the modal class.
An investor holds three assets: €10,000 in Asset A (return 8%), €20,000 in Asset B (return 5%), and €5,000 in Asset C (return 12%). Compute the portfolio’s weighted mean return.
Solution. =SUMPRODUCT({10000,20000,5000},{0.08,0.05,0.12})/SUM({10000,20000,5000}) = 7.00%. The portfolio return is the weighted mean of the individual returns, with weights equal to the amounts invested.
Using the Harbour 30-day class table from the worked example (five classes of width 100, midpoints 450, 550, 650, 750, 850), compute the grouped mean and compare it to the raw-data mean (628.34).
Solution. =SUMPRODUCT({450,550,650,750,850},{6,6,8,7,3})/SUM({6,6,8,7,3}) = 19,000/30 = 633.33. The grouped mean (633.33) is close to the raw mean (628.34) — the difference (about €5) is the approximation error from using class midpoints instead of the exact 30 values.
In wandering-fork.xlsx, filter franchise to “Campus” and compute the mean, median, and mode of daily_revenue_eur for all 120 days. How do they compare to Harbour’s values?
Solution. Campus 120-day: mean = 455.03, median = 390.39, mode (no unique mode). Over the same 120 days, Harbour has mean = 413.49 and median = 391.47. The medians are almost equal (390 vs 391), but the Campus mean sits much further above its median (gap €65 vs €22), confirming the stronger right skew seen in Module 7 (Campus skewness = 1.00 vs Harbour’s 0.47).
Common mistakes
- Using the mean for skewed data. The payroll example shows how a single extreme value can make the mean misleading; the median is safer.
- Forgetting to lock the total in a weighted mean. The denominator
SUM(weights)must be absolute or a structured reference, otherwise copying the formula shifts the total. - Confusing grouped mean with raw mean. The grouped mean is an approximation; always state it as such and prefer the raw mean when available.
- Using
MODE.SNGLon continuous data. Raw continuous data rarely has exact repeats; use the grouped mode formula instead. - Treating ordinal numbers as ratio.
satisfaction_ratingis 1–5, but=AVERAGEon it is questionable; use the median.
Further reading
| Source | Where | What it adds | Time | Verified |
|---|---|---|---|---|
| OpenStax IBS 2e | §2.3 Measures of the Center of the Data (book pages 74–78) | Mean, median, mode, grouped mean | ~4 pages | checked against the PDF contents (T046) |
| OpenStax IBS 2e | §2.4 Sigma Notation and Calculating the Arithmetic Mean (book pages 78–79) | Summation notation, mean formula | ~1 page | checked against the PDF contents (T046) |
| OpenStax IBS 2e | §2.5 Geometric Mean (book pages 79–80) | Geometric mean for growth rates (enrichment) | ~1 page | checked against the PDF contents (T046) |
| Khan Academy | Unit Summarizing quantitative data in the Statistics and probability course; look for lessons on mean, median, mode, and weighted mean | Short videos and practice on central tendency measures | ~30–45 min (estimate) | link check pending (T048) |
OpenStax IBS 2e: Alexander Holmes, Barbara Illowsky and Susan Dean, Introductory Business Statistics 2e, OpenStax, Rice University, https://openstax.org/details/books/introductory-business-statistics-2e, licensed under CC BY-NC-SA 4.0. Sections are cited, not reproduced.