Module 7 — Shape and univariate review
Data Description | Part II — Summarising one variable
The business question. “The Wandering Fork” has a mean daily revenue of €628.34 and a standard deviation of €137.09 (Harbour, 30 days) — but knowing how much the data spreads out says nothing about how it spreads out: is it balanced evenly around its center, or lopsided with a longer tail in one direction? Skewness indicators answer exactly that question, and they close the loop on what the mean, median, and mode (Module 4) were quietly hinting at all along. This module also consolidates the entire univariate toolkit (Modules 3–6) into one complete workflow.
Learning objectives
By the end of this module, you will be able to:
- Explain what a symmetric, right-skewed, and left-skewed distribution look like, and how the mean, median, and mode typically order themselves in each case.
- Compute and interpret the Fisher-Pearson sample skewness coefficient (Excel’s
SKEW), the primary skewness measure used in this course. - Compute the Pearson skewness coefficient and the Yule-Kendall skewness coefficient as alternative, simpler measures, and recognize that different formulas can disagree in sign or magnitude for near-symmetric data.
- Read skewness directly off a histogram or boxplot.
- Build a complete descriptive-statistics summary for a quantitative variable in Excel, from raw data to shape indicators, without external guidance.
Time plan
| Block | Minutes | What you do |
|---|---|---|
| Theory | 60 | Symmetry vs skewness, Fisher-Pearson skewness, Pearson and Yule-Kendall coefficients, reading skew from charts |
| Excel lab | 50 | SKEW, histogram overlay, full univariate workflow |
| Exercises | 70 | Three by-hand and three Excel exercises (including full workflow) |
| Total | 180 |
Theory
Symmetry and skewness
A distribution is symmetric if its left and right halves, measured from the center, are mirror images of each other; in that case the mean, median, and mode coincide. A distribution is skewed when one tail is longer or fatter than the other. Right-skewed (positive skew) means a longer tail toward high values, typically pulling the mean above the median above the mode (\(\bar{x} > Me > Mo\)). Left-skewed (negative skew) is the mirror case, with a longer tail toward low values (\(\bar{x} < Me < Mo\)).
In plain words: symmetric = balanced; right-skewed = tail drags to the right (a few exceptionally high values); left-skewed = tail drags to the left (a few exceptionally low values).
Fisher-Pearson sample skewness coefficient
The primary skewness measure used in this course is the adjusted Fisher-Pearson standardized moment coefficient, which is what Excel’s SKEW function computes: \[
SKEW = \frac{n}{(n-1)(n-2)} \sum_{i=1}^{n}
\left(\frac{x_i-\bar{x}}{s}\right)^3.
\] This standardizes every deviation from the mean (dividing by \(s\), so the result is unit-free), cubes it (so, unlike variance, positive and negative deviations do not cancel and large deviations dominate), and averages the result with a small-sample correction factor. A value near \(0\) signals near-symmetry; positive means right-skewed; negative means left-skewed.
In plain words: cube the standardized deviations (so direction matters), average them with a small-sample correction. The correction factor \(\frac{n}{(n-1)(n-2)}\) matters most for small samples and becomes negligible for large ones.
Alternative asymmetry coefficients
Two simpler, complementary measures are also common. The Pearson skewness coefficient compares the mean to the mode, rescaled by the standard deviation. The Yule-Kendall skewness coefficient instead compares how the median sits between \(Q_1\) and \(Q_3\), using only the five-number summary from Module 5 — useful when the mode is hard to pin down or the data is only available as a boxplot.
\[ \text{Pearson asymmetry} = \frac{\bar{x}-Mo}{s}, \qquad\qquad YK = \frac{Q_3+Q_1-2\,Me}{Q_3-Q_1}. \] When the mode is not well defined (raw continuous data, no repeated value), use Pearson’s median form, \(3(\bar{x}-Me)/s\), which swaps the mode for the median. The Pearson coefficient asks “how many standard deviations separate the mean from the mode (or median)?” The Yule-Kendall coefficient asks “is the median closer to \(Q_1\) or to \(Q_3\)?” — if \(Me\) sits closer to \(Q_1\) than to \(Q_3\), the upper quarter of the data is more spread out, which is itself a sign of right skew, and vice versa.
Different skewness formulas can genuinely disagree in sign or magnitude, especially for near-symmetric data where all the values involved are close to each other and small estimation differences get amplified by division. Always state which formula you used when reporting a skewness value — “the Fisher-Pearson skewness is +0.034” is a precise, checkable claim; “the skewness is positive” is not, unless the formula is named.
Kurtosis (conceptual only)
Kurtosis describes the “tailedness” and peak of a distribution relative to a normal distribution: leptokurtic (heavy tails, sharp peak), mesokurtic (normal-like), platykurtic (light tails, flat peak). In this course, kurtosis is treated conceptually only — names and pictures, no formula, no Excel function.
Worked example
Using “The Wandering Fork”’s daily_revenue_eur (\(n=30\)), with \(\bar{x}=628.34\), \(Me=643.63\), grouped mode \(Mo=666.67\) (Module 4), \(s=137.09\) (Module 6), and \(Q_1=506.61\), \(Q_3=729.99\), \(IQR=223.39\) (Module 5), all verified in data-model.md:
| Statistic | Value | Interpretation |
|---|---|---|
| Fisher-Pearson sample skewness | +0.034 | near-symmetric, very mild right skew |
| Pearson skewness coefficient | −0.280 | mildly left, per this formula |
| Pearson median form \(3(\bar{x}-Me)/s\) | −0.335 | mildly left, per this formula |
| Yule-Kendall skewness coefficient | −0.227 | mildly left, per this formula |
Step 1 — the headline measure. The Fisher-Pearson sample skewness, computed with Excel’s SKEW on all 30 raw observations, is \(+0.034\). Being so close to zero, this is the course’s official conclusion: “The Wandering Fork”’s daily revenue is essentially symmetric, with only a very slight lean toward high values.
Step 2 — Pearson asymmetry. \((\bar{x}-Mo)/s = (628.34-666.67)/137.09 \approx -0.280\): comparing the mean to the grouped-estimated mode instead of using every raw value gives a mildly negative reading. The median form, \(3(628.34-643.63)/137.09 \approx -0.335\), agrees in sign.
Step 3 — Yule-Kendall. \((Q_3+Q_1-2Me)/(Q_3-Q_1) = (729.99+506.61-2\times643.63)/223.39 \approx -0.227\): comparing only the five-number summary also leans mildly negative.
Step 4 — reconciling the disagreement. Steps 2 and 3 lean slightly negative while Step 1 leans slightly positive. This is not a contradiction, it is a lesson: the Fisher-Pearson coefficient uses every one of the 30 raw daily revenues, while the Pearson coefficient depends on the grouped mode (\(Mo=666.67\)), itself only an estimate from five class frequencies, and the Yule-Kendall coefficient uses just four numbers (\(Q_1, Me, Q_3\), and the range between the first two). For a distribution this close to symmetric, such small formula-specific approximations are enough to flip the sign — which is itself consistent with the headline conclusion that the true skew is very small in absolute terms either way. The value to report and remember for “The Wandering Fork” is the Fisher-Pearson \(+0.034\).
Campus franchise (right-skewed case)
For Campus 120-day revenue: Fisher-Pearson skewness = +1.00, Pearson median form = +0.63 (Campus revenue has no unique mode, so the median form is used), Yule-Kendall = +0.15. All three are positive and agree on right skew — the event-driven revenue spikes create a long upper tail. Yule-Kendall is the weakest because it only looks at the middle half of the data and ignores the tail beyond \(Q_3\). This is the pattern Module 6’s high \(CV\) (67.6%) and Module 5’s three outliers were already signaling.
Using Excel
The complete list of function names, with their French equivalents, is on the Excel functions page.
| Concept | Excel function / steps |
|---|---|
| Fisher-Pearson sample skewness | =SKEW(range) |
| Pearson skewness coefficient | No direct function — =(AVERAGE(range)-Mo)/STDEV.S(range), with \(Mo\) computed manually (Module 4) |
| Yule-Kendall skewness coefficient | No direct function — =(QUARTILE.EXC(range,3)+QUARTILE.EXC(range,1)-2*MEDIAN(range))/(QUARTILE.EXC(range,3)-QUARTILE.EXC(range,1)) |
For the full table (120 days per franchise) use the Excel Table tblFork and add the franchise as one more condition, for example =SKEW(FILTER(tblFork[daily_revenue_eur], tblFork[franchise]="Campus")).
Proof / derivation
SKEW includes a bias-correction factor (click to expand)
The conceptually simplest skewness measure is the third standardized moment, \[ g_1 = \frac{\frac{1}{n}\sum_{i=1}^{n}(x_i-\bar{x})^3} {\left(\frac{1}{n}\sum_{i=1}^{n}(x_i-\bar{x})^2\right)^{3/2}} = \frac{m_3}{m_2^{3/2}}, \] the average cubed deviation (\(m_3\)), rescaled by the standard deviation cubed (via \(m_2^{3/2}\)) so the result is unit-free. Cubing (rather than squaring, as in variance) is what makes this measure sensitive to direction: a value far above the mean and one equally far below contribute equal-magnitude but opposite-sign terms, so they do not cancel the way squared deviations would.
However, \(g_1\) computed directly from a sample is a biased estimator of the true population skewness — it systematically under- or overstates it, especially for small \(n\). Excel’s SKEW applies the standard small-sample correction, replacing \(\frac{1}{n}\) with the factor \(\frac{n}{(n-1)(n-2)}\) and using the sample standard deviation \(s\) (with its own \(n-1\) correction, Module 6) in place of the population-style \(m_2^{1/2}\): \[
SKEW = \frac{n}{(n-1)(n-2)}\sum_{i=1}^n\left(\frac{x_i-\bar x}{s}\right)^3.
\] As \(n\) grows large, \(\frac{n}{(n-1)(n-2)} \to \frac{1}{n}\) and \(s \to
m_2^{1/2}\), so \(SKEW \to g_1\): the correction matters most for small samples (like “The Wandering Fork”’s \(n=30\)) and becomes negligible for very large ones.
Visual intuition
On a histogram, skewness is the direction of the longer tail: a right-skewed histogram has its bulk of bars on the left and a thin tail trailing off to the right (think: most days are ordinary, a few are exceptionally good); left-skewed is the mirror image. On a boxplot, skewness shows up as an off-center median line inside the box and unequal whisker lengths: if the whisker toward \(Q_3\) is longer than the one toward \(Q_1\), that is visual right skew. For “The Wandering Fork”, the near-equal spacing seen in Module 5’s boxplot (median roughly centered between \(Q_1=506.61\) and \(Q_3=729.99\)) is the visual signature of the near-zero skewness confirmed numerically here.
Interactive demo
TODO: an interactive skewness visualizer (editable sample, live histogram + skewness value) is a candidate for a future Shinylive demo.
Exercises
A distribution has a Fisher-Pearson sample skewness of \(-0.85\). What does the sign and rough magnitude tell you about its shape, and about the typical ordering of its mean and median?
Solution. The negative sign means the distribution is left-skewed: it has a longer or fatter tail toward low values. A magnitude of \(0.85\) (using the common rule of thumb: roughly \(|SKEW|<0.5\) near-symmetric, \(0.5\)–\(1\) moderate skew, \(>1\) strong skew) indicates a moderate left skew — noticeable, but not extreme. In this case, the typical ordering is \(\bar{x} < Me < Mo\): the few low outlying values pull the mean below the median.
“The Wandering Fork”’s mean (628.34), median (643.63), and grouped mode (666.67) follow the order \(\bar{x} < Me < Mo\), which is the classic textbook signature of left skew — yet the Fisher-Pearson coefficient reports \(+0.034\), a very mild right skew. How do you reconcile this?
Solution. The “\(\bar{x}<Me<Mo\) implies left skew” rule is only a rough heuristic, not a guarantee, and it is especially unreliable here for two reasons: (1) the gaps between the three values are tiny relative to the spread of the data (\(s=137.09\)), so their exact order can flip with negligible changes in the underlying data; and (2) the mode used is a grouped estimate (\(Mo=666.67\), from five class frequencies), not an exact value computed from all 30 raw observations the way \(\bar{x}\), \(Me\), and the Fisher-Pearson coefficient are. The Fisher-Pearson value, using every raw observation, is the more precise and authoritative measure here, and at \(+0.034\) it confirms the distribution is essentially symmetric — close enough to zero that a heuristic based on three closely-spaced, partly-estimated landmarks should not be trusted to override it.
A dataset has \(Q_1=40\), \(Me=50\), \(Q_3=55\). Compute the Yule-Kendall skewness coefficient and interpret its sign.
Solution. \[ YK = \frac{Q_3+Q_1-2Me}{Q_3-Q_1} = \frac{55+40-2\times50}{55-40} = \frac{95-100}{15} = -\frac{5}{15} \approx -0.33. \] The negative sign indicates mild left skew by this measure: the median (50) sits closer to \(Q_3\) (55) than to \(Q_1\) (40), meaning the lower quarter of the data is more spread out than the upper quarter — a longer effective “reach” on the low side.
In wandering-fork.xlsx, filter franchise to “Campus” and compute =SKEW for daily_revenue_eur (120 days). Also compute the Pearson (median form, since there is no unique mode) and Yule-Kendall coefficients. Do all three agree on the direction?
Solution. Campus 120-day: SKEW = 1.00, Pearson median form = \(3(455.03-390.39)/307.43 \approx 0.63\), Yule-Kendall = \((598.60+236.07-2\times390.39)/362.54 \approx 0.15\). All three are positive, confirming right skew — consistent with the event-driven revenue spikes and the three outliers found in Module 5. Yule-Kendall is much smaller because the long tail lies beyond \(Q_3\), where it cannot see.
Open payroll-skew.csv. Compute the Fisher-Pearson skewness for the 11 salaries. What does the value tell you?
Solution. SKEW = 3.12. This is strong right skew — the CEO’s €180,000 salary creates an extremely long upper tail. The mean (€47,231) is far above the median (€38,050), and the mode (no unique mode) would be even lower. This is the most skewed dataset in the course.
Using wandering-fork.xlsx filtered to “Campus” (120 days), build a complete descriptive summary for daily_revenue_eur in one continuous Excel workflow: class table (Module 3), quartiles and boxplot (Module 5), mean/median/mode (Module 4), variance/std/CV (Module 6), skewness (this module). Verify internal consistency (e.g., \(Q_1 \le Me \le Q_3\), mean ≈ median for symmetric data).
Solution. Campus 120-day summary: - Class table: 8 classes of width 200 (Module 3, Exercise 4), modal class [200,400[ (37 days), long tail up to [1400,1600[ - Quartiles: Q1=236.07, Me=390.39, Q3=598.60, IQR=362.54, fences -307.74 / 1142.40, three high outliers (1335.48, 1362.93, 1407.59) - Mean=455.03, median=390.39, no unique mode - Variance=94,511.80, std=307.43, CV=67.56% - Skewness=1.00 (right-skewed) - Consistency: Q1 (236) ≤ Me (390) ≤ Q3 (599) ✓; mean > median ✓ (right skew); CV > 35% ✓ (high variability)
Common mistakes
- Reporting “the skewness” without naming the formula. Always specify: Fisher-Pearson (Excel’s
SKEW), Pearson, or Yule-Kendall. - Trusting the mean/median/mode ordering heuristic for near-symmetric data. The gaps can be tiny; the Fisher-Pearson coefficient using all raw data is the authoritative measure.
- Using the grouped mode in the Pearson coefficient without stating it is an estimate. The grouped mode is an approximation; the Pearson coefficient computed with it inherits that approximation error.
- Confusing kurtosis with skewness. Skewness = asymmetry (lopsided); kurtosis = tailedness/peak (heavy vs light tails). This course covers skewness numerically; kurtosis conceptually only.
- Forgetting the small-sample correction in
SKEW. Excel’sSKEWincludes it automatically; if you implement the formula manually, you must include \(\frac{n}{(n-1)(n-2)}\).
Further reading
| Source | Where | What it adds | Time | Verified |
|---|---|---|---|---|
| OpenStax IBS 2e | §2.6 Skewness and the Mean, Median, and Mode (book pages 80–82) | Skewness definition, mean/median/mode ordering, Pearson coefficient | ~2 pages | checked against the PDF contents (T046) |
| Khan Academy | Unit Modeling data distributions in the Statistics and probability course; look for lessons on skewness and comparing distributions | Short videos and practice on skewness and distribution shape | ~30–45 min (estimate) | link check pending (T048) |
OpenStax IBS 2e: Alexander Holmes, Barbara Illowsky and Susan Dean, Introductory Business Statistics 2e, OpenStax, Rice University, https://openstax.org/details/books/introductory-business-statistics-2e, licensed under CC BY-NC-SA 4.0. Sections are cited, not reproduced.