Session 7 — Dispersion

Data Description | Part II

Session 6 pinned down “The Wandering Fork”’s central tendency — mean €628.34, median €643.63, mode €666.67 — three closely-clustered landmarks that describe the “typical” day. But two businesses can share the exact same mean and still be very different to run — one reliably steady, the other wildly erratic. Dispersion indicators measure exactly that variability, and they are what turns “the average day” into “the average day, ± how much”.

Learning objectives

  • Compute the range and the interquartile range.
  • Compute the sample variance and sample standard deviation, using both the definitional and the shortcut (computational) formula.
  • Compute and interpret the coefficient of variation.
  • Use the interactive calculator below to see how editing a sample changes its dispersion indicators live.

Theory

Definition — range and interquartile range

The range is \(\max - \min\): simple, but sensitive to a single extreme value. The interquartile range \(IQR = Q_3-Q_1\) (Session 5) measures the spread of only the central 50% of the data, making it more robust to extreme values.

Formula — sample variance and standard deviation

For a sample of \(n\) values \(x_1,\dots,x_n\) with mean \(\bar{x}\), the sample variance is \[ s^2 = \frac{1}{n-1}\sum_{i=1}^{n} (x_i-\bar{x})^2, \] i.e. the average squared distance of each observation from the mean (with \(n-1\), not \(n\), correcting for the fact that \(\bar{x}\) is itself estimated from the same sample). The sample standard deviation \(s=\sqrt{s^2}\) returns to the original unit (€, not €²).

Formula — coefficient of variation

\[ CV = \frac{s}{\bar{x}} \times 100\%. \] Because it is unit-free, \(CV\) is the only dispersion indicator that lets you compare the relative variability of two variables measured on different scales or units (e.g., revenue in € vs. customer count).

Worked example

Using “The Wandering Fork”’s daily_revenue_eur (\(n=30\), \(\bar{x}=628.34\)), verified in data-model.md:

Statistic Value
Range 478.97
Sample variance \(s^2\) 18,793.33
Sample std. dev. \(s\) 137.09
Coefficient of variation \(CV\) 21.82%

Step 1 — range. \(886.56 - 407.59 = 478.97\): over the 30-day period, the best day out-earned the worst by almost €480.

Step 2 — variance and standard deviation. Squaring each day’s deviation from \(\bar{x}=628.34\), summing, and dividing by \(n-1=29\) gives \(s^2 = 18{,}793.33\) €²; taking the square root gives \(s = 137.09\) €. On a typical day, revenue is about €137 away from the €628.34 average, in either direction.

Step 3 — coefficient of variation. \(CV = 137.09/628.34 \times 100\% = 21.82\%\) — using the guide from session-06.qmd’s central tendency companion (below \(35\%\)), a moderate level of relative variability: noticeably variable, but not wildly erratic day to day.

Using Excel

Concept Excel function / steps
Range =MAX(range)-MIN(range)
Sample variance =VAR.S(range)
Sample standard deviation =STDEV.S(range)
Coefficient of variation =STDEV.S(range)/AVERAGE(range)
Interquartile range =QUARTILE.EXC(range,3)-QUARTILE.EXC(range,1)

Proof / derivation

Starting from the definitional formula and expanding the square: \[ \sum_{i=1}^{n}(x_i-\bar{x})^2 = \sum_{i=1}^n x_i^2 - 2\bar{x}\sum_{i=1}^n x_i + n\bar{x}^2. \] Since \(\sum_{i=1}^n x_i = n\bar{x}\), the middle term becomes \(-2\bar{x}(n\bar{x}) = -2n\bar{x}^2\), so \[ \sum_{i=1}^{n}(x_i-\bar{x})^2 = \sum_{i=1}^n x_i^2 - 2n\bar{x}^2 + n\bar{x}^2 = \sum_{i=1}^n x_i^2 - n\bar{x}^2. \] Dividing by \(n-1\) gives the shortcut formula \[ s^2 = \frac{1}{n-1}\left(\sum_{i=1}^{n} x_i^2 - n\bar{x}^2\right), \] which needs only \(\sum x_i^2\) and \(\bar{x}\) — one pass over the data instead of first computing every deviation \((x_i-\bar{x})\) and then squaring it. Both formulas give the identical result, \(s^2=18{,}793.33\), for “The Wandering Fork”; the shortcut form is simply faster to compute by hand or with SUMPRODUCT/SUMSQ in Excel.

Visual intuition

Variance is literally the average squared distance from the mean: picture each of the 30 days as a point on a number line at its revenue value, and \(\bar{x}=628.34\) as a fixed pivot. Variance answers “on average, how far (squared) do the points sit from the pivot?” — a wide, flat spread of points gives a large variance; a tight cluster around the pivot gives a small one.

Interactive demo — central tendency & dispersion calculator

Edit the six values below and watch the mean, variance, standard deviation, and coefficient of variation update live. Starting values are a 6-day sample of “The Wandering Fork”’s daily_revenue_eur.

#| standalone: true
#| components: [viewer]
#| viewerHeight: 560

from shiny import App, render, ui
import matplotlib
matplotlib.use('Agg')
import matplotlib.pyplot as plt
import numpy as np

BG = '#1C2E22'
FG = '#D2CCC0'

app_ui = ui.page_fluid(
    ui.tags.style(f"""
        body {{ background-color: {BG}; color: {FG}; padding: 12px; font-family: sans-serif; margin:0; }}
        .form-label {{ color: {FG} !important; }}
        .form-range {{ accent-color: steelblue; width: 100%; }}
        .info-box {{
            background: #152119; border-radius: 6px; padding: 10px 14px;
            margin: 8px 0; border-left: 3px solid steelblue;
            font-family: monospace; font-size: 0.9em;
        }}
    """),
    ui.h5("🧮 Central tendency & dispersion calculator", style=f"color:{FG}"),
    ui.p(
        "Move the sliders to change a 6-day revenue sample; the mean, "
        "variance, standard deviation, and CV recompute live.",
        style=f"color:{FG}; font-size:0.88em; margin-bottom:6px;",
    ),
    ui.input_slider("d1", "Day 1 (€):", min=300, max=950, value=460.48, step=1),
    ui.input_slider("d2", "Day 2 (€):", min=300, max=950, value=500.02, step=1),
    ui.input_slider("d3", "Day 3 (€):", min=300, max=950, value=418.65, step=1),
    ui.input_slider("d4", "Day 4 (€):", min=300, max=950, value=785.79, step=1),
    ui.input_slider("d5", "Day 5 (€):", min=300, max=950, value=719.66, step=1),
    ui.input_slider("d6", "Day 6 (€):", min=300, max=950, value=886.56, step=1),
    ui.output_ui("info"),
    ui.output_plot("plot", height="280px"),
)

def server(input, output, session):
    def sample():
        return np.array([input.d1(), input.d2(), input.d3(), input.d4(), input.d5(), input.d6()], dtype=float)

    @output
    @render.ui
    def info():
        x = sample()
        mean = x.mean()
        median = float(np.median(x))
        var = x.var(ddof=1)
        std = x.std(ddof=1)
        cv = std / mean * 100 if mean else 0
        return ui.HTML(f"""
            <div class='info-box'>
                Mean = <strong style='color:darkorange'>{mean:.2f}</strong> &nbsp;|&nbsp;
                Median = <strong style='color:darkorange'>{median:.2f}</strong><br>
                Variance (s²) = {var:.2f} &nbsp;|&nbsp;
                Std. dev. (s) = {std:.2f} &nbsp;|&nbsp;
                CV = {cv:.2f}%
            </div>
        """)

    @output
    @render.plot
    def plot():
        x = sample()
        mean = x.mean()

        fig, ax = plt.subplots(figsize=(7, 3.2))
        fig.patch.set_facecolor(BG)
        ax.set_facecolor(BG)
        for spine in ax.spines.values():
            spine.set_edgecolor(FG)
        ax.tick_params(colors=FG)

        days = np.arange(1, 7)
        ax.scatter(days, x, color='steelblue', zorder=5, s=70)
        ax.axhline(mean, color='darkorange', lw=1.8, linestyle='--', label=f'Mean = {mean:.2f}')
        for d, v in zip(days, x):
            ax.plot([d, d], [mean, v], color=FG, lw=1, alpha=0.5)
        ax.set_xticks(days)
        ax.set_xlabel("Day", color=FG)
        ax.set_ylabel("Revenue (€)", color=FG)
        ax.set_title("Deviations from the mean", color=FG)
        ax.legend(fontsize=8.5, facecolor=BG, edgecolor='none', labelcolor=FG)
        ax.grid(alpha=0.15, color=FG)
        plt.tight_layout()
        return fig

app = App(app_ui, server)

Exercises

A sample of \(n=5\) values has \(\sum x_i = 100\) and \(\sum x_i^2 = 2140\). Find the sample variance.

Solution. \(\bar{x}=100/5=20\). Shortcut formula: \(s^2 = \frac{1}{n-1}\left(\sum x_i^2 - n\bar{x}^2\right) = \frac{1}{4}(2140 - 5\times400) = \frac{1}{4}(2140-2000) = 35\).

“The Wandering Fork”’s items_sold has mean \(74.40\) and sample standard deviation \(16.22\) (from data-model.md). Which variable — daily_revenue_eur or items_sold — has higher relative variability?

Solution. \(CV_{\text{revenue}} = 137.09/628.34 = 21.82\%\); \(CV_{\text{items}} = 16.22/74.40 = 21.80\%\). The two are almost identical (about 21.8%) even though the raw units (€ vs. items sold) are completely different — exactly the kind of comparison only \(CV\), not \(s\) alone, can make fairly.

If one exceptionally bad day of \(\text{daily\_revenue\_eur}=150\) replaced the current minimum (407.59), how would the range change, and would the standard deviation be affected in the same direction?

Solution. New range \(= 886.56 - 150 = 736.56\) (up from 478.97) — the range, using only two values, reacts strongly to a single extreme point. The standard deviation would also increase (an extra-large deviation from the mean adds to \(\sum(x_i-\bar{x})^2\)), but by a much smaller relative amount, since it accounts for all 30 days rather than just the two most extreme.