Glossary

Data Description

Every mathematical symbol and named formula introduced across Modules 1–10 is listed below, with its notation, a plain-language definition, a pronunciation guide (for Greek letters only), and the module where it is first introduced.

Conventions used throughout

Question Convention Excel
Quartiles and percentiles Position \((n+1)\times p\) with interpolation QUARTILE.EXC, PERCENTILE.EXC
Variance and standard deviation Sample formulas (\(n-1\)) by default VAR.S, STDEV.S
Skewness Adjusted Fisher–Pearson coefficient SKEW
Kurtosis Conceptual only: no formula in this course β€”

French names for each function are on the Excel functions page.

Percentages and totals (Module 1–2)

Notation Name Plain-language definition Pronunciation (Greek only) Introduced in
\(p_i\) Percentage distribution Share of a total that item \(i\) represents, \(x_i/T\times100\%\) β€” Module 1
\(\Delta\%\) Percentage change Relative change between an old and a new value β€” Module 1
\(T\) Grand total Sum of all values in a range or all group totals β€” Module 1

Univariate position (Modules 2–5)

Notation Name Plain-language definition Pronunciation (Greek only) Introduced in
\(n\) Sample size Number of observations/individuals β€” Module 2
\(k\) Number of modalities Number of distinct categories/classes of a variable β€” Module 2
\(n_i\) Absolute frequency Count of observations in modality/class \(i\) β€” Module 2
\(f_i\) Relative frequency Share of observations in modality/class \(i\), \(n_i/n\) β€” Module 2
\(h\) Class width Width of a class interval in a grouped frequency table β€” Module 3
\([L_i, L_i+h[\) Half-open class Class that includes its lower bound \(L_i\) but not its upper bound, so each value belongs to exactly one class β€” Module 3
\(k\) (classes) Number of classes Number of intervals in a grouped table; Sturges’ rule suggests \(k \approx 1 + 3.3\log_{10}n\) β€” Module 3
\(m_i\) Class midpoint Centre of class \(i\), \((L_i + L_i + h)/2\); where the frequency polygon plots the class β€” Module 3
\(n_i/h_i\) Density Count divided by class width; the bar height of a histogram when class widths are unequal β€” Module 3
\(N_i^{+}\) Increasing cumulative frequency (count) Running total of counts up to and including class \(i\) β€” Module 3
\(F_i^{+}\) Increasing cumulative frequency (share) Running total of relative frequencies up to class \(i\), \(N_i^{+}/n\) β€” Module 3
Ogive Cumulative frequency curve Line joining (class upper bound, \(F_i^{+}\)); read it to find shares and percentiles β€” Module 3
\(Me\) Median The value splitting an ordered series into two equal halves β€” Module 5
\(Q_1, Q_2, Q_3\) Quartiles Values splitting an ordered series into four equal quarters (\(Q_2=Me\)) β€” Module 5
\(P_p\) Percentile Value such that approximately \(p\%\) of observations are \(\leq\) it β€” Module 5
\(IQR\) Interquartile range \(Q_3-Q_1\); spread of the central 50% of the data β€” Module 5
\((n+1)p\) Quantile position Rank of the \(p\)-th quantile in the ordered data, interpolated between neighbours (QUARTILE.EXC) β€” Module 5
\(Q_1-1.5\,IQR\), \(Q_3+1.5\,IQR\) Outlier fences Values beyond these limits are flagged as outliers on a boxplot β€” Module 5

Univariate center, dispersion, and shape (Modules 4–7)

Notation Name Plain-language definition Pronunciation (Greek only) Introduced in
\(\bar{x}\) Sample mean Arithmetic average of all observations, \(\frac{1}{n}\sum x_i\) β€œx-bar” Module 4
\(Mo\) Mode Most frequent value, or the center of the most frequent class for grouped data β€” Module 4
\(\bar{x}_w\) Weighted mean \(\sum w_i x_i / \sum w_i\); mean where each value counts according to its weight β€œx-bar w” Module 4
\(w_i\) Weight Importance given to value \(x_i\) in a weighted mean (e.g. amount invested, credits) β€” Module 4
\(L, f_0, f_1, f_2\) Grouped-mode inputs Modal class lower bound, and the frequencies of the preceding/modal/following classes β€” Module 4
\(s^2\) Sample variance Average squared distance of observations from the mean, \(\frac{1}{n-1}\sum(x_i-\bar{x})^2\) β€” Module 6
\(s\) Sample standard deviation \(\sqrt{s^2}\); dispersion in the original unit β€” Module 6
\(\sigma^2\), \(\sigma\) Population variance and standard deviation Same as \(s^2\), \(s\) but dividing by \(n\); used only when the data are the whole population (VAR.P, STDEV.P) β€œsigma” Module 6
\(CV\) Coefficient of variation Unit-free relative dispersion, \(s/\bar{x}\times100\%\) β€” Module 6
\(SKEW\) / \(g_1\) Fisher-Pearson sample skewness Bias-corrected, unit-free measure of asymmetry around the mean (Excel’s SKEW) β€” Module 7
Pearson asymmetry coefficient Pearson skewness coefficient \((\bar{x}-Mo)/s\); compares the mean to the mode β€” Module 7
\(3(\bar{x}-Me)/s\) Pearson median skewness Compares the mean to the median; usable when there is no clear mode β€” Module 7
\(YK\) Yule-Kendall skewness coefficient \((Q_3+Q_1-2Me)/(Q_3-Q_1)\); asymmetry from the five-number summary β€” Module 7

Bivariate (Modules 8–9)

Notation Name Plain-language definition Pronunciation (Greek only) Introduced in
\(n_{ij}\) Contingency table cell count Count of individuals with modality \(i\) of one variable and modality \(j\) of the other β€” Module 8
\(n_{i\cdot}\), \(n_{\cdot j}\) Row / column margin Row or column total of a contingency table; recovers each variable’s own frequency table β€” Module 8
\(\text{Cov}(X,Y)\) Sample covariance Average product of deviations from each variable’s mean, \(\frac{1}{n-1}\sum(x_i-\bar{x})(y_i-\bar{y})\) β€” Module 9
\(r\) Correlation coefficient Unit-free strength/direction of the linear relationship between \(X\) and \(Y\), \(-1\le r\le1\) β€” Module 9
\(r^2\) Coefficient of determination Share of \(Y\)’s variability explained by a linear link with \(X\) β€” Module 9
\(a, b\) Regression intercept and slope Coefficients of the least-squares line \(\hat{y}=a+bx\) β€” Module 9
\(\hat{y}\) Predicted / fitted value The value of \(y\) predicted by the regression line for a given \(x\) β€” Module 9

Synthesis (Module 10)

Notation Name Plain-language definition Pronunciation (Greek only) Introduced in
\(n_j\) Group size Number of observations in group \(j\) (e.g. one franchise or outlet) β€” Module 10
\(\bar{x}_j\) Group mean Mean of the observations in group \(j\) β€œx-bar j” Module 10
\(\bar{x}=\sum_j \frac{n_j}{n}\bar{x}_j\) Overall mean from group means The overall mean is the weighted mean of the group means, with weights \(n_j/n\) β€” Module 10

Synthesis and review (Module 10)

Notation Name Plain-language definition Pronunciation (Greek only) Introduced in
\(\bar{x}_w\) Weighted mean Mean of group means weighted by group sizes, \(\sum w_j \bar{x}_j\) with \(w_j = n_j/n\) β€” Module 10
Mock exam Original assessment Newly authored MCQ + computational questions mapping to M1–M9 β€” Module 10