✍ 02: Descriptive Statistics — Exercise¶
Tip
Practice — try each question first, then expand the answer to check your reasoning.
Part A uses the Friends dataset. Part B uses a small paired (x, y) sample to practice covariance and correlation by hand.
Friends dataset (Part A):
| Friend | Max temp | Weight | Height | Years | Gender | Company |
|---|---|---|---|---|---|---|
| Andrew | 25 | 77 | 175 | 10 | M | Good |
| Bernhard | 31 | 110 | 195 | 12 | M | Good |
| Carolina | 15 | 70 | 172 | 2 | F | Bad |
| Dennis | 20 | 85 | 180 | 16 | M | Good |
| Eve | 10 | 65 | 168 | 0 | F | Bad |
| Fred | 12 | 75 | 173 | 6 | M | Good |
| Gwyneth | 16 | 75 | 180 | 3 | F | Bad |
| Hayden | 26 | 63 | 165 | 2 | F | Bad |
| Irene | 15 | 55 | 158 | 5 | F | Bad |
| James | 21 | 66 | 163 | 14 | M | Good |
| Kevin | 30 | 95 | 190 | 1 | M | Bad |
| Lea | 13 | 72 | 172 | 11 | F | Good |
| Marcus | 8 | 83 | 185 | 3 | F | Bad |
| Nigel | 12 | 115 | 192 | 15 | M | Good |
📊 Q1. (Part A) Build the absolute, relative, absolute-cumulative, and relative-cumulative frequency table for Weight.¶
Show answer
n = 14. Every weight value in this sample is unique except 75 (appears twice). | Weight | Absolute | Relative | Abs. Cumulative | Rel. Cumulative | | --- | --- | --- | --- | --- | | 55 | 1 | 7.14% | 1 | 7.14% | | 63 | 1 | 7.14% | 2 | 14.29% | | 65 | 1 | 7.14% | 3 | 21.43% | | 66 | 1 | 7.14% | 4 | 28.57% | | 70 | 1 | 7.14% | 5 | 35.71% | | 72 | 1 | 7.14% | 6 | 42.86% | | 75 | 2 | 14.29% | 8 | 57.14% | | 77 | 1 | 7.14% | 9 | 64.29% | | 83 | 1 | 7.14% | 10 | 71.43% | | 85 | 1 | 7.14% | 11 | 78.57% | | 95 | 1 | 7.14% | 12 | 85.71% | | 110 | 1 | 7.14% | 13 | 92.86% | | 115 | 1 | 7.14% | 14 | 100% | Relative frequency = absolute / 14; cumulative columns are running totals of the columns to their left, in sorted order.📈 Q2. (Part A) Find the mode, median, 1st quartile (Q1), and 3rd quartile (Q3) for Years.¶
Show answer
Sorted **Years**: 0, 1, 2, 2, 3, 3, 5, 6, 10, 11, 12, 14, 15, 16 (n = 14) - **Mode** — values 2 and 3 each occur twice (every other value occurs once) → **bimodal: 2 and 3**. - **Median** — n is even, so median = average of the 7th and 8th sorted values = (5 + 6) / 2 = **5.5**. - **Q1** — median of the lower half {0, 1, 2, 2, 3, 3, 5} (7 values, odd) → the 4th value = **2**. - **Q3** — median of the upper half {6, 10, 11, 12, 14, 15, 16} (7 values, odd) → the 4th value = **12**.Part B dataset (paired samples):
| x | 2 | −1 | 0 | 1 | −2 | −3 |
|---|---|---|---|---|---|---|
| y | −1 | 1 | −2 | 0 | 1 | 2 |
📊 Q3. (Part B) Compute the covariance and Pearson correlation coefficient between x and y.¶
Show answer
n = 6, x̄ = −0.5, ȳ = 0.1667 | x−x̄ | 2.5 | −0.5 | 0.5 | 1.5 | −1.5 | −2.5 | | --- | --- | --- | --- | --- | --- | --- | | y−ȳ | −1.1667 | 0.8333 | −2.1667 | −0.1667 | 0.8333 | 1.8333 | Σ(x−x̄)(y−ȳ) = −10.5, Σ(x−x̄)² = 17.5, Σ(y−ȳ)² = 10.8333 - **Cov(x, y)** = −10.5 / (n−1) = −10.5 / 5 = **−2.1** - **Pearson r** = −10.5 / √(17.5 × 10.8333) = −10.5 / 13.769 ≈ **−0.76** A strong negative linear relationship — as x increases, y tends to decrease.📈 Q4. (Part B) Compute Spearman's rank correlation for the same data.¶
Show answer
Rank x (ascending, 1 = smallest): rx = 6, 3, 4, 5, 2, 1 (mean 3.5) Rank y (ties averaged): ry = 2, 4.5, 1, 3, 4.5, 6 (mean 3.5) Σ(rx−rx̄)(ry−rȳ) = −14, Σ(rx−rx̄)² = 17.5, Σ(ry−rȳ)² = 17 **Spearman ρ** = −14 / √(17.5 × 17) = −14 / 17.249 ≈ **−0.81** Spearman agrees in sign and magnitude with Pearson r here, confirming a consistent negative monotonic relationship.📚 All Exercises · Next: Chapter 03 — Multivariate Analysis ➡️