What's Inside this Chapter?
What’s Inside the Chapter? (After Subscription)
1. Mean, Median & Mode
1.1. Central Tendency and Mathematical Averages
1.2. Positional Averages and Partition Values
1.3. Mode and Distributional Relationships
1.4. Comparative Properties and Economic Suitability
2. Skewness
3. Quartile Deviation, Average Deviation and Standard Deviation
3.1. Quartile Deviation
3.2. Average Deviation
3.3. Standard Deviation
3.4. Relationship Between the Measures
4. Correlation
4.1. Conceptual Foundations and Classification
4.2. Karl Pearson’s Coefficient of Correlation
4.3. Spearman’s Rank Correlation and Concurrent Deviation
5. Probability Distribution Sampling
5.1. Introduction
5.2. Basic Terminology
5.3. Types of Sampling
5.4. Sampling Distribution
5.5. Other Important Sampling Distributions
5.6. Properties of a Good Estimator
5.7. Key Formulas
Access This Topic With Any Subscription Below:
- CUET PG Economics
- CUET PG Economics + Book Notes
Statistical Methods in Economics
CUET PG ECONOMICS
Mean, Median & Mode
Central Tendency and Mathematical Averages
Positional Averages and Partition Values
Mode and Distributional Relationships
Comparative Properties and Economic Suitability
| Criterion / Feature | Arithmetic Mean | Median | Mode |
| Type of Measure | Mathematical Average | Positional Average | Positional Average |
| Sensitivity to Outliers | High (Extreme values distort value) | Robust (Insensitive to extreme values) | Robust (Insensitive to extreme values) |
| Open-Ended Classes | Cannot be computed without assumptions | Can be readily computed | Can be readily computed |
| Algebraic Flexibility | High (Combined mean, regression, sampling) | None | None |
| Graphical Determination | Cannot be located graphically | Determined using Ogives | Determined using Histogram |
| Sampling Stability | Highly stable across samples | Moderately stable | Least stable |
| Primary Economic Applications | Per capita income, national accounts | Wealth distribution, poverty lines | Consumer preferences, market demand |
Skewness
Skewness is a statistical measure that describes the degree and direction of asymmetry in a frequency distribution around its central value. A distribution is symmetric when the observations are distributed equally on both sides of a central point. When the distribution is not symmetric, it is said to be skewed. Skewness therefore indicates whether the distribution has a longer or heavier tail toward the right or toward the left, and also indicates the extent of such asymmetry. In a perfectly symmetric distribution, the two sides of the distribution are mirror images of each other and the measure of skewness is zero. Skewness is a measure of the shape of a distribution, whereas measures such as mean, median and mode primarily describe its location.
A distribution may be positively skewed, negatively skewed, or approximately symmetric. In a positively skewed distribution, the tail extends farther toward the higher values. It is also called right-skewed distribution. In a negatively skewed distribution, the tail extends farther toward the lower values. It is also called left-skewed distribution. The terms positive and negative refer to the direction of asymmetry, not to whether the observations themselves are positive or negative.
In a symmetric distribution, the left and right portions of the distribution are approximately equal in shape. If the distribution is unimodal and perfectly symmetric, then:
$$
\text{Mean}=\text{Median}=\text{Mode}
$$
and the skewness is:
$$
Sk=0
$$
A normal distribution is the most important example of a perfectly symmetric distribution. In a symmetric distribution, the mean, median and mode coincide at the centre of the distribution. For a distribution that is only approximately symmetric, these three measures may be very close to one another even if they are not exactly equal.
In a positively skewed distribution, a relatively small number of observations with very high values pull the arithmetic mean toward the right-hand tail. The usual relationship is:
$$
\text{Mode}<\text{Median}<\text{Mean}
$$
Thus, the mean is generally greater than the median, and the median is generally greater than the mode. The right tail is longer than the left tail. Examples can include distributions of income, wealth, property prices, or waiting times where a small number of observations have exceptionally high values. The existence of a positive skew does not require every observation to be positive; positive skewness concerns the shape and direction of the tail.
In a negatively skewed distribution, a relatively small number of observations with very low values pull the arithmetic mean toward the left-hand tail. The usual relationship is:
$$
\text{Mean}<\text{Median}<\text{Mode}
$$
The left tail is longer than the right tail. Thus, the mean is generally smaller than the median, while the mode is generally larger than the median.
The basic distinction can be represented as follows:
| Distribution | Direction of longer tail | Usual relationship |
|---|---|---|
| Symmetric | Neither | Mean = Median = Mode |
| Positively skewed | Right | Mode < Median < Mean |
| Negatively skewed | Left | Mean < Median < Mode |
The mean, median and mode relationship is particularly important in identifying the direction of skewness. The mean is more sensitive to extreme observations than the median and mode. Therefore, extreme observations tend to pull the mean in the direction of the tail. This is the principal reason for the usual ordering of mean, median and mode in skewed distributions.
Skewness differs from dispersion. Dispersion measures the extent to which observations are spread around a central value, whereas skewness measures asymmetry. Standard deviation, variance and range are measures of dispersion. Skewness does not directly measure how widely observations are scattered; it measures whether the distribution is balanced or unbalanced around its centre.
Skewness also differs from kurtosis. Skewness concerns asymmetry, whereas kurtosis concerns the shape of the distribution in relation to the concentration of observations and the tails. A distribution can have zero skewness but still differ substantially from the normal distribution in its kurtosis.
Measures of Skewness:
Skewness can be measured through several statistical coefficients. The important measures include Karl Pearson’s coefficient of skewness, Bowley’s coefficient of skewness, Kelly’s coefficient of skewness, and moment coefficient of skewness. These measures are dimensionless coefficients and allow distributions to be compared even when they are measured in different units.
Karl Pearson’s Coefficient:
Karl Pearson’s coefficient of skewness is one of the most commonly used measures of skewness. When the mode is known, it is calculated as:
$$
Sk=\frac{\bar X-\text{Mode}}{\sigma}
$$
where:
-
\(\bar X\) = arithmetic mean
-
\(\text{Mode}\) = mode
-
\(\sigma\) = standard deviation
The numerator \(\bar X-\text{Mode}\) measures the difference between the mean and mode, while division by standard deviation converts the measure into a unit-free coefficient.
If:
$$
\bar X>\text{Mode}
$$
then:
$$
Sk>0
$$
and the distribution is positively skewed.
If:
$$
\bar X<\text{Mode}
$$
then:
$$
Sk<0
$$
and the distribution is negatively skewed.
If:
$$
\bar X=\text{Mode}
$$
then:
$$
Sk=0
$$
provided the distribution satisfies the relevant conditions.
In many practical distributions, the mode may not be clearly defined or may not be reliable. In such cases, Pearson’s second coefficient is used. It is:
$$
Sk=\frac{3(\bar X-\text{Median})}{\sigma}
$$
This formula is derived from the empirical relationship among mean, median and mode:
$$
\text{Mode}\approx3\text{Median}-2\bar X
$$
Rearranging:
$$
\bar X-\text{Mode}
\bar X-(3\text{Median}-2\bar X)
$$
Therefore:
$$
\bar X-\text{Mode}
3(\bar X-\text{Median})
$$
and hence:
$$
Sk=\frac{3(\bar X-\text{Median})}{\sigma}
$$
The Pearson coefficient is particularly useful for moderately skewed distributions. Its value is generally interpreted as positive, negative or zero according to the direction of asymmetry. The coefficient is not measured in the original units of the variable.
Bowley’s Coefficient:
Bowley’s coefficient of skewness is based on quartiles and the median. It is particularly useful when the distribution contains extreme observations because quartiles are less affected by extreme values than the arithmetic mean.
Bowley’s coefficient is:
$$
Sk_B=
\frac{Q_3+Q_1-2M}{Q_3-Q_1}
$$
where:
-
\(Q_1\) = first quartile
-
\(Q_2=M\) = median
-
\(Q_3\) = third quartile
The numerator:
$$
Q_3+Q_1-2M
$$
measures the asymmetry between the two halves of the distribution around the median.
The denominator:
$$
Q_3-Q_1
$$
is the interquartile range.
Therefore:
$$
Sk_B=
\frac{(Q_3-M)-(M-Q_1)}
{(Q_3-M)+(M-Q_1)}
$$
This form makes the interpretation particularly clear. If the distance from the median to \(Q_3\) is greater than the distance from \(Q_1\) to the median, the distribution is positively skewed. If the distance from \(Q_1\) to the median is greater, the distribution is negatively skewed.
If:
$$
Q_3-M>M-Q_1
$$
then:
$$
Sk_B>0
$$
and the distribution is positively skewed.
If:
$$
Q_3-M<M-Q_1
$$
then:
$$
Sk_B<0
$$
and the distribution is negatively skewed.
If:
$$
Q_3-M=M-Q_1
$$
then:
$$
Sk_B=0
$$
and the distribution is symmetric with respect to the quartiles and median.
An important property of Bowley’s coefficient is that it is less affected by extreme values because it uses \(Q_1\), median and \(Q_3\) rather than the mean and standard deviation. It is therefore useful for distributions containing open-ended class intervals, where calculation of the mean or standard deviation may be difficult or inappropriate.
Bowley’s coefficient is particularly useful for ordinal data or distributions in which the median and quartiles are more meaningful than the arithmetic mean. It does not utilise all observations directly.
The coefficient can also be expressed using the semi-interquartile range:
$$
Q.D.=\frac{Q_3-Q_1}{2}
$$
but Bowley’s coefficient itself is normally written as:
$$
Sk_B=
\frac{Q_3+Q_1-2Q_2}{Q_3-Q_1}
$$
since \(Q_2=M\).
Kelly’s Coefficient:
Kelly’s coefficient of skewness uses deciles or percentiles and therefore incorporates more information from the distribution than Bowley’s quartile-based measure.
Using deciles, Kelly’s coefficient is:
$$
Sk_K=
\frac{D_9+D_1-2D_5}{D_9-D_1}
$$
where:
-
\(D_1\) = first decile
-
\(D_5\) = fifth decile, which is the median
-
\(D_9\) = ninth decile
Since:
$$
D_5=M
$$
the formula can also be written as:
$$
Sk_K=
\frac{D_9+D_1-2M}{D_9-D_1}
$$
Using percentiles, it can be expressed as:
$$
Sk_K=
\frac{P_{90}+P_{10}-2P_{50}}
{P_{90}-P_{10}}
$$
where:
$$
P_{50}=M
$$
Kelly’s coefficient is useful when information about deciles or percentiles is available. Compared with Bowley’s coefficient, it uses a wider range of the distribution.
The principal difference is that Bowley’s measure concentrates on the central 50 percent of observations through \(Q_1\), \(M\), and \(Q_3\), whereas Kelly’s measure uses the central 80 percent through \(D_1\), \(D_5\), and \(D_9\).
