Range & Interquartile Range (IQR)

When analyzing data, measures of central tendency like the mean or median tell us about the 'typical' value. However, they don't tell us how spread out or varied the data points are. Understanding variability is crucial because two datasets can have the same mean but vastly different distributions. For instance, two investment portfolios might have the same average return, but one could have much higher risk (variability) than the other.

This lesson explores two fundamental measures of spread: the Range and the Interquartile Range (IQR). We'll learn how to calculate them, interpret what they tell us about a dataset, and understand their respective strengths and weaknesses, especially concerning the presence of outliers.

Understanding the Range

The Range is the simplest measure of variability. It quantifies the total spread of a dataset by calculating the difference between the highest and lowest values. A larger range indicates greater variability, meaning the data points are more spread out, while a smaller range suggests data points are clustered closer together.

While easy to calculate and understand, the range provides a quick, initial sense of spread. It's particularly useful for small datasets or when you need a very quick estimate of the data's extent.

pythonCalculating the Range of Customer Ages
📐 Range Formula

Range = Maximum Value - Minimum Value

This formula applies to any numerical dataset.

A significant limitation of the range is its extreme sensitivity to outliers. Since it only considers the two most extreme values, a single unusually high or low data point can drastically inflate the range, making it unrepresentative of the typical spread of the majority of the data. This makes the range a non-robust statistic, meaning it's easily distorted by extreme values.

Consider a dataset of employee salaries where one executive earns significantly more than everyone else. The range would be very large, but it wouldn't accurately reflect the salary spread for the vast majority of employees.

Impact of Outliers on Range
Compares the range of two datasets: one without an outlier and one with a single extreme outlier. The bars represent the calculated range value for each dataset.
Loading chart...
Key Insight: A single outlier can dramatically increase the range, making it a less reliable measure of typical spread.
Check Your Understanding
Why is the Range considered a non-robust measure of variability?

Introducing Quartiles

To overcome the range's sensitivity to outliers, we often turn to measures that focus on the spread of the central portion of the data. This brings us to quartiles. Quartiles divide a dataset into four equal parts, each containing 25% of the data points, after the data has been ordered from smallest to largest.

There are three main quartiles: Q1 (First Quartile), Q2 (Second Quartile), and Q3 (Third Quartile). Q2 is equivalent to the median, which divides the data into two equal halves. Q1 is the median of the lower half of the data, and Q3 is the median of the upper half.

📌 What Quartiles Represent
  • Q1 (25th Percentile): 25% of the data falls below this value.
  • Q2 (50th Percentile / Median): 50% of the data falls below this value.
  • Q3 (75th Percentile): 75% of the data falls below this value.
pythonCalculating Quartiles for Customer Spending

The Interquartile Range (IQR)

The Interquartile Range (IQR) is a measure of statistical dispersion, or spread, that describes the middle 50% of values when ordered from lowest to highest. It is calculated as the difference between the third quartile (Q3) and the first quartile (Q1). The IQR is a more robust measure of spread than the range because it is not affected by extreme outliers.

By focusing on the central portion of the data, the IQR provides a clearer picture of the typical spread, making it particularly useful for skewed distributions or datasets containing outliers. It tells us how spread out the 'bulk' of the data is, ignoring the most extreme 25% on either end.

pythonCalculating IQR for Customer Spending
📐 IQR Formula

IQR = Q3 - Q1

Where Q3 is the 75th percentile and Q1 is the 25th percentile of the data.

Visualizing Quartiles and IQR with a Box Plot
A box plot illustrating the distribution of customer spending, showing the minimum, Q1, median (Q2), Q3, maximum, and an outlier. The box itself represents the IQR.
Loading chart...
Key Insight: The box in a box plot visually represents the IQR, showing the spread of the central 50% of the data, with whiskers extending to the min/max (excluding outliers) and individual points for outliers.

Interpreting the IQR is straightforward: a larger IQR indicates a wider spread in the middle 50% of the data, while a smaller IQR suggests that the central data points are more tightly clustered. Unlike the range, the IQR is not influenced by extreme values, making it a robust measure of spread. This robustness is why the IQR is often preferred in statistical analysis, especially when dealing with data that may contain outliers or is not symmetrically distributed.

For example, if we're looking at house prices, a few extremely expensive mansions won't skew the IQR as much as they would the range, giving a more realistic picture of typical house price variability.

Check Your Understanding
You are analyzing the salaries of employees in a company. The dataset includes the CEO's salary, which is significantly higher than everyone else's. Which measure of variability would best represent the typical spread of salaries for the majority of employees?

Range vs. IQR: When to Use Which

Comparing Range and Interquartile Range (IQR)
FeatureRangeInterquartile Range (IQR)
CalculationMax Value - Min ValueQ3 - Q1
Robustness to OutliersNot robust (highly sensitive)Robust (less sensitive)
Data CoverageAll data points (from min to max)Middle 50% of data
Ease of UnderstandingVery easyModerate
When to UseQuick, initial overview; small datasets; no expected outliers; symmetrical data.Detailed analysis; skewed data; presence of outliers; comparing distributions.
A summary of the key differences and appropriate use cases for Range and IQR.
Key Takeaways
  • The Range is the difference between the maximum and minimum values, providing the total spread of a dataset.

  • The Range is highly sensitive to outliers, meaning a single extreme value can significantly distort its representation of spread.

  • Quartiles (Q1, Q2, Q3) divide an ordered dataset into four equal parts, with Q2 being the median.

  • The Interquartile Range (IQR) is the difference between the third quartile (Q3) and the first quartile (Q1), representing the spread of the middle 50% of the data.

  • The IQR is a robust measure of variability because it is not affected by extreme outliers, making it ideal for skewed distributions or data with unusual values.

  • Use the Range for a quick, general idea of spread in small, outlier-free datasets. Use the IQR for a more reliable and representative measure of spread, especially in larger or outlier-prone datasets.

← All lessons in Descriptive Statistics

Ready to keep this from fading?

Bitelrn turns lessons like this into a full course — quizzes, a knowledge map, and spaced review.

Get started free