Variance & Standard Deviation
When analyzing data, the mean tells you the central tendency, but it doesn't reveal how spread out the individual data points are. Two datasets can have the exact same mean yet look completely different: one might have all values clustered tightly, while the other has values scattered widely. To truly understand a dataset, you need to quantify this spread, or variability.
Quantifying Spread: Deviation from the Mean
The most intuitive way to measure how spread out data is to look at how far each data point deviates from the average. This difference, (for a population) or (for a sample), is called the deviation from the mean. A large positive deviation means the data point is much higher than the average, while a large negative deviation means it's much lower.
A key property of the mean is that the sum of all deviations from the mean is always zero. This is why simply averaging the deviations won't work as a measure of spread; positive and negative values will cancel each other out, always resulting in zero.
The Concept of Variance
To prevent positive and negative deviations from canceling out, we square each deviation before summing them. Squaring ensures all values are non-negative, and it also gives greater weight to larger deviations, emphasizing outliers. The average of these squared deviations is called the variance. Variance provides a measure of how much the data points are spread out from the mean, but its units are squared (e.g., dollars squared), making direct interpretation difficult.
For a population of data points with mean , the variance (denoted , sigma squared) is:
Sample Variance: The Correction
When you calculate variance from a sample of data (rather than an entire population), you typically divide the sum of squared deviations by instead of . This is known as Bessel's correction. The reason for this adjustment is that a sample mean () is usually a slightly biased estimator of the true population mean (). Using in the denominator provides an unbiased estimate of the population variance, meaning it's more likely to be closer to the true population variance on average.
For a sample of data points with sample mean , the variance (denoted ) is:
Introducing Standard Deviation
Variance, while mathematically useful, is often hard to interpret because its units are squared. To bring the measure of spread back to the original units of the data, we take the square root of the variance. This value is called the standard deviation. Standard deviation is the most widely used measure of variability because it's directly comparable to the data points themselves and the mean, making it much easier to interpret in practical terms.
For a population:
For a sample:
Notice that standard deviation is simply the square root of the corresponding variance formula.
Interpreting Standard Deviation
Standard deviation tells you the typical distance a data point is from the mean. A small standard deviation indicates that data points are clustered closely around the mean, implying low variability. A large standard deviation suggests that data points are spread out over a wider range, indicating high variability. For many real-world datasets that approximate a normal distribution, about 68% of the data falls within one standard deviation of the mean, and about 95% falls within two standard deviations.
temperatures dataset from the previous example, calculate the population standard deviation. Then, add a new temperature reading, 15, to the dataset and recalculate both the population mean and population standard deviation. Observe how the new outlier affects these measures.Variance vs. Standard Deviation
| Feature | Variance ( or ) | Standard Deviation ( or ) |
|---|---|---|
| Definition | Average of the squared differences from the mean. | Square root of the variance. |
| Units | Squared units of the original data (e.g., , ). | Same units as the original data (e.g., , ). |
| Interpretability | Less intuitive due to squared units; primarily used for mathematical properties. | Highly intuitive; represents typical distance from the mean. |
| Sensitivity to Outliers | Highly sensitive, as deviations are squared. | Highly sensitive, as it's derived from variance. |
| Use Cases | Statistical inference, ANOVA, regression analysis (as a component). | Descriptive statistics, quality control, comparing variability across datasets. |
Impact of Outliers
Because both variance and standard deviation involve squaring the deviations from the mean, they are particularly sensitive to outliers. A single data point far from the mean will have a very large squared deviation, disproportionately increasing the overall variance and standard deviation. This sensitivity means that while these measures are powerful, it's crucial to inspect your data for extreme values that might skew your interpretation of spread. In such cases, alternative measures like the Interquartile Range (IQR) might offer a more robust view of typical variability.
Variance ( or ) is the average of the squared deviations from the mean, providing a measure of data spread in squared units.
Standard Deviation ( or ) is the square root of variance, returning the measure of spread to the original units of the data, making it more interpretable.
For population variance, divide by ; for sample variance, divide by (Bessel's correction) to get an unbiased estimate.
Standard deviation quantifies the typical distance a data point is from the mean; a larger value indicates greater data dispersion.
Both measures are highly sensitive to outliers due to the squaring of deviations, which can disproportionately inflate their values.
Use standard deviation for direct interpretation of spread in real-world units, and variance for mathematical properties in statistical models.
Try it yourself
Distribution Explorer
Drag the parameters and watch the shape, mean, and spread respond. Sweep the plot to read off the probability of landing at or below a point.
Open the lab