Sample Size Impact

Even with a population of highly variable incomes, a sample of just 30 individuals can yield a surprisingly normal distribution of sample means. This phenomenon is a cornerstone of statistical inference, allowing us to draw reliable conclusions about large populations from smaller subsets. This lesson explores how and why increasing your sample size transforms the behavior of sample means, making them more predictable and precise.

Why Individual Samples Can Mislead

When you take a single, small sample from a population, its mean might not accurately reflect the true population mean. Imagine surveying only five people about their income; their average might be very different from the overall average income of the city. This variability means that if you were to take many such small samples, their means would likely be spread out, making it difficult to pinpoint the true population parameter with confidence.

Sampling Distribution: Small Samples (n=5)

Sampling Distribution of Mean Income (n=5)
This chart shows the distribution of sample means when taking many small samples of size 5 from a highly right-skewed population (like income). The distribution of these sample means still exhibits some right-skewness and is relatively wide, indicating high variability.
Loading chart...
Key Insight: With small samples (n=5), the sampling distribution of the mean remains somewhat skewed and has a wide spread, reflecting the parent population's characteristics.

Looking at the distribution for n=5n=5, we observe a shape that still clearly resembles the original skewed population. The peak is shifted to the left, and a long tail extends to the right. This wide spread signifies that individual sample means can vary considerably from one sample to the next. Such high variability makes it challenging to use any single sample mean as a reliable estimate for the true population mean.

Sampling Distribution: Moderate Samples (n=30)

Sampling Distribution of Mean Income (n=30)
This chart illustrates the distribution of sample means for samples of size 30, drawn from the same skewed population. The distribution is now much more bell-shaped and significantly narrower than for n=5, demonstrating the Central Limit Theorem's effect.
Loading chart...
Key Insight: As sample size increases to n=30, the sampling distribution of the mean becomes approximately normal and much narrower, indicating reduced variability.

With a moderate sample size of n=30n=30, a significant transformation occurs. The sampling distribution of the mean now appears much more bell-shaped and symmetric, closely approximating a normal distribution. Crucially, the spread of the distribution has noticeably tightened around the population mean. This reduction in variability means that individual sample means are now much closer to the true population mean, making them more reliable estimators.

Sampling Distribution: Large Samples (n=100)

Sampling Distribution of Mean Income (n=100)
This chart shows the distribution of sample means for large samples of size 100, from the same skewed population. The curve is now very narrow and clearly normal, demonstrating further reduction in spread and increased precision.
Loading chart...
Key Insight: For large samples (n=100), the sampling distribution of the mean is highly concentrated around the population mean and nearly perfectly normal, providing very precise estimates.

When the sample size increases further to n=100n=100, the sampling distribution becomes even narrower and more concentrated around the population mean. It is now almost perfectly normal. This extreme reduction in spread means that nearly all sample means will be very close to the true population mean. Such large samples provide highly reliable and precise estimates, minimizing the impact of random sampling variability.

The Standard Error: Quantifying Sampling Variability

To formally quantify the spread of the sampling distribution of the mean, we use the standard error (SE). The standard error is essentially the standard deviation of the sample means. A smaller standard error indicates that the sample means are clustered more tightly around the population mean, implying greater precision in our estimates. It directly measures how much sample means are expected to vary from the true population mean due to random sampling.

📐 Standard Error Formula

SE=σnSE = \frac{\sigma}{\sqrt{n}}

Where:
- SESE is the standard error of the mean
- σ\sigma is the population standard deviation
- nn is the sample size

The formula for standard error reveals a critical relationship: it is inversely proportional to the square root of the sample size, n\sqrt{n}. This means that to halve the standard error (and thus double the precision), you need to quadruple the sample size. While increasing nn always reduces the standard error, the returns diminish. For example, going from n=100n=100 to n=400n=400 reduces the standard error by the same factor as going from n=1n=1 to n=4n=4.

Check Your Understanding
If you double your sample size, how does the standard error of the mean change?

The 'Large Enough' Heuristic (n ≥ 30)

A widely used heuristic in statistics is that a sample size of n30n \ge 30 is generally considered 'large enough' for the Central Limit Theorem to apply. This means that even if the parent population is not normally distributed, the sampling distribution of its means will be approximately normal once the sample size reaches 30 or more. This guideline is a practical rule of thumb, but it's important to remember it's a heuristic; for extremely skewed populations, a larger sample size might be needed to achieve a truly normal sampling distribution.

Key Takeaways
  • Increasing sample size significantly reduces the spread of the sampling distribution of the mean.

  • Larger samples cause the sampling distribution to become more bell-shaped and approximate a normal distribution, regardless of the parent population's shape.

  • The standard error quantifies this spread, measuring the typical deviation of sample means from the population mean.

  • The standard error decreases proportionally to 1/n1/\sqrt{n}, meaning precision improves with the square root of the sample size.

  • The n30n \ge 30 heuristic provides a practical guideline for when the Central Limit Theorem's promise of normality generally holds, making estimates from moderate samples reliable.

← All lessons in Central Limit Theorem

Ready to keep this from fading?

Bitelrn turns lessons like this into a full course — quizzes, a knowledge map, and spaced review.

Get started free