Bayes' Theorem Fundamentals

A friend calls you, worried. They just received a positive result on a screening test for a rare medical condition. The test boasts an impressive 99% accuracy rate: it correctly identifies the condition 99% of the time in those who have it, and correctly identifies no condition 99% of the time in those who don't. Based on these numbers, your friend is convinced they almost certainly have the condition.

Is your friend's 99% certainty justified, or is there a crucial piece of information missing from their calculation? Understanding how to update our beliefs with new evidence, especially when base rates are involved, is precisely what Bayes' Theorem helps us do.

The Trap of Intuition: P(A|B) vs. P(B|A)

Our intuition often leads us astray when dealing with conditional probabilities. In the scenario, the test's accuracy tells us the probability of a positive test given the disease, or P(Positive Test | Disease). However, what your friend truly wants to know is the probability of having the disease given a positive test result, or P(Disease | Positive Test).

These two probabilities, P(A|B) and P(B|A), are not interchangeable. Confusing them is a common logical fallacy, especially when one event (like a rare disease) has a very low prior probability (or base rate) in the general population.

Breakdown of Positive Test Results (Population of 100,000)
This pie chart illustrates the composition of all positive test results in a hypothetical population of 100,000, where the disease prevalence is 0.1% (1 in 1000) and the test accuracy (sensitivity and specificity) is 99%. It shows how many positive tests are true positives versus false positives.
Loading chart...
Key Insight: Despite the test's high accuracy, the vast majority of positive test results come from individuals who do NOT have the rare condition, due to the low prevalence of the disease.

The visualization reveals a critical insight: even with a 99% accurate test, if the disease is very rare (e.g., 0.1% prevalence), a positive test result doesn't mean a 99% chance of having the disease. In our example, out of 100,000 people, only 100 have the disease. The 99% accurate test correctly identifies 99 of them (True Positives).

However, among the 99,900 healthy individuals, 1% will incorrectly test positive (False Positives), which amounts to 999 people. So, for every 99 true positives, there are 999 false positives. This means that a positive test result is far more likely to be a false alarm than an indication of the actual disease, drastically lowering the actual probability of having the disease given a positive test.

Enter Bayes' Theorem: Updating Beliefs with Evidence

This is where Bayes' Theorem comes to our rescue. It provides a formal, mathematical framework for updating our beliefs about a hypothesis in light of new evidence. Instead of relying on intuition, which can be misleading, Bayes' Theorem allows us to systematically combine our initial belief (the prior probability) with the strength of the new evidence (the likelihood) to arrive at a revised, more accurate belief (the posterior probability).

Essentially, it helps us 'flip' conditional probabilities. If we know P(Evidence | Hypothesis), Bayes' Theorem enables us to calculate P(Hypothesis | Evidence), which is often the probability we are truly interested in.

📐 Bayes' Theorem Formula

P(H∣E)=P(E∣H)⋅P(H)P(E)P(H|E) = \frac{P(E|H) \cdot P(H)}{P(E)}

Posterior Probability, P(H∣E)P(H|E)
The probability of the hypothesis (H) being true given the evidence (E) has been observed. This is the updated belief after considering the evidence.
Example: In our scenario, P(Disease∣Positive Test)P(\text{Disease}|\text{Positive Test}) is the posterior probability – the actual chance your friend has the disease given their positive test result.
Likelihood, P(E∣H)P(E|H)
The probability of observing the evidence (E) given that the hypothesis (H) is true. This measures how well the evidence supports the hypothesis.
Example: The test's accuracy: P(Positive Test∣Disease)P(\text{Positive Test}|\text{Disease}) is the likelihood – the probability of testing positive if one truly has the disease.
Prior Probability, P(H)P(H)
The initial probability of the hypothesis (H) being true before any new evidence (E) is considered. It represents our pre-existing belief or the base rate.
Example: The prevalence of the disease in the general population, P(Disease)P(\text{Disease}), is the prior probability.
Marginal Likelihood (Evidence), P(E)P(E)
The overall probability of observing the evidence (E), regardless of whether the hypothesis is true or false. It acts as a normalizing constant.
Example: The overall probability of anyone testing positive, P(Positive Test)P(\text{Positive Test}), considering both those with and without the disease.

Solving the Dilemma: Applying Bayes' Theorem

Calculating the Probability of Disease Given a Positive Test
1
Identify Known Probabilities
Let D be the event of having the disease, and Pos be the event of a positive test result. We are given: - Prior Probability, P(D)P(D): Prevalence of the disease. Let's assume 0.1% or 0.001. - Likelihood, P(Pos∣D)P(\text{Pos}|D): Test sensitivity (correctly identifies disease). Given as 99% or 0.99. - Specificity, P(Neg∣not D)P(\text{Neg}|\text{not }D): Test correctly identifies no disease. Given as 99% or 0.99. From specificity, we can derive the probability of a false positive: - P(Pos∣not D)P(\text{Pos}|\text{not }D): Probability of testing positive when not having the disease. This is 1−P(Neg∣not D)=1−0.99=0.011 - P(\text{Neg}|\text{not }D) = 1 - 0.99 = 0.01. - P(not D)P(\text{not }D): Probability of not having the disease. This is 1−P(D)=1−0.001=0.9991 - P(D) = 1 - 0.001 = 0.999.
2
Calculate the Marginal Likelihood, P(Pos)P(\text{Pos})
The marginal likelihood is the overall probability of a positive test result, which can happen in two ways: a true positive (having the disease and testing positive) or a false positive (not having the disease and testing positive). P(Pos)=P(Pos∣D)⋅P(D)+P(Pos∣not D)⋅P(not D)P(\text{Pos}) = P(\text{Pos}|D) \cdot P(D) + P(\text{Pos}|\text{not }D) \cdot P(\text{not }D) P(Pos)=(0.99⋅0.001)+(0.01⋅0.999)P(\text{Pos}) = (0.99 \cdot 0.001) + (0.01 \cdot 0.999) P(Pos)=0.00099+0.00999P(\text{Pos}) = 0.00099 + 0.00999 P(Pos)=0.01098P(\text{Pos}) = 0.01098
3
Apply Bayes' Theorem
Now we can plug these values into Bayes' Theorem to find the posterior probability, P(D∣Pos)P(D|\text{Pos}): P(D∣Pos)=P(Pos∣D)⋅P(D)P(Pos)P(D|\text{Pos}) = \frac{P(\text{Pos}|D) \cdot P(D)}{P(\text{Pos})} P(D∣Pos)=0.99⋅0.0010.01098P(D|\text{Pos}) = \frac{0.99 \cdot 0.001}{0.01098} P(D∣Pos)=0.000990.01098P(D|\text{Pos}) = \frac{0.00099}{0.01098} P(D∣Pos)≈0.09016P(D|\text{Pos}) \approx 0.09016

After applying Bayes' Theorem, we find that the probability of your friend actually having the disease, given a positive test result, is approximately 0.09016, or about 9.02%. This is a stark contrast to the initial intuitive belief of 99%.

The low prior probability of the disease (0.1%) plays a dominant role here. Even a highly accurate test produces many false positives when the condition is rare, diluting the significance of a positive result. This calculation demonstrates how crucial it is to consider the base rate when interpreting diagnostic tests or any evidence.

Check Your Understanding
If the prevalence of the disease were much higher, say 1 in 100 people (1%), how would the posterior probability P(Disease∣Positive Test)P(\text{Disease}|\text{Positive Test}) change?

Beyond Medical Tests: Where Bayes' Theorem Shines

The power of Bayes' Theorem extends far beyond medical diagnostics. It is a cornerstone of modern statistical inference and machine learning, providing a principled way to update beliefs and make decisions under uncertainty across numerous domains.

For instance, spam filters use Bayesian principles to calculate the probability that an email is spam given the words it contains. Machine learning algorithms, such as Naive Bayes classifiers, leverage it for tasks like text categorization and sentiment analysis. In A/B testing, it helps determine the probability that one version of a product is better than another given observed user behavior. Even in legal reasoning, it can be used to update the probability of guilt given new evidence. Its versatility makes it an indispensable tool for anyone working with data and uncertainty.

Key Takeaways
  • Bayes' Theorem provides a formal method to update our beliefs (hypotheses) based on new evidence.

  • It allows us to calculate P(H∣E)P(H|E) (posterior) from P(E∣H)P(E|H) (likelihood), which are often confused.

  • The theorem's components are Posterior Probability, Likelihood, Prior Probability, and Marginal Likelihood.

  • The Prior Probability (or base rate) is crucial; neglecting it can lead to highly inaccurate conclusions, especially for rare events.

  • Bayes' Theorem is widely applied in fields like medical diagnostics, spam filtering, machine learning, and A/B testing.

  • It helps make more informed decisions by systematically combining initial knowledge with observed data.

  • In the diagnostic dilemma, your friend's 99% certainty was misplaced; Bayes' Theorem revealed a much lower actual probability of disease due to its rarity.

← All lessons in Probability Basics

Ready to keep this from fading?

Bitelrn turns lessons like this into a full course — quizzes, a knowledge map, and spaced review.

Get started free