What is the difference between selection bias and sampling bias?
Caroline Morton
September 30, 2026
I touched on selection bias in my explainer on bias in health data research. In this explainer, I want to separate it more carefully from sampling bias, as the two issues can appear similar even though they can cause different problems in your study.
As a brief reminder before we dive in:
- Selection bias occurs when the way people enter or remain in a study is connected to the exposure or outcome you are measuring which can distort the association.
- With sampling bias, the people in your dataset do not properly represent the wider population you want to describe.
The difference between these terms is important because selection bias can affect whether the result is correct within the study, while sampling bias mainly affects how far we can extrapolate that result to other groups.
What do selection and sampling bias look like in health research?
Let’s look at an example for selection bias first. Say you are trying to assess whether a painkiller is linked to stomach bleeding, so you compare patients admitted with a stomach bleed with controls from the same hospital’s orthopaedic ward. From the outside this seems like a reasonable control group. However we know that orthopaedic patients are more likely to use painkillers than the wider population. They might for example, be the most likely to have taken the painkiller recently or more likely to have taken it regularly over a long period for an injury that have sustained (which is why they are in hospital in the first place). The people in the control group are therefore not representative of the general population, introducing selection bias. It is not a fair comparison to use, and once this bias has been introduced it can be very difficult to correct later in the analysis.
Now, in another example, let’s say you want to know how common frailty is in all people aged over 75, so you use records from a handful of GP practices in a comfortable commuter town. The patients are real people and the measurements are taken perfectly, however, the issue is that those particular GP practices may not represent the wider over-75 population. These GP surgeries might not have any nursing home residents among their patients, or they might cater to a more affluent demographic that is healthier than the general population. This is sampling bias because while the study is internally sound, the estimate really describes a narrower group than the researcher intended. Working out which of the two you are dealing with is what the rest of this explainer is about.
Comparing selection and sampling biases
| Selection bias | Sampling bias | |
|---|---|---|
| What it is | The people included in the study are selected in a way that creates an unfair comparison between groups | The people in your sample aren’t a fair representation of the population you want to understand |
| Where it happens | When the study population is defined, including how controls are chosen and who is included or excluded | When the sample is selected, or when people choose to take part |
| What it affects | Whether the comparison between groups is valid | Whether the results can be generalised to the wider population |
| How you can spot it | Assess how each group was recruited, and whether that process is related to the exposure or outcome | Compare your sample against an outside source, e.g. census or registry data |
| What you can do about it | Aim to prevent it at the design stage, as it can be difficult to correct afterwards | Weight or standardise the sample |
As noted in the table, sampling bias can sometimes be managed in the analysis, as long as you know it’s there and have external information to compare or weight against, whereas selection bias is often much harder to correct once it has been introduced as it might require a complete redesign of your study. You can imagine in the example above with the orthopaedic patients as the control group, the best option might be to pick a different control group!
How can selection bias affect your results?
I’ll go through another example to show how selection bias can occur. Imagine you want to know whether heavy coffee drinking is associated with pancreatic cancer, so you decide to compare 300 patients with the disease and 300 controls. In order to illustrate how selection bias happens I’ve represented two different control options in the table below to show how it affects the results. The first set of controls are patients being seen by the same consultants for other digestive complaints, which makes them easy to recruit and superficially well matched, while the second set of controls were drawn from the general population. This example is based on a real case-control study published by MacMahon and colleagues in 1981, though the numbers used here are illustrative.
| Groups (total number) | Heavy coffee drinkers | Odds ratio vs cases |
|---|---|---|
| Pancreatic cancer cases (300) | 180 (60%) | |
| Digestive clinic patients (300) | 105 (35%) | 2.8 |
| General population (300) | 165 (55%) | 1.2 |
Among the patients with pancreatic cancer, 60% (180/300) were heavy coffee drinkers, whereas among the digestive-clinic controls (who do not have pancreatic cancer), 35% (105/300) were heavy coffee drinkers. If we don’t investigate further, we might reasonably assume that heavy coffee drinking is associated with pancreatic cancer, however, the second control group shows us why that interpretation may be misleading. At 55% prevalence we can see that heavy coffee drinking is nearly as common in the general population as it is among pancreatic cancer cases (60%).

The problem in this example is that the digestive-clinic controls were not a fair comparison group. Patients with gastrointestinal problems may already have reduced their coffee intake because of their symptoms or dietary advice, so coffee consumption can appear artificially low among these controls. In this example, the apparent association is created by the choice of control group. If the people selected as controls already have unusually high or low levels of the exposure you are studying, the comparison can give you a misleading result. Since the distortion comes from who was selected into the control group, adjusting the analysis for other variables does not necessarily remove it.
What is the impact of sampling bias?
Let’s look at an example to show how sampling bias can occur. One common way it arises is through non-participation, when the people who take part in a study are different from the people who don’t. Countries with extensive national health registers, such as Finland, can be particularly useful for studying these kinds of sampling bias because survey records can be linked to information about both participants and non-participants. McMinn and colleagues used this kind of linkage to examine whether people who took part in health surveys differed from those who did not.

The example below uses illustrative numbers to show how this can affect the results. In this example, people who smoke, have recently been hospitalised, or left education earlier are less likely to take part in the survey. As a result, the people who do participate look healthier or more advantaged than the population that was originally invited.
| Took part | Didn’t take part | Total patients invited | |
|---|---|---|---|
| Number of patients | 6,200 | 3,800 | 10,000 |
| Current smokers | 18% | 39% | 26% |
| Hospitalised in the past year | 9% | 24% | 15% |
| Left education at 16 | 22% | 41% | 29% |
The differences between participants and non-participants are very important because the survey estimate was based only on the people who took part. In this example, 18% of participants are current smokers compared with 39% of non-participants, meaning that if we only looked at the people who responded to the survey, we would estimate smoking prevalence at 18%, even though it is 26% across everyone invited. Estimates across hospitalisations and education level would similarly be biased.
This is sampling bias because the people who took part are not representative of the wider population the survey was intended to describe. While the measurements collected from participants can be completely accurate, the overall estimate can still be misleading if some groups are systematically more or less likely to participate. Increasing the sample size would not necessarily fix the problem either, because if the same kinds of people continue to be less likely to take part, you can end up with a larger and more precise study while still producing a biased estimate.
What should I take away from this?
A useful place to start is to determine what kind of bias you’re dealing with, which you can do by asking what exactly is wrong with the study population.
-
For selection bias, look at how people entered the study or how comparison groups were chosen. Ask whether that process could be related to the exposure, the outcome, or both. In a case-control study, for example, would the controls have had the same chance of being selected regardless of their exposure?
-
For sampling bias, compare the people in your dataset with the wider population you want to describe. Are some groups more or less likely to be included? Do the participants differ from non-participants on characteristics that matter to the question you are asking?
A study can have both biases, so it is worth checking for each separately as identifying one does not rule out the other.
Further reading
If you want to read more about some of the ideas in this explainer, I recommend starting with my explainer on bias in health data research, which looks at selection bias alongside several other forms of bias that can affect health data studies. I also have an explainer on confounding in health data research which covers how confounding can distort associations between exposures and outcomes.
My articles on electronic health records and codelists are also useful background for understanding how patients end up in a dataset and how study populations are defined. If you want to go deeper into the formal treatment of selection bias and causal inference, I also recommend Hernán and Robins’ Causal Inference: What If.
If you work with health data and have examples of selection or sampling bias in practice, it would be great to hear from you. I teach courses on epidemiology and health data analysis, and I am always looking for real-world examples to discuss with my students. If you want me to come and give a guest lecture or talk about your experiences, please get in touch via the contact page.
Enjoyed this? Subscribe to my newsletter.
I write about open science, research code, and building better tools for researchers.