By Dr. Maddie Swannack

Next Lesson - Models of Health

Population & Social Science


Contents

Abstract

  • There are four main types of observational study, two descriptive (ecological and cross-sectional) and two analytical (cohort and case control). Each of these study designs has flaws specific to that design (such as sampling bias in cross sectional studies) and flaws common to all study designs (such as lack of generalisability or risk of random error). 
  • Ecological studies split a population by characteristic and look at the number of cases within each population. 
  • Cross-sectional studies are a snapshot of the number of cases in a population at any time. It does not make comments on the causes of this disease.
  • Cohort studies split one population of people based on their exposure status (independent variable) and investigate their disease status (dependent variable). They can be prospective or retrospective. 
  • Case control studies identify cases (those with the disease) and controls (those without the disease) and look back to compare previous exposure between the groups. 
  • Confidence intervals give a range of values that are compatible with the data from the study. A 95% confidence interval means that, if the study were repeated many times, 95% of such intervals would contain the true value. Wide confidence intervals indicate that the estimate is less precise. The null hypothesis of a study is indicated either by the number 1 or the value zero (depending on the type of data being compared). If the number 1 (for relative measures) or 0 (for absolute measures) lies within the confidence interval, there is insufficient evidence to reject the null hypothesis, meaning no statistically significant association is shown.

Core

 

See our summary and glossary article which explains the terms used in this article. 

Types of Studies 

There are four main types of observational studies used in the study of populations. 

They are split into descriptive (where the outcome is a qualitative or wordy description of the findings) and analytical (where the outcome is a quantitative or numerical value representing the findings). 

The two types of descriptive study discussed in this article are ecological and cross-sectional studies, and the two analytical studies discussed are cohort and case-control studies. 

Throughout this article, disease is used as the focus in every study design, and asthma is used as every example. This is done to make the study designs easier to compare. However, these study designs could be used for a wide variety of epidemiological hypotheses, and their use is not limited to disease investigations.

A “Case” in this article refers to a study participant that has the disease. A “Control” is a study participant that does not have the disease/ has not been exposed to a factor but is included in the study for comparison. 

Any numbers in this article are purely hypothetical and do not represent the true results of any known studies. 

 

Common Flaws with Study Designs 

All study designs have common flaws, relevant to all types. These are grouped together here so that they don’t have to be explained four times, as they apply to all four study designs. 

Random error: It is important to remember that the most common error in all study design is random error, aka chance. It is impossible to completely remove this element, but many things can be done to help to reduce it and make its effect known, such as declaring confidence intervals in the results, increasing sample size or randomising which participants are in each of the study groups. Increasing sample size is a very useful tool to reduce random error as it dilutes the effects of chance. 

Lack of general applicability/ generalisability: Lack of generalisability applies to any study where a sample of the population is chosen rather than studying the whole population. If the designer has taken 1000 people from the city of London to do their study, they cannot necessarily say that those 1000 people represent the 8.2 million other people in London. This means that their study would have low generalisability. This could be improved by increasing the sample size (number of people in the study) or sample pool (taking a group from each borough) but this is expensive.

Confounding variables: Confounding variables are external variables that influence individuals within the study that the study organisers may not necessarily know about. Confounders can affect both the dependent and independent variable. This means that the confounding variable could cause an inappropriate association between the independent and dependent variable (a false result). A relationship could be found between two variables which does not exist in isolation but is a direct result of the confounding variable. An example of this would be: the amount that a loaded bow is pulled back and how far the arrow flies. In theory, the more force that is applied, the further the arrow would fly in one direction. However, the confounding variable in this case would be the wind on that day - if the wind blows in different directions the flight of the arrow will be affected. Once identified this confounder can be controlled for e.g. trying the experiment on a windless day. Only then can a true association between distance drawn back and distance be established.

 

Ecological Studies

Ecological studies are the easiest study design to understand and an example of a descriptive study. They involve collecting a population of people, separating that group by a characteristic (usually geographical), and then looking at the occurrence of the disease (number of cases) in each group.  As a result, they are fast to complete and cheap. 

For example, the population of a town could be the target population for this study, and the characteristics that the population is being separated by could be distance from a motorway. If the disease investigated was asthma, the results of the study would be a set number, for example, “There are 4000 cases of asthma in residents less than a mile from the motorway compared to 3000 cases of asthma in residents more than a mile from the motorway.”

There are a number of flaws with the ecological study design:

  • Confounding variables are common and difficult to expose.
  • Lack of generalisability is common.
  • Ecological fallacy - as data is gathered about the large groups or populations it can’t accurately be applied to individuals.

 

Cross Sectional Studies

Cross-sectional studies are a second type of descriptive study commonly used. To perform this study, the researcher must find the number of cases at a certain time within a group of the population. This ‘snapshot’ provides good information for prevalence but does not study exposure or any causal link. This study also is less helpful when studying rare diseases as it is likely that individuals will be missed in the sample size.

For example, “In 2015, there were 250,000 people with asthma in England.” 

This study is commonly carried out via door-to-door surveying and other similar methods.

This study design also has multiple new terms associated with it:

  • Prevalence - the number of people with the disease in the population divided by the total number of people in the population at a given time. 

There are some flaws with the cross sectional study design:

  • Sampling bias - this occurs when some members of the population are more likely to be selected to participate in the study. This could happen, for example, if the study takes place during the day. This would then exclude individuals who work during the day, and would therefore skew the results of the study because these people cannot be included. This would then mean that the study would not accurately represent the population.
  • Participant bias - this is when the participant acts in a certain way because they think that is how the researchers wants them to behave. For example, if a patient knows whether they have been given the drug or the placebo in a drug trial, they might exaggerate or play down their symptoms. This can be unintentional on the part of the participant.

Cross sectional studies can still be used to test hypotheses, including a null hypothesis, for example when comparing prevalence between groups. However, because exposure and outcome are measured at the same time, they cannot establish temporal sequence or causality. 

 

Cohort Studies

Cohort studies are a common type of analytical study. They involve classifying the healthy participants based on their exposure and then investigating how many of each type go on to develop disease. 

For example, if investigating the influence of mould in a home on the development of asthma, the participants would be classified into whether they had lived in a house with mould (the independent variable), and then the development of asthma would be the outcome of interest (the dependent variable). The study would then follow both classes or groups throughout their life to see if and when they develop asthma. This is useful as it can be used to demonstrate associations between exposures (mould) and conditions (asthma), and can support causal inference.

Cohort studies can come in two types:

  • Prospective cohort studies take participants who have just been exposed and investigate their future development of the disease. 
  • Retrospective cohort studies take participants who have been exposed in the past and investigate the development of the disease in the present. This is still classified as a cohort study because the participants are put in groups based on their exposure rather then based on their disease state. 

Cohort studies also use some specific language terms;

Person years - the total number of years that all individuals in the study are observed for. A person is no longer observed when they die or when they leave the study. Person years is simply a measure of time. If 100 people are observed for 2 years, that is 200 person years.

Incidence rate - The incidence rate of the exposed population is worked out by dividing the number of people who developed the disease in the exposed group by the total person years at risk. For example, if 100 exposed individuals each contribute an average of 5 years at risk, this gives 500 person years. If 10 people develop asthma, the incidence rate is 10 / 500, which is 0.02 cases per person year at risk. While incidence rates are based on person time rather than headcount, the sample size still affects the precision of the estimate.

Incidence rate ratio - this is the ratio between the incidence rate of the exposed and the incidence rate of the unexposed. This shows how much more likely it is for a participant to have developed the disease if they have been exposed.

Cohort studies also have many issues:

  • Loss to follow-up - outcomes may be unavailable when participants leave a prospective cohort or when records are incomplete in a retrospective cohort. Fewer observed outcomes reduce precision. Loss can also bias the exposure–outcome comparison if the reasons for missing outcomes are related to exposure and disease risk. Record the reasons for loss and assess their relationship to both exposure and outcome. Examples include:
    • Survivor bias - restricting analysis to participants who survive or remain available can select a healthier group. This may distort the exposure–outcome comparison, rather than merely make the sample less representative.
  • Differential loss - the frequency or reasons for missing outcomes may differ between exposure groups. If loss is also related to the outcome, the estimated association may be biased. Even loss that does not introduce bias reduces the information available. Observation bias - this error comes from a miscalculation error. If participants are asked to provide measurement or data there can be issues with incorrect reporting of results.

Cohort studies are usually expensive as they involve large sample sizes to mitigate loss to follow up and require the resources to follow up these large groups many years into the future. However cohort studies are very helpful for producing a large amount of data and allowing for prevalence and incidence to be calculated.

The following considerations should be made when designing a successful cohort study:

The acronym PICO can be used to remember what needs to be established when designing a cohort study.

 

Population - who will be sampled from?

Intervention - what is the exposure of interest

Comparison - who is being compared against, who is the control group from the sample population (those who lack exposure).

Outcome - what is the end condition being determined.

 

In our example

P - Children who live in x city.

I - the presence of mould in the family home.

C - children who live in the city in non mouldy houses.

O - development of asthma in 20 years.

 

Quiz

Preview the Epidemiological Studies quiz