2 Epidemiology Basics

Learn how epidemiology measures disease frequency, compares groups, uses study designs, and evaluates evidence about possible causes.

What measures

examines how health-related conditions are distributed across populations, what factors influence them, and how this knowledge can help prevent or control health problems. It studies patterns by person, place, and time, and compares groups to investigate possible causes.

A frequency measure needs a clearly defined population, outcome, and time period. Its numerator counts the people or events of interest; its denominator represents the population in which those events could occur.

and existing cases

describes how widespread a condition is by measuring existing cases, including both new and pre-existing cases. refers to a specified time; refers to a specified period.

Prevalence=existing casespopulation×100%\text{Prevalence} = \frac{\text{existing cases}}{\text{population}} \times 100\%

For example, if 4040 of 800800 people have asthma on a survey date, is 5%5\%. is useful for estimating the burden of ongoing conditions and planning services. It can be high because new cases occur frequently, because people live with the condition for a long time, or both.

and new cases

measures the occurrence of new disease among people initially at risk. An proportion, also called risk, is the proportion of an at-risk group that develops the condition over a specified period. It requires a defined period and a population whose members can be followed.

If 1212 of 200200 initially healthy people develop an illness during one year, the one-year risk is 6%6\%.

An divides new cases by the total time people were observed and at risk. It uses person-time, so participants can contribute different amounts of observation time. For example, 1212 new cases over 1,0001{,}000 person-years gives an of 1212 cases per 1,0001{,}000 person-years.

measures new disease, whereas measures how widespread a condition is.

Comparing disease frequency

Measures of association compare disease frequency between groups. Suppose 2020 of 200200 exposed people and 1010 of 200200 unexposed people develop a condition during the same period. The risk is 10%10\% in the exposed group and 5%5\% in the unexposed group.

  • Risk ratio: 10%÷5%=210\% \div 5\% = 2. The exposed group had twice the risk.

  • : 10%−5%=510\% - 5\% = 5 percentage points. There were five additional cases per 100100 people in the exposed group over the period.

Relative measures such as the risk ratio describe proportional differences; absolute measures such as the show the size of the difference in population terms. In studies, researchers select participants based on outcome status, so they generally use an odds ratio rather than directly calculating risk.

Common study designs

Study design determines what can be measured and how confidently results can be interpreted. Designs differ in how they select participants and measure exposure and outcome; none is automatically reliable without attention to how it was conducted.

  • : Measures exposure and health status at one point or over a short period. It can estimate and identify patterns, but the time order of exposure and outcome may be unclear because they are measured together.

  • : Groups people by exposure and follows them to compare outcomes. Data may be collected prospectively or from past records. studies can estimate and clarify whether exposure preceded outcome, but follow-up can be costly and differences between groups may distort comparisons.

  • : Selects people with an outcome as cases and people without it as controls, then compares prior exposures. It is efficient for rare outcomes or outcomes with long latency. It usually estimates an odds ratio and can be affected by participant selection or inaccurate recall of past exposure.

  • : Assigns an intervention by chance and compares outcomes between groups. Randomization helps balance other causes of the outcome, supporting causal conclusions. Trials may be impractical or unethical for some exposures, and their results may not apply to every population.

  • : Compares exposure and outcome data summarized for groups rather than individuals. It is useful for population-level patterns, but a group-level association may not hold for individuals; this problem is called the .

Interpreting epidemiologic evidence

An observed association means exposure and outcome vary together in the data; by itself, it does not prove that the exposure caused the outcome. Interpretation requires considering possible sources of error and alternative explanations.

  • is systematic error. It can arise from participant selection, loss to follow-up, or inaccurate measurement of exposure or outcome.

  • occurs when a third factor is related to both exposure and outcome, making their association appear stronger, weaker, or different from the true relationship. For example, age could affect both an exposure’s and disease risk.

  • Chance and precision: Estimates from small samples are often less precise. A shows a range of values compatible with the data under the statistical model; a wide interval signals greater uncertainty. alone does not establish practical importance or causation.

  • Time order and alternatives: Consider whether exposure preceded the outcome and whether other causes could explain the pattern. Randomization can help address , but observational evidence remains essential when trials are infeasible or unethical.

  • Applicability: Assess whether the participants, setting, and follow-up resemble the population and circumstances to which the findings will be applied.

A careful interpretation considers the study question, design, measurements, possible sources of error, the size and precision of the association, and consistency with other evidence.