01 Foundations of Statistics and Data Collection

A structured guide to defining statistical questions, collecting reliable data, choosing study designs, recognizing bias, and protecting research participants.

From Questions to Evidence

Statistics begins with a question about a group, process, or phenomenon. A well-designed study connects the research question to the of interest, the variables to be measured, the method of data collection, and the analysis plan.

A typical investigation follows this progression:

  1. Define the research question and target .

  2. Identify the variables that will be measured.

  3. Choose a or decide to conduct a census.

  4. Select an observational or experimental design.

  5. Collect data using consistent procedures.

  6. Evaluate , uncertainty, and ethical risks.

  7. Analyze and communicate results with appropriate limitations.

Good statistical conclusions depend on more than calculations. If the question is vague, the measurements are poorly defined, or the recruitment process excludes important groups, technically correct calculations may still produce misleading conclusions.

Takeaway: Statistical reasoning begins with sound planning before any data are analyzed.

Populations, Samples, and Sampling

A is the complete group relevant to a question, while a is the subset actually observed. The target is the group to which the researcher intends to generalize.

For a study of average daily screen time among university students:

  • The target is all students at the university.

  • The is the students who provide usable responses.

  • The parameter is the true average for all students.

  • The statistic is the average calculated from the respondents.

A census collects information from every member of the target . A is usually faster and less expensive, but the observed units may differ from the as a whole. The sampling frame is the list or procedure used to identify eligible units. If the frame omits important parts of the target , even a large may be unrepresentative.

In probability sampling, each unit has a known, nonzero probability of selection. Common methods include:

  • Simple random sampling: every possible of a given size has the same chance of selection.

  • Systematic sampling: units are selected at a regular interval after a random starting point.

  • Stratified sampling: the is divided into relevant subgroups, and units are sampled within each subgroup.

  • Cluster sampling: groups such as schools or neighborhoods are selected, and units within them are studied.

  • Multistage sampling: selection occurs through several stages.

Convenience samples and voluntary-response samples are easy to obtain but generally do not support reliable -wide conclusions because participation may be related to the measured variables.

Takeaway: Generalization depends on how well the represents the target , not simply on how large the is.

Variables and Measurement

A is a characteristic that can differ across observational units. Researchers should define how each will be measured before collecting data.

Variables can be classified in several ways:

  • A categorical places observations into groups or labels, such as eye color, insurance type, or treatment group.

  • A quantitative records numerical amounts for which arithmetic comparisons are meaningful, such as height, income, or number of errors.

  • A discrete quantitative has countable values, usually whole numbers, such as number of hospital visits.

  • A continuous quantitative can take values across an interval, such as weight, time, or temperature.

The measurement scale indicates what comparisons are meaningful:

  • Nominal: categories have no inherent order, such as blood type or brand.

  • Ordinal: categories are ordered, but the spacing between categories is unequal or unknown, such as satisfaction levels.

  • Interval: values have meaningful equal differences but no meaningful zero, such as temperature in Celsius.

  • Ratio: values have equal differences and a meaningful zero, so ratios are meaningful, such as mass, length, or elapsed time.

A numerical code does not automatically make a quantitative. Coding political parties as 1, 2, and 3 creates labels, not measurements that should be averaged.

Operationalization converts an abstract concept into an observable rule. For example, academic achievement could be measured by grade-point average, a standardized test score, or successful course completion. Different operational definitions can lead to different results.

Takeaway: The type and measurement scale of each determine which summaries and analyses are appropriate.

Study Designs and Causal Reasoning

A study design specifies how units are selected, how variables are measured, whether conditions are assigned, and how results will be compared.

A cross-sectional study measures variables at one point or during a short period. It can describe current conditions or prevalence, but it often cannot establish which event occurred first. A longitudinal study follows units over time and can examine change and temporal relationships.

A prospective study identifies participants before the outcome occurs and follows them forward. A retrospective study uses existing records or asks participants about past exposures or outcomes. Either timing can occur in observational or experimental research.

In an , researchers measure conditions as they naturally occur and do not assign the main exposure or treatment. Common forms include:

  • A cohort study follows groups formed according to an exposure or characteristic and compares later outcomes.

  • A case-control study compares units with an outcome, called cases, with units without it, called controls, and examines previous exposures.

  • A cross-sectional study measures exposure and outcome at approximately the same time.

An deliberately assigns an intervention, treatment, or exposure and measures its effect. Units receiving the intervention are commonly compared with a control group or another comparison condition.

In an , uses chance to place units into treatment groups. It helps balance measured and unmeasured characteristics across groups. This differs from random sampling, which is used to help make a representative of a .

Experiments may also use placebos, blinding, or a standard-of-care comparison. In a double-blind design, neither participants nor personnel assessing outcomes know the assigned treatment when practical.

Takeaway: Observational studies reveal naturally occurring relationships, while experiments can provide stronger evidence about causation when their design and implementation are sound.

, Error, and Confounding

is a systematic tendency for a study procedure or estimate to deviate from the truth. differs from ordinary random variation: a larger can reduce random but does not necessarily correct systematic problems.

Important forms of include:

  • Selection : recruitment or retention makes some parts of the target more likely to be included.

  • Undercoverage: important members of the are missing from the sampling frame.

  • Nonresponse : selected units do not participate, and nonparticipants differ from participants.

  • Volunteer : people who choose to participate differ systematically from those who do not.

  • Attrition : dropout differs across treatment groups or according to participants' outcomes.

  • Measurement : recorded values systematically differ from the intended measurements.

  • Recall : participants in different groups remember past events with different accuracy.

  • Interviewer : an interviewer's wording, behavior, or expectations influence responses or measurements.

A is related to both an explanatory and an outcome. For example, an association between carrying umbrellas and traffic accidents could be explained partly by rain: rain increases umbrella use and can also reduce visibility and road safety.

Researchers can address confounding through , restriction, matching, stratification, statistical adjustment, or careful interpretation. Standardized protocols, validated instruments, interviewer training, blinding, and pilot testing can reduce measurement problems.

is the natural variation among samples from the same . Nonsampling error includes undercoverage, nonresponse, inaccurate answers, coding mistakes, and processing errors. Increasing size generally reduces but does not automatically eliminate nonsampling error or .

Takeaway: Reliable results require attention to recruitment, measurement, follow-up, and possible alternative explanations—not just a large .

Ethics and an Integrated Example

Ethical research protects participants and preserves the integrity of evidence. Three central principles are:

  • Respect for persons: treat individuals as autonomous decision-makers and provide meaningful .

  • Beneficence: minimize possible harms and reasonably balance risks against potential benefits.

  • Justice: select participants fairly and avoid placing research burdens disproportionately on vulnerable groups.

should explain the study's purpose, procedures, risks, potential benefits, privacy protections, and voluntary nature of participation. Researchers should also:

  • protect privacy and maintain data confidentiality;

  • collect only data necessary for the research question;

  • secure identifiable records and control access;

  • provide additional safeguards for vulnerable groups;

  • disclose conflicts of interest and funding sources;

  • report methods, results, and limitations honestly;

  • avoid fabricating, falsifying, selectively omitting, or deceptively presenting results.

Institutional review boards evaluate whether risks are minimized, participant selection is equitable, consent is appropriate, and privacy and confidentiality are protected. Ethical responsibilities continue after data collection through responsible storage, sharing, analysis, and communication.

Applying the framework

Suppose a school wants to determine whether an online tutoring program improves algebra performance. A careful plan would:

  1. Define the as students enrolled in the school's algebra courses.

  2. Select a using a specified sampling plan.

  3. Measure tutoring participation, prior achievement, study time, and final exam score.

  4. Distinguish categorical, quantitative, and ordinal measurements.

  5. Compare an observational design with an experimental design.

  6. Use in the when appropriate to strengthen causal conclusions.

  7. Apply the same exam and scoring rules, track nonresponse and dropout, and blind scorers when possible.

  8. Protect student records and ensure that the comparison condition does not unfairly deny necessary educational support.

An observational comparison may be confounded because motivated students could be more likely to seek tutoring. A randomized can provide stronger evidence about causation, while the sampling plan determines how confidently the findings can be generalized.

Takeaway: Ethical safeguards and transparent limitations are essential parts of valid statistical research, not additions made after analysis.