1. Statistical Questions and Data
A foundational guide to identifying statistical questions, defining observations and variables, classifying data, and matching data sources to appropriate conclusions.
1. Asking Statistical Questions
A statistical investigation begins by asking whether the situation calls for a study of variation.
A anticipates different answers across cases. Examples include:
How many hours do college students sleep on a typical weeknight?
What proportion of voters in a city support a proposed policy?
Do students who exercise regularly have different average stress scores from students who do not?
A has one definite answer for the particular case being considered. Examples include asking whether one person owns a bicycle or asking for the boiling point of water under a specified condition.
The same topic can produce either kind of question. “How tall is Maya?” usually concerns one person and is deterministic. “How tall are students in Maya’s class?” concerns a group in which different heights are expected, so it is statistical.
A useful way to rewrite a statistically is to identify the and ask about a distribution, proportion, typical value, or comparison. For example:
“How long does one student take to commute?” becomes “How long do students at this school typically take to commute?”
“Does one patient recover after treatment?” becomes “What proportion of patients recover after treatment?”
“Which team won the game?” becomes “How does the team’s win rate compare with its opponents’ win rates?”
Takeaway: A question is statistical when meaningful variability is expected in the observations needed to answer it.
2. Defining Observations and Variables
Once the question is clear, specify what each observation represents and what is measured.
An is the person, object, event, or other entity described by one observation. In a school survey, the units might be individual students. In other investigations, they might be households, hospitals, days, or manufactured products.
A is a characteristic recorded for each that can take different values. For a survey of 200 students about transportation, possible variables include:
transportation method;
commute time in minutes;
grade level; and
number of school days missed.
The school might be a constant if every student attends the same school. A is the larger group to which the investigation is intended to apply. For example, the 200 surveyed students may be a sample from all students enrolled at the school.
A well-organized data table usually has one row per and one column per . The entries in a row describe one unit, while the entries in a column show the values of one across units.
Takeaway: Before calculating or graphing, identify the units, the variables, and the of interest.
3. Classifying Data
The type of determines which summaries and graphs are appropriate.
A records labels or group membership. Examples include transportation method, eye color, political party identification, and whether a patient experienced a side effect. Some categorical variables are ordered, such as satisfaction levels from very dissatisfied to very satisfied, but the gaps between categories should not automatically be treated as equal.
A records numerical amounts for which ordering and differences are meaningful. Examples include age, commute time, number of absences, and body mass.
Quantitative variables can be described further:
Discrete variables are usually counts with separate, countable values, such as number of pets or clinic visits.
Continuous variables are measurements that can, in principle, take any value within an interval, such as temperature, time, height, or mass.
Classify a by what its values mean, not merely by whether digits appear. A student identification number and a ZIP code contain digits but function as labels, so they are categorical rather than quantitative. Similarly, “number of hours studied” is quantitative when respondents enter values such as 6.5, but it becomes categorical when the response choices are “none,” “a little,” “some,” or “a lot.”
Use this checklist:
Is the value a label or a measured amount?
If it is a label, do the categories have a meaningful order?
If it is numerical, do differences or averages make sense?
Is it a count or a measurement on a continuum?
Takeaway: Meaning and measurement determine a ’s type, not the visual appearance of its entries.
4. Matching Data Sources to Conclusions
Data sources affect what a study can reasonably conclude.
are collected directly for the current investigation. Common methods include surveys, interviews, direct measurements, observations, controlled experiments, and sensors. These data can be tailored to the research question, but collection may require substantial planning and resources. Wording, sampling, measurement procedures, and nonresponse can influence the results.
were collected previously and are reused for a new analysis. Examples include government records, school or hospital administrative records, public scientific datasets, and earlier surveys. can save time and money, but the analyst must check:
who or what was included;
what the variables actually measure;
when and where the data were collected;
whether missing values or exclusions are present; and
whether the data represent the in the new question.
A is the complete group to which a conclusion is intended to apply. A census attempts to collect information from every member of a defined . A sample survey collects information from only a subset. An observational study records variables without assigning treatments, whereas an experiment deliberately assigns conditions or treatments and measures outcomes.
Example: commute times
Suppose a city wants to know whether residents who use public transportation have longer commutes than residents who drive. A sound plan identifies:
Question: Do commute times differ between public-transit users and drivers in the city?
: City residents who regularly commute to work.
Observational units: Individual commuters.
Variables: Primary commute method, commute time, work schedule, and possibly distance from home to work.
Data source: A survey of residents or an existing transportation dataset.
Analysis: Compare the distributions or typical commute times for the two groups.
If the sample includes only people who use a city transit app, it may not represent all commuters. Also, if residents choose their own transportation, an observed difference is an association and is not automatically evidence that transportation method causes the difference.
Takeaway: Match the question, , units, variables, and data source before interpreting results.
5. Investigation Planning Checklist
A complete description of a statistical investigation should answer six questions:
What is being asked?
To whom or to what should the conclusion apply?
What entities provide the individual observations?
What variables are measured on each unit?
Were the data collected directly or obtained from an existing source?
What limitations might the source impose on the conclusion?
This framework prevents common errors, such as treating labels as numerical measurements, applying a result to a that was not represented, or interpreting an association as proof of causation. It also clarifies which summaries and graphs are appropriate.
The central progression is:
begin with a question that can be answered using data;
identify the and observational units;
define the variables and classify their types;
examine how the data were collected; and
limit the conclusion to what the design and source support.
Final takeaway: Good statistical reasoning starts before computation. Clear definitions and an appropriate data source are necessary for meaningful analysis.