How should a dataset’s rows and columns be organized?
Each row represents one observation, and each column represents a variable recorded about that observation.
Study 1 Descriptive Statistics and Data Visualization with 12 free online flashcards. Review key terms, definitions, and concepts with this interactive flashcard deck.
How should a dataset’s rows and columns be organized?
Each row represents one observation, and each column represents a variable recorded about that observation.
What is a line chart suited to showing?
A line chart shows change over time. Put time in chronological order on the horizontal axis.
Why should missing values not be treated as zero?
A missing value means the value was not recorded; zero is an observed value. Keep them distinct when summarizing data.
How do categorical and quantitative variables differ?
Categorical variables describe groups or labels, while quantitative variables represent numerical values. Counts and proportions suit categories; means and standard deviations require quantitative values.
What does relative frequency represent?
Relative frequency is a category’s count divided by the total number of observations; it can be reported as a proportion or percentage.
When is a bar chart appropriate, and how are its bars arranged?
A bar chart compares counts or percentages across distinct categories. Its bars are separated.
What does a histogram show, and why do its bars touch?
A histogram displays the distribution of a quantitative variable by grouping values into intervals. Its bars touch because the intervals form a continuous numerical scale.
What can a scatterplot reveal, and what can it not establish by itself?
A scatterplot displays the relationship between two quantitative variables. A visible association alone does not show that one variable caused the other.
What features does a boxplot summarize?
A boxplot compactly summarizes a quantitative distribution using its median, quartiles, and spread. Side-by-side boxplots help compare groups.
Which four features help describe a quantitative distribution?
Describe shape, center, spread, and unusual observations. Together, these features explain the distribution more fully than a single summary.
How is the mean calculated, and how can an extreme value affect it?
The mean is the sum of the values divided by the number of observations. Because it uses every value, unusually high or low observations can pull it toward the tail.
How is the median found, and why can it suit skewed data?
The median is the middle value after sorting the data; with an even number of observations, it is the average of the two middle values. It is less affected by extremes than the mean.