Free Practice Quiz Question List

5. Cleaning and Preparing Datasets Online Quiz Questions

Use this free practice quiz with 20 questions to review 5. Cleaning and Preparing Datasets, test your knowledge, and prepare for your next test or exam.

20 questions
01
Choose one
1 point

How does a comma-separated values (CSV) file represent tabular data?

  1. A

    As a single undifferentiated text block

  2. B

    As rows and fields separated by delimiters

  3. C

    As one image per observation

  4. D

    As a list containing only field names

02
Choose one
1 point

What is the most appropriate first step in a reliable dataset-cleaning workflow?

  1. A

    Delete the source file after importing it

  2. B

    Replace all missing values immediately

  3. C

    Keep an unchanged copy of the source file

  4. D

    Remove unusual values before profiling

03
True or false
1 point

True or false: Converting every text value to lowercase is always an appropriate standardization step.

  1. A

    True

  2. B

    False

04
Written response
1 point

What is the standard term for documentation that defines a dataset's field meanings, expected types, units, permitted values, and relationships?

05
Fill in the blank
1 point

In the dataset type system, a score such as 82 should have the type, while a field containing only true or false should have the type.

06
Choose all
1 point

Which checks are valid examples of data-validation rules? Select all that apply.

  1. A

    Completeness checks for required fields

  2. B

    Uniqueness checks for fields that should be unique

  3. C

    A requirement that every value have the same visual color

  4. D

    Relationship checks between related fields

07
Choose one
1 point

A dataset has missing income values, and income differs substantially between relevant groups. If the original values cannot be recovered, which action is most defensible?

  1. A

    Replace every missing income with zero

  2. B

    Use a justified median from a relevant group and document that it is an estimate

  3. C

    Delete all records containing a missing income without assessing the effect

  4. D

    Treat a blank cell as proof that the income was zero

08
True or false
1 point

True or false: Every repeated identifier in a dataset proves that one of the corresponding records should be deleted.

  1. A

    True

  2. B

    False

09
Written response
1 point

What database term describes an identifier that has a unique value for each valid record?

10
Fill in the blank
1 point

A validation test that checks whether a return date comes after a purchase date is a check.

11
Choose all
1 point

Which practices support ethical dataset preparation? Select all that apply.

  1. A

    Retain only data necessary for the stated purpose

  2. B

    Keep all direct identifiers permanently, regardless of need

  3. C

    Assess whether a cleaning rule disproportionately removes a group

  4. D

    Document transformations, assumptions, exclusions, and remaining limitations

12
Choose one
1 point

In the worked school-data example, what preparation rule would make the attendance values consistently represented?

  1. A

    Convert every attendance value to a letter grade

  2. B

    Leave the percentage and decimal formats mixed

  3. C

    Standardize attendance to a proportion between 0 and 1

  4. D

    Replace all attendance values with the overall mean

13
Open ended
1 point

Design a reproducible plan for preparing a dataset for analysis. Explain how you would preserve and understand the original data, handle missing and inconsistent values, investigate duplicates and outliers, validate the result, document changes, and assess ethical consequences.

14
Choose one
1 point

In a tabular dataset, what is a field?

  1. A

    A complete case or observation in the dataset

  2. B

    One characteristic such as score or program

  3. C

    A rule used to validate a value

  4. D

    A unique identifier used as a primary key

15
Choose one
1 point

A team is beginning to clean a dataset and wants its work to remain auditable. Which action best follows the recommended workflow?

  1. A

    Delete the original after creating a cleaned file

  2. B

    Keep an unchanged copy of the source file

  3. C

    Replace all unusual values with the median

  4. D

    Remove every record containing a blank

16
Choose one
1 point

A measurement is missing, and the original record cannot be recovered. Which response is most appropriate when any plausible replacement would be highly speculative?

  1. A

    Replace it with zero because the field has no recorded value

  2. B

    Replace it with the overall average in every case

  3. C

    Leave it missing when guessing would introduce more error

  4. D

    Delete the entire dataset so that no missing values remain

17
Choose one
1 point

A dataset contains two rows with the same customer identifier, but the unit of observation is a purchase. What should the analyst do before removing either row?

  1. A

    Remove every row sharing an identifier immediately

  2. B

    Merge all rows from the same customer into one record

  3. C

    Investigate the records and decide using the dataset's unit and rules

  4. D

    Keep all repeated rows without checking their contents

18
True or false
1 point

True or false: If direct identifiers have been removed, a dataset can no longer create a privacy risk.

  1. A

    True

  2. B

    False

19
Written response
1 point

In the worked school dataset, attendance is recorded as 95%. If attendance must be standardized to a proportion between 0 and 1, what numeric value should be entered?

20
Written response
1 point

What term names the documentation that records a dataset's field meanings, expected types, units, permitted values, and relationships?