1. Data, Information, and the Data Life Cycle
A practical introduction to how computers represent, organize, process, preserve, and responsibly use data throughout its life cycle.
From recorded data to meaningful
Computers work with recorded symbols, measurements, observations, and instructions. These records may represent numbers, words, images, sounds, locations, transactions, or events.
A useful distinction is:
Data are recorded facts or symbols, such as
72,Alex, or2026-09-15.is data interpreted in a meaningful context.
Knowledge is a conclusion or understanding developed from .
Data do not automatically become useful . Meaning depends on context, definitions, organization, quality, and interpretation. For example, the separate values 184, 207, and 231 become meaningful when a table identifies them as daily visits to Library A and supplies the dates and units.
Takeaway: Data are recorded inputs; is contextualized meaning; knowledge is understanding or a conclusion derived from that .
How computers represent data
Digital computers represent with binary digits, or bits. A bit has one of two conventional values, 0 or 1. Eight bits make one byte, which can represent one of possible patterns.
Binary is a base-two number system. For example, the bit pattern 01011010 represents:
The same bit pattern can mean different things depending on the encoding rules, , and . For example, 01000001 may represent the number , the character A under a particular character encoding, or part of an image.
A field’s determines what kind of value it contains and which operations are appropriate. A ZIP code such as 02139 is generally stored as a text string because its leading zero matters and arithmetic is not normally performed on it. Common types include integers, decimal numbers, text strings, Boolean values, dates and times, categorical values, binary objects, and geospatial data.
Takeaway: Binary patterns store data, but encoding rules and data types provide the context needed to interpret them.
Organization, , and context
Data organization affects how easily can be searched, validated, combined, and analyzed.
follow a defined schema, such as a relational table with a known column for each field.
contain labels or key-value pairs but allow records to have different fields; JSON, XML, and application logs are examples.
lack a regular tabular schema; essays, photographs, videos, scanned documents, and social-media posts are examples.
describe other data. They may identify a file’s creation date, size, format, creator, units, column meanings, collection location, processing procedures, license, or change history. also support provenance, which records where data came from and how they were changed.
Without , a value such as 12.4 may be impossible to interpret: it could represent meters, dollars, kilograms, or degrees. make data easier to find, understand, evaluate, combine, and reuse.
Takeaway: Organization makes data manageable, while supplies the meaning and history needed for responsible use.
Managing data across its life cycle
The describes how data move through interconnected stages rather than a permanently fixed sequence:
Plan: Define the purpose, required data, quality requirements, formats, storage, security, ethics, and retention period.
Generate or acquire: Collect observations, create data through computation, or obtain data from another source while recording relevant .
Store and protect: Use suitable files, databases, or repositories with access controls, backups, versioning, and integrity checks.
Clean and prepare: Check types, standardize formats, handle missing values, investigate duplicates and outliers, and document transformations.
Analyze and interpret: Apply calculations, queries, statistical methods, models, spreadsheets, or visualizations to answer questions.
Share, use, and reuse: Provide appropriate data, documentation, and to authorized users or the public, with licenses and limitations.
Preserve or dispose: Retain data that must remain available, migrate them when necessary, and securely delete data that no longer have a legitimate purpose.
Consider a school energy project. The school might plan to compare electricity use by building and month, acquire meter readings, store them in a controlled database, convert all readings to the same unit, identify impossible values, calculate monthly totals, and create line charts. It could then share a documented summary with facility managers, retain useful aggregated results, and delete unnecessary sensitive details.
improves consistency, but it requires judgment. Blank, unknown, not-applicable, and suppressed values should not automatically be treated as zero. Removing all incomplete records can also disproportionately exclude observations from a particular community. A responsible analyst asks whether missingness itself contains useful .
Takeaway: A life cycle connects purpose, collection, protection, preparation, analysis, sharing, and responsible retention or disposal.
Analysis, visualization, and efficient storage
Tables, formulas, and visual tools help transform rows of data into summaries and patterns. Useful operations include sorting, filtering, calculating totals and averages, grouping records with pivot tables, and comparing values across time or categories.
Choose a visualization that matches the question:
A line chart is generally suitable for change over time.
A bar chart is useful for comparing categories.
A scatter plot can help examine relationships between numerical variables.
A map can show geographic patterns.
Clear titles, labels, units, scales, and notes help prevent misleading interpretations. A trend or correlation does not by itself prove causation. Conclusions also depend on data quality, completeness, and representativeness.
reduces the number of bits needed to store or transmit data. preserves every original value, whereas permanently removes some to achieve a smaller representation. Lossless methods are appropriate for text, programs, spreadsheets, scientific measurements, and archival files; lossy methods are often used for images, audio, and video when some detail can be sacrificed.
Takeaway: Tools can reveal patterns efficiently, but interpretation must account for the question, the visualization, the data quality, and any lost through compression.
Responsible and ethical data practice
Data collection can improve medical research, transportation, public services, scientific discovery, and business operations. It can also create risks when is collected without meaningful choice, used for unexpected purposes, retained unnecessarily, exposed through security failures, or used to classify people unfairly.
Responsible practice asks:
Purpose: Why is the data being collected, and is that purpose clear?
Necessity: Is each field needed, or is more being collected than required?
Consent and control: Do people understand the collection and have meaningful choices?
Access and security: Who can view, change, share, or delete the data?
Accuracy: Can people correct incorrect records?
Fairness: Could the dataset or its use disadvantage particular groups?
Retention: How long should the data be kept?
Reidentification: Could supposedly anonymous records be combined with other data to identify individuals?
means collecting and retaining only what is needed for a clear, legitimate purpose. It should be paired with appropriate access controls, documentation, quality checks, retention limits, and evaluation of possible effects on individuals and communities. Privacy and ethical considerations belong at every stage of the , not only after a system has been built.
Takeaway: Good data practice balances usefulness with privacy, security, fairness, transparency, purpose limitation, and respect for affected people.