3. Digital Data Formats and Storage
A progressive guide to how digital information is represented as binary patterns, organized into files, encoded as text, images, and audio, compressed, and selected for technical and social purposes.
1. Bits, Bytes, and Binary Capacity
Digital information is stored as patterns of bits, each of which has one of two values, commonly or . A contains eight bits. The same pattern can represent a number, character, color, or instruction depending on the rules applied to it.
Binary values and capacity
uses base two. Each position represents a power of two, so the sequence means in decimal.
Storage quantities require attention to prefixes. A kilobyte, written kB, is bytes, while a kibibyte, written KiB, is bytes. Similarly, GB uses a decimal definition of bytes, whereas GiB uses a binary definition of bytes. Confusing these units can lead to incorrect estimates of file or drive capacity.
A file becomes larger when it contains more values or represents them with greater precision. More pixels, samples, color levels, channels, or metadata generally require more storage.
Takeaway: Bits and bytes provide the physical representation, but interpretation rules give stored patterns their meaning.
2. File Structure and Interpretation
A digital file commonly includes several organized components:
A header identifies the format and records information needed for interpretation.
Metadata describes properties such as dimensions, creator, date, language, sample rate, or color space.
The payload contains the actual text, pixels, audio samples, or other content.
An index or directory can help software locate pieces of a large or compressed file.
Checksums or other integrity data can help detect accidental corruption.
A specifies how these parts are arranged and interpreted. A packages one or more streams with metadata, while a encodes or decodes particular content. For example, WAVE can act as a for linear PCM audio; the structure and the audio encoding are therefore distinct.
File extensions such as .txt, .png, .wav, and .mp3 provide useful clues, but software should also inspect format signatures and internal headers. Renaming an extension does not convert the underlying data.
Takeaway: Reliable interpretation depends on internal structure and documented rules, not merely on a filename extension.
3. and
Text is stored by assigning numbers to characters and then representing those numbers as binary. A defines both the relationship between characters and numeric code points and the way those code points become bytes.
ASCII represents a limited set of characters. provides a shared system for characters used by languages and technical writing around the world. Common encoding forms include:
UTF-8 uses one to four 8- code units per character and preserves ASCII characters' familiar one- values.
UTF-16 uses one or two 16- code units per character.
UTF-32 uses one 32- code unit per character.
The letter A has code point U+0041 and is represented by the single UTF-8 41 in hexadecimal. A character such as é requires multiple UTF-8 bytes. Thus, one character and one are not interchangeable concepts.
If a program reads UTF-8 bytes as though they used another encoding, the result may be garbled text called mojibake. Correct text processing therefore requires knowing which encoding was used.
Takeaway: Text meaning depends on a shared character system and the correct conversion between code points and bytes.
4. Digital Images and Color Information
A is a rectangular grid of pixels. Each pixel stores a numerical color value, and the amount of data depends on several factors:
Resolution is the width and height of the pixel grid.
Color channels describe components such as red, green, and blue.
determines how many levels each channel or pixel can represent.
An alpha channel can store transparency information.
Compression can reduce the number of stored bits.
For an uncompressed RGB image with width , height , and 8 bits per channel, an approximate size is
A RGB image therefore requires about bytes before headers, metadata, and compression. An 8- channel has possible levels. Three 8- RGB channels can represent up to , or about million, color combinations.
Image formats make different tradeoffs:
JPEG is usually lossy and is effective for photographs, but repeated saving can introduce artifacts.
PNG is lossless and useful for graphics, text, transparency, and sharp edges.
TIFF can store high-quality images and extensive metadata.
SVG stores geometric descriptions rather than a fixed pixel grid, allowing shapes to scale without becoming pixelated.
Higher native resolution and greater retain more information but increase storage and processing requirements.
Takeaway: Image size and quality are shaped jointly by pixel count, channel information, , and compression.
5. Digital Audio: Sampling and Quantization
Digital audio converts continuous sound into a sequence of numerical measurements. Two major properties determine its resolution:
A is the number of samples recorded per second, measured in hertz.
is the number of bits used for each sample and determines the number of possible amplitude levels.
For uncompressed PCM audio, the approximate data rate is
Stereo audio at samples per second and bits per sample requires
This is approximately kilobits per second before headers and other overhead. One minute at this rate requires roughly MB of uncompressed audio data.
A higher can represent higher-frequency components, while a higher provides finer amplitude resolution. Both increase storage, bandwidth, and processing requirements.
Common choices include WAVE/PCM, which is typically uncompressed and faithful; FLAC, which uses ; and MP3 or AAC, which use to create smaller files.
Takeaway: Audio quality and size depend mainly on how frequently sound is sampled and how precisely each sample is represented.
6. Compression, Size, and
preserves all original information. After decompression, the result is identical to the input. It works by representing repeated or predictable patterns more efficiently, but its compression ratio depends on the data. Highly random or already compressed data may shrink very little.
removes information judged less noticeable or less important to human perception. It can create much smaller files, but decompression cannot reconstruct the exact original. JPEG, MP3, and AAC are common examples.
The central tradeoff is . Increasing generally reduces quality and may cause image blocking, ringing, blurring, or audio artifacts. Repeatedly opening and resaving a lossy file can compound the damage.
The appropriate choice depends on purpose:
An archival master prioritizes maximum and preservation information.
An editing source should maintain high quality and minimize generation loss.
A web image may prioritize small size with acceptable visual quality.
Streaming audio may prioritize a low data rate and efficient decoding.
Scientific or financial data generally requires exact values and lossless storage.
Compression can also shift the cost from storage to processing: a smaller file may require more computation to encode, decode, open, or edit.
Takeaway: Choose lossless methods when exact recovery matters; choose lossy methods only when reduced size justifies discarded information.
7. Choosing Formats and Interpreting Data
Format selection should follow the data's purpose rather than only its extension or apparent file size. Evaluate the following questions:
: Must every original value be preserved?
Size: Is storage or network bandwidth limited?
Processing: Can the target device decode and edit the format efficiently?
Interoperability: Is the format supported by multiple programs and platforms?
Metadata: Can it preserve descriptions, timestamps, color information, or technical settings?
Sustainability: Is the specification documented and likely to remain usable?
Security and privacy: Could the file contain hidden metadata or sensitive information?
A practical workflow often keeps a high-quality master and creates smaller derivative files for distribution. For example, an organization might preserve an uncompressed or losslessly compressed image and generate JPEG copies for a website.
Software extracts meaning by applying format rules. A spreadsheet interprets bytes as rows, columns, dates, and numbers. A media player reads information, selects a , reconstructs samples, and sends them to a display or speaker.
For large datasets, analysts should verify the intended , identify missing or duplicated values, use consistent units and date formats, filter irrelevant records, document transformations and assumptions, and distinguish genuine patterns from artifacts of compression, sampling, or collection methods.
Digital storage also has social consequences. Searchable records can improve research and services, but they can enable surveillance, expose sensitive information, or reproduce bias when data is incomplete or collected unfairly. Responsible use requires purpose limitation, appropriate security, transparency, retention limits, and attention to who may be harmed.
Takeaway: A technically suitable format balances preservation, access, efficiency, compatibility, metadata, security, and responsible use.