4. Data Compression
A structured guide to how data compression exploits redundancy, how lossless and lossy methods differ, and how to choose an appropriate balance among size, quality, speed, and usability.
The purpose and central idea of data compression
Data compression represents information with fewer bits than an uncompressed representation. It can reduce storage requirements and transmission time, but a useful method must preserve the information required by users and applications.
The central idea is to identify patterns that do not need to be represented independently. A file may contain repeated symbols, predictable symbol frequencies, similar neighboring values, similar consecutive frames, or details that people are unlikely to notice. A compressor replaces these patterns with shorter codes, references, transformations, or controlled approximations.
Two broad strategies are possible:
removes representational while preserving every part of the original.
also removes information considered less important, allowing a greater reduction in size.
A compressed file is useful only when it remains suitable for its purpose. Smaller size may improve storage and delivery, while exact recovery, perceptual quality, compatibility, and processing speed may impose limits.
Takeaway: Compression is the efficient representation of information, based on and the acceptable level of change.
patterns and techniques
Different kinds of call for different techniques:
Repetition: The same value occurs many times. can represent
AAAAABBBas5A3B.Statistical : Some symbols are more frequent than others. can give common symbols shorter codes.
Spatial : Nearby pixels often have similar colors. Storing differences between neighboring pixels may require fewer bits than storing every color independently.
Temporal : Consecutive audio or video frames are often similar. A system can store one complete frame and record only changes in later frames.
Perceptual : Some visual or auditory details are difficult to notice, such as small color changes or sounds masked by louder sounds.
Compression frequently uses multiple stages. can replace a repeated sequence with a reference, prediction can produce a residual difference, and can then represent frequent residuals efficiently. A reversible filter may also reorganize data so that patterns become easier to compress without permanently changing the underlying information.
Takeaway: The best compression technique depends on the type of present in the data.
How preserves information
allows the original data to be reconstructed exactly. It exploits without discarding information, which makes it essential when even a small change could affect meaning or operation.
Common techniques include:
: Stores a value together with the number of consecutive repetitions.
: Replaces repeated strings with references to earlier entries or sequences.
: Assigns shorter codes to more frequent symbols.
Prediction and residual coding: Predicts a value from nearby values and stores the usually smaller difference between the prediction and the actual value.
is appropriate for source code, executable programs, legal and financial records, scientific measurements, spreadsheets, databases, text documents, and archival masters. It is also useful when a file will be edited extensively or compressed again.
The possible size reduction depends on the data. Already-compressed files, encrypted files, and random-looking data contain little detectable , so a lossless compressor may produce little reduction or may make the file slightly larger because of headers and other overhead.
Takeaway: Choose whenever exact reconstruction is required.
How trades fidelity for size
produces a representation that is sufficiently similar to the original for its intended use, but it does not preserve every detail. It commonly uses a model of human visual or auditory perception to decide which information can be reduced or removed.
A lossy image compressor may transform pixel data, apply , and encode the remaining values efficiently. At a high-quality setting, changes may be difficult to notice. At a lower-quality setting, artifacts such as blockiness, ringing, color banding, or blurred detail may appear.
is often suitable for photographs, streaming music and video, video conferencing, online games, thumbnails, and previews. These uses may value fast delivery and small files more than exact reconstruction.
Repeatedly opening and saving a lossy file can cause cumulative degradation because each generation may discard additional information. A high-quality or lossless master should therefore be retained when future editing or preservation matters.
Takeaway: trades exact fidelity for substantially smaller files, so its quality setting and intended use must be considered together.
Comparing lossless and lossy methods
The difference between the two approaches can be summarized as follows:
Reconstruction: restores an identical original; restores an approximation.
Information discarded: Lossless methods discard none; lossy methods discard some information.
Typical size reduction: Lossless reduction is moderate and data-dependent; lossy reduction is often much greater for media.
Quality control: preserves exact quality; allows quality to be exchanged for size.
Repeated editing: Lossless editing is safe for data integrity; repeated lossy editing may accumulate visible or audible degradation.
Common uses: Lossless methods suit documents, source code, archives, ZIP files, and PNG images; lossy methods suit photographs, streaming media, and previews.
Some formats support both modes. JPEG 2000, for example, defines both bit-preserving lossless methods and lossy methods for continuous-tone images. The correct choice depends on whether exact recovery or a controlled approximation is more important.
Takeaway: The lossless-versus-lossy decision is primarily a decision about whether any information may be discarded.
Choosing a method for a real task
A practical choice should balance quality, size, and usability rather than optimizing only one measure.
Determine whether exact recovery is required. If text, code, measurements, records, or archival data must remain unchanged, use .
Identify the data type. Text, numerical tables, images, audio, and video contain different forms of and respond differently to compression.
Set an acceptable quality level. For lossy media, compare several quality settings instead of assuming that the smallest file is best.
Measure performance as well as size. Compare the , time, decoding time, and memory use. The is calculated as:
Check compatibility and access needs. Confirm that the intended software can decode the format and determine whether random access is needed or sequential decompression is acceptable.
Consider future use. Repeated editing, searching, analysis, error recovery, legal requirements, and scientific preservation may favor a lossless master.
Keep an original when information may be valuable. Smaller delivery copies can be produced while preserving a high-quality or lossless version.
For example, a small JPEG image may be appropriate for rapid web delivery, while PNG may be preferable for a diagram with sharp text, flat colors, or transparency. A ZIP archive is generally better for a spreadsheet because it preserves every cell exactly.
Takeaway: Select a method by matching the compression tradeoff to the data, workflow, quality requirement, and receiving system.