02 — Information and Representation
A progressive guide to how computers encode numbers, text, images, and sound as binary data, including the trade-offs and limits of digital representation.
Binary Patterns and Units
Computers store, process, and transmit information as patterns of bits. A can have the value or , but those values do not carry meaning by themselves. Meaning comes from an agreed representation that tells a system how to interpret the pattern.
Binary place values
The uses base . Each position represents a power of , beginning with at the rightmost position. For example:
A sequence of bits can represent up to different patterns. For instance, bits provide patterns. Electronic components commonly use two distinguishable states, such as high and low voltage or on and off, which makes binary representation practical.
Bytes and storage units
A normally contains bits. Its patterns can represent unsigned values from through . Storage units use either decimal or binary prefixes: bytes, while bytes.
Takeaway: Binary data consists of patterns; an agreed interpretation gives those patterns meaning.
Representing Numbers
The same bits can represent different kinds of numbers depending on the numerical format. The number of bits controls the available range and precision.
Whole numbers
An represents only zero and positive whole numbers. With bits, its range is:
Thus, an - ranges from to , and a - ranges from to .
Negative numbers require a signed representation. is a common method. With bits, it usually provides the range:
An - signed value therefore ranges from through . A pattern must be interpreted as signed or unsigned; the interpretation changes its meaning.
Fractional values
Numbers such as are often stored using . A floating-point value includes information comparable to a sign, significant digits, and an exponent or scale. This format supports a broad range, but it cannot represent every real number exactly. A stored value intended to be may be slightly above or below it, and repeated calculations can accumulate rounding differences.
Takeaway: More bits can expand range or precision, but every numerical format has defined limits.
Representing Text
Text becomes digital data when characters are assigned numerical codes and those codes are stored as bytes. A specifies this mapping and the way the resulting values are represented.
ASCII and
ASCII assigns codes to common English letters, digits, punctuation marks, and control characters. For example, the character A has the value , which is . ASCII is compact, but it does not contain enough characters for all writing systems.
is designed to represent characters from modern and historical writing systems, as well as symbols and punctuation. Each character is assigned a code point, such as U+0041 for A.
A code point is not itself a sequence. encodes code points as sequences of one to four bytes. ASCII characters use one in , while many other characters use multiple bytes. is therefore both broad enough for many writing systems and compatible with ASCII.
Four related ideas
A character is an abstract item of text, such as
é.A code point is the number assigned to that character in a character standard such as .
An encoding is the sequence used to store or transmit the code point.
A font or glyph is the visual design used to display the character.
If bytes are decoded using the wrong encoding, the bytes may remain unchanged while the displayed text becomes corrupted. This kind of garbled text is sometimes called mojibake.
Takeaway: Correct text handling requires agreement about characters, code points, encodings, and visual display.
Representing Images
A raster image represents a picture as a rectangular grid of pixels. Each pixel stores numerical information about brightness and color.
Resolution and pixels
describes the number of pixels used to represent an image. An image measuring pixels by pixels contains:
More pixels can preserve more fine detail, but they also require more storage.
Color values and storage
is the number of bits used for a pixel or color channel. A channel with bits can represent up to values. An - grayscale channel provides brightness levels. A - RGB image uses bits each for red, green, and blue, giving:
possible color combinations.
A simple uncompressed size estimate is:
For a -by- image with bits per pixel:
Actual files can differ because of compression, headers, metadata, or extra channels such as transparency.
Raster and vector graphics
Raster graphics store individual pixels, so enlarging them can make edges appear jagged or blurry. Vector graphics describe shapes with lines, curves, and filled regions. They can often be enlarged without the same loss of sharpness.
Takeaway: Image quality and size depend on pixel dimensions, , representation type, and compression.
Representing Sound
Physical sound is a continuous pattern of air-pressure changes over time. Digital systems commonly convert this waveform into a sequence of numerical samples using pulse-code modulation.
Sampling over time
The is the number of measurements taken per second, measured in hertz. A rate of records samples per second. A higher rate can represent faster changes and a wider range of frequencies when the recording and reconstruction system are designed appropriately.
If the is too low, high-frequency information can be misrepresented as lower-frequency information. This distortion is called aliasing.
Quantizing amplitude
Each sample must be stored as a number. The number of bits used for each sample is the sample , also called quantization resolution. With bits, a sample can have up to possible levels. More levels allow smaller amplitude differences to be represented.
Rounding a continuous measurement to one of these finite levels creates . Increasing the sample generally reduces this error, although it increases the amount of data.
For a mono signal sampled at with - samples, the uncompressed data rate is:
Stereo requires twice as much sample data because it stores two channels.
Takeaway: Digital sound depends on both how often the waveform is sampled and how precisely each sample's amplitude is recorded.
Limits and Trade-offs
Digital representation is a useful model, but a fixed collection of bits cannot capture unlimited range, precision, or detail. Understanding these constraints helps explain why data formats make trade-offs.
Range and precision
A fixed number of bits supports only a finite range. For example, an - unsigned value cannot represent or without a different representation or more storage. Discrete numerical levels also leave gaps between representable values, and floating-point calculations can accumulate rounding differences.
Sampling and quantization
Images sample space into pixels, while sound samples time into measurements. Detail that falls between samples may be lost. Increasing or can reduce this loss, but it also increases storage and processing requirements. Quantization introduces a difference between a continuous value and the selected digital level.
Compression
reduces file size while preserving the ability to recover the exact original data. Lossy compression achieves smaller files by discarding some information. It can be effective for images, audio, and video, but excessive or repeated lossy compression can create visible or audible artifacts.
Interpretation and compatibility
A sequence is useful only when the sender and receiver agree about its interpretation. Compatibility can depend on character encodings, number formats, order, image layouts, audio parameters, and file structure. The same pattern may therefore be valid data under one convention and incorrect or meaningless data under another.
Takeaway: Digital systems balance accuracy, range, storage, processing cost, and compatibility. Every representation is powerful within its intended rules and limited outside them.