02 — Information and Representation

A progressive guide to how computers encode numbers, text, images, and sound as binary data, including the trade-offs and limits of digital representation.

Binary Patterns and Units

Computers store, process, and transmit information as patterns of bits. A can have the value 00 or 11, but those values do not carry meaning by themselves. Meaning comes from an agreed representation that tells a system how to interpret the pattern.

Binary place values

The uses base 22. Each position represents a power of 22, beginning with 202^0 at the rightmost position. For example:

10112=1×23+0×22+1×21+1×20=11101011_2 = 1\times 2^3 + 0\times 2^2 + 1\times 2^1 + 1\times 2^0 = 11_{10}

A sequence of nn bits can represent up to 2n2^n different patterns. For instance, 88 bits provide 28=2562^8 = 256 patterns. Electronic components commonly use two distinguishable states, such as high and low voltage or on and off, which makes binary representation practical.

Bytes and storage units

A normally contains 88 bits. Its 256256 patterns can represent unsigned values from 00 through 255255. Storage units use either decimal or binary prefixes: 1 kB=1,0001\ \text{kB} = 1{,}000 bytes, while 1 KiB=1,024=2101\ \text{KiB} = 1{,}024 = 2^{10} bytes.

Takeaway: Binary data consists of patterns; an agreed interpretation gives those patterns meaning.

Representing Numbers

The same bits can represent different kinds of numbers depending on the numerical format. The number of bits controls the available range and precision.

Whole numbers

An represents only zero and positive whole numbers. With nn bits, its range is:

0 through 2n−10\ \text{through}\ 2^n - 1

Thus, an 88- ranges from 00 to 255255, and a 1616- ranges from 00 to 65,53565{,}535.

Negative numbers require a signed representation. is a common method. With nn bits, it usually provides the range:

−2n−1 through 2n−1−1-2^{n-1}\ \text{through}\ 2^{n-1}-1

An 88- signed value therefore ranges from −128-128 through 127127. A pattern must be interpreted as signed or unsigned; the interpretation changes its meaning.

Fractional values

Numbers such as 3.143.14 are often stored using . A floating-point value includes information comparable to a sign, significant digits, and an exponent or scale. This format supports a broad range, but it cannot represent every real number exactly. A stored value intended to be 0.10.1 may be slightly above or below it, and repeated calculations can accumulate rounding differences.

Takeaway: More bits can expand range or precision, but every numerical format has defined limits.

Representing Text

Text becomes digital data when characters are assigned numerical codes and those codes are stored as bytes. A specifies this mapping and the way the resulting values are represented.

ASCII and

ASCII assigns codes to common English letters, digits, punctuation marks, and control characters. For example, the character A has the value 651065_{10}, which is 01000001201000001_2. ASCII is compact, but it does not contain enough characters for all writing systems.

is designed to represent characters from modern and historical writing systems, as well as symbols and punctuation. Each character is assigned a code point, such as U+0041 for A.

A code point is not itself a sequence. encodes code points as sequences of one to four bytes. ASCII characters use one in , while many other characters use multiple bytes. is therefore both broad enough for many writing systems and compatible with ASCII.

Four related ideas

  • A character is an abstract item of text, such as é.

  • A code point is the number assigned to that character in a character standard such as .

  • An encoding is the sequence used to store or transmit the code point.

  • A font or glyph is the visual design used to display the character.

If bytes are decoded using the wrong encoding, the bytes may remain unchanged while the displayed text becomes corrupted. This kind of garbled text is sometimes called mojibake.

Takeaway: Correct text handling requires agreement about characters, code points, encodings, and visual display.

Representing Images

A raster image represents a picture as a rectangular grid of pixels. Each pixel stores numerical information about brightness and color.

Resolution and pixels

describes the number of pixels used to represent an image. An image measuring 1,9201{,}920 pixels by 1,0801{,}080 pixels contains:

1,920×1,080=2,073,600 pixels1{,}920 \times 1{,}080 = 2{,}073{,}600\ \text{pixels}

More pixels can preserve more fine detail, but they also require more storage.

Color values and storage

is the number of bits used for a pixel or color channel. A channel with nn bits can represent up to 2n2^n values. An 88- grayscale channel provides 256256 brightness levels. A 2424- RGB image uses 88 bits each for red, green, and blue, giving:

256×256×256=16,777,216256 \times 256 \times 256 = 16{,}777{,}216

possible color combinations.

A simple uncompressed size estimate is:

width×height×bits per pixel÷8\text{width} \times \text{height} \times \text{bits per pixel} \div 8

For a 1,0001{,}000-by-1,0001{,}000 image with 2424 bits per pixel:

1,000×1,000×24÷8=3,000,000 bytes1{,}000 \times 1{,}000 \times 24 \div 8 = 3{,}000{,}000\ \text{bytes}

Actual files can differ because of compression, headers, metadata, or extra channels such as transparency.

Raster and vector graphics

Raster graphics store individual pixels, so enlarging them can make edges appear jagged or blurry. Vector graphics describe shapes with lines, curves, and filled regions. They can often be enlarged without the same loss of sharpness.

Takeaway: Image quality and size depend on pixel dimensions, , representation type, and compression.

Representing Sound

Physical sound is a continuous pattern of air-pressure changes over time. Digital systems commonly convert this waveform into a sequence of numerical samples using pulse-code modulation.

Sampling over time

The is the number of measurements taken per second, measured in hertz. A rate of 44,100 Hz44{,}100\ \text{Hz} records 44,10044{,}100 samples per second. A higher rate can represent faster changes and a wider range of frequencies when the recording and reconstruction system are designed appropriately.

If the is too low, high-frequency information can be misrepresented as lower-frequency information. This distortion is called aliasing.

Quantizing amplitude

Each sample must be stored as a number. The number of bits used for each sample is the sample , also called quantization resolution. With nn bits, a sample can have up to 2n2^n possible levels. More levels allow smaller amplitude differences to be represented.

Rounding a continuous measurement to one of these finite levels creates . Increasing the sample generally reduces this error, although it increases the amount of data.

For a mono signal sampled at 44,100 Hz44{,}100\ \text{Hz} with 1616- samples, the uncompressed data rate is:

44,100 samples/second×2 bytes/sample=88,200 bytes/second44{,}100\ \text{samples/second} \times 2\ \text{bytes/sample} = 88{,}200\ \text{bytes/second}

Stereo requires twice as much sample data because it stores two channels.

Takeaway: Digital sound depends on both how often the waveform is sampled and how precisely each sample's amplitude is recorded.

Limits and Trade-offs

Digital representation is a useful model, but a fixed collection of bits cannot capture unlimited range, precision, or detail. Understanding these constraints helps explain why data formats make trade-offs.

Range and precision

A fixed number of bits supports only a finite range. For example, an 88- unsigned value cannot represent 256256 or −1-1 without a different representation or more storage. Discrete numerical levels also leave gaps between representable values, and floating-point calculations can accumulate rounding differences.

Sampling and quantization

Images sample space into pixels, while sound samples time into measurements. Detail that falls between samples may be lost. Increasing or can reduce this loss, but it also increases storage and processing requirements. Quantization introduces a difference between a continuous value and the selected digital level.

Compression

reduces file size while preserving the ability to recover the exact original data. Lossy compression achieves smaller files by discarding some information. It can be effective for images, audio, and video, but excessive or repeated lossy compression can create visible or audible artifacts.

Interpretation and compatibility

A sequence is useful only when the sender and receiver agree about its interpretation. Compatibility can depend on character encodings, number formats, order, image layouts, audio parameters, and file structure. The same pattern may therefore be valid data under one convention and incorrect or meaningless data under another.

Takeaway: Digital systems balance accuracy, range, storage, processing cost, and compatibility. Every representation is powerful within its intended rules and limited outside them.