2. Binary and Digital Representation
A progressive guide to binary notation and the rules used to represent numbers, text, images, and sound as digital data.
Binary foundations and storage units
Computers represent information as patterns of bits. A has one of two possible values, 0 or 1. These values can correspond to physical states such as low or high voltage, off or on, or two magnetic orientations.
A pattern of bits has no meaning by itself. A system must interpret it as a number, character, image value, sound sample, instruction, or another kind of information. This distinction between physical storage and logical meaning is the foundation of digital representation.
Binary place values
uses base , whereas the decimal system uses base . In a binary number, the rightmost position has value , the next position has value , and each position to the left doubles the place value. A 1 means that a place value is included; a 0 means that it is not.
For example, the binary number 1011₂ represents:
The subscripts identify the bases: 1011₂ is binary and 11₁₀ is decimal.
Bits, bytes, and storage units
A is one binary digit. A contains eight bits, so it has possible patterns. An unsigned represents decimal values from through .
Common powers-of-two storage units include bytes for a kibibyte, bytes for a mebibyte, and bytes for a gibibyte. The terms kilobyte, megabyte, and gigabyte may also refer to decimal quantities based on powers of .
Takeaway: Binary provides the patterns, but an interpretation rule is needed to give each pattern meaning.
Converting and storing binary numbers
Binary-to-decimal conversion uses the place values marked with 1. Multiply each binary digit by its corresponding power of , then add the results.
For example:
Therefore, .
To convert decimal to binary, repeatedly divide by and record each remainder. Read the remainders from bottom to top. For , the successive remainders are 1, 0, 0, 1, and 1, which produce 11001₂ when read upward.
A second method is to select powers of that add to the target value. Since
the positions from through contain 1, 1, 0, 0, and 1.
Fixed-width storage
In a fixed-width representation, a value is stored using a predetermined number of bits. The value can be written as 101₂, but an eight- stores it as 00000101₂. The leading zeros do not change the value; they fill the available positions.
A pattern does not automatically reveal whether it represents a positive integer, a negative integer, or a fraction. The representation rules must be known. For example, an unsigned eight- value ranges from through , while signed conventions such as two's complement use some patterns for negative values.
Takeaway: Conversion depends on place values, and storage depends on the width and numerical convention used.
Numbers, precision, and interpretation
Digital systems use different rules to represent different kinds of numerical values. For nonnegative integers, each binary position contributes a power of . Fractions use positions after a binary point, such as , , and .
Signed integers use conventions such as two's complement to assign negative meanings to some patterns. Real-number approximations commonly use floating-point formats with fields for a sign, an exponent, and a fraction. Floating-point formats cover a wide range of values, but they cannot represent every real number exactly.
The number of available bits limits both range and precision. More bits provide more distinct patterns, but they also require more storage. Choosing a representation therefore involves balancing the values that must be represented, the required accuracy, and the storage available.
Representation rules
A defines how a pattern should be interpreted. Its rules may specify the data type, layout, encoding, size, order, and required metadata. Without these rules, the same pattern can support several valid interpretations.
For example, the same eight bits could represent an unsigned decimal number, a character, a grayscale value, part of a color value, part of an audio sample, or an instruction. The bits alone do not determine which meaning is intended.
Takeaway: Digital representation is a contract between stored patterns and the rules used to interpret them.
Text and image representation
Text is represented by mapping characters to numbers and then storing those numbers as bits. An encoding system must specify which number corresponds to each character and how that number is stored.
ASCII represents a limited set of characters using numeric codes. assigns a code point to characters from many writing systems, punctuation marks, symbols, and emoji. can be stored using encoding forms such as , UTF-16, and UTF-32.
uses one to four bytes for a character. Characters in the ASCII range use one , while many other characters use two, three, or four bytes. For example, A has code point U+0041 and is represented by the single 41 in hexadecimal.
The same sequence of bits can produce different text if it is decoded using the wrong encoding. Text files and communication protocols therefore need an agreed-upon character encoding.
Images as numerical data
A raster image is a rectangular grid of pixels. Each stores numerical information about brightness or color. A black-and-white image may use one per . A grayscale image may use several bits per ; with bits, it can represent shades.
Color images commonly use separate red, green, and blue channels. If each channel uses bits, the image uses bits per and can represent up to:
RGB color combinations before transparency and other color-management details are considered. Greater spatial resolution means more pixels, while greater means more possible brightness or color values per . Both generally increase the amount of data.
Vector images use mathematical descriptions of shapes, lines, and curves rather than storing every . The display system calculates pixels when the image is rendered.
Takeaway: Text and images both use bits, but their formats define different mappings between numerical patterns and visible characters or colors.
Digital sound and the common representation model
Sound begins as a continuously varying analog signal. To store it digitally, a computer repeatedly measures the signal through sampling. Each measurement is assigned a numerical value through and stored in binary.
The is the number of measurements taken per second. A rate of means samples per second. is the number of bits used for each sample; a larger allows more possible amplitude values and a wider range of signal levels.
Digital audio may contain multiple channels, such as left and right stereo channels. For one second of uncompressed stereo audio at samples per second, bits per sample, and channels, the amount of data is:
This equals bytes before file headers and other metadata. The stored amount depends on , , number of channels, and duration.
Sampling and are approximations. Increasing the can capture changes that occur more quickly, while increasing provides finer measurement levels. Compression can reduce stored size, sometimes by removing information judged less noticeable.
One principle across all media
Numbers, text, images, and sound are all ultimately stored as patterns. Their meanings differ because each format supplies different interpretation rules. A format may need to identify dimensions for an image, sampling information for audio, or character encoding for text.
Takeaway: Digital media differ in their encoding methods, but every format connects binary storage to a defined interpretation.