1 Binary Representation and Data

Learn how bits are grouped, interpreted as numbers, and encoded to represent text and other kinds of data.

Bits: the building blocks

Computers store information as patterns of bits, each of which is either 00 or 11. A is one digit, and a groups 88 bits together. A pattern does not have an inherent meaning: a program or data format must interpret it.

Reading and

The number system is base 22. Each position in a number represents a power of 22, just as each decimal position represents a power of 1010. For example:

1011012=1×25+0×24+1×23+1×22+0×21+1×20=4510\texttt{101101}_2 = 1\times 2^5 + 0\times 2^4 + 1\times 2^3 + 1\times 2^2 + 0\times 2^1 + 1\times 2^0 = 45_{10}

The number system is base 1616, using digits 00–99 and letters A–F, where A–F represent 1010–1515. Each digit corresponds to exactly four bits, so can be converted by grouping bits into chunks of four:

  • 1101 10102=DA16=21810\texttt{1101\ 1010}_2 = \texttt{DA}_{16} = 218_{10}.

  • 0010 11112=2F16\texttt{0010\ 1111}_2 = \texttt{2F}_{16}.

is therefore a compact way to write patterns. The prefix 0x\texttt{0x} is often used to indicate a value, as in 0x2F\texttt{0x2F}.

Takeaway: place values are powers of 22, and each digit neatly summarizes four bits.

Interpreting integer patterns

The same number of bits can represent different numerical ranges depending on the convention. An uses all positions for nonnegative place values. For nn bits, its range is:

0 through 2n−10\ \text{through}\ 2^n-1

Thus, an 88- ranges from 00 to 255255.

Computers commonly use to represent signed integers. In an nn- two’s-complement pattern, the leftmost has place value −2n−1-2^{n-1}, while the other positions retain positive powers of 22. The range is:

−2n−1 through 2n−1−1-2^{n-1}\ \text{through}\ 2^{n-1}-1

For 88 bits, that range is −128-128 to 127127. For example, the pattern 11111011\texttt{11111011} represents −5-5:

−128+64+32+16+8+2+1=−5-128 + 64 + 32 + 16 + 8 + 2 + 1 = -5

A common way to form a negative value in is to invert the bits of its positive counterpart and add 11. In this way, 00000101\texttt{00000101}, representing 55, becomes 11111011\texttt{11111011}, representing −5-5.

The pattern 11111111\texttt{11111111} illustrates why the interpretation convention matters: it represents 255255 as an unsigned 88- integer and −1-1 as an 88- two’s-complement integer.

Takeaway: a pattern alone does not determine a number; the signed or unsigned representation convention does.

Representing text

Text is stored by assigning numeric values to characters and encoding those values as bits. assigns values to a limited set of characters; uppercase A has the value 6565, written in as 0x41\texttt{0x41}.

assigns code points to characters. A is written in forms such as U+0041\texttt{U+0041} for uppercase A. The identifies the character, but an encoding form determines the bytes or units used to store or transmit it.

For example, uses between 11 and 44 bytes per . UTF-16 uses one or two 1616- code units, while UTF-32 uses one 3232- code unit. preserves the values for the first 128128 code points, so A is encoded as the single 0x41\texttt{0x41}.

Takeaway: a identifies a character; its UTF encoding determines its stored representation.

From patterns to data

A data format gives stored patterns their meaning. The same 88 bits might represent an , a signed integer, or a character, depending on how they are interpreted. Other data types follow their own conventions: a Boolean value may indicate false or true, while image and audio data can be represented by numerical pixel or sound samples. Some formats also specify the order in which multiple bytes are stored or transmitted.

Key idea: bits are the physical representation; the agreed encoding or data type assigns meaning to the pattern.