1 Binary Representation and Data
Learn how bits are grouped, interpreted as numbers, and encoded to represent text and other kinds of data.
Bits: the building blocks
Computers store information as patterns of bits, each of which is either or . A is one digit, and a groups bits together. A pattern does not have an inherent meaning: a program or data format must interpret it.
Reading and
The number system is base . Each position in a number represents a power of , just as each decimal position represents a power of . For example:
The number system is base , using digits – and letters A–F, where A–F represent –. Each digit corresponds to exactly four bits, so can be converted by grouping bits into chunks of four:
.
.
is therefore a compact way to write patterns. The prefix is often used to indicate a value, as in .
Takeaway: place values are powers of , and each digit neatly summarizes four bits.
Interpreting integer patterns
The same number of bits can represent different numerical ranges depending on the convention. An uses all positions for nonnegative place values. For bits, its range is:
Thus, an - ranges from to .
Computers commonly use to represent signed integers. In an - two’s-complement pattern, the leftmost has place value , while the other positions retain positive powers of . The range is:
For bits, that range is to . For example, the pattern represents :
A common way to form a negative value in is to invert the bits of its positive counterpart and add . In this way, , representing , becomes , representing .
The pattern illustrates why the interpretation convention matters: it represents as an unsigned - integer and as an - two’s-complement integer.
Takeaway: a pattern alone does not determine a number; the signed or unsigned representation convention does.
Representing text
Text is stored by assigning numeric values to characters and encoding those values as bits. assigns values to a limited set of characters; uppercase A has the value , written in as .
assigns code points to characters. A is written in forms such as for uppercase A. The identifies the character, but an encoding form determines the bytes or units used to store or transmit it.
For example, uses between and bytes per . UTF-16 uses one or two - code units, while UTF-32 uses one - code unit. preserves the values for the first code points, so A is encoded as the single .
Takeaway: a identifies a character; its UTF encoding determines its stored representation.
From patterns to data
A data format gives stored patterns their meaning. The same bits might represent an , a signed integer, or a character, depending on how they are interpreted. Other data types follow their own conventions: a Boolean value may indicate false or true, while image and audio data can be represented by numerical pixel or sound samples. Some formats also specify the order in which multiple bytes are stored or transmitted.
Key idea: bits are the physical representation; the agreed encoding or data type assigns meaning to the pattern.