6 Data Representation

A structured guide to how computers encode, store, transmit, and interpret numbers, logical values, characters, and text as patterns of bits.

Bits, Bytes, and Meaning

Computers store every kind of data as a pattern of bits. A has the value 00 or 11, and a contains eight bits. The bits do not inherently mean “number,” “letter,” or “instruction”; software supplies the interpretation.

A therefore has 28=2562^8 = 256 possible patterns. For example, the pattern 01000001\mathtt{01000001} can represent the unsigned integer 6565, the ASCII character A, or part of another data structure. This is the central idea of : the same physical bits can acquire different meanings under different rules.

Takeaway: Storage consists of patterns; meaning comes from the rules used to interpret those patterns.

Binary Numbers and Order

Binary uses base 22, so each position represents a power of two. In the eight- pattern 00101101\mathtt{00101101}, the positions containing 11 contribute 3232, 88, 44, and 11:

001011012=32+8+4+1=4510\mathtt{00101101}_2 = 32 + 8 + 4 + 1 = 45_{10}

More generally, an nn- pattern can represent 2n2^n distinct combinations. When the pattern is treated as an unsigned integer, the largest value is 2n−12^n - 1. Thus, eight unsigned bits cover 00 through 255255, while sixteen unsigned bits cover 00 through 65,53565{,}535.

A nibble is four bits. Larger values are commonly grouped into sixteen-, thirty-two-, or sixty-four- quantities. When a value occupies multiple bytes, determines their order: big-endian places the most significant first, whereas little-endian places the least significant first.

Takeaway: Binary place values explain numerical conversion, while order determines how multi- values are laid out.

Integers and Fractional Values

An unsigned integer represents zero and positive whole numbers. With eight bits, 00000000\mathtt{00000000} represents 00, 00000001\mathtt{00000001} represents 11, and 11111111\mathtt{11111111} represents 255255. The general maximum is:

2n−12^n - 1

Signed integers must also represent negative values. is commonly used because ordinary binary addition can then work for both positive and negative integers. To create −5-5 in eight bits, begin with 00000101\mathtt{00000101}, invert the bits to obtain 11111010\mathtt{11111010}, and add 11, producing 11111011\mathtt{11111011}.

For an eight- signed two’s-complement value, the usual range is −128-128 through 127127. One pattern cannot be used to represent every possible integer: if a result exceeds the available range, a language or operation may wrap around, report an error, or use a larger representation.

Fractional values require another approach. stores a sign, a significant portion, and a scale or exponent. It can cover a wide range, but its precision is limited, so a decimal value such as 0.10.1 may not have an exact binary representation and calculations can produce small rounding errors. Fixed-point representation instead keeps the binary-point position fixed.

Takeaway: Integer formats trade range against the number of bits, while floating-point formats trade exact precision for range and flexibility.

Logical Values and Data Types

A Boolean value represents a logical condition with two states: true and false. It may be stored as one , commonly using 00 for false and 11 for true, although a programming language or file format may use a whole or another fixed-size unit.

A program can use Boolean values in decisions, such as testing whether a user is logged in or has permission. The stored pattern alone does not guarantee a Boolean interpretation; the program must apply a rule that defines how the value is used.

A connects a stored pattern with valid operations. Common data types include integers, floating-point numbers, Booleans, characters, strings, and raw sequences. The type determines whether software should perform arithmetic, logical tests, character processing, or some other operation.

Takeaway: A is an interpretation contract: it identifies the kind of value and the operations that make sense for it.

Characters, Unicode, and Text

A character is an abstract symbol such as A, 7, ?, or é; a string is an ordered sequence of characters. Computers store numeric codes rather than letters directly. A defines how character identities become bytes.

ASCII assigns codes to common English letters, digits, punctuation, and control characters. For example, uppercase A has decimal code 6565, which is binary 01000001\mathtt{01000001}. ASCII does not provide enough codes for all writing systems and symbols used around the world.

Unicode supplies a shared repertoire and assigns each character a . Encodings such as , UTF-16, and UTF-32 specify how those code points are represented in bytes. is variable-length: characters in the ASCII range use one , while other characters use two, three, or four bytes. Consequently, a string’s number of characters is not always equal to its number of bytes.

If bytes encoded as are incorrectly interpreted using another encoding, the result can be unreadable text called mojibake. The sender and receiver must therefore agree on the encoding when text is saved or transmitted.

Takeaway: Characters, code points, encodings, and bytes are related but distinct: a character is the symbol, a identifies it, and an encoding determines its representation.

From Stored Bits to Interpreted Data

A typical data-handling process connects four stages:

  1. Representation: A value is converted into a pattern of bits.

  2. Storage: The bits are placed in memory, a file, or another storage device.

  3. Transmission: The pattern may be sent between components or across a network.

  4. Interpretation: Software applies a , encoding, or file format to determine what the bits mean.

For example, 01000001\mathtt{01000001} can be interpreted as the unsigned integer 6565 or as the ASCII character A. It could also be part of image data or part of a machine instruction, depending on the surrounding rules. The pattern itself does not decide which interpretation is correct.

This principle explains why formats and protocols are necessary. Text requires an agreed encoding; an image requires rules for colors and pixel arrangement; and program data requires rules for distinguishing instructions from values. Fixed-size formats also impose limits: an eight- unsigned value cannot represent a result greater than 255255 without more bits or a different handling rule.

Takeaway: Correct interpretation depends on shared context, including the type, encoding, format, and storage conventions.