6 Data Representation
A structured guide to how computers encode, store, transmit, and interpret numbers, logical values, characters, and text as patterns of bits.
Bits, Bytes, and Meaning
Computers store every kind of data as a pattern of bits. A has the value or , and a contains eight bits. The bits do not inherently mean “number,” “letter,” or “instruction”; software supplies the interpretation.
A therefore has possible patterns. For example, the pattern can represent the unsigned integer , the ASCII character A, or part of another data structure. This is the central idea of : the same physical bits can acquire different meanings under different rules.
Takeaway: Storage consists of patterns; meaning comes from the rules used to interpret those patterns.
Binary Numbers and Order
Binary uses base , so each position represents a power of two. In the eight- pattern , the positions containing contribute , , , and :
More generally, an - pattern can represent distinct combinations. When the pattern is treated as an unsigned integer, the largest value is . Thus, eight unsigned bits cover through , while sixteen unsigned bits cover through .
A nibble is four bits. Larger values are commonly grouped into sixteen-, thirty-two-, or sixty-four- quantities. When a value occupies multiple bytes, determines their order: big-endian places the most significant first, whereas little-endian places the least significant first.
Takeaway: Binary place values explain numerical conversion, while order determines how multi- values are laid out.
Integers and Fractional Values
An unsigned integer represents zero and positive whole numbers. With eight bits, represents , represents , and represents . The general maximum is:
Signed integers must also represent negative values. is commonly used because ordinary binary addition can then work for both positive and negative integers. To create in eight bits, begin with , invert the bits to obtain , and add , producing .
For an eight- signed two’s-complement value, the usual range is through . One pattern cannot be used to represent every possible integer: if a result exceeds the available range, a language or operation may wrap around, report an error, or use a larger representation.
Fractional values require another approach. stores a sign, a significant portion, and a scale or exponent. It can cover a wide range, but its precision is limited, so a decimal value such as may not have an exact binary representation and calculations can produce small rounding errors. Fixed-point representation instead keeps the binary-point position fixed.
Takeaway: Integer formats trade range against the number of bits, while floating-point formats trade exact precision for range and flexibility.
Logical Values and Data Types
A Boolean value represents a logical condition with two states: true and false. It may be stored as one , commonly using for false and for true, although a programming language or file format may use a whole or another fixed-size unit.
A program can use Boolean values in decisions, such as testing whether a user is logged in or has permission. The stored pattern alone does not guarantee a Boolean interpretation; the program must apply a rule that defines how the value is used.
A connects a stored pattern with valid operations. Common data types include integers, floating-point numbers, Booleans, characters, strings, and raw sequences. The type determines whether software should perform arithmetic, logical tests, character processing, or some other operation.
Takeaway: A is an interpretation contract: it identifies the kind of value and the operations that make sense for it.
Characters, Unicode, and Text
A character is an abstract symbol such as A, 7, ?, or é; a string is an ordered sequence of characters. Computers store numeric codes rather than letters directly. A defines how character identities become bytes.
ASCII assigns codes to common English letters, digits, punctuation, and control characters. For example, uppercase A has decimal code , which is binary . ASCII does not provide enough codes for all writing systems and symbols used around the world.
Unicode supplies a shared repertoire and assigns each character a . Encodings such as , UTF-16, and UTF-32 specify how those code points are represented in bytes. is variable-length: characters in the ASCII range use one , while other characters use two, three, or four bytes. Consequently, a string’s number of characters is not always equal to its number of bytes.
If bytes encoded as are incorrectly interpreted using another encoding, the result can be unreadable text called mojibake. The sender and receiver must therefore agree on the encoding when text is saved or transmitted.
Takeaway: Characters, code points, encodings, and bytes are related but distinct: a character is the symbol, a identifies it, and an encoding determines its representation.
From Stored Bits to Interpreted Data
A typical data-handling process connects four stages:
Representation: A value is converted into a pattern of bits.
Storage: The bits are placed in memory, a file, or another storage device.
Transmission: The pattern may be sent between components or across a network.
Interpretation: Software applies a , encoding, or file format to determine what the bits mean.
For example, can be interpreted as the unsigned integer or as the ASCII character A. It could also be part of image data or part of a machine instruction, depending on the surrounding rules. The pattern itself does not decide which interpretation is correct.
This principle explains why formats and protocols are necessary. Text requires an agreed encoding; an image requires rules for colors and pixel arrangement; and program data requires rules for distinguishing instructions from values. Fixed-size formats also impose limits: an eight- unsigned value cannot represent a result greater than without more bits or a different handling rule.
Takeaway: Correct interpretation depends on shared context, including the type, encoding, format, and storage conventions.