10 — Data, Systems, and Applications
A progression through how data is organized, searched, exposed through software interfaces, and managed through abstraction layers in modern computer systems.
From Raw Data to Usable Information
A computer system becomes useful not merely by storing values, but by giving those values structure and meaning.
Data consists of recorded values such as numbers, characters, images, sounds, measurements, or symbols. The same values become information when they are labeled and interpreted in context. For example, the values 42, Maya, and 91 are more useful when identified as a student ID, a name, and an exam score.
Common levels of organization include:
Bits and bytes, which provide binary representation.
Values, such as numbers, dates, characters, and Boolean values.
Records, which group related values describing an object or event.
Files or documents, which collect records or other content.
Tables, which organize records around common fields.
Databases, which manage related collections of data.
Indexes, which provide additional structures for faster searching.
Representation and meaning are not identical. A date might be stored internally as an integer, displayed as September 15, 2026, and sent through an interface as 2026-09-15. These representations differ while referring to the same underlying meaning.
Takeaway: Organization and interpretation turn raw values into usable information.
Databases, Tables, and Queries
A is an organized collection of data designed for efficient storage, retrieval, and modification. A supplies the software that creates, queries, protects, and maintains that data.
In a , data is represented in related tables. Each table contains rows, also called records or tuples, and columns, also called attributes or fields. A schema describes the table structure, data types, relationships, and constraints.
Keys connect related records:
A uniquely identifies each row in its table.
A refers from one table to a uniquely identified row in another table.
Constraints help ensure that references and values remain valid.
For example, a student table might contain StudentID, Name, and Major. A course table might contain CourseID, Title, and Credits. An enrollment table can connect students and courses through StudentID and CourseID, along with details such as semester and grade.
A query asks the DBMS to select or transform data. The programmer describes what information is needed, while the DBMS chooses how to find it. It may use indexes, different join algorithms, caching, or parallel execution. This separation keeps applications from needing to manage the physical storage layout directly.
Takeaway: Schemas organize data, keys express relationships, constraints protect correctness, and queries describe needed results.
Consistency, , and Transactions
Poorly organized data often repeats the same facts in multiple places. If a student's name is copied into every enrollment record, changing the name can require many updates and may leave inconsistent copies.
is a design technique that separates related data into tables to reduce unnecessary duplication and update problems. Student information is stored once, course information is stored once, and enrollment records connect them with keys.
involves trade-offs. Separating data can improve consistency and storage efficiency, but it may require more joins when information is retrieved. design therefore balances:
Consistency
Storage efficiency
Query speed
Simplicity
Flexibility
A groups operations into one logical operation. Consider a transfer between two accounts: subtracting from one account and adding to another should be treated as a single unit. If the operation fails partway through, the system should be able to discard the incomplete changes. This all-or-nothing behavior is called atomicity.
Transactions are especially important when many users or programs access the same data concurrently. They help prevent intermediate or conflicting states from producing incorrect results.
Takeaway: Good design reduces duplication, while transactions protect multi-step changes.
and Search Ranking
focuses on finding relevant material in a large collection. Unlike a query that may ask for records satisfying exact conditions, often works with text, images, audio, or web pages and estimates which items best match a user's need.
An maps each term to the documents in which it occurs. For example, the term applications might map to documents 2 and 3, while the term databases might map to documents 1 and 2. The associated document identifiers form postings lists.
An index can also store:
How often a term appears in a document.
The positions where the term appears.
How many documents contain the term.
Information used to calculate relevance scores.
A typical retrieval process includes:
Tokenization: split content into terms.
: handle case, punctuation, spelling, or word forms.
Indexing: build term-to-document mappings.
Matching: find items containing query terms.
Scoring: estimate relevance.
Ranking: present the most useful results first.
Search quality can be described using two complementary measures. is the proportion of returned results that are relevant. is the proportion of all relevant results that are returned. Returning only one highly relevant result may produce high but low ; returning nearly everything may produce high but low .
Takeaway: Search systems use indexes and ranking to find useful material efficiently, balancing against .
Software Interfaces and APIs
A software interface is a defined way for one component to interact with another. It specifies available operations, accepted inputs, produced outputs, and possible errors or side effects.
An is an interface intended for programs. A library-search application, for example, might provide operations to retrieve a student, create an enrollment, or delete an enrollment. A client follows the API's contract without needing to know how the server stores data or performs authorization.
A useful interface contract can document:
Operation names
Input types and required fields
Output formats and data types
Error conditions
Authentication and authorization requirements
Versioning and compatibility rules
Performance expectations or rate limits
Interfaces occur at many levels. A function interface defines how code calls a function. A file-system interface lets programs create, read, update, and delete files. A interface accepts queries and returns results. A network protocol defines how systems exchange messages. A user interface lets people interact with software.
Good interfaces are clear, consistent, stable, minimal, compositional, portable, and secure. An interface such as findStudent(id) is generally safer and more maintainable than allowing every application to read files directly.
Takeaway: Interfaces are contracts that let components cooperate without exposing every internal detail.
Abstraction Layers: Benefits and Limits
An abstraction represents a complex system through a simpler model. An provides services to the layer above while using services from the layer below.
A typical application stack can be viewed as:
User interface
Application logic
API or service interface
management system
Operating system and file system
Hardware
Each layer has a focused responsibility. The user interface presents information and accepts actions. Application logic applies rules. The API defines communication between components. The DBMS manages queries, transactions, and storage structures. The operating system manages files, memory, processes, and devices. Hardware performs physical computation and stores bits.
Layering improves manageability, replaceability, portability, security, reuse, and complexity control. An application can request a file from an operating system without knowing which physical disk sector contains it. Similarly, a browser can display search results without knowing the schema.
Abstraction has costs. Layers can add processing overhead, memory use, communication delays, and difficulty diagnosing failures. A high-level interface is often easier to use but may provide less control; a low-level interface can provide more control but requires more detailed knowledge.
occurs when users must understand details that the abstraction was intended to hide. A query may become slow because the programmer needs to understand indexes or join strategies. Such leakage cannot always be eliminated, but details should be exposed when they matter for correctness, performance, security, or reliability.
Takeaway: Layers reduce complexity by separating responsibilities, but their boundaries can introduce costs and may sometimes reveal hidden details.
Putting the System Together
Consider a library-search application that combines structured data, search, interfaces, and layers.
Books, authors, borrowers, and loans are stored in related tables.
A query checks which books are currently available.
An searches titles, descriptions, and subject terms.
An API receives a search request, validates it, and returns structured results.
A user interface displays titles, filters, and availability.
The underlying layers handle storage, files, memory, processes, and hardware.
A request for the terms robotics ethics can move through the system as follows:
The user enters the terms in a web interface.
The interface creates a search request.
The API validates the request.
The search service consults its .
The checks availability and loan status.
The API returns structured results.
The user interface displays ranked books.
The system works because each component has a defined responsibility and communicates through interfaces. The browser does not need to know the schema, and the does not need to know how results will be displayed.
Final takeaway: Modern applications make information usable by combining organized data, targeted queries, relevance-based search, contractual interfaces, and layered abstractions.