11 Testing and Debugging

A practical guide to distinguishing program errors, debugging systematically, designing effective tests, and reasoning about correctness from specifications.

and Fundamentals

and work together, but they answer different questions. asks whether observed behavior matches expected behavior for selected inputs. investigates why a discrepancy occurred and removes its underlying cause.

A program can pass several tests and still be incorrect. Tests provide evidence about the cases they cover; they do not normally prove correctness for every possible input. Test expectations should come from a precise specification rather than from assumptions about how the implementation happens to be written.

A useful distinction is:

  • finds a discrepancy between expected and actual behavior.

  • explains the discrepancy and fixes the defect.

  • Re- checks that the correction works and has not damaged related behavior.

Takeaway: Begin with the intended behavior, not with the line that happens to display the wrong result.

Recognizing Program Errors

Program errors differ in when they appear and how they affect behavior.

A violates the programming language’s grammar. For example, a conditional statement missing its required colon usually prevents the program from starting.

A occurs during execution. Dividing by zero, accessing an invalid array index, opening a missing file, or using an inappropriate value type can all cause runtime failures.

A is more subtle: the program runs but computes or returns the wrong result. If an adult is defined as someone whose age is at least 1818, a test for age greater than 1818 incorrectly excludes the boundary value 1818. The correction must follow the intended rule.

When diagnosing a failure, ask not only, “Which line failed?” but also, “What behavior was intended, and where did the actual behavior first diverge from it?” The place where an error becomes visible may differ from the place where the incorrect value was created.

Takeaway: Classify the failure, then search backward from the first incorrect behavior to the defect that caused it.

A Systematic Process

Use a repeatable process instead of changing unrelated lines until the program appears to work.

  1. Reproduce the problem. Find a small input that reliably demonstrates the failure. Record the input, expected result, actual result, error message, traceback, and reproduction steps.

  2. Read diagnostic information. An error message often identifies the type and detection location. A traceback shows the sequence of calls that led there, but the defect may be in an earlier caller.

  3. Form one specific hypothesis. For example: “The loop skips the final item because its upper bound is exclusive.” A precise hypothesis suggests a focused inspection or test.

  4. Inspect relevant state. Examine variables, indexes, object fields, return values, and stack frames immediately before and after the suspicious operation.

  5. Make the smallest justified change. Correct the code that violates the intended rule, not merely the code nearest the visible symptom.

  6. Re-run the failing test and related tests. Confirm the original failure is fixed, then check nearby behavior and previously working functionality.

  7. Preserve the failure as a . Keep a test that would have failed before the correction so future changes can detect a recurrence.

A debugger can pause at breakpoints, execute one statement at a time, inspect expressions, and show the call stack. Step over a called method when its internals are not the focus; step into it when the called method may contain the defect; step out when its inspection is complete.

Takeaway: Reproduce, hypothesize, inspect, change minimally, and verify broadly.

Tracing Execution and Loop Boundaries

A records the important state of a program as it executes. For a loop that counts even values, record the current value, the condition result, and the count after each iteration. With input values [3,8,5,10][3, 8, 5, 10], the count changes from 00 to 11 when 88 is processed and from 11 to 22 when 1010 is processed.

When tracing a loop, record:

  • the initial state;

  • the state at each iteration;

  • the condition result;

  • every update to an accumulator, index, or object field;

  • the state after the loop ends.

Tracing can reveal a variable that is never updated, an update that happens too early or too late, a condition that is always true or false, a loop that executes one time too many or too few, an unexpected mutation, or a method that returns before its intended final step.

For a loop processing nn items, explicitly consider n=0n = 0, n=1n = 1, and a typical larger value. Check whether indexing begins at 00 or 11, and whether the final bound is inclusive or exclusive. These checks expose many off-by-one defects.

Takeaway: A turns vague suspicion into a sequence of observable state changes.

Designing Effective Tests

A checks a small unit, such as a function, method, or class operation, in isolation. A normal test case has setup, an action, and an assertion that compares actual behavior with the expected result. Tests should check the specification rather than depend on an internal technique such as a particular loop, sorting method, or data structure.

Build a test set that covers several categories:

  • Normal cases: common valid inputs.

  • Boundary cases: values at the limits of valid behavior.

  • Empty cases: empty strings, arrays, collections, or files.

  • Singleton cases: inputs containing exactly one item.

  • Large cases: inputs near expected capacity limits.

  • Invalid cases: malformed, out-of-range, or unacceptable inputs.

  • Special-value cases: negative numbers, zero, duplicates, null-like values, and unusual characters.

  • Interaction cases: combinations of conditions that may reveal defects.

For a method accepting ages from 00 through 120120, useful boundary inputs include −1-1, 00, 11, 119119, 120120, and 121121. For an array method, consider an empty array, a one-item array, two items, repeated values, an already sorted array, and a reverse-sorted array.

For a search operation, test a target at the first position, last position, and middle; a missing target; an empty array; and duplicates. For a loop expected to process nn items, verify zero, one, and a typical larger input.

Takeaway: Strong tests cover meaningful input categories and expected claims, not merely lines of code.

Boundaries and

An lies near an unusual or limiting condition. Boundary analysis is especially valuable because incorrect comparison operators and loop limits often behave correctly for ordinary inputs but fail at the edges.

makes invalid states difficult to create and detects them early. Apply it by:

  1. validating ranges, required values, formats, and collection sizes;

  2. stating clear requirements before an operation runs;

  3. checking the results of risky file, conversion, network, or search operations;

  4. failing clearly with an appropriate error or documented failure result;

  5. preserving object invariants through controlled updates;

  6. providing meaningful error messages when it is safe to identify the invalid value or required condition;

  7. avoiding accidental dependence on mutable shared arrays or objects.

Assertions are useful for internal assumptions and invariants. For example, after computing an average, an assertion may check that the result lies between the minimum and maximum values. However, assertions can be removed when Python runs with optimization, so they should not replace validation of untrusted user input or required production checks.

Takeaway: Reject or handle invalid states explicitly, and distinguish normal input validation from checks of internal programming assumptions.

Reasoning About Program Correctness

A program is correct when its observable behavior satisfies its specification for every input in the specified domain. Evaluation should consider both results and side effects, including changes to collections, object fields, files, or other external state.

A describes what must be true before an operation. A describes what must be true afterward. For an operation removing an item from a collection, the might require a valid position; the might require one fewer item and preservation of the remaining order.

An remains true at a defined point in execution. Suppose a loop processes the first ii elements of an array. A useful is that, after each iteration, the accumulated total equals the sum of those first ii processed elements. To support correctness, show that the is true before the loop, remains true after each iteration, and implies the desired result when the loop ends.

A correctness review should ask:

  • Does the program handle every valid input category?

  • Does it reject or safely handle invalid input?

  • Are return values and side effects correct and documented?

  • Are array indexes and loop limits valid?

  • Do object fields preserve their class invariants?

  • Are inherited methods valid for derived objects?

  • Does the program terminate when required?

  • Does it satisfy performance and resource constraints?

  • Are failure behavior and error messages appropriate?

Takeaway: Combine tests with preconditions, postconditions, invariants, and reasoning about the full input domain.

An Integrated and Workflow

A practical workflow combines the ideas throughout this guide:

  1. Define expected behavior precisely.

  2. Write a small test that demonstrates the suspected problem.

  3. Reproduce the failure consistently.

  4. Read the error message and traceback.

  5. execution and locate the first incorrect value.

  6. Form and test one hypothesis at a time.

  7. Make the smallest justified correction.

  8. Re-run the failing test.

  9. Run boundary, invalid-input, interaction, and regression tests.

  10. Keep the new test and document important assumptions.

An off-by-one example illustrates the method. If a summing loop processes indexes from 00 through n−2n - 2, it omits the final index n−1n - 1. A with the values [4,7,2][4, 7, 2] shows that indexes 00 and 11 are visited while index 22 is not. Correct the loop so every required element is processed, then test an empty array, a one-element array, and a multi-element array. When indexing is unnecessary, iterating directly over values reduces opportunities for index errors.

The final verification should include the original failing case, nearby boundary cases, invalid inputs, and previously passing behavior. Keep the new test as a and record any assumption that future maintainers must preserve.

Takeaway: A reliable workflow turns a one-time fix into durable evidence that the specification continues to be met.