6 File Input and Output

Learn how to open, read, write, organize, and safely persist text and structured data with Python file I/O.

Opening Files and Choosing Modes

File input and output lets a Python program exchange data with persistent files. Reading brings stored data into memory, while writing saves results for later use.

Python's open() function returns a file object. Its usual form is open(path, mode="r", =None). The controls the operation:

  • "r" reads an existing file and raises if it is missing.

  • "w" writes by creating a file or replacing its existing contents.

  • "a" appends to the end of a file and creates it if necessary.

  • "x" creates a new file and fails if the file already exists.

  • Modes such as "rb" and "wb" use bytes instead of text.

Text mode works with strings. Binary mode works with bytes and is appropriate for images, audio, compressed files, and other non-text data. Do not supply an when using binary mode.

Takeaway: Select the mode deliberately, especially before using "w", because it can erase existing contents.

Managing File Lifetime Safely

A coordinates setup and cleanup for a resource. For files, the preferred pattern is a with statement:

with open("notes.txt", "r", ="utf-8") as file:

The indented block can read or write through file; when the block ends, Python closes the file automatically. This cleanup also occurs when an exception leaves the block, so buffered data is flushed and the operating-system resource is released.

Manual cleanup is possible with try and finally, but it is longer and easier to get wrong. Use with open(...) or Path.open() as the normal approach.

Takeaway: Keep file operations inside a with block so closure is predictable during both normal execution and failures.

Reading and Converting Text

A file object offers several ways to read text:

  • read() returns the remaining contents as one string. It is convenient for small files, but loading a very large file at once can consume substantial memory.

  • readline() returns one line at a time.

  • readlines() returns the remaining lines as a list.

  • Iterating directly over the file processes one line at a time and is generally the most memory-efficient pattern for large files.

Lines normally include their trailing newline. Use rstrip() or rstrip("\\n") when that newline should be removed. Use strip() cautiously because it also removes leading and trailing spaces, which may be meaningful.

File contents arrive as strings, so convert them when the program needs another type. For example, a list of integer values can be built by applying int(line.strip()) to each line. An average can then be computed with sum(numbers) / len(numbers).

Takeaway: Iterate over large files, remove only the whitespace you intend to remove, and convert input strings explicitly.

Writing Text and Working with Paths

Writing is also performed through a file object. write() accepts one string and returns the number of characters written; it does not add a newline automatically. Include "\\n" when separate output lines are needed. writelines() writes an iterable of strings, so each string must already contain any required newline.

Opening in "w" mode replaces the previous contents. Opening in "a" mode preserves existing contents and adds new text at the end. For example, appending "Program started\\n" is suitable for a log entry.

A such as data/input.txt is interpreted from the current working directory. An gives the complete location, such as /home/user/data/input.txt on many Unix-like systems. The class provides a portable way to construct paths:

path = Path("data") / "input.txt"

A Path can report whether a location exists, whether it is a regular file, and what its parent, name, or suffix is. Its read_text() and write_text() methods provide concise text operations and automatically open and close the file. write_text() replaces an existing file.

Takeaway: Use "a" for additions, treat "w" as destructive replacement, and prefer Path for portable path construction.

Encodings, Newlines, and Exceptions

An defines how characters become bytes and how bytes become characters again. Explicitly using ="utf-8" makes text handling more predictable across computers, especially when files contain non-ASCII characters. If the reading and writing encodings do not match, Python may raise UnicodeDecodeError or produce incorrect text.

In ordinary text mode, Python translates platform-specific line endings into "\\n" while reading and performs the corresponding translation while writing. Binary mode avoids this text translation and preserves bytes. When using the csv module, open the file with newline="" so the module can handle line-ending conventions correctly.

Expected file failures should be handled specifically. can indicate that a required file is absent, PermissionError can indicate insufficient access, and UnicodeDecodeError can indicate that the data is not valid in the expected . Avoid a bare except: because it can hide programming errors and interrupt signals.

Takeaway: Make text explicit, distinguish text from binary data, and catch only failures the program can address.

Persisting Structured Data with

For simple persistence of lists and dictionaries, provides a standard text representation. .dump() writes a Python value to an open file, while .load() reads the file and reconstructs a corresponding Python value.

A typical persistence workflow is:

  1. Create a Path for the data file.

  2. Open it in UTF-8 text mode.

  3. Use .dump(..., indent=2) to save readable structured data.

  4. Open it for reading later and call .load().

  5. Handle a missing file or malformed when an empty default is an appropriate recovery.

supports common values such as dictionaries, lists, strings, numbers, booleans, and None, but it does not automatically serialize every Python object. For important data, validation, backups, and safer update strategies are advisable because a program that stops during a direct write can leave incomplete contents.

Takeaway: Use for basic structured persistence, and plan for missing, invalid, or interrupted data when reliability matters.

Putting File I/O Patterns Together

Reliable file-processing programs combine safe resource handling, incremental processing, validation, and deliberate error handling. A typical transformation opens a source and destination with context managers, iterates over the source line by line, validates or transforms each line, and writes accepted results to the destination.

For example, a filtering task can copy only lines for which line.strip() is nonempty. A transformation task can remove surrounding whitespace, convert a name with .upper(), and write the result followed by "\\n". A numeric summary can use enumerate(file, start=1) to retain line numbers, skip blank lines, attempt float(text), and report invalid values with their locations.

These patterns scale because they avoid unnecessary full-file loading and make bad input visible without treating every failure as a programming error. Before finalizing a program, check that it uses a , specifies UTF-8 for text when appropriate, chooses the correct mode, constructs paths portably, validates untrusted input, and avoids accidentally replacing valuable files.

Final checklist:

  • Use with open(...) or Path.open().

  • Specify ="utf-8" for UTF-8 text.

  • Choose "r", "w", or "a" according to the intended operation.

  • Use binary modes for non-text data.

  • Iterate over large files.

  • Handle expected exceptions specifically.

  • Use and where they simplify the design.