5 Collections and String Processing

A practical guide to Python strings, sequences, mappings, sets, iteration, comprehensions, and choosing the right collection for a task.

The Collection Landscape

Python collections organize groups of values. The main built-in choices are strings, lists, tuples, dictionaries, and sets. The best choice depends on whether order, mutability, duplicates, positional access, key-based lookup, or uniqueness matters.

A useful first distinction is between sequences and mappings or sets:

  • Sequences keep values in an order and support positional operations. Strings, lists, and tuples are sequences.

  • Dictionaries find values through keys rather than numeric positions.

  • Sets store distinct values without providing positional access.

Before choosing a collection, ask:

  1. Does order matter?

  2. Must the data be changed after creation?

  3. Are duplicate values meaningful?

  4. Will values be found by position, by a key, or by membership?

  5. Is required?

Takeaway: Choose a collection according to the operations the program needs, not merely according to the shape of the sample data.

Strings and Text Processing

Strings represent textual data and are immutable sequences of Unicode characters. They can be indexed, sliced, iterated over, and tested for membership.

Common operations include:

  • lower() and upper() for case conversion.

  • strip(), lstrip(), and rstrip() for surrounding whitespace.

  • replace() for substitution.

  • split() for turning text into a of pieces.

  • join() for combining strings with a separator.

  • find() for locating a substring, returning -1 when it is absent.

  • count() for counting non-overlapping occurrences.

  • startswith() and endswith() for prefix and suffix checks.

For example, splitting "Ada,Grace,Linus" at commas produces a of names, and joining that with " | " produces formatted text. Formatted literals, also called f-strings, make it convenient to place values inside readable output.

Because strings are immutable, an operation such as text.upper() returns a new . To retain the result, assign it back to a variable or store it elsewhere.

Takeaway: Treat a as text that can be inspected and transformed, but not edited in place.

Editable and Fixed Sequences

Lists and tuples are both ordered sequences, so both support , , iteration, membership tests, and length checks. Their central difference is mutability.

A is appropriate for an editable sequence. Methods such as append(), extend(), insert(), remove(), pop(), sort(), and reverse() operate on a . Lists can also act as stacks: add items with append() and remove the most recently added item with pop().

A is appropriate for a fixed group of related values. For example, a coordinate can be represented as (3, 4), and its values can be unpacked into separate variables. elements cannot be replaced after creation.

Use a when:

  • order and position matter;

  • the collection must change;

  • duplicates should remain; or

  • items will be added or removed.

Use a when:

  • the group should remain fixed;

  • the values form one logical record; or

  • an immutable, hashable group is needed as a key or element, provided all contents are hashable.

Takeaway: Lists represent changeable ordered data; tuples represent fixed ordered data.

Key-Based Lookup and Unique Values

Dictionaries and sets solve different lookup problems. A associates each unique, hashable key with a value. A stores distinct, hashable values without associated values.

Dictionaries are useful for records and lookups. A student record might map "name" to a name and "age" to an age. Assigning to an existing key updates its value, while assigning to a new key adds a pair. Accessing a missing key with subscription raises KeyError; get() can provide a default instead. Use keys(), values(), and items() when iterating explicitly over contents.

A common counting pattern starts with an empty and updates each item using get() with a default count. This is useful for counting letters, grouping records, or tallying categories.

Sets are useful when duplicates should disappear or when membership testing and relationships are central. The union combines values from two sets, the intersection keeps shared values, the difference keeps values present in one but not the other, and the symmetric difference keeps values present in exactly one . Sets do not support or .

Use remove() when a missing item should be treated as an error, and discard() when missing is acceptable.

Takeaway: Use a to map keys to values and a to represent unique membership.

Iteration, Comprehensions, and Selection

Iteration provides a common way to process strings, lists, tuples, dictionaries, sets, ranges, and many other iterable objects. A for loop visits each item in turn.

Use enumerate() when both a position and its value are needed. Use zip() when corresponding items from multiple iterables should be processed together. For dictionaries, iterate over keys by default, over values with values(), or over key–value pairs with items().

Comprehensions make common transformations concise. A can transform or filter items, a creates unique results, and a creates key–value pairs. For example, a can create squares from a range of numbers, while a can collect the distinct lengths of several words.

Prefer a regular loop when a would require several nested loops, complicated conditions, or side effects. Readability is more important than compactness.

A final selection guide is:

  • Use str for textual data.

  • Use for an ordered, editable sequence.

  • Use for an ordered, fixed sequence.

  • Use dict for lookup by key or structured associations.

  • Use for unique values and membership operations.

  • Use collections.deque when efficient additions and removals at both ends are required.

Takeaway: Clear iteration patterns and an appropriate collection make transformations easier to read, maintain, and reason about.