JEPA4Japan · tutorials

Iterators and Generators

1,090 words 5 min read #Python

Process streams lazily and understand the protocol behind Python's for loop.

Course progress Course outline 24 of 24 lessons available

A shelf and a bookmark are different

A list is a reusable shelf of values. An iterable is anything that can hand out a bookmark. An iterator is that stateful bookmark: it remembers where the next value lives.

  1. Iterable a shelf you can visit
  2. iter() make a bookmark
  3. next() move and get one value
The shelf can make bookmarks; each bookmark owns one position.
topics = ["iterables", "iterators", "generators"]
bookmark = iter(topics)

print(next(bookmark))
print(next(bookmark))
print(list(bookmark))
print(list(bookmark))

Output:

iterables
iterators
['generators']
[]

The iterator is single-use, so the second list(bookmark) is empty. The original list is reusable because another iter(topics) makes a fresh iterator. Files, strings, ranges, and many other objects are iterable too.

StopIteration means “the trail ends”

When no value remains, next() raises StopIteration. A for loop quietly handles that normal signal for you.

  1. iter(iterable) get one iterator
  2. next(iterator) receive one value
  3. Repeat ask again
  4. StopIteration finish the loop
This small protocol is hiding behind every for loop.
bookmark = iter([10, 20])

while True:
    try:
        value = next(bookmark)
    except StopIteration:
        print("done")
        break
    else:
        print(value)

Output:

10
20
done

Write ordinary for loops for ordinary work. Use explicit next() when the exact boundary matters. next(bookmark, "empty") can return a default instead, but choose a marker that cannot be confused with real data.

Do not raise StopIteration inside a generator body. Use return or reach the end; an escaping StopIteration becomes RuntimeError.

yield is a pause button

A function containing yield is a generator function. Calling it creates a generator but does not run its body yet. The first pull starts it; each yield hands out one value and freezes local state.

  1. Run work until yield
  2. Pause remember every local value
  3. Give send one value outward
The next pull resumes immediately after the paused yield.
def count_to(limit):
    print("started")
    for number in range(1, limit + 1):
        yield number


numbers = count_to(3)
print("created")
print(next(numbers))
print(list(numbers))

Output:

created
started
1
[2, 3]

A list comprehension builds everything now; a generator expression such as (n * n for n in range(5)) produces values only when pulled. Choose a list for repeated traversal, length, indexing, or a snapshot. Choose a lazy stream when one pass is enough and delayed work or lower peak memory matters. Lazy does not automatically mean faster.

Join small streams carefully

yield from passes through every value from another iterable:

def whole_course():
    yield from ["basics", "collections"]
    yield from ["generators", "testing"]


print(list(whole_course()))

Output:

['basics', 'collections', 'generators', 'testing']

Lazy work also delays errors. A bad line may fail far from the call that created the generator, so test consumption, not only creation. If a generator opens a file, it may keep that file open while paused. Consume it inside the resource’s with block, or give ownership and closing a clear try/finally policy. Catch only exceptions that the stream can actually classify.

generator.send(value) enables a two-way conversation, but keep it at the edge of your toolbox. A new generator must first reach its initial yield with next(generator); sending a non-None value before that raises TypeError. For normal pipelines, function arguments in and yielded values out are much clearer.

Build a lazy reading conveyor

Save this complete standard-library project as reading_conveyor.py:

from dataclasses import dataclass
from itertools import islice


@dataclass(frozen=True)
class Reading:
    topic: str
    minutes: int


def joined(*groups):
    for group in groups:
        yield from group


def parse_readings(lines):
    for line_number, line in enumerate(lines, start=1):
        parts = [part.strip() for part in line.split("|")]
        if len(parts) != 2 or not parts[0]:
            raise ValueError(f"line {line_number}: expected topic|minutes")
        try:
            minutes = int(parts[1])
        except ValueError as cause:
            raise ValueError(f"line {line_number}: invalid minutes") from cause
        if minutes <= 0:
            raise ValueError(f"line {line_number}: minutes must be positive")
        yield Reading(parts[0], minutes)


def at_least(readings, minimum):
    for reading in readings:
        if reading.minutes >= minimum:
            yield reading


raw_lines = joined(
    ["Iterators | 25", "Generators | 40"],
    ["Broken line"],
)
selected = at_least(parse_readings(raw_lines), minimum=20)

first_two = list(islice(selected, 2))
assert first_two == [Reading("Iterators", 25), Reading("Generators", 40)]
print("First pull:", first_two)

try:
    next(selected)
except ValueError as error:
    print("Next pull:", error)

Run python3 reading_conveyor.py:

First pull: [Reading(topic='Iterators', minutes=25), Reading(topic='Generators', minutes=40)]
Next pull: line 3: expected topic|minutes

islice() asks for only two values, so the broken third line is untouched at first. The next pull reaches it and raises the delayed error. The assertion verifies order and contents without turning the whole source into a list.

Three tiny missions

  1. Make one iterator, call next() twice, then convert its remainder to a list twice.
  2. Add a third group to joined() and predict the combined order before running it.
  3. Change islice(selected, 2) to islice(selected, 1) and explain which line the next pull reaches.

Ready for Chapter 18?

  • I can distinguish an iterable from an iterator.
  • I can explain iter(), next(), and StopIteration behind a for loop.
  • I know when a generator starts and what yield pauses.
  • I can choose between a reusable list and a single-pass lazy stream.
  • I can join stages with yield from without hiding resource ownership.
  • I ran the Reading Conveyor and saw the error arrive only when pulled.

Next, you will put useful policy around a function call with decorators and around a code block with context managers.