JEPA4Japan · tutorials

Threads, Processes, and Parallel Work

1,320 words 6 min read #Python

Match I/O-bound and CPU-bound jobs to safe concurrency primitives.

Course progress Course outline 24 of 24 lessons available

Overlapping is not the same as simultaneous

Concurrency means several jobs are in progress during the same stretch of time. Parallelism means jobs truly run at the same instant, often on different CPU cores.

  1. Sequential finish one dish, then start one
  2. Concurrent switch while a dish waits
  3. Parallel two cooks work now
Overlapping lifetimes and simultaneous execution answer different questions.

Blocking file or network work is often I/O-bound: it spends much of its life waiting. Long pure-Python calculation is CPU-bound. Threads often help overlap blocking waits; processes can run independent CPU jobs on separate cores. Scheduling and communication cost time, so measure before claiming a speedup.

Threads share one room

Threads in one process share the same address space, globals, and heap objects. Each has its own call stack, but two references can point to one mutable object.

In the standard CPython 3.12 build, the GIL allows only one thread in a process to execute Python bytecode at a time. CPython releases it around many blocking I/O waits, and some native extensions release it. Thus threads can overlap I/O, but several threads do not normally make pure-Python CPU bytecode run across cores. Other interpreters or specialized builds can differ.

  1. One process one shared memory room
  2. Threads share objects in the room
  3. GIL one runs Python bytecode
  4. I/O wait another thread may progress
The GIL is an interpreter boundary, not a lock for your business rules.

The GIL does not protect a multi-step invariant. Use a Lock around the complete check-and-change operation:

from concurrent.futures import ThreadPoolExecutor
from threading import Lock

balance = 3
balance_lock = Lock()


def spend_one():
    global balance
    with balance_lock:
        if balance > 0:
            balance -= 1


with ThreadPoolExecutor(max_workers=4) as pool:
    list(pool.map(lambda _: spend_one(), range(10)))

print(balance)
0

The lock protects “never spend below zero.” Keep critical sections small. Better still, let workers return values and combine them in one owner when sharing is unnecessary.

Executors own workers; Futures carry outcomes

ThreadPoolExecutor reuses a bounded group of threads. map() returns results in input order, even if workers finish differently. submit() returns a Future:

from concurrent.futures import ThreadPoolExecutor


def parse_minutes(text):
    minutes = int(text)
    if minutes < 0:
        raise ValueError("negative minutes")
    return minutes


with ThreadPoolExecutor(max_workers=1) as pool:
    future = pool.submit(parse_minutes, "-2")
    try:
        future.result()
    except ValueError as error:
        print(f"Worker failed: {error}")
Worker failed: negative minutes

result() returns the value or re-raises the worker’s exception in the coordinating thread. Always observe futures. as_completed() exposes completion order, which is intentionally unstable; attach identifiers and sort before public output.

Processes use separate rooms

Worker processes have separate interpreters and memory. In standard CPython they can execute independent Python CPU work on different cores. Inputs, callables, and results must cross the boundary through serialization, usually pickle.

  1. Process A its own memory
  2. Serialized data travels between rooms
  3. Process B may use another core
Process isolation enables CPU parallelism but adds startup, memory, and transfer costs.

Process tasks should be importable top-level functions. Arguments and results should be reasonably small and serializable; open files, locks, live connections, lambdas, and nested functions are poor task values. Put orchestration behind the portable startup guard:

def main():
    run_process_pool()


if __name__ == "__main__":
    main()

Without the guard, a spawned child may import the module and recursively create another pool.

Bounds and stopping have honest limits

More workers can increase disk contention, memory use, service rejection, and serialization overhead. Choose max_workers from real capacity. For many tiny CPU jobs, batch work; for an endless stream, use finite batches or a bounded producer queue. A fixed worker count alone does not limit how many futures have been submitted.

Stopping APIs have narrower meanings than their names may suggest:

  • future.result(timeout=1) limits how long the caller waits; the work may continue.
  • future.cancel() succeeds only before the call starts.
  • shutdown(wait=True) stops accepting work and waits for running and pending work.
  • shutdown(wait=False) returns sooner, but calls continue and Python still waits before process exit.
  • cancel_futures=True cancels pending futures, not those already running.

Use native I/O timeouts and cooperative stop signals. Python cannot safely force an arbitrary running thread to unwind.

Tiny project: Local Reading Digest

Create reading_digest.py. It is offline and standard-library only. A bounded thread pool reads temporary files; a bounded process pool performs independent text analysis. Sorting makes the output deterministic.

from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
from pathlib import Path
from tempfile import TemporaryDirectory


def load_article(path):
    return path.name, path.read_text(encoding="utf-8")


def analyze_article(record):
    filename, text = record
    words = text.split()
    mentions = sum(
        word.strip(".,:;!?()").casefold() == "python" for word in words
    )
    return filename, len(words), mentions


def build_digest(paths):
    paths = list(paths)
    if not paths:
        return []
    with ThreadPoolExecutor(max_workers=min(2, len(paths))) as pool:
        loaded = list(pool.map(load_article, paths))
    with ProcessPoolExecutor(max_workers=min(2, len(loaded))) as pool:
        analyzed = list(pool.map(analyze_article, loaded))
    return sorted(analyzed)


def main():
    with TemporaryDirectory() as temporary:
        directory = Path(temporary)
        samples = {
            "a.txt": "Python waits. Python reads files.",
            "b.txt": "Processes analyze independent Python text.",
        }
        paths = []
        for name, text in samples.items():
            path = directory / name
            path.write_text(text, encoding="utf-8")
            paths.append(path)

        items = build_digest(reversed(paths))
        print("Reading digest")
        for filename, words, mentions in items:
            print(f"- {filename}: {words} words, python={mentions}")

        with ThreadPoolExecutor(max_workers=1) as pool:
            missing = pool.submit(load_article, directory / "missing.txt")
            try:
                missing.result()
            except FileNotFoundError:
                print("Missing file surfaced: True")


if __name__ == "__main__":
    main()

Save it as a file and run python3 reading_digest.py:

Reading digest
- a.txt: 5 words, python=2
- b.txt: 5 words, python=1
Missing file surfaced: True

The sample is intentionally too small to prove a speedup. It proves process startup, serialization, exception delivery, worker bounds, and stable output. A sequential analyzer may be faster for tiny input.

Three tiny missions and the Chapter 23 checklist

  1. Remove sharing. Replace the locked balance with workers that return local counts and combine them centrally.
  2. Keep order. Use submit() and as_completed(), store an input index, then print by that index.
  3. Measure honestly. Compare sequential and process analysis for larger local texts and record worker count and data size.

You are ready for Chapter 23 when:

  • you can separate concurrency from parallelism;
  • you can state what threads share and the standard CPython 3.12 GIL boundary;
  • you can protect a compound invariant with a lock;
  • you observe every Future and surface its exception;
  • you use top-level, serializable process tasks behind a main guard;
  • you bound workers and submitted work from resource limits;
  • you know timeout, cancel, and shutdown do not kill arbitrary running work;
  • you can run the digest and reproduce its sorted output.

Next, one event loop will coordinate many I/O waits with coroutines, tasks, deadlines, and cooperative cancellation.