Course progress Course outline 24 of 24 lessons available
Python Foundations
Data and Collections
Building Reliable Programs
Modeling with Objects
Professional Python
Advanced Python
Bring a ruler, not a wish
“Make it faster” is a wish. Freeze behavior, measure real work, change one thing, and prove correctness again.
- Contract freeze visible behavior
- Measure find the real slow part
- Change one bounded idea
- Prove same answers, new evidence
The contract freezes inputs, outputs, order, errors, side effects, and supported Python. A checkable goal names workload, boundary, runner, and budget: “Rank 10,000 local records under 200 ms on this runner.”
A microbenchmark omits files, networks, warm-up, memory pressure, and other machines. Keep each claim as small as its evidence.
Three tools ask three questions
- timeit how long is one focused job?
- cProfile where does the full run spend time?
- tracemalloc where do Python allocations grow?
timeit.repeat() runs the same focused callable several times:
from timeit import repeat
data = tuple(range(1_000))
samples = repeat(lambda: sum(data), number=1_000, repeat=5)
print(len(samples))
print(all(sample > 0 for sample in samples))
5
True
The values vary. Keep every sample plus Python, platform, input, setup, number, and repeat. Compare candidates inside the same boundary.
When the slow part is unknown, profile the complete command:
python3 -m cProfile -s cumulative digest_evidence.py
tottime excludes callees; cumtime includes them. Profiling adds overhead, so use it to locate cost, not promise exact latency.
tracemalloc watches allocations made by Python:
import tracemalloc
tracemalloc.start()
numbers = [number * 2 for number in range(10_000)]
current, peak = tracemalloc.get_traced_memory()
print(len(numbers))
print(peak >= current)
tracemalloc.stop()
10000
True
This is traced Python memory, not the process’s complete operating-system memory.
Big-O is a growth map, not a stopwatch
Complexity predicts how work grows. A measurement describes selected points on one machine.
from collections import Counter
tokens = ["ai", "agents", "ai", "python"]
keywords = ("ai", "python")
distinct_matches = len(set(tokens) & set(keywords))
occurrence_matches = sum(Counter(tokens)[word] for word in keywords)
print(distinct_matches)
print(occurrence_matches)
2
3
A set counts distinct matching words; a Counter totals occurrences. Changing 3 into 2 changes behavior. Repeated tokens.count(word) costs about O(k × t); one Counter plus lookups averages O(t + k). Setup can still lose on tiny inputs, so measure representative sizes.
Put a guardrail around every faster idea.
Keep a clear reference implementation while testing a candidate.
- Reference simple and trusted
- Candidate possibly faster
- Edge cases empty, repeats, ties
- Equivalent? only then compare speed
from collections import Counter
def reference_score(tokens, keywords):
return sum(tokens.count(word) for word in keywords)
def counter_score(tokens, keywords):
counts = Counter(tokens)
return sum(counts[word] for word in keywords)
cases = [
([], ("ai",)),
(["ai", "ai", "python"], ("ai", "python")),
(["other"], ("ai",)),
]
for tokens, keywords in cases:
assert reference_score(tokens, keywords) == counter_score(tokens, keywords)
print("Equivalent on all cases")
Equivalent on all cases
Test empty data, repeats, normalization, ordering ties, invalid input, and the representative maximum before trusting a candidate.
Tiny project: Reading Digest Evidence
Save this as digest_evidence.py. It checks equivalent scorers, keeps samples, then ranks by score, newest date, and case-folded title.
from collections import Counter
from dataclasses import dataclass
from datetime import date
import re
from timeit import repeat
TOKEN = re.compile(r"[A-Za-z0-9]+")
@dataclass(frozen=True)
class Article:
title: str
summary: str
published: date
def tokens(article):
text = f"{article.title} {article.summary}"
return [match.group(0).casefold() for match in TOKEN.finditer(text)]
def reference_score(article, keywords):
words = tokens(article)
return sum(words.count(keyword) for keyword in keywords)
def counter_score(article, keywords):
counts = Counter(tokens(article))
return sum(counts[keyword] for keyword in keywords)
def rank(articles, keywords):
scored = []
for article in articles:
score = counter_score(article, keywords)
if score:
scored.append((score, article))
return sorted(
scored,
key=lambda pair: (
-pair[0],
-pair[1].published.toordinal(),
pair[1].title.casefold(),
),
)
articles = [
Article(
"Reliable AI agents",
"Agent evaluation makes AI systems safer.",
date(2026, 8, 10),
),
Article(
"Python for model evaluation",
"Python tools compare AI outputs offline.",
date(2026, 8, 11),
),
Article("Hardware update", "New accelerators arrived.", date(2026, 8, 12)),
]
keywords = ("ai", "python")
for article in articles:
assert reference_score(article, keywords) == counter_score(article, keywords)
reference_samples = repeat(
lambda: [reference_score(article, keywords) for article in articles],
number=1_000,
repeat=5,
)
counter_samples = repeat(
lambda: [counter_score(article, keywords) for article in articles],
number=1_000,
repeat=5,
)
print("AI reading digest")
for number, (score, article) in enumerate(rank(articles, keywords), start=1):
print(f"{number}. {article.title} ({score})")
print("Equivalent: True")
print(f"Samples per version: {len(reference_samples)}")
print(f"All samples positive: {all(reference_samples + counter_samples)}")
Run python3 digest_evidence.py and verify:
AI reading digest
1. Python for model evaluation (3)
2. Reliable AI agents (2)
Equivalent: True
Samples per version: 5
All samples positive: True
The ASCII token rule is a versioned contract, not universal language processing. Supporting phrases or Japanese text needs a named tokenizer and new equivalence tests.
Ship the tested box
- Freeze and test behavior plus evidence
- Build and inspect wheel, metadata, no secrets
- Install cleanly test away from source
- Release safely monitor and keep rollback
For the package from Chapter 20, a release rehearsal can include:
python3 -m unittest discover -s tests -v
python3 -m build
python3 -m zipfile -l dist/your_package-1.0.0-py3-none-any.whl
python3 -m venv .release-venv
.release-venv/bin/python -m pip install --no-deps dist/your_package-1.0.0-py3-none-any.whl
.release-venv/bin/your-command --help
build is an external frontend; install it deliberately. On Windows, executables are under .release-venv\Scripts\.
Build reviewed source and inspect contents, metadata, and secrets. Test that exact wheel away from the source tree. Publish those immutable bytes under a new version. Monitor the user goal, and keep the previous artifact, rollback command, and owner.
Three tiny missions
- Write a real goal. Name a workload, machine, measured boundary, repetitions, and budget. List what the goal does not cover.
- Catch a semantic change. Compare occurrence scoring with distinct-word scoring on
['ai', 'ai', 'python']and explain3versus2. - Find before fixing. Profile
digest_evidence.py, identify the highest cumulative paths, and propose one change guarded by equivalence tests.
You finished the course when…
- you define behavior before measuring or changing it;
- you choose
timeit,cProfile, ortracemallocfor the question you actually have; - you keep full samples and measurement context;
- you use complexity as a growth model, not a stopwatch;
- you prove a candidate equivalent on normal and edge cases;
- you build, inspect, clean-install, and smoke-test an artifact;
- you publish immutable tested bytes and can name the rollback path;
- the Reading Digest Evidence project prints the expected report.
You now have the whole Python loop: model the problem, write clear code, validate boundaries, test behavior, choose concurrency carefully, measure real costs, and ship something another person can trust.