JEPA4Japan · 教程

测量、优化与交付

2,184字 7分钟阅读 #Python

分析真正的性能瓶颈,并交付一个经过测试的 AI 阅读摘要命令行应用程序。

课程进度 课程大纲 已发布 24/24 课

带上尺子,而不是愿望

“让它更快”只是一个愿望。先固定行为,测量真实工作负载,只改变一件事,然后再次证明正确性。

  1. 契约 固定可见行为
  2. 测量 找到真正缓慢的部分
  3. 改变 一个范围明确的想法
  4. 证明 答案相同,证据更新
正确 → 测量 → 改变 → 再次测试和测量。

契约会固定输入、输出、顺序、错误、副作用以及所支持的 Python 版本。一个可检验的目标要说明工作负载、测量边界、运行环境和预算:“在这个运行环境中,用不到 200 ms 对 10,000 条本地记录进行排名。”

微基准测试不会涵盖文件、网络、预热、内存压力和其他机器。每项结论都不应超出其证据所能支持的范围。

三种工具提出三个问题

  1. timeit 一项明确的任务要花多长时间?
  2. cProfile 完整运行的时间花在了哪里?
  3. tracemalloc Python 的内存分配在哪里增长?
时钟、地图和内存透镜回答不同的问题。

timeit.repeat() 会多次运行同一个专注于特定任务的可调用对象:

from timeit import repeat

data = tuple(range(1_000))
samples = repeat(lambda: sum(data), number=1_000, repeat=5)

print(len(samples))
print(all(sample > 0 for sample in samples))
5
True

这些数值会有所变化。请保留每个样本,以及 Python 版本、平台、输入、准备过程、number 和 repeat。在相同的测量边界内比较候选方案。

如果还不知道缓慢的部分在哪里,就分析完整命令:

python3 -m cProfile -s cumulative digest_evidence.py

tottime 不包含被调用函数的耗时;cumtime 包含它们。性能分析本身会增加开销,所以要用它定位成本,而不是承诺精确的延迟。

tracemalloc 会监视 Python 进行的内存分配:

import tracemalloc

tracemalloc.start()
numbers = [number * 2 for number in range(10_000)]
current, peak = tracemalloc.get_traced_memory()

print(len(numbers))
print(peak >= current)
tracemalloc.stop()
10000
True

这里测量的是被追踪的 Python 内存,而不是进程在操作系统层面的完整内存占用。

Big-O 是增长地图,不是秒表

复杂度预测工作量如何增长。测量结果描述的是某台机器上选定的几个数据点。

from collections import Counter

tokens = ["ai", "agents", "ai", "python"]
keywords = ("ai", "python")

distinct_matches = len(set(tokens) & set(keywords))
occurrence_matches = sum(Counter(tokens)[word] for word in keywords)

print(distinct_matches)
print(occurrence_matches)
2
3

集合计算的是不同匹配单词的数量;Counter 计算的是出现总次数。把 3 变成 2 就改变了程序行为。重复调用 tokens.count(word) 的成本约为 O(k × t);创建一次 Counter 再进行查找,平均成本为 O(t + k)。对于很小的输入,准备成本仍可能让它更慢,因此要测量有代表性的数据规模。

为每一个更快的想法设置护栏。

测试候选方案时,保留一个清晰的参考实现。

  1. 参考实现 简单且可信
  2. 候选方案 可能更快
  3. 边界情况 空数据、重复项、并列
  4. 等价吗? 只有等价后才比较速度
在一个想法进入基准测试之前,先用正确性测试把关。
from collections import Counter


def reference_score(tokens, keywords):
    return sum(tokens.count(word) for word in keywords)


def counter_score(tokens, keywords):
    counts = Counter(tokens)
    return sum(counts[word] for word in keywords)


cases = [
    ([], ("ai",)),
    (["ai", "ai", "python"], ("ai", "python")),
    (["other"], ("ai",)),
]

for tokens, keywords in cases:
    assert reference_score(tokens, keywords) == counter_score(tokens, keywords)

print("Equivalent on all cases")
Equivalent on all cases

在信任候选方案之前,要测试空数据、重复项、规范化、排序并列、无效输入以及具有代表性的最大规模。

小项目:阅读摘要证据

将下面的内容保存为 digest_evidence.py。它会检查评分函数是否等价、保留测量样本,然后依次按照分数、最新日期和经过大小写折叠的标题进行排名。

from collections import Counter
from dataclasses import dataclass
from datetime import date
import re
from timeit import repeat


TOKEN = re.compile(r"[A-Za-z0-9]+")


@dataclass(frozen=True)
class Article:
    title: str
    summary: str
    published: date


def tokens(article):
    text = f"{article.title} {article.summary}"
    return [match.group(0).casefold() for match in TOKEN.finditer(text)]


def reference_score(article, keywords):
    words = tokens(article)
    return sum(words.count(keyword) for keyword in keywords)


def counter_score(article, keywords):
    counts = Counter(tokens(article))
    return sum(counts[keyword] for keyword in keywords)


def rank(articles, keywords):
    scored = []
    for article in articles:
        score = counter_score(article, keywords)
        if score:
            scored.append((score, article))
    return sorted(
        scored,
        key=lambda pair: (
            -pair[0],
            -pair[1].published.toordinal(),
            pair[1].title.casefold(),
        ),
    )


articles = [
    Article(
        "Reliable AI agents",
        "Agent evaluation makes AI systems safer.",
        date(2026, 8, 10),
    ),
    Article(
        "Python for model evaluation",
        "Python tools compare AI outputs offline.",
        date(2026, 8, 11),
    ),
    Article("Hardware update", "New accelerators arrived.", date(2026, 8, 12)),
]
keywords = ("ai", "python")

for article in articles:
    assert reference_score(article, keywords) == counter_score(article, keywords)

reference_samples = repeat(
    lambda: [reference_score(article, keywords) for article in articles],
    number=1_000,
    repeat=5,
)
counter_samples = repeat(
    lambda: [counter_score(article, keywords) for article in articles],
    number=1_000,
    repeat=5,
)

print("AI reading digest")
for number, (score, article) in enumerate(rank(articles, keywords), start=1):
    print(f"{number}. {article.title} ({score})")
print("Equivalent: True")
print(f"Samples per version: {len(reference_samples)}")
print(f"All samples positive: {all(reference_samples + counter_samples)}")

运行 python3 digest_evidence.py 并验证:

AI reading digest
1. Python for model evaluation (3)
2. Reliable AI agents (2)
Equivalent: True
Samples per version: 5
All samples positive: True

ASCII 词元规则是一份有版本记录的契约,而不是通用的语言处理方式。若要支持短语或日语文本,就需要指定一个分词器并编写新的等价性测试。

发布经过测试的盒子

  1. 固定并测试 行为加证据
  2. 构建并检查 wheel、元数据、没有机密信息
  3. 干净安装 离开源代码进行测试
  4. 安全发布 监控并保留回滚方案
发布你测试过的完全相同的字节,并保留上一个正常的盒子。

对于第 20 章中的软件包,一次发布演练可以包括:

python3 -m unittest discover -s tests -v
python3 -m build
python3 -m zipfile -l dist/your_package-1.0.0-py3-none-any.whl
python3 -m venv .release-venv
.release-venv/bin/python -m pip install --no-deps dist/your_package-1.0.0-py3-none-any.whl
.release-venv/bin/your-command --help

build 是一个外部前端工具;请有意识地安装它。在 Windows 上,可执行文件位于 .release-venv\Scripts\ 下。

从经过审查的源代码进行构建,并检查内容、元数据和机密信息。离开源代码目录树测试那个完全相同的 wheel。以一个新版本发布这些不可变的字节。监控用户目标,并保留上一个构件、回滚命令和负责人信息。

三个小任务

  1. 写出一个真实目标。 说明工作负载、机器、测量边界、重复次数和预算。列出这个目标不涵盖的内容。
  2. 发现语义变化。 在 ['ai', 'ai', 'python'] 上比较按出现次数评分和按不同单词评分,并解释 3 与 2 的区别。
  3. 先找到问题,再修复。 分析 digest_evidence.py 的性能,找出累计耗时最高的路径,并提出一项由等价性测试保护的改动。

当你做到这些时,就完成了本课程……

  • 你在测量或改变行为之前先定义行为;
  • 你会根据真正要回答的问题选择 timeit、cProfile 或 tracemalloc;
  • 你会保留完整的样本和测量上下文;
  • 你把复杂度当作增长模型,而不是秒表;
  • 你会证明候选方案在正常情况和边界情况下都与参考实现等价;
  • 你会构建、检查、干净安装并对构件进行冒烟测试;
  • 你会发布经过测试的不可变字节,并且能够说出回滚路径;
  • 阅读摘要证据项目会打印预期的报告。

现在,你已经掌握了完整的 Python 循环:为问题建模、编写清晰的代码、验证边界、测试行为、谨慎选择并发方式、测量真实成本,并发布出其他人可以信任的成果。