课程进度 课程大纲 已发布 24/24 课
Python 基础
数据与集合
构建可靠的程序
使用对象建模
专业 Python
高级 Python
Python 语法会发送协议请求
特殊方法并不是什么魔法装饰。它们是具有明确约定的钩子。普通语法会向对象请求某种行为。
-
len(box)
请求
__len__ -
a == b
通过
__eq__协商 - for item in box 请求一个迭代器
-
box[0]
请求
__getitem__
只实现你的领域确实能够支持的协议。例如,__len__ 必须返回一个非负整数,而不能返回文本 "2"。
为值提供两副面孔,并进行公平的比较
repr(value) 是供开发者使用的稳定诊断形式。str(value) 是供用户阅读的形式。如果缺少 __str__,Python 会退回使用 __repr__。
class ReadingItem:
def __init__(self, title, minutes):
self.title = title
self.minutes = minutes
def __repr__(self):
return f"ReadingItem({self.title!r}, {self.minutes!r})"
def __str__(self):
return f"{self.title} ({self.minutes} min)"
!r 会让引号边界清晰可见。绝不要在 repr 中放入秘密信息;对象表示可能会出现在日志和回溯信息中。
相等性比较是一场协商。对于不支持的类型,应返回单例 NotImplemented,而不是 False,也不是异常 NotImplementedError:
def __eq__(self, other):
if type(other) is not type(self):
return NotImplemented
return (self.title, self.minutes) == (other.title, other.minutes)
- 询问左侧 你能与这种类型比较吗?
- NotImplemented 让 Python 询问另一侧
- 最终结果 相等性比较可能为 False;顺序比较可能报错
只有当领域中存在一种有意义的顺序时,才定义 <。用 (minutes, title) 表示阅读优先级可能是合理的;但说一个电话号码“小于”另一个电话号码通常没有意义。当不同场景需要不同的排序方式时,调用者应该传入排序键。
对于普通值,相等性应该保持自反性、对称性和传递性。应将参与比较的字段视为领域约定的一部分来选择,而不能仅仅因为这些字段恰好存在就使用它们。
只对稳定的相等性状态进行哈希
哈希容器要求遵守一条定律:
如果
a == b,那么hash(a) == hash(b)。
只对相等性比较所使用的稳定字段和类型进行哈希:
def __hash__(self):
return hash((type(self), self.title, self.minutes))
当对象用作字典键或集合成员时,这些字段不能发生变化。定义了值相等性但仍然可变的类应该是不可哈希的。如果定义了 __eq__ 却没有定义 __hash__,Python 通常会设置 __hash__ = None;请保留这个安全的默认行为。不要在字段相等性旁边恢复基于对象身份的哈希,因为这样可能导致相等的对象拥有不同的哈希值。
让容器回答问题,而不泄露内部容器
- len / bool 大小;零为假
- in 含义明确的成员检查
- iter 每次都返回一个新迭代器
- [0] 和 [1:] 元素和安全切片
如果没有 __bool__,Python 会使用 __len__:零为假,正数为真。__contains__ 为 in 提供支持。每次调用 __iter__ 都必须返回一个新迭代器,这样嵌套遍历就不会共享位置。__getitem__ 应该支持整数索引、负数索引、切片以及正常的 IndexError 行为。切片可以返回元组或一个新容器,但绝不能直接返回内部的可变列表。
构建一个符合 Python 风格的阅读队列
将下面这个完整的标准库项目保存为 reading_queue.py:
class ReadingItem:
__slots__ = ("_title", "_minutes", "_sealed")
def __init__(self, title, minutes):
title = title.strip()
if not title:
raise ValueError("title must not be empty")
if isinstance(minutes, bool) or not isinstance(minutes, int) or minutes < 0:
raise ValueError("minutes must be a non-negative integer")
object.__setattr__(self, "_title", title)
object.__setattr__(self, "_minutes", minutes)
object.__setattr__(self, "_sealed", True)
def __setattr__(self, name, value):
if getattr(self, "_sealed", False):
raise AttributeError("ReadingItem is immutable")
object.__setattr__(self, name, value)
@property
def title(self):
return self._title
@property
def minutes(self):
return self._minutes
def __repr__(self):
return f"ReadingItem({self.title!r}, {self.minutes!r})"
def __str__(self):
return f"{self.title} ({self.minutes} min)"
def __eq__(self, other):
if type(other) is not type(self):
return NotImplemented
return (self.title, self.minutes) == (other.title, other.minutes)
def __lt__(self, other):
if type(other) is not type(self):
return NotImplemented
return (self.minutes, self.title) < (other.minutes, other.title)
def __hash__(self):
return hash((type(self), self.title, self.minutes))
class ReadingQueue:
__hash__ = None
def __init__(self, items=()):
self._items = []
for item in items:
self.add(item)
def add(self, item):
if not isinstance(item, ReadingItem):
raise TypeError("item must be a ReadingItem")
self._items.append(item)
def __len__(self):
return len(self._items)
def __contains__(self, item):
return item in self._items
def __iter__(self):
return iter(tuple(self._items))
def __getitem__(self, index):
if isinstance(index, slice):
return tuple(self._items[index])
return self._items[index]
def main():
python = ReadingItem(" Python data model ", 30)
same = ReadingItem("Python data model", 30)
async_item = ReadingItem("Async I/O", 45)
assert repr(python) == "ReadingItem('Python data model', 30)"
assert python == same and hash(python) == hash(same)
assert len({python, same}) == 1
assert python != ("Python data model", 30)
assert sorted([async_item, python]) == [python, async_item]
queue = ReadingQueue([python, async_item])
left, right = iter(queue), iter(queue)
assert next(left) == next(right) == python
assert len(queue) == 2 and bool(queue) and same in queue
assert queue[-1] == async_item and queue[1:] == (async_item,)
try:
hash(queue)
except TypeError:
pass
else:
raise AssertionError("mutable queue was hashable")
print("Tests passed.")
print(f"Queue: {len(queue)} items")
for item in queue:
print(f"- {item}")
if __name__ == "__main__":
main()
输出:
Tests passed.
Queue: 2 items
- Python data model (30 min)
- Async I/O (45 min)
这个队列特意没有哈希,迭代时使用一个全新的快照迭代器,切片返回元组而不是其私有列表。
三个小任务与下一步
- 两副面孔。 构建
Bookmark(title, page),为它提供用于诊断的repr和面向读者的str;使用一个包含引号的标题进行测试。 - 比较握手。 让
Duration.__eq__和__lt__在遇到整数时返回NotImplemented。验证Duration(5) == 5为假,而Duration(5) < 5会引发TypeError。 - 独立书架。 基于私有存储支持长度、成员检查、两个同时使用的迭代器、负数索引和元组切片。
容易踩坑的地方:直接调用特殊方法会绕过公开语法所表达的意图,NotImplementedError 不是比较操作的哨兵值,凭空设计的顺序会歪曲领域含义,可变的哈希状态会破坏查找,而返回内部存储会泄露修改能力。
- 我能将公开语法对应到协议请求。
- 我能区分面向开发者的
repr和面向读者的str。 - 对于不支持的比较类型,我会返回
NotImplemented。 - 我只会为确实存在排序规则的领域定义顺序。
- 我会遵守相等性与哈希定律,或者让可变值保持不可哈希。
- 我能实现大小、真假判断、成员检查、全新迭代、索引和切片。
- 我的队列能通过所有断言,并且不会暴露内部存储列表。
- 我完成了三个小任务。
接下来,你将区分并发与并行,并根据不同工作能够安全执行的方式选择线程或进程。