元组与具名元组:把「第 3 个字段」变成「price」
写作约定:本文所有代码都在真实解释器 上运行过,输出是从终端直接复制的。除特别标注外,输出均来自 Python 3.14.7;涉及版本差异的地方会单独标出实测版本。凡是没能实测验证的说法,本文不写。
配套代码:
examples/01_sales_report.py------ 用具名元组做销售汇总examples/02_config.py------ 用具名元组做不可变配置verify/------ 每个结论的验证脚本,07_final_matrix.py可跨版本重跑
一、先把元组本身搞清楚
具名元组是元组的子类,所以先把元组那几个容易被误解的点确认清楚。
1.1 决定元组的是逗号,不是括号
python
t1 = (1)
t2 = (1,)
t3 = 1, 2
print("type((1)) ->", type(t1).__name__, "| (1) == 1 ->", t1 == 1)
print("type((1,)) ->", type(t2).__name__, "| len((1,)) ->", len(t2))
print("1, 2 ->", repr(t3))
输出:
text
type((1)) -> int | (1) == 1 -> True
type((1,)) -> tuple | len((1,)) -> 1
1, 2 -> (1, 2)
(1) 就是整数 1,括号只是分组符号。一个元素也要写逗号,这是新手最常见的 bug 来源:
python
def get_user():
return "alice", # 有逗号:返回元组
# return "alice" # 没逗号:返回字符串
1.2 「不可变」指的是元组这一层
元组本身不能增删改元素,但它引用的对象该变还是变:
python
inner = [1, 2]
t = (inner, "x")
t[0].append(3)
print("元组内容被间接修改后:", t)
try:
t[0] = [9]
except TypeError as e:
print("t[0] = [9] ->", type(e).__name__, ":", e)
text
元组内容被间接修改后: ([1, 2, 3], 'x')
t[0] = [9] -> TypeError : 'tuple' object does not support item assignment
1.3 可哈希是有条件的
元组能当字典键,前提是它里面的每个元素都可哈希:
python
print("{(1, 2): 'a'} ->", {(1, 2): "a"}[1, 2])
try:
{([1, 2],): "a"}
except TypeError as e:
print("含 list 的元组作键 ->", type(e).__name__, ":", e)
text
{(1, 2): 'a'} -> a
含 list 的元组作键 -> TypeError : cannot use 'tuple' as a dict key (unhashable type: 'list')
上面这条报错是 3.14.7 的原文。3.12.10 上同样的代码只说
TypeError: unhashable type: 'list'。报错格式会随版本变,别拿它做字符串匹配。
这个规则原样传递给了具名元组:字段里有 list/dict/set,整个具名元组就不可哈希(见 3.6 节)。
1.4 函数「返回多个值」其实是返回元组
python
def min_max(nums):
return min(nums), max(nums)
lo, hi = min_max([3, 1, 4, 1, 5]) # 解包
head, *tail = (1, 2, 3, 4) # 星号解包
print("min_max([3,1,4,1,5]) ->", (lo, hi))
print("head, *tail = (1,2,3,4) ->", head, tail)
text
min_max([3,1,4,1,5]) -> (1, 5)
head, *tail = (1,2,3,4) -> 1 [2, 3, 4]
记住这一点,后面就能理解为什么具名元组可以无缝用在 for a, b in pairs 这类地方------它就是元组。
二、具名元组解决什么问题
看一段很典型的代码:
python
def total(rows):
return sum(r[2] * r[3] for r in rows) # r[2]、r[3] 是什么?
r[2] * r[3] 能跑,但没人知道它算的是哪两个字段。三个月后回来改,只能去翻上游的解析代码。具名元组把「位置」变成「名字」:
python
class Sale(NamedTuple):
order_id: str
sku: str
qty: int
unit_price: float
def total(rows):
return sum(r.qty * r.unit_price for r in rows)
同时它还是元组:长度、索引、解包、比较、字典键,一样都没丢。
Python 提供了两条路:
collections.namedtuple |
typing.NamedTuple |
|
|---|---|---|
| 写法 | 工厂函数 + 字段名 | class 语法 + 类型注解 |
| 加方法 | 需要再继承一次 | 直接写在 class 里 |
| 支持类型注解 | 否 | 是 |
| 默认值 | defaults=(...) 按尾部对齐 |
字段后写 = value |
| 保留 docstring | 否(工厂函数会给一个自动文档) | 是 |
| 能力上限 | 低 | 高(推荐) |
两者生成的类都是 tuple 的子类,运行期结构完全一样,区别只在于「怎么写」和「能写多复杂」。
三、collections.namedtuple 逐个 API 拆解
3.1 定义与访问
python
from collections import namedtuple
Point = namedtuple("Point", ["x", "y"])
p = Point(3, 4)
print("p ->", p)
print("repr(p) ->", repr(p))
print("p.x, p.y ->", p.x, p.y)
print("p[0], p[1] ->", p[0], p[1])
print("tuple(p) ->", tuple(p))
print("isinstance(p, tuple) ->", isinstance(p, tuple))
print("p == (3, 4) ->", p == (3, 4))
print("len(p) ->", len(p))
text
p -> Point(x=3, y=4)
repr(p) -> Point(x=3, y=4)
p.x, p.y -> 3 4
p[0], p[1] -> 3 4
tuple(p) -> (3, 4)
isinstance(p, tuple) -> True
p == (3, 4) -> True
len(p) -> 2
注意最后几行:具名元组和普通元组相等 ,Point(3, 4) == (3, 4) 是 True,而且 hash(Point(3, 4)) == hash((3, 4))。这是它和 dataclass 的关键区别(第十一节有对照)。
字段名既可以用列表,也可以用一个字符串(空格或逗号分隔):
python
P2 = namedtuple("P2", "x y")
P3 = namedtuple("P3", "x,y")
print("P2._fields ->", P2._fields, "| P3._fields ->", P3._fields)
text
P2._fields -> ('x', 'y') | P3._fields -> ('x', 'y')
字符串形式内部是 field_names.replace(',', ' ').split(),所以 "x y"、"x,y"、"x, y" 都可以。字段多的时候用列表更清楚。
3.2 _fields 与 defaults
python
Point3 = namedtuple("Point3", "x y z", defaults=(0, 0))
print("Point3._field_defaults ->", Point3._field_defaults)
print("Point3(1) ->", Point3(1))
text
Point3._field_defaults -> {'y': 0, 'z': 0}
Point3(1) -> Point3(x=1, y=0, z=0)
defaults 是从右往左 对齐的,所以 defaults=(0, 0) 落在 y 和 z 上。参数校验的报错信息也很直接:
python
try:
Point3(1, 2, 3, 4)
except TypeError as e:
print("多余位置参数 ->", type(e).__name__, ":", e)
try:
Point3(1, 2, 3, z=9)
except TypeError as e:
print("同时给 z 位置和关键字 ->", type(e).__name__, ":", e)
try:
Point3(1, 2, 3, extra=1)
except TypeError as e:
print("未知关键字 ->", type(e).__name__, ":", e)
text
多余位置参数 -> TypeError : Point3.__new__() takes from 2 to 4 positional arguments but 5 were given
同时给 z 位置和关键字 -> TypeError : Point3.__new__() got multiple values for argument 'z'
未知关键字 -> TypeError : Point3.__new__() got an unexpected keyword argument 'extra'
3.3 _make:从任意序列构造(带长度校验)
python
print("p._asdict() ->", p._asdict(), "| type ->", type(p._asdict()).__name__)
print("Point._make([7, 8]) ->", Point._make([7, 8]))
print("Point(*[7, 8]) 等价 ->", Point(*[7, 8]))
text
p._asdict() -> {'x': 3, 'y': 4} | type -> dict
Point._make([7, 8]) -> Point(x=7, y=8)
Point(*[7, 8]) 等价 -> Point(x=7, y=8)
_make 比 Point(*seq) 多一层长度校验:长度不对会抛 TypeError。解析 CSV 行的时候,用 _make 能把「表头字段数和数据列数不一致」这类错误当场拦住。
3.4 _replace:派生新实例,不是原地修改
python
p2 = p._replace(x=10)
print("p._replace(x=10) ->", p2, "| 原对象不变 ->", p)
text
p._replace(x=10) -> Point(x=10, y=4) | 原对象不变 -> Point(x=3, y=4)
_replace 内部走的是 _make,所以未知字段会报错。这个错误类型在 3.13 变了(实测):
| 版本 | p._replace(z=1) |
|---|---|
| 3.12.10 | ValueError: Got unexpected field names: ['z'] |
| 3.14.7 | TypeError: Got unexpected field names: ['z'] |
如果你写 except ValueError 来兜住拼错的字段名,升级到 3.13+ 会漏掉。
_replace 是浅拷贝,可变字段仍然共享:
python
Box = namedtuple("Box", "items name")
b = Box([1, 2], "a")
c = b._replace(name="b")
c.items.append(3)
print("b ->", b, "| c ->", c, "| b.items is c.items ->", b.items is c.items)
text
b -> Box(items=[1, 2, 3], name='a') | c -> Box(items=[1, 2, 3], name='b') | b.items is c.items -> True
只改了 name,b 的 items 也跟着变成 [1, 2, 3]。要真正隔离就自己 copy.deepcopy。
3.5 rename:字段名非法时的兜底
python
try:
namedtuple("Bad", ["a", "class", "a", "1x"])
except ValueError as e:
print("默认 rename=False ->", type(e).__name__, ":", e)
Bad = namedtuple("Bad", ["a", "class", "a", "1x"], rename=True)
print("rename=True ->", Bad._fields)
print("Bad(1,2,3,4) ->", Bad(1, 2, 3, 4))
text
默认 rename=False -> ValueError : Type names and field names cannot be a keyword: 'class'
rename=True -> ('a', '_1', '_2', '_3')
Bad(1,2,3,4) -> Bad(a=1, _1=2, _2=3, _3=4)
非法字段(关键字、非标识符、以下划线开头、重复)会被替换成 _下标。注意 rename 只按真值判断 ,传 "yes"、1 都会生效:
text
rename='yes' -> OK, _fields=('a', '_1')
rename=1 -> OK, _fields=('a', '_1')
rename=0 -> ValueError: Type names and field names cannot be a keyword: 'class'
建议显式传 True / False,别传字符串。
3.6 哈希、排序、去重
python
print("hash(Point(3,4)) == hash((3,4)) ->", hash(Point(3, 4)) == hash((3, 4)))
print("{Point(3,4): 'v'} 可用 ->", {Point(3, 4): "v"}[(3, 4)])
print("len({Point(1,2), Point(1,2), (1,2)}) ->", len({Point(1, 2), Point(1, 2), (1, 2)}))
print("sorted ->", sorted([Point(2, 0), Point(1, 9)]))
text
hash(Point(3,4)) == hash((3,4)) -> True
{Point(3,4): 'v'} 可用 -> v
len({Point(1,2), Point(1,2), (1,2)}) -> 1
sorted -> [Point(x=1, y=9), Point(x=2, y=0)]
字段里放可变对象就不可哈希,但仍然可以比较相等:
text
字段含 list 时 hash -> TypeError : unhashable type: 'list'
字段含 list 时作字典键 -> TypeError : cannot use 'Shared' as a dict key (unhashable type: 'list')
但仍可用 == 比较 -> True
(这两条报错是 3.14.7 的原文,3.12.10 只打印 unhashable type: 'list'。)
四、typing.NamedTuple:带类型和方法
4.1 基本类语法
python
from typing import NamedTuple
class Point(NamedTuple):
"""二维平面上的一个点。"""
x: int
y: int = 0 # 默认值
def dist_to_origin(self) -> float:
return (self.x**2 + self.y**2) ** 0.5
p = Point(3, 4)
print("p ->", p)
print("type(p) ->", type(p).__name__, "| isinstance tuple ->", isinstance(p, tuple))
print("p.dist_to_origin() ->", p.dist_to_origin())
print("__annotations__ ->", Point.__annotations__)
print("_fields ->", Point._fields)
print("_field_defaults ->", Point._field_defaults)
print("Point(5) ->", Point(5))
print("实例有没有 __dict__ ->", hasattr(p, "__dict__"), "| __slots__ ->", Point.__slots__)
print("Point.__doc__ ->", repr(Point.__doc__))
print("Point.__mro__[:3] ->", [c.__name__ for c in Point.__mro__[:3]])
text
p -> Point(x=3, y=4)
type(p) -> Point | isinstance tuple -> True
p.dist_to_origin() -> 5.0
__annotations__ -> {'x': <class 'int'>, 'y': <class 'int'>}
_fields -> ('x', 'y')
_field_defaults -> {'y': 0}
Point(5) -> Point(x=5, y=0)
实例有没有 __dict__ -> False | __slots__ -> ()
Point.__doc__ -> '二维平面上的一个点。'
Point.__mro__[:3] -> ['Point', 'tuple', 'object']
三个信息量很大的点:
__slots__是(),实例没有__dict__------所以每个实例的内存开销和同长度元组一样(第 8 节实测)。__mro__里第二项就是tuple,它真的是元组。- docstring 保留 ,这对
help()和文档生成很重要。
方法、@property、普通类属性都可以写在类里,非注解的类属性不会被当成字段:
python
class Circle(NamedTuple):
radius: float
pi = 3.141592653589793 # 普通类属性,不是字段
@property
def area(self) -> float:
return Circle.pi * self.radius**2
def scaled(self, k: float) -> "Circle":
return self._replace(radius=self.radius * k)
c = Circle(2.0)
print("_fields ->", Circle._fields)
print("Circle.pi ->", Circle.pi, "| c.area ->", c.area)
print("c.scaled(3) ->", c.scaled(3))
text
_fields -> ('radius',)
Circle.pi -> 3.141592653589793 | c.area -> 12.566370614359172
c.scaled(3) -> Circle(radius=6.0)
4.2 不能覆盖那几个特殊属性
text
覆盖 __new__ -> AttributeError: Cannot overwrite NamedTuple attribute __new__
覆盖 __init__ -> AttributeError: Cannot overwrite NamedTuple attribute __init__
显式 __slots__ -> AttributeError: Cannot overwrite NamedTuple attribute __slots__
想自定义构造逻辑,正确做法是写类方法:
python
class Sale(NamedTuple):
qty: int
unit_price: float
@classmethod
def from_row(cls, qty, price):
return cls(int(qty), float(price))
4.3 默认值陷阱:可变默认值会被共享
typing.NamedTuple 的默认值只求值一次 ,而且不会像 dataclass 那样拦截可变默认值(后者会直接报 ValueError: mutable default <class 'list'> for field tags is not allowed: use default_factory):
python
class Cfg(NamedTuple):
name: str
tags: list = []
a, b = Cfg("a"), Cfg("b")
a.tags.append("x")
print("b.tags ->", b.tags, "(两个实例共享同一个默认 list)")
print("默认值对象的身份相同 ->", Cfg("a").tags is Cfg("b").tags)
text
b.tags -> ['x'] (两个实例共享同一个默认 list)
默认值对象的身份相同 -> True
只改 a,b 也被污染了。字段默认值只用不可变对象 ,需要 list 就传 None 再在方法里兜底,或者干脆改用 dataclasses.field(default_factory=list)。
4.4 函数式写法与泛型
python
Pair = NamedTuple("Pair", [("key", str), ("value", int)])
print("Pair('a', 1) ->", Pair("a", 1), "| _fields ->", Pair._fields)
print("Pair.__annotations__ ->", Pair.__annotations__)
text
Pair('a', 1) -> Pair(key='a', value=1) | _fields -> ('key', 'value')
Pair.__annotations__ -> {'key': <class 'str'>, 'value': <class 'int'>}
泛型也支持:
python
from typing import Generic, TypeVar
T = TypeVar("T")
class Page(NamedTuple, Generic[T]):
items: list
total: int
def is_empty(self) -> bool:
return self.total == 0
pg = Page([1, 2, 3], 3)
print("pg ->", pg, "| is_empty ->", pg.is_empty())
text
pg -> Page(items=[1, 2, 3], total=3) | is_empty -> False
注意:
NamedTuple("X", a=int, b=str)这种关键字写法在 3.13+ 已弃用 (3.12 无告警,3.13/3.14 实测会抛DeprecationWarning,提示 3.15 将移除):
textDeprecationWarning: Creating NamedTuple classes using keyword arguments is deprecated and will be disallowed in Python 3.15. Use the class-based or functional syntax instead.新代码用 class 语法或
[("a", int), ...]列表形式。
五、模式匹配与 copy.replace()(新版本专属)
5.1 结构化模式匹配(3.10+)
namedtuple 生成的类带 __match_args__,所以可以按字段名匹配:
python
class Pt(NamedTuple):
x: int
y: int
def describe(p):
match p:
case Pt(x=0, y=0):
return "原点"
case Pt(x=0, y=y):
return f"y 轴上, y={y}"
case Pt(x=x, y=0):
return f"x 轴上, x={x}"
case Pt(x, y) if x == y:
return f"对角线上 ({x}, {y})"
case Pt(x, y):
return f"普通点 ({x}, {y})"
case _:
return "不是点(没有匹配到 Pt 的模式)"
for q in [Pt(0,0), Pt(0,5), Pt(5,0), Pt(3,3), Pt(1,2), (1,2)]:
print(f"{q!r:>14} -> {describe(q)}")
text
Pt(x=0, y=0) -> 原点
Pt(x=0, y=5) -> y 轴上, y=5
Pt(x=5, y=0) -> x 轴上, x=5
Pt(x=3, y=3) -> 对角线上 (3, 3)
Pt(x=1, y=2) -> 普通点 (1, 2)
(1, 2) -> 不是点(没有匹配到 Pt 的模式)
最后一行值得注意:Pt(1, 2) == (1, 2) 是 True,但类模式要求 isinstance 成立 ,所以普通元组不会匹配 case Pt(...)。想同时接受普通元组,就再加一条 case (x, y):。
case Pt(x, y) 这种位置模式依赖 __match_args__,它的存在性有版本差异(第十节表格):3.9 没有,3.10+ 有。
5.2 __replace__ 与 copy.replace()(3.13+)
python
import copy
print("copy.replace(Pt(1,2), x=9) ->", copy.replace(Pt(1, 2), x=9))
print("Pt(1,2).__replace__(y=9) ->", Pt(1, 2).__replace__(y=9))
try:
copy.replace(Pt(1, 2), z=9)
except TypeError as e:
print("copy.replace 未知字段 ->", type(e).__name__, ":", e)
text
copy.replace(Pt(1,2), x=9) -> Pt(x=9, y=2)
Pt(1,2).__replace__(y=9) -> Pt(x=1, y=9)
copy.replace 未知字段 -> TypeError : Got unexpected field names: ['z']
3.12 实测没有 copy.replace,也没有 __replace__,只能用 _replace。_replace 在所有版本都能用,所以写兼容库时优先用它。
六、实战一:销售明细汇总
完整代码见 examples/01_sales_report.py,下面是真实输出。
python
class Sale(NamedTuple):
"""一条销售明细。order_id 是订单号,qty 是数量,unit_price 是单价。"""
order_id: str
sku: str
qty: int
unit_price: float
@property
def amount(self) -> float:
return round(self.qty * self.unit_price, 2)
RAW = [
("order_id", "sku", "qty", "unit_price"),
("A-1001", "KB-01", 2, 199.0),
# ...省略
]
header, *rows = RAW
assert tuple(header) == Sale._fields, "表头和字段不一致,先改字段定义"
sales = [Sale._make(r) for r in rows]
text
== 明细 ==
A-1001 KB-01 2 × 199.00 = 398.00
A-1001 MS-02 1 × 89.50 = 89.50
A-1002 KB-01 1 × 199.00 = 199.00
A-1002 HD-03 3 × 45.00 = 135.00
A-1003 MS-02 4 × 89.50 = 358.00
A-1003 HD-03 2 × 45.00 = 90.00
_replace 派生打折记录,原记录不动:
text
原记录 -> Sale(order_id='A-1001', sku='KB-01', qty=2, unit_price=199.0) 金额 398.0
折后记录-> Sale(order_id='A-1001', sku='KB-01', qty=2, unit_price=159.2) 金额 318.4
原记录没有被改动 -> True
分组汇总:
text
SKU 订单数 件数 金额
HD-03 2 5 225.00
KB-01 2 3 597.00
MS-02 2 5 447.50
转 JSON 时必须显式 _asdict(),否则字段名会丢(见 9.4 节):
json
{
"by_sku": [
{"sku": "HD-03", "orders": 2, "qty": 5, "revenue": 225.0},
{"sku": "KB-01", "orders": 2, "qty": 3, "revenue": 597.0},
{"sku": "MS-02", "orders": 2, "qty": 5, "revenue": 447.5}
],
"total_revenue": 1269.5
}
这个例子里具名元组承担了三个角色:表结构声明 (_fields 校验表头)、只读记录 (打折靠 _replace)、可哈希的分组键。
七、实战二:不可变配置对象
完整代码见 examples/02_config.py。
python
Timeouts = namedtuple("Timeouts", "connect read write", defaults=(5.0, 30.0, 30.0))
print("Timeouts() ->", Timeouts())
print("Timeouts(1.5) ->", Timeouts(1.5))
print("Timeouts(connect=2.0) ->", Timeouts(connect=2.0))
text
Timeouts() -> Timeouts(connect=5.0, read=30.0, write=30.0)
Timeouts(1.5) -> Timeouts(connect=1.5, read=30.0, write=30.0)
Timeouts(connect=2.0) -> Timeouts(connect=2.0, read=30.0, write=30.0)
派生配置:
python
base = Timeouts(connect=1.0)
slow_net = base._replace(connect=10.0, read=60.0)
print("base ->", base)
print("slow_net ->", slow_net)
print("base 未被修改 ->", base.connect == 1.0)
text
base -> Timeouts(connect=1.0, read=30.0, write=30.0)
slow_net -> Timeouts(connect=10.0, read=60.0, write=30.0)
base 未被修改 -> True
配置对象可以安全地当默认参数、当字典键:
python
def fetch(url, timeouts=Timeouts()):
return f"GET {url} (connect={timeouts.connect}s, read={timeouts.read}s)"
print(fetch("https://example.com"))
print(fetch("https://slow.example.com", timeouts=slow_net))
cache = {Timeouts(): "默认配置的结果", slow_net: "慢网络的结果"}
print("用配置对象作缓存键 ->", cache[Timeouts(5.0, 30.0, 30.0)])
text
GET https://example.com (connect=5.0s, read=30.0s)
GET https://slow.example.com (connect=10.0s, read=60.0s)
用配置对象作缓存键 -> 默认配置的结果
(因为它可哈希且不可变,不会出现「传进去被下游改掉」的问题。)
八、性能与内存:实测数据
环境:macOS / Apple Silicon,同一份脚本 verify/06_bench.py,timeit.repeat(..., number=1_000_000, repeat=5) 取最小值。不同机器和版本会有波动,看量级关系即可。
8.1 构造与访问(ns/次)
| 操作 | 3.12.10 | 3.14.7 |
|---|---|---|
tuple((x, y)) |
15.8 | 17.9 |
Point(x, y)(具名元组) |
117.2 | 119.7 |
DPoint(x, y)(frozen dataclass) |
181.2 | 147.9 |
MPoint(x, y)(普通 dataclass) |
76.8 | 50.2 |
{"x": x, "y": y}(dict) |
45.1 | 42.6 |
p.x(具名元组属性) |
10.0 | 10.7 |
dp.x(frozen dataclass 属性) |
5.0 | 6.0 |
d["x"](dict 取值) |
9.1 | 9.3 |
tup[0](元组索引) |
6.1 | 6.5 |
a, b = p(解包) |
19.2 | 17.3 |
p._asdict() |
155.0 | 157.0 |
json.dumps(p) |
702.2 | 724.5 |
json.dumps(p._asdict()) |
885.3 | 986.8 |
三个结论,都有点反直觉:
- 构造具名元组比构造裸元组慢约 7 倍 (117.2 / 15.8 ≈ 7.4)------因为
tuple.__new__走的是 C,而具名元组生成的__new__是一个 Python 层 lambda。如果你的热路径是「每秒构造几百万条记录」,这一点要实测再决定。 - 属性访问并不比
d["x"]快(两个版本上都是大约慢 1 ns)。所以「用具名元组替代 dict 更快」这个说法在近几个版本上不成立,别把它当性能优化手段。 json.dumps(p)比json.dumps(p._asdict())快------因为前者按数组处理,只输出[3, 4],省掉了字段名。代价是丢了语义。
8.2 内存:这才是具名元组的杀手锏
python
sys.getsizeof(...) # 只算容器对象本身
| 字段数 | tuple | namedtuple | slots dataclass | 普通 dataclass | dict |
|---|---|---|---|---|---|
| 2 | 64 | 64 | 48 | 48 | 184 |
| 3 | 72 | 72 | 56 | 48 | 184 |
| 5 | 88 | 88 | 72 | 48 | 184 |
| 8 | 112 | 112 | 96 | 48 | 272 |
(Python 3.14.7 实测。3.12.10 上同一张表的 tuple / namedtuple 两列是 56 / 64 / 80 / 104,每行比 3.14 少 8 字节;slots dataclass、dataclass、dict 三列在两个版本上完全相同。也就是说,跨版本唯一稳定且关键的规律是:namedtuple 和 tuple 逐格相等。)
结论很清楚:
- 具名元组 = 裸元组的内存 :字段没有存成 dict,就是 tuple 里那几个槽位。换成
dict大约多占 3 倍。 - 普通 dataclass 反而"看起来"小 (48 B),因为字段值存在
__dict__里,而__dict__是另一个对象,getsizeof没算进去。算上字典就大了:tracemalloc实测 10 万个两字段实例,dict 每实例约 192 B,普通 dataclass 约 96 B,dataclass(slots=True)约 56 B,具名元组约 80 B,裸元组约 72 B。 - 想让 dataclass 在内存上超过具名元组,得加
slots=True。
(tracemalloc 的数字包含分配器账目,所以具名元组每实例比裸元组多记了约 8 B;以 sys.getsizeof 为准,两者对象大小完全相同。)
九、坑清单(全部有实测)
9.1 字段名规则
python
for names in (["_x"], ["x", "_y"], ["x", "x"], ["class"]):
try:
namedtuple("C", names)
print(names, "-> OK")
except ValueError as e:
print(names, "-> ValueError:", e)
print(["self"], "-> OK _fields=" + str(namedtuple("C", ["self"])._fields))
text
['_x'] -> ValueError: Field names cannot start with an underscore: '_x'
['x', '_y'] -> ValueError: Field names cannot start with an underscore: '_y'
['x', 'x'] -> ValueError: Encountered duplicate field name: 'x'
['class'] -> ValueError: Type names and field names cannot be a keyword: 'class'
['self'] -> OK _fields=('self',)
以下划线开头是被禁的(避免和 _fields、_make、_replace、_asdict 这些内部 API 撞名);关键字和重复名也不行。self 这种名字反而合法------虽然不建议。
9.2 字段名会覆盖 tuple 自带的 count / index
python
T = namedtuple("T", "count index")
t = T(10, 20)
print("T(10,20) ->", t, "| t.count ->", t.count, "| t.index ->", t.index)
try:
t.count(10)
except TypeError as e:
print("t.count(10) ->", type(e).__name__, ":", e)
print("tuple.count 被字段访问器覆盖了吗 ->", type(t).count is not tuple.count)
text
T(10,20) -> T(count=10, index=20) | t.count -> 10 | t.index -> 20
t.count(10) -> TypeError : 'int' object is not callable
tuple.count 被字段访问器覆盖了吗 -> True
字段访问器把 tuple.count / tuple.index 顶掉了,t.count(10) 直接崩。避开这两个名字。
9.3 给 _fields 赋值不会改变行为,只会让它说谎
_fields 只是一个普通类属性,可以赋值 ,但 __new__、__repr__、_replace 用的是类创建时就固定下来的字段名,只有 _asdict() 会去读 self._fields:
python
Q = namedtuple("Q", "x y")
Q._fields = ("a", "b")
q = Q(1, 2)
print("repr(q) ->", repr(q))
print("q._asdict() ->", q._asdict())
text
repr(q) -> Q(x=1, y=2)
q._asdict() -> {'a': 1, 'b': 2}
repr 和 _asdict() 互相矛盾,_replace(a=...) 还会报错。3.7 到 3.14 实测行为一致(3.7 的 _asdict 是 OrderedDict([('a', 1), ('b', 2)]))。结论:永远不要给 _fields 赋值。
9.4 JSON 序列化会丢掉字段名;dict() 也不是转字典
python
import json
print("json.dumps(p) ->", json.dumps(p))
print("json.dumps(p._asdict()) ->", json.dumps(p._asdict()))
text
json.dumps(p) -> [3, 4]
json.dumps(p._asdict()) -> {"x": 3, "y": 4}
因为它是元组,json 按数组序列化。别指望用 default= 钩子救回来------钩子只对「无法序列化」的对象调用,而元组是能被序列化的:
python
json.dumps(p, default=lambda o: o._asdict())
text
[3, 4]
default 压根没被调用。要保留字段名就只能先 _asdict() 再序列化。
同样地,dict(p) 不是 转字典,它把元组的每个元素当成 (键, 值):
python
P = namedtuple("P", "x y")
try:
dict(P(1, 2))
except TypeError as e:
print("dict(P(1, 2)) ->", type(e).__name__, ":", e)
print("dict(P(('a',1),('b',2))) ->", dict(P(("a", 1), ("b", 2))))
print("P(('a',1),('b',2))._asdict() ->", P(("a", 1), ("b", 2))._asdict())
text
dict(P(1, 2)) -> TypeError : object is not iterable
dict(P(("a",1),("b",2))) -> {'a': 1, 'b': 2}
P(("a",1),("b",2))._asdict() -> {'x': ('a', 1), 'y': ('b', 2)}
要转字典永远用 _asdict()。
9.5 pickle 的真实规则
流传的说法是「函数里定义的具名元组不能 pickle」------不准确 。pickle 通过 __module__ + __qualname__ 回查类对象,所以规则是「模块命名空间里有没有这个名字,且指向同一个类」:
python
def make_same():
return namedtuple("Same", "x y")
def make_other():
return namedtuple("Other", "x y")
def make_ghost():
return namedtuple("Ghost", "x y")(1, 2)
Same = make_same() # A:类名 Same,绑定的名字也叫 Same
Renamed = make_other() # B:类名 Other,绑定的名字却叫 Renamed
ghost = make_ghost() # C:类只在函数内存在,模块里查不到
for label, obj in [("A 绑定同名(模块里能查到 Same) ", Same(1, 2)),
("B 绑定异名(类叫 Other,绑定为 Renamed)", Renamed(1, 2)),
("C 从未绑定(只在函数内存在) ", ghost)]:
try:
pickle.dumps(obj)
print(f"{label} -> 成功")
except Exception as e:
print(f"{label} -> {type(e).__name__}: {e}")
text
A 绑定同名(模块里能查到 Same) -> 成功
B 绑定异名(类叫 Other,绑定为 Renamed) -> PicklingError: Can't pickle <class '__main__.Other'>: it's not found as __main__.Other
C 从未绑定(只在函数内存在) -> PicklingError: Can't pickle <class '__main__.Ghost'>: it's not found as __main__.Ghost
所以:把具名元组定义在模块顶层 。放在函数里创建、只通过返回值暴露,就会在跨进程(多进程 / 缓存 / 消息队列)时炸掉。module= 参数也救不了,它只是改 __module__,名字没绑定照样失败。
顺带一提,copy.copy 和 copy.deepcopy 行为不同:
python
import copy
plain = (1, 2) # 必须用变量,否则常量折叠会让 is 比较失去意义
named = Point(3, 4)
box = Box([1], "a")
print("copy.copy(普通元组) is 原对象 ->", copy.copy(plain) is plain, " (元组不可变,原样返回)")
print("copy.copy(具名元组) is 原对象 ->", copy.copy(named) is named, " (走 __getnewargs__,new 一个新对象)")
print("copy.deepcopy 是否共享可变字段 ->", copy.deepcopy(box).items is box.items)
text
copy.copy(普通元组) is 原对象 -> True (元组不可变,原样返回)
copy.copy(具名元组) is 原对象 -> False (走 __getnewargs__,new 一个新对象)
copy.deepcopy 是否共享可变字段 -> False
9.6 子类化时字段怎么变(两种写法不一样)
collections.namedtuple 的子类字段会拼接(父类的字段在前):
python
import sys
Base = namedtuple("Base", "a b")
class Child(Base):
__slots__ = ()
class Child2(Base): # 不写 __slots__
pass
print("Child._fields ->", Child._fields)
print("Child 实例有 __dict__ 吗 ->", hasattr(Child(1, 2), "__dict__"))
print("Child2 实例有 __dict__ 吗 ->", hasattr(Child2(1, 2), "__dict__"))
print("两者实例大小 ->", sys.getsizeof(Child(1, 2)), sys.getsizeof(Child2(1, 2)))
text
Child._fields -> ('a', 'b')
Child 实例有 __dict__ 吗 -> False
Child2 实例有 __dict__ 吗 -> True
两者实例大小 -> 64 80
这正是「内存和元组一样」的前提:子类里写 __slots__ = () 。不写的话,子类会给每个实例加一个 __dict__(上面实测 64 B → 80 B,还没算字典本身),具名元组的零开销优势就没了。注意 Base.__slots__ 虽然是 (),但它只是被继承,不会阻止子类创建 __dict__。
但 typing.NamedTuple 的子类不会新增字段,注解被静默忽略:
python
class A(NamedTuple):
a: int
class B(A):
b: int # 不会变成字段
print("B._fields ->", B._fields) # ('a',)
print("B(1) ->", B(1)) # B(a=1)
print("hasattr(B(1), 'b') ->", hasattr(B(1), "b")) # False
text
B._fields -> ('a',) | B(1) -> B(a=1) | hasattr b -> False
3.8 / 3.12 / 3.14 实测一致。想要新字段,就重新定义一个完整的 NamedTuple。
9.7 两个具名元组做多重继承:不报错,但字段会丢
python
class A(NamedTuple):
a: int
class B(NamedTuple):
b: int
class C(A, B):
pass
text
C.__mro__ -> ['C', 'A', 'B', 'tuple', 'object']
C._fields -> ('a',) # B 的 b 完全丢失
C(1) -> C(a=1)
C(1, 2) -> TypeError : A.__new__() takes 2 positional arguments but 3 were given
C(a=1,b=2) -> TypeError : A.__new__() got an unexpected keyword argument 'b'
(以上为 3.14.7 输出;3.12.10 实测 C._fields 同样是 ('a',)、C(1) 同样得到 C(a=1),只是 TypeError 的措辞不同。)
静默丢掉 B 的字段,只在构造时才报错。别把两个具名元组混在一起继承。
9.8 其他小坑
typing.NamedTuple里写__slots__ = ()会报AttributeError(它自己已经设好了)。_asdict()每次返回新字典 ,改它不影响实例(实测:改了之后p仍然是Point(x=3, y=4))。_replace()返回新对象,但原对象和它共享可变字段(3.4 节)。- 字段默认值用可变对象会被所有实例共享(4.3 节)。
十、版本差异速查表(3.7 → 3.14 实测)
用 verify/07_final_matrix.py 在 8 个解释器上跑出来的原始结果:
| 特性 | 3.7 | 3.8 | 3.9 | 3.10 | 3.11 | 3.12 | 3.13 | 3.14 |
|---|---|---|---|---|---|---|---|---|
_asdict() 返回类型 |
OrderedDict |
dict |
dict |
dict |
dict |
dict |
dict |
dict |
_field_defaults / defaults= |
有 | 有 | 有 | 有 | 有 | 有 | 有 | 有 |
__match_args__(模式匹配) |
无 | 无 | 无 | 有 | 有 | 有 | 有 | 有 |
__replace__ / copy.replace() |
无 | 无 | 无 | 无 | 无 | 无 | 有 | 有 |
_replace() 未知字段的异常 |
ValueError |
ValueError |
ValueError |
ValueError |
ValueError |
ValueError |
TypeError |
TypeError |
NamedTuple("X", a=int) |
可用 | 可用 | 可用 | 可用 | 可用 | 可用 | 弃用告警 | 弃用告警 |
| 类语法保留 docstring | 是 | 是 | 是 | 是 | 是 | 是 | 是 | 是 |
(3.13/3.14 的告警原文:Creating NamedTuple classes using keyword arguments is deprecated and will be disallowed in Python 3.15.)
如果库要兼容 3.9 及更早:不要用模式匹配的类模式,不要用 copy.replace,_replace 的错误处理要同时兼容 ValueError 和 TypeError。
十一、怎么选:一张决策表
| 需求 | 推荐 |
|---|---|
| 就是一组固定长度的值,要快、要省内存 | tuple |
| 字段名固定、需要按名字访问、要能当字典键、要能解包 | namedtuple / NamedTuple |
字段名固定,但需要可变、需要 default_factory、需要继承体系 |
@dataclass |
| 字段多且不固定、需要动态增删、与 JSON 直接对应 | dict |
| 需要运行期类型校验 / 从外部数据反序列化 | pydantic.BaseModel 等 |
再补几条实际判断标准:
- 要放进 set / 当 dict 键 / 做缓存 key → 具名元组(dataclass 默认不可哈希,需要
frozen=True)。 - 要和普通元组互操作 (比如和已有接口对接、
==比较、for a, b in ...)→ 具名元组。dataclass实例== (1, 2)是False,也不支持len()和下标。 - 字段有几十个、还需要默认值和校验 →
dataclass+ 类型检查工具。 - 构造频率极高 (每秒百万级)→ 实测:具名元组构造比裸元组慢约 7 倍,这种场景考虑裸元组或
slotsdataclass。
具名元组 vs dataclass 的语义差异,实测对照:
python
from dataclasses import dataclass
P = namedtuple("P", "x y")
@dataclass(frozen=True)
class DP:
x: int
y: int
p, dp = P(1, 2), DP(1, 2)
print("P(1,2) == (1,2) ->", p == (1, 2), " # namedtuple")
print("DP(1,2) == (1,2) ->", dp == (1, 2), " # frozen dataclass")
print("hash(P(1,2)) == hash((1,2)) ->", hash(p) == hash((1, 2)))
print("P 有 __len__ ->", len(p), "| DP 有 __len__ ->", hasattr(dp, "__len__"))
print("P 可索引 p[0] ->", p[0], "| DP 可索引 ->", end=" ")
try:
dp[0]
except TypeError as e:
print(f"TypeError: {e}")
text
P(1,2) == (1,2) -> True # namedtuple
DP(1,2) == (1,2) -> False # frozen dataclass
hash(P(1,2)) == hash((1,2)) -> True
P 有 __len__ -> 2 | DP 有 __len__ -> False
P 可索引 p[0] -> 1 | DP 可索引 -> TypeError: 'DP' object is not subscriptable
十二、速查表
python
from collections import namedtuple
from typing import NamedTuple
# ---- 定义 ----
Point = namedtuple("Point", "x y") # 字符串
Point = namedtuple("Point", ["x", "y"]) # 列表
Point = namedtuple("Point", "x y", defaults=(0,)) # 默认值(从右往左)
Point = namedtuple("Point", ["class", "y"], rename=True) # 非法名自动改 _下标
class Point(NamedTuple): # 推荐:带类型
x: int
y: int = 0
def dist(self) -> float: ...
# ---- 构造 ----
Point(3, 4) # 位置
Point(x=3, y=4) # 关键字
Point._make(seq) # 从任意序列(带长度校验)
Point(**d) # 从字典
# ---- 访问 ----
p.x / p[0] / x, y = p / tuple(p)
# ---- 元信息 ----
Point._fields # ('x', 'y')
Point._field_defaults # {'y': 0}
Point.__annotations__ # {'x': int, 'y': int}
Point.__match_args__ # ('x', 'y') (3.10+)
# ---- 转换 ----
p._asdict() # dict(3.8+;3.7 及以前是 OrderedDict)
p._replace(x=10) # 派生新实例(浅拷贝)
copy.replace(p, x=10) # 3.13+
Point(**p._asdict()) # dict 转回来
json.dumps(p._asdict()) # 要 JSON 对象就显式转;default= 钩子无效
# dict(p) 不是转字典!它把每个元素当 (键, 值) 解析,请用 p._asdict()
# ---- 限制与注意事项 ----
# pickle:类必须绑定在模块命名空间的同名位置(__module__ + __qualname__)
# 不可哈希:只要有任一字段不可哈希
# 不可覆盖:__new__ / __init__ / __slots__(typing.NamedTuple)
# 不要赋值:_fields / _field_defaults(赋了会让 _asdict() 与 repr() 不一致)
# 子类化:collections.namedtuple 用 __slots__ = ();typing.NamedTuple 不支持加字段
# 可变默认值会被所有实例共享
结语
具名元组的定位其实非常清晰:它是「元组 + 字段名」,不是「轻量级类」。
- 内存和裸元组完全相同、可哈希、可解包、可比较,这四点让它在「值对象 / 记录 / 复合键」场景几乎没有对手;
- 代价是不能改、构造偏慢、字段名必须合法且固定;
- 需要可变、需要
default_factory、需要更自由的继承时,dataclass更合适。