元组与具名元组:把「第 3 个字段」变成「price」

元组与具名元组:把「第 3 个字段」变成「price」

写作约定:本文所有代码都在真实解释器 上运行过,输出是从终端直接复制的。除特别标注外,输出均来自 Python 3.14.7;涉及版本差异的地方会单独标出实测版本。凡是没能实测验证的说法,本文不写。

配套代码:

  • examples/01_sales_report.py ------ 用具名元组做销售汇总
  • examples/02_config.py ------ 用具名元组做不可变配置
  • verify/ ------ 每个结论的验证脚本,07_final_matrix.py 可跨版本重跑

一、先把元组本身搞清楚

具名元组是元组的子类,所以先把元组那几个容易被误解的点确认清楚。

1.1 决定元组的是逗号,不是括号

python 复制代码
t1 = (1)
t2 = (1,)
t3 = 1, 2
print("type((1)) ->", type(t1).__name__, "| (1) == 1 ->", t1 == 1)
print("type((1,)) ->", type(t2).__name__, "| len((1,)) ->", len(t2))
print("1, 2 ->", repr(t3))

输出:

text 复制代码
type((1)) -> int | (1) == 1 -> True
type((1,)) -> tuple | len((1,)) -> 1
1, 2 -> (1, 2)

(1) 就是整数 1,括号只是分组符号。一个元素也要写逗号,这是新手最常见的 bug 来源:

python 复制代码
def get_user():
    return "alice",        # 有逗号:返回元组
    # return "alice"       # 没逗号:返回字符串

1.2 「不可变」指的是元组这一层

元组本身不能增删改元素,但它引用的对象该变还是变:

python 复制代码
inner = [1, 2]
t = (inner, "x")
t[0].append(3)
print("元组内容被间接修改后:", t)
try:
    t[0] = [9]
except TypeError as e:
    print("t[0] = [9] ->", type(e).__name__, ":", e)
text 复制代码
元组内容被间接修改后: ([1, 2, 3], 'x')
t[0] = [9] -> TypeError : 'tuple' object does not support item assignment

1.3 可哈希是有条件的

元组能当字典键,前提是它里面的每个元素都可哈希:

python 复制代码
print("{(1, 2): 'a'} ->", {(1, 2): "a"}[1, 2])
try:
    {([1, 2],): "a"}
except TypeError as e:
    print("含 list 的元组作键 ->", type(e).__name__, ":", e)
text 复制代码
{(1, 2): 'a'} -> a
含 list 的元组作键 -> TypeError : cannot use 'tuple' as a dict key (unhashable type: 'list')

上面这条报错是 3.14.7 的原文。3.12.10 上同样的代码只说 TypeError: unhashable type: 'list'。报错格式会随版本变,别拿它做字符串匹配。

这个规则原样传递给了具名元组:字段里有 list/dict/set,整个具名元组就不可哈希(见 3.6 节)。

1.4 函数「返回多个值」其实是返回元组

python 复制代码
def min_max(nums):
    return min(nums), max(nums)

lo, hi = min_max([3, 1, 4, 1, 5])       # 解包
head, *tail = (1, 2, 3, 4)              # 星号解包
print("min_max([3,1,4,1,5]) ->", (lo, hi))
print("head, *tail = (1,2,3,4) ->", head, tail)
text 复制代码
min_max([3,1,4,1,5]) -> (1, 5)
head, *tail = (1,2,3,4) -> 1 [2, 3, 4]

记住这一点,后面就能理解为什么具名元组可以无缝用在 for a, b in pairs 这类地方------它就是元组。


二、具名元组解决什么问题

看一段很典型的代码:

python 复制代码
def total(rows):
    return sum(r[2] * r[3] for r in rows)   # r[2]、r[3] 是什么?

r[2] * r[3] 能跑,但没人知道它算的是哪两个字段。三个月后回来改,只能去翻上游的解析代码。具名元组把「位置」变成「名字」:

python 复制代码
class Sale(NamedTuple):
    order_id: str
    sku: str
    qty: int
    unit_price: float

def total(rows):
    return sum(r.qty * r.unit_price for r in rows)

同时它还是元组:长度、索引、解包、比较、字典键,一样都没丢。

Python 提供了两条路:

collections.namedtuple typing.NamedTuple
写法 工厂函数 + 字段名 class 语法 + 类型注解
加方法 需要再继承一次 直接写在 class 里
支持类型注解 否 是
默认值 defaults=(...) 按尾部对齐 字段后写 = value
保留 docstring 否(工厂函数会给一个自动文档) 是
能力上限 低 高(推荐)

两者生成的类都是 tuple 的子类,运行期结构完全一样,区别只在于「怎么写」和「能写多复杂」。


三、collections.namedtuple 逐个 API 拆解

3.1 定义与访问

python 复制代码
from collections import namedtuple

Point = namedtuple("Point", ["x", "y"])
p = Point(3, 4)
print("p ->", p)
print("repr(p) ->", repr(p))
print("p.x, p.y ->", p.x, p.y)
print("p[0], p[1] ->", p[0], p[1])
print("tuple(p) ->", tuple(p))
print("isinstance(p, tuple) ->", isinstance(p, tuple))
print("p == (3, 4) ->", p == (3, 4))
print("len(p) ->", len(p))
text 复制代码
p -> Point(x=3, y=4)
repr(p) -> Point(x=3, y=4)
p.x, p.y -> 3 4
p[0], p[1] -> 3 4
tuple(p) -> (3, 4)
isinstance(p, tuple) -> True
p == (3, 4) -> True
len(p) -> 2

注意最后几行:具名元组和普通元组相等 ,Point(3, 4) == (3, 4) 是 True,而且 hash(Point(3, 4)) == hash((3, 4))。这是它和 dataclass 的关键区别(第十一节有对照)。

字段名既可以用列表,也可以用一个字符串(空格或逗号分隔):

python 复制代码
P2 = namedtuple("P2", "x y")
P3 = namedtuple("P3", "x,y")
print("P2._fields ->", P2._fields, "| P3._fields ->", P3._fields)
text 复制代码
P2._fields -> ('x', 'y') | P3._fields -> ('x', 'y')

字符串形式内部是 field_names.replace(',', ' ').split(),所以 "x y"、"x,y"、"x, y" 都可以。字段多的时候用列表更清楚。

3.2 _fields 与 defaults

python 复制代码
Point3 = namedtuple("Point3", "x y z", defaults=(0, 0))
print("Point3._field_defaults ->", Point3._field_defaults)
print("Point3(1) ->", Point3(1))
text 复制代码
Point3._field_defaults -> {'y': 0, 'z': 0}
Point3(1) -> Point3(x=1, y=0, z=0)

defaults 是从右往左 对齐的,所以 defaults=(0, 0) 落在 y 和 z 上。参数校验的报错信息也很直接:

python 复制代码
try:
    Point3(1, 2, 3, 4)
except TypeError as e:
    print("多余位置参数 ->", type(e).__name__, ":", e)
try:
    Point3(1, 2, 3, z=9)
except TypeError as e:
    print("同时给 z 位置和关键字 ->", type(e).__name__, ":", e)
try:
    Point3(1, 2, 3, extra=1)
except TypeError as e:
    print("未知关键字 ->", type(e).__name__, ":", e)
text 复制代码
多余位置参数 -> TypeError : Point3.__new__() takes from 2 to 4 positional arguments but 5 were given
同时给 z 位置和关键字 -> TypeError : Point3.__new__() got multiple values for argument 'z'
未知关键字 -> TypeError : Point3.__new__() got an unexpected keyword argument 'extra'

3.3 _make:从任意序列构造(带长度校验)

python 复制代码
print("p._asdict() ->", p._asdict(), "| type ->", type(p._asdict()).__name__)
print("Point._make([7, 8]) ->", Point._make([7, 8]))
print("Point(*[7, 8]) 等价 ->", Point(*[7, 8]))
text 复制代码
p._asdict() -> {'x': 3, 'y': 4} | type -> dict
Point._make([7, 8]) -> Point(x=7, y=8)
Point(*[7, 8]) 等价 -> Point(x=7, y=8)

_make 比 Point(*seq) 多一层长度校验:长度不对会抛 TypeError。解析 CSV 行的时候,用 _make 能把「表头字段数和数据列数不一致」这类错误当场拦住。

3.4 _replace:派生新实例,不是原地修改

python 复制代码
p2 = p._replace(x=10)
print("p._replace(x=10) ->", p2, "| 原对象不变 ->", p)
text 复制代码
p._replace(x=10) -> Point(x=10, y=4) | 原对象不变 -> Point(x=3, y=4)

_replace 内部走的是 _make,所以未知字段会报错。这个错误类型在 3.13 变了(实测):

版本 p._replace(z=1)
3.12.10 ValueError: Got unexpected field names: ['z']
3.14.7 TypeError: Got unexpected field names: ['z']

如果你写 except ValueError 来兜住拼错的字段名,升级到 3.13+ 会漏掉。

_replace 是浅拷贝,可变字段仍然共享:

python 复制代码
Box = namedtuple("Box", "items name")
b = Box([1, 2], "a")
c = b._replace(name="b")
c.items.append(3)
print("b ->", b, "| c ->", c, "| b.items is c.items ->", b.items is c.items)
text 复制代码
b -> Box(items=[1, 2, 3], name='a') | c -> Box(items=[1, 2, 3], name='b') | b.items is c.items -> True

只改了 name,b 的 items 也跟着变成 [1, 2, 3]。要真正隔离就自己 copy.deepcopy。

3.5 rename:字段名非法时的兜底

python 复制代码
try:
    namedtuple("Bad", ["a", "class", "a", "1x"])
except ValueError as e:
    print("默认 rename=False ->", type(e).__name__, ":", e)
Bad = namedtuple("Bad", ["a", "class", "a", "1x"], rename=True)
print("rename=True ->", Bad._fields)
print("Bad(1,2,3,4) ->", Bad(1, 2, 3, 4))
text 复制代码
默认 rename=False -> ValueError : Type names and field names cannot be a keyword: 'class'
rename=True -> ('a', '_1', '_2', '_3')
Bad(1,2,3,4) -> Bad(a=1, _1=2, _2=3, _3=4)

非法字段(关键字、非标识符、以下划线开头、重复)会被替换成 _下标。注意 rename 只按真值判断 ,传 "yes"、1 都会生效:

text 复制代码
rename='yes' -> OK, _fields=('a', '_1')
rename=1 -> OK, _fields=('a', '_1')
rename=0 -> ValueError: Type names and field names cannot be a keyword: 'class'

建议显式传 True / False,别传字符串。

3.6 哈希、排序、去重

python 复制代码
print("hash(Point(3,4)) == hash((3,4)) ->", hash(Point(3, 4)) == hash((3, 4)))
print("{Point(3,4): 'v'} 可用 ->", {Point(3, 4): "v"}[(3, 4)])
print("len({Point(1,2), Point(1,2), (1,2)}) ->", len({Point(1, 2), Point(1, 2), (1, 2)}))
print("sorted ->", sorted([Point(2, 0), Point(1, 9)]))
text 复制代码
hash(Point(3,4)) == hash((3,4)) -> True
{Point(3,4): 'v'} 可用 -> v
len({Point(1,2), Point(1,2), (1,2)}) -> 1
sorted -> [Point(x=1, y=9), Point(x=2, y=0)]

字段里放可变对象就不可哈希,但仍然可以比较相等:

text 复制代码
字段含 list 时 hash -> TypeError : unhashable type: 'list'
字段含 list 时作字典键 -> TypeError : cannot use 'Shared' as a dict key (unhashable type: 'list')
但仍可用 == 比较 -> True

(这两条报错是 3.14.7 的原文,3.12.10 只打印 unhashable type: 'list'。)


四、typing.NamedTuple:带类型和方法

4.1 基本类语法

python 复制代码
from typing import NamedTuple

class Point(NamedTuple):
    """二维平面上的一个点。"""

    x: int
    y: int = 0                      # 默认值

    def dist_to_origin(self) -> float:
        return (self.x**2 + self.y**2) ** 0.5


p = Point(3, 4)
print("p ->", p)
print("type(p) ->", type(p).__name__, "| isinstance tuple ->", isinstance(p, tuple))
print("p.dist_to_origin() ->", p.dist_to_origin())
print("__annotations__ ->", Point.__annotations__)
print("_fields ->", Point._fields)
print("_field_defaults ->", Point._field_defaults)
print("Point(5) ->", Point(5))
print("实例有没有 __dict__ ->", hasattr(p, "__dict__"), "| __slots__ ->", Point.__slots__)
print("Point.__doc__ ->", repr(Point.__doc__))
print("Point.__mro__[:3] ->", [c.__name__ for c in Point.__mro__[:3]])
text 复制代码
p -> Point(x=3, y=4)
type(p) -> Point | isinstance tuple -> True
p.dist_to_origin() -> 5.0
__annotations__ -> {'x': <class 'int'>, 'y': <class 'int'>}
_fields -> ('x', 'y')
_field_defaults -> {'y': 0}
Point(5) -> Point(x=5, y=0)
实例有没有 __dict__ -> False | __slots__ -> ()
Point.__doc__ -> '二维平面上的一个点。'
Point.__mro__[:3] -> ['Point', 'tuple', 'object']

三个信息量很大的点:

  1. __slots__ 是 (),实例没有 __dict__------所以每个实例的内存开销和同长度元组一样(第 8 节实测)。
  2. __mro__ 里第二项就是 tuple,它真的是元组。
  3. docstring 保留 ,这对 help() 和文档生成很重要。

方法、@property、普通类属性都可以写在类里,非注解的类属性不会被当成字段:

python 复制代码
class Circle(NamedTuple):
    radius: float
    pi = 3.141592653589793        # 普通类属性,不是字段

    @property
    def area(self) -> float:
        return Circle.pi * self.radius**2

    def scaled(self, k: float) -> "Circle":
        return self._replace(radius=self.radius * k)

c = Circle(2.0)
print("_fields ->", Circle._fields)
print("Circle.pi ->", Circle.pi, "| c.area ->", c.area)
print("c.scaled(3) ->", c.scaled(3))
text 复制代码
_fields -> ('radius',)
Circle.pi -> 3.141592653589793 | c.area -> 12.566370614359172
c.scaled(3) -> Circle(radius=6.0)

4.2 不能覆盖那几个特殊属性

text 复制代码
覆盖 __new__ -> AttributeError: Cannot overwrite NamedTuple attribute __new__
覆盖 __init__ -> AttributeError: Cannot overwrite NamedTuple attribute __init__
显式 __slots__ -> AttributeError: Cannot overwrite NamedTuple attribute __slots__

想自定义构造逻辑,正确做法是写类方法:

python 复制代码
class Sale(NamedTuple):
    qty: int
    unit_price: float

    @classmethod
    def from_row(cls, qty, price):
        return cls(int(qty), float(price))

4.3 默认值陷阱:可变默认值会被共享

typing.NamedTuple 的默认值只求值一次 ,而且不会像 dataclass 那样拦截可变默认值(后者会直接报 ValueError: mutable default <class 'list'> for field tags is not allowed: use default_factory):

python 复制代码
class Cfg(NamedTuple):
    name: str
    tags: list = []

a, b = Cfg("a"), Cfg("b")
a.tags.append("x")
print("b.tags ->", b.tags, "(两个实例共享同一个默认 list)")
print("默认值对象的身份相同 ->", Cfg("a").tags is Cfg("b").tags)
text 复制代码
b.tags -> ['x'] (两个实例共享同一个默认 list)
默认值对象的身份相同 -> True

只改 a,b 也被污染了。字段默认值只用不可变对象 ,需要 list 就传 None 再在方法里兜底,或者干脆改用 dataclasses.field(default_factory=list)。

4.4 函数式写法与泛型

python 复制代码
Pair = NamedTuple("Pair", [("key", str), ("value", int)])
print("Pair('a', 1) ->", Pair("a", 1), "| _fields ->", Pair._fields)
print("Pair.__annotations__ ->", Pair.__annotations__)
text 复制代码
Pair('a', 1) -> Pair(key='a', value=1) | _fields -> ('key', 'value')
Pair.__annotations__ -> {'key': <class 'str'>, 'value': <class 'int'>}

泛型也支持:

python 复制代码
from typing import Generic, TypeVar

T = TypeVar("T")

class Page(NamedTuple, Generic[T]):
    items: list
    total: int

    def is_empty(self) -> bool:
        return self.total == 0

pg = Page([1, 2, 3], 3)
print("pg ->", pg, "| is_empty ->", pg.is_empty())
text 复制代码
pg -> Page(items=[1, 2, 3], total=3) | is_empty -> False

注意:NamedTuple("X", a=int, b=str) 这种关键字写法在 3.13+ 已弃用 (3.12 无告警,3.13/3.14 实测会抛 DeprecationWarning,提示 3.15 将移除):

text 复制代码
DeprecationWarning: Creating NamedTuple classes using keyword arguments is deprecated
and will be disallowed in Python 3.15. Use the class-based or functional syntax instead.

新代码用 class 语法或 [("a", int), ...] 列表形式。


五、模式匹配与 copy.replace()(新版本专属)

5.1 结构化模式匹配(3.10+)

namedtuple 生成的类带 __match_args__,所以可以按字段名匹配:

python 复制代码
class Pt(NamedTuple):
    x: int
    y: int

def describe(p):
    match p:
        case Pt(x=0, y=0):
            return "原点"
        case Pt(x=0, y=y):
            return f"y 轴上, y={y}"
        case Pt(x=x, y=0):
            return f"x 轴上, x={x}"
        case Pt(x, y) if x == y:
            return f"对角线上 ({x}, {y})"
        case Pt(x, y):
            return f"普通点 ({x}, {y})"
        case _:
            return "不是点(没有匹配到 Pt 的模式)"

for q in [Pt(0,0), Pt(0,5), Pt(5,0), Pt(3,3), Pt(1,2), (1,2)]:
    print(f"{q!r:>14} -> {describe(q)}")
text 复制代码
  Pt(x=0, y=0) -> 原点
  Pt(x=0, y=5) -> y 轴上, y=5
  Pt(x=5, y=0) -> x 轴上, x=5
  Pt(x=3, y=3) -> 对角线上 (3, 3)
  Pt(x=1, y=2) -> 普通点 (1, 2)
        (1, 2) -> 不是点(没有匹配到 Pt 的模式)

最后一行值得注意:Pt(1, 2) == (1, 2) 是 True,但类模式要求 isinstance 成立 ,所以普通元组不会匹配 case Pt(...)。想同时接受普通元组,就再加一条 case (x, y):。

case Pt(x, y) 这种位置模式依赖 __match_args__,它的存在性有版本差异(第十节表格):3.9 没有,3.10+ 有。

5.2 __replace__ 与 copy.replace()(3.13+)

python 复制代码
import copy
print("copy.replace(Pt(1,2), x=9) ->", copy.replace(Pt(1, 2), x=9))
print("Pt(1,2).__replace__(y=9) ->", Pt(1, 2).__replace__(y=9))
try:
    copy.replace(Pt(1, 2), z=9)
except TypeError as e:
    print("copy.replace 未知字段 ->", type(e).__name__, ":", e)
text 复制代码
copy.replace(Pt(1,2), x=9) -> Pt(x=9, y=2)
Pt(1,2).__replace__(y=9) -> Pt(x=1, y=9)
copy.replace 未知字段 -> TypeError : Got unexpected field names: ['z']

3.12 实测没有 copy.replace,也没有 __replace__,只能用 _replace。_replace 在所有版本都能用,所以写兼容库时优先用它。


六、实战一:销售明细汇总

完整代码见 examples/01_sales_report.py,下面是真实输出。

python 复制代码
class Sale(NamedTuple):
    """一条销售明细。order_id 是订单号,qty 是数量,unit_price 是单价。"""

    order_id: str
    sku: str
    qty: int
    unit_price: float

    @property
    def amount(self) -> float:
        return round(self.qty * self.unit_price, 2)


RAW = [
    ("order_id", "sku", "qty", "unit_price"),
    ("A-1001", "KB-01", 2, 199.0),
    # ...省略
]

header, *rows = RAW
assert tuple(header) == Sale._fields, "表头和字段不一致,先改字段定义"
sales = [Sale._make(r) for r in rows]
text 复制代码
== 明细 ==
A-1001  KB-01  2 × 199.00 =  398.00
A-1001  MS-02  1 ×  89.50 =   89.50
A-1002  KB-01  1 × 199.00 =  199.00
A-1002  HD-03  3 ×  45.00 =  135.00
A-1003  MS-02  4 ×  89.50 =  358.00
A-1003  HD-03  2 ×  45.00 =   90.00

_replace 派生打折记录,原记录不动:

text 复制代码
原记录  -> Sale(order_id='A-1001', sku='KB-01', qty=2, unit_price=199.0) 金额 398.0
折后记录-> Sale(order_id='A-1001', sku='KB-01', qty=2, unit_price=159.2) 金额 318.4
原记录没有被改动 -> True

分组汇总:

text 复制代码
SKU       订单数    件数        金额
HD-03       2     5    225.00
KB-01       2     3    597.00
MS-02       2     5    447.50

转 JSON 时必须显式 _asdict(),否则字段名会丢(见 9.4 节):

json 复制代码
{
  "by_sku": [
    {"sku": "HD-03", "orders": 2, "qty": 5, "revenue": 225.0},
    {"sku": "KB-01", "orders": 2, "qty": 3, "revenue": 597.0},
    {"sku": "MS-02", "orders": 2, "qty": 5, "revenue": 447.5}
  ],
  "total_revenue": 1269.5
}

这个例子里具名元组承担了三个角色:表结构声明 (_fields 校验表头)、只读记录 (打折靠 _replace)、可哈希的分组键。

七、实战二:不可变配置对象

完整代码见 examples/02_config.py。

python 复制代码
Timeouts = namedtuple("Timeouts", "connect read write", defaults=(5.0, 30.0, 30.0))

print("Timeouts()        ->", Timeouts())
print("Timeouts(1.5)     ->", Timeouts(1.5))
print("Timeouts(connect=2.0) ->", Timeouts(connect=2.0))
text 复制代码
Timeouts()        -> Timeouts(connect=5.0, read=30.0, write=30.0)
Timeouts(1.5)     -> Timeouts(connect=1.5, read=30.0, write=30.0)
Timeouts(connect=2.0) -> Timeouts(connect=2.0, read=30.0, write=30.0)

派生配置:

python 复制代码
base = Timeouts(connect=1.0)
slow_net = base._replace(connect=10.0, read=60.0)
print("base     ->", base)
print("slow_net ->", slow_net)
print("base 未被修改 ->", base.connect == 1.0)
text 复制代码
base     -> Timeouts(connect=1.0, read=30.0, write=30.0)
slow_net -> Timeouts(connect=10.0, read=60.0, write=30.0)
base 未被修改 -> True

配置对象可以安全地当默认参数、当字典键:

python 复制代码
def fetch(url, timeouts=Timeouts()):
    return f"GET {url} (connect={timeouts.connect}s, read={timeouts.read}s)"

print(fetch("https://example.com"))
print(fetch("https://slow.example.com", timeouts=slow_net))
cache = {Timeouts(): "默认配置的结果", slow_net: "慢网络的结果"}
print("用配置对象作缓存键 ->", cache[Timeouts(5.0, 30.0, 30.0)])
text 复制代码
GET https://example.com (connect=5.0s, read=30.0s)
GET https://slow.example.com (connect=10.0s, read=60.0s)
用配置对象作缓存键 -> 默认配置的结果

(因为它可哈希且不可变,不会出现「传进去被下游改掉」的问题。)


八、性能与内存:实测数据

环境:macOS / Apple Silicon,同一份脚本 verify/06_bench.py,timeit.repeat(..., number=1_000_000, repeat=5) 取最小值。不同机器和版本会有波动,看量级关系即可。

8.1 构造与访问(ns/次)

操作 3.12.10 3.14.7
tuple((x, y)) 15.8 17.9
Point(x, y)(具名元组) 117.2 119.7
DPoint(x, y)(frozen dataclass) 181.2 147.9
MPoint(x, y)(普通 dataclass) 76.8 50.2
{"x": x, "y": y}(dict) 45.1 42.6
p.x(具名元组属性) 10.0 10.7
dp.x(frozen dataclass 属性) 5.0 6.0
d["x"](dict 取值) 9.1 9.3
tup[0](元组索引) 6.1 6.5
a, b = p(解包) 19.2 17.3
p._asdict() 155.0 157.0
json.dumps(p) 702.2 724.5
json.dumps(p._asdict()) 885.3 986.8

三个结论,都有点反直觉:

  1. 构造具名元组比构造裸元组慢约 7 倍 (117.2 / 15.8 ≈ 7.4)------因为 tuple.__new__ 走的是 C,而具名元组生成的 __new__ 是一个 Python 层 lambda。如果你的热路径是「每秒构造几百万条记录」,这一点要实测再决定。
  2. 属性访问并不比 d["x"] 快(两个版本上都是大约慢 1 ns)。所以「用具名元组替代 dict 更快」这个说法在近几个版本上不成立,别把它当性能优化手段。
  3. json.dumps(p) 比 json.dumps(p._asdict()) 快------因为前者按数组处理,只输出 [3, 4],省掉了字段名。代价是丢了语义。

8.2 内存:这才是具名元组的杀手锏

python 复制代码
sys.getsizeof(...)   # 只算容器对象本身
字段数 tuple namedtuple slots dataclass 普通 dataclass dict
2 64 64 48 48 184
3 72 72 56 48 184
5 88 88 72 48 184
8 112 112 96 48 272

(Python 3.14.7 实测。3.12.10 上同一张表的 tuple / namedtuple 两列是 56 / 64 / 80 / 104,每行比 3.14 少 8 字节;slots dataclass、dataclass、dict 三列在两个版本上完全相同。也就是说,跨版本唯一稳定且关键的规律是:namedtuple 和 tuple 逐格相等。)

结论很清楚:

  • 具名元组 = 裸元组的内存 :字段没有存成 dict,就是 tuple 里那几个槽位。换成 dict 大约多占 3 倍。
  • 普通 dataclass 反而"看起来"小 (48 B),因为字段值存在 __dict__ 里,而 __dict__ 是另一个对象,getsizeof 没算进去。算上字典就大了:tracemalloc 实测 10 万个两字段实例,dict 每实例约 192 B,普通 dataclass 约 96 B,dataclass(slots=True) 约 56 B,具名元组约 80 B,裸元组约 72 B。
  • 想让 dataclass 在内存上超过具名元组,得加 slots=True。

(tracemalloc 的数字包含分配器账目,所以具名元组每实例比裸元组多记了约 8 B;以 sys.getsizeof 为准,两者对象大小完全相同。)


九、坑清单(全部有实测)

9.1 字段名规则

python 复制代码
for names in (["_x"], ["x", "_y"], ["x", "x"], ["class"]):
    try:
        namedtuple("C", names)
        print(names, "-> OK")
    except ValueError as e:
        print(names, "-> ValueError:", e)
print(["self"], "-> OK _fields=" + str(namedtuple("C", ["self"])._fields))
text 复制代码
['_x'] -> ValueError: Field names cannot start with an underscore: '_x'
['x', '_y'] -> ValueError: Field names cannot start with an underscore: '_y'
['x', 'x'] -> ValueError: Encountered duplicate field name: 'x'
['class'] -> ValueError: Type names and field names cannot be a keyword: 'class'
['self'] -> OK _fields=('self',)

以下划线开头是被禁的(避免和 _fields、_make、_replace、_asdict 这些内部 API 撞名);关键字和重复名也不行。self 这种名字反而合法------虽然不建议。

9.2 字段名会覆盖 tuple 自带的 count / index

python 复制代码
T = namedtuple("T", "count index")
t = T(10, 20)
print("T(10,20) ->", t, "| t.count ->", t.count, "| t.index ->", t.index)
try:
    t.count(10)
except TypeError as e:
    print("t.count(10) ->", type(e).__name__, ":", e)
print("tuple.count 被字段访问器覆盖了吗 ->", type(t).count is not tuple.count)
text 复制代码
T(10,20) -> T(count=10, index=20) | t.count -> 10 | t.index -> 20
t.count(10) -> TypeError : 'int' object is not callable
tuple.count 被字段访问器覆盖了吗 -> True

字段访问器把 tuple.count / tuple.index 顶掉了,t.count(10) 直接崩。避开这两个名字。

9.3 给 _fields 赋值不会改变行为,只会让它说谎

_fields 只是一个普通类属性,可以赋值 ,但 __new__、__repr__、_replace 用的是类创建时就固定下来的字段名,只有 _asdict() 会去读 self._fields:

python 复制代码
Q = namedtuple("Q", "x y")
Q._fields = ("a", "b")
q = Q(1, 2)
print("repr(q)      ->", repr(q))
print("q._asdict()  ->", q._asdict())
text 复制代码
repr(q)      -> Q(x=1, y=2)
q._asdict()  -> {'a': 1, 'b': 2}

repr 和 _asdict() 互相矛盾,_replace(a=...) 还会报错。3.7 到 3.14 实测行为一致(3.7 的 _asdict 是 OrderedDict([('a', 1), ('b', 2)]))。结论:永远不要给 _fields 赋值。

9.4 JSON 序列化会丢掉字段名;dict() 也不是转字典

python 复制代码
import json
print("json.dumps(p) ->", json.dumps(p))
print("json.dumps(p._asdict()) ->", json.dumps(p._asdict()))
text 复制代码
json.dumps(p) -> [3, 4]
json.dumps(p._asdict()) -> {"x": 3, "y": 4}

因为它是元组,json 按数组序列化。别指望用 default= 钩子救回来------钩子只对「无法序列化」的对象调用,而元组是能被序列化的:

python 复制代码
json.dumps(p, default=lambda o: o._asdict())
text 复制代码
[3, 4]

default 压根没被调用。要保留字段名就只能先 _asdict() 再序列化。

同样地,dict(p) 不是 转字典,它把元组的每个元素当成 (键, 值):

python 复制代码
P = namedtuple("P", "x y")
try:
    dict(P(1, 2))
except TypeError as e:
    print("dict(P(1, 2)) ->", type(e).__name__, ":", e)
print("dict(P(('a',1),('b',2))) ->", dict(P(("a", 1), ("b", 2))))
print("P(('a',1),('b',2))._asdict() ->", P(("a", 1), ("b", 2))._asdict())
text 复制代码
dict(P(1, 2)) -> TypeError : object is not iterable
dict(P(("a",1),("b",2))) -> {'a': 1, 'b': 2}
P(("a",1),("b",2))._asdict() -> {'x': ('a', 1), 'y': ('b', 2)}

要转字典永远用 _asdict()。

9.5 pickle 的真实规则

流传的说法是「函数里定义的具名元组不能 pickle」------不准确 。pickle 通过 __module__ + __qualname__ 回查类对象,所以规则是「模块命名空间里有没有这个名字,且指向同一个类」:

python 复制代码
def make_same():
    return namedtuple("Same", "x y")

def make_other():
    return namedtuple("Other", "x y")

def make_ghost():
    return namedtuple("Ghost", "x y")(1, 2)

Same = make_same()          # A:类名 Same,绑定的名字也叫 Same
Renamed = make_other()      # B:类名 Other,绑定的名字却叫 Renamed
ghost = make_ghost()        # C:类只在函数内存在,模块里查不到

for label, obj in [("A 绑定同名(模块里能查到 Same)   ", Same(1, 2)),
                   ("B 绑定异名(类叫 Other,绑定为 Renamed)", Renamed(1, 2)),
                   ("C 从未绑定(只在函数内存在)       ", ghost)]:
    try:
        pickle.dumps(obj)
        print(f"{label} -> 成功")
    except Exception as e:
        print(f"{label} -> {type(e).__name__}: {e}")
text 复制代码
A 绑定同名(模块里能查到 Same)    -> 成功
B 绑定异名(类叫 Other,绑定为 Renamed) -> PicklingError: Can't pickle <class '__main__.Other'>: it's not found as __main__.Other
C 从未绑定(只在函数内存在)        -> PicklingError: Can't pickle <class '__main__.Ghost'>: it's not found as __main__.Ghost

所以:把具名元组定义在模块顶层 。放在函数里创建、只通过返回值暴露,就会在跨进程(多进程 / 缓存 / 消息队列)时炸掉。module= 参数也救不了,它只是改 __module__,名字没绑定照样失败。

顺带一提,copy.copy 和 copy.deepcopy 行为不同:

python 复制代码
import copy

plain = (1, 2)              # 必须用变量,否则常量折叠会让 is 比较失去意义
named = Point(3, 4)
box = Box([1], "a")
print("copy.copy(普通元组) is 原对象 ->", copy.copy(plain) is plain, "     (元组不可变,原样返回)")
print("copy.copy(具名元组) is 原对象 ->", copy.copy(named) is named, "    (走 __getnewargs__,new 一个新对象)")
print("copy.deepcopy 是否共享可变字段 ->", copy.deepcopy(box).items is box.items)
text 复制代码
copy.copy(普通元组) is 原对象 -> True      (元组不可变,原样返回)
copy.copy(具名元组) is 原对象 -> False     (走 __getnewargs__,new 一个新对象)
copy.deepcopy 是否共享可变字段 -> False

9.6 子类化时字段怎么变(两种写法不一样)

collections.namedtuple 的子类字段会拼接(父类的字段在前):

python 复制代码
import sys

Base = namedtuple("Base", "a b")

class Child(Base):
    __slots__ = ()

class Child2(Base):        # 不写 __slots__
    pass

print("Child._fields ->", Child._fields)
print("Child 实例有 __dict__ 吗 ->", hasattr(Child(1, 2), "__dict__"))
print("Child2 实例有 __dict__ 吗 ->", hasattr(Child2(1, 2), "__dict__"))
print("两者实例大小 ->", sys.getsizeof(Child(1, 2)), sys.getsizeof(Child2(1, 2)))
text 复制代码
Child._fields -> ('a', 'b')
Child 实例有 __dict__ 吗 -> False
Child2 实例有 __dict__ 吗 -> True
两者实例大小 -> 64 80

这正是「内存和元组一样」的前提:子类里写 __slots__ = () 。不写的话,子类会给每个实例加一个 __dict__(上面实测 64 B → 80 B,还没算字典本身),具名元组的零开销优势就没了。注意 Base.__slots__ 虽然是 (),但它只是被继承,不会阻止子类创建 __dict__。

但 typing.NamedTuple 的子类不会新增字段,注解被静默忽略:

python 复制代码
class A(NamedTuple):
    a: int

class B(A):
    b: int          # 不会变成字段

print("B._fields ->", B._fields)     # ('a',)
print("B(1) ->", B(1))               # B(a=1)
print("hasattr(B(1), 'b') ->", hasattr(B(1), "b"))   # False
text 复制代码
B._fields -> ('a',) | B(1) -> B(a=1) | hasattr b -> False

3.8 / 3.12 / 3.14 实测一致。想要新字段,就重新定义一个完整的 NamedTuple。

9.7 两个具名元组做多重继承:不报错,但字段会丢

python 复制代码
class A(NamedTuple):
    a: int
class B(NamedTuple):
    b: int
class C(A, B):
    pass
text 复制代码
C.__mro__ -> ['C', 'A', 'B', 'tuple', 'object']
C._fields -> ('a',)          # B 的 b 完全丢失
C(1) -> C(a=1)
C(1, 2) -> TypeError : A.__new__() takes 2 positional arguments but 3 were given
C(a=1,b=2) -> TypeError : A.__new__() got an unexpected keyword argument 'b'

(以上为 3.14.7 输出;3.12.10 实测 C._fields 同样是 ('a',)、C(1) 同样得到 C(a=1),只是 TypeError 的措辞不同。)

静默丢掉 B 的字段,只在构造时才报错。别把两个具名元组混在一起继承。

9.8 其他小坑

  • typing.NamedTuple 里写 __slots__ = () 会报 AttributeError(它自己已经设好了)。
  • _asdict() 每次返回新字典 ,改它不影响实例(实测:改了之后 p 仍然是 Point(x=3, y=4))。
  • _replace() 返回新对象,但原对象和它共享可变字段(3.4 节)。
  • 字段默认值用可变对象会被所有实例共享(4.3 节)。

十、版本差异速查表(3.7 → 3.14 实测)

用 verify/07_final_matrix.py 在 8 个解释器上跑出来的原始结果:

特性 3.7 3.8 3.9 3.10 3.11 3.12 3.13 3.14
_asdict() 返回类型 OrderedDict dict dict dict dict dict dict dict
_field_defaults / defaults= 有 有 有 有 有 有 有 有
__match_args__(模式匹配) 无 无 无 有 有 有 有 有
__replace__ / copy.replace() 无 无 无 无 无 无 有 有
_replace() 未知字段的异常 ValueError ValueError ValueError ValueError ValueError ValueError TypeError TypeError
NamedTuple("X", a=int) 可用 可用 可用 可用 可用 可用 弃用告警 弃用告警
类语法保留 docstring 是 是 是 是 是 是 是 是

(3.13/3.14 的告警原文:Creating NamedTuple classes using keyword arguments is deprecated and will be disallowed in Python 3.15.)

如果库要兼容 3.9 及更早:不要用模式匹配的类模式,不要用 copy.replace,_replace 的错误处理要同时兼容 ValueError 和 TypeError。


十一、怎么选:一张决策表

需求 推荐
就是一组固定长度的值,要快、要省内存 tuple
字段名固定、需要按名字访问、要能当字典键、要能解包 namedtuple / NamedTuple
字段名固定,但需要可变、需要 default_factory、需要继承体系 @dataclass
字段多且不固定、需要动态增删、与 JSON 直接对应 dict
需要运行期类型校验 / 从外部数据反序列化 pydantic.BaseModel 等

再补几条实际判断标准:

  • 要放进 set / 当 dict 键 / 做缓存 key → 具名元组(dataclass 默认不可哈希,需要 frozen=True)。
  • 要和普通元组互操作 (比如和已有接口对接、== 比较、for a, b in ...)→ 具名元组。dataclass 实例 == (1, 2) 是 False,也不支持 len() 和下标。
  • 字段有几十个、还需要默认值和校验 → dataclass + 类型检查工具。
  • 构造频率极高 (每秒百万级)→ 实测:具名元组构造比裸元组慢约 7 倍,这种场景考虑裸元组或 slots dataclass。

具名元组 vs dataclass 的语义差异,实测对照:

python 复制代码
from dataclasses import dataclass

P = namedtuple("P", "x y")

@dataclass(frozen=True)
class DP:
    x: int
    y: int

p, dp = P(1, 2), DP(1, 2)
print("P(1,2) == (1,2)      ->", p == (1, 2), "       # namedtuple")
print("DP(1,2) == (1,2)     ->", dp == (1, 2), "      # frozen dataclass")
print("hash(P(1,2)) == hash((1,2)) ->", hash(p) == hash((1, 2)))
print("P 有 __len__ ->", len(p), "| DP 有 __len__ ->", hasattr(dp, "__len__"))
print("P 可索引 p[0] ->", p[0], "| DP 可索引 ->", end=" ")
try:
    dp[0]
except TypeError as e:
    print(f"TypeError: {e}")
text 复制代码
P(1,2) == (1,2)      -> True        # namedtuple
DP(1,2) == (1,2)     -> False       # frozen dataclass
hash(P(1,2)) == hash((1,2)) -> True
P 有 __len__ -> 2 | DP 有 __len__ -> False
P 可索引 p[0] -> 1 | DP 可索引 -> TypeError: 'DP' object is not subscriptable

十二、速查表

python 复制代码
from collections import namedtuple
from typing import NamedTuple

# ---- 定义 ----
Point = namedtuple("Point", "x y")                 # 字符串
Point = namedtuple("Point", ["x", "y"])            # 列表
Point = namedtuple("Point", "x y", defaults=(0,))  # 默认值(从右往左)
Point = namedtuple("Point", ["class", "y"], rename=True)   # 非法名自动改 _下标

class Point(NamedTuple):                           # 推荐:带类型
    x: int
    y: int = 0
    def dist(self) -> float: ...

# ---- 构造 ----
Point(3, 4)              # 位置
Point(x=3, y=4)          # 关键字
Point._make(seq)         # 从任意序列(带长度校验)
Point(**d)               # 从字典

# ---- 访问 ----
p.x  /  p[0]  /  x, y = p  /  tuple(p)

# ---- 元信息 ----
Point._fields            # ('x', 'y')
Point._field_defaults    # {'y': 0}
Point.__annotations__    # {'x': int, 'y': int}
Point.__match_args__     # ('x', 'y')  (3.10+)

# ---- 转换 ----
p._asdict()                          # dict(3.8+;3.7 及以前是 OrderedDict)
p._replace(x=10)                     # 派生新实例(浅拷贝)
copy.replace(p, x=10)                # 3.13+
Point(**p._asdict())                 # dict 转回来
json.dumps(p._asdict())              # 要 JSON 对象就显式转;default= 钩子无效
# dict(p) 不是转字典!它把每个元素当 (键, 值) 解析,请用 p._asdict()

# ---- 限制与注意事项 ----
# pickle:类必须绑定在模块命名空间的同名位置(__module__ + __qualname__)
# 不可哈希:只要有任一字段不可哈希
# 不可覆盖:__new__ / __init__ / __slots__(typing.NamedTuple)
# 不要赋值:_fields / _field_defaults(赋了会让 _asdict() 与 repr() 不一致)
# 子类化:collections.namedtuple 用 __slots__ = ();typing.NamedTuple 不支持加字段
# 可变默认值会被所有实例共享

结语

具名元组的定位其实非常清晰:它是「元组 + 字段名」,不是「轻量级类」。

  • 内存和裸元组完全相同、可哈希、可解包、可比较,这四点让它在「值对象 / 记录 / 复合键」场景几乎没有对手;
  • 代价是不能改、构造偏慢、字段名必须合法且固定;
  • 需要可变、需要 default_factory、需要更自由的继承时,dataclass 更合适。
相关推荐
___波子 Pro Max.1 小时前
pip 管库,pipx 管工具:Python 依赖隔离的正确姿势
python
我的xiaodoujiao3 小时前
Django 基础知识详细图文教程 12-Django 模型定义与使用 2
后端·python·django
happylifetree11 小时前
Python09:核心语法-数据存储与运算-字面量
python
心之语歌11 小时前
Tkinter 画布基本梳理
运维·服务器·python
Wx-bishekaifayuan12 小时前
django个性化旅游路线推荐平台49005-计算机课程设计、毕业设计
spring boot·后端·python·django·课程设计·express·旅游
今晚打老虎12 小时前
c++之提高A(前缀和)(第三课)
数据结构·c++·算法
小白快快跑哦13 小时前
python-字符串全解(六):正则表达式-量词
python·正则表达式·字符串
禹凕13 小时前
机器学习之Selenium(Machina Learning about Selenium)
爬虫·python·selenium·测试工具·机器学习
不会就选b13 小时前
算法日常・每日刷题--<贪心>26
数据结构·算法