1. 引言
Python 是一门「一切皆对象」的语言。无论是整数、字符串、函数,还是类本身,在 Python 中都是对象。理解 Python 的对象模型(Object Model)与数据模型(Data Model),是深入掌握这门语言的关键分水岭。
很多开发者在使用 Python 时,停留在「调用 API」的层面,而真正的高手则能理解「API 背后的机制」。本文将从三个核心主题展开:
- 魔术方法(Magic Methods) :如何通过
__xxx__方法让自定义对象融入 Python 语法。 - 描述符(Descriptors) :属性访问的底层协议,
property、classmethod、staticmethod的实现基石。 - 元类(Metaclasses):类的「类」,掌控类创建过程的终极武器。
读完本文,你将理解 Python 对象系统的底层运作逻辑,并能写出更优雅、更 Pythonic 的代码。
2. 对象模型:一切皆对象
2.1 对象的基本概念
在 Python 中,对象是数据与行为的封装体。每个对象都有三个核心属性:
- 身份(Identity) :对象在内存中的唯一标识,可通过
id()获取。 - 类型(Type) :对象所属的类,可通过
type()获取。 - 值(Value):对象承载的数据内容。
python
>>> a = 42
>>> id(a)
140736518948720
>>> type(a)
<class 'int'>
2.2 类也是对象
这是 Python 对象模型中最反直觉、也最重要的一点:类本身也是对象 。当你定义一个类时,Python 会创建一个类型为 type 的对象:
python
class Dog:
def bark(self):
return "Woof!"
print(type(Dog)) # <class 'type'>
Dog 是 type 的实例,这意味着类可以被赋值给变量、作为参数传递、动态创建------这正是元类机制的基石。
2.3 对象的创建与销毁
对象的生命周期由 __new__、__init__、__del__ 三个魔术方法协作管理:
python
class Person:
def __new__(cls, name):
print("1. __new__ 分配内存")
instance = super().__new__(cls)
return instance
def __init__(self, name):
print("2. __init__ 初始化")
self.name = name
def __del__(self):
print("3. __del__ 销毁前清理")
p = Person("Alice")
# 输出:
# 1. __new__ 分配内存
# 2. __init__ 初始化
del p
# 输出:
# 3. __del__ 销毁前清理
3. 数据模型:魔术方法
3.1 什么是魔术方法
魔术方法(Magic Methods),又称双下划线方法(Dunder Methods),是 Python 数据模型的核心。它们以双下划线开头和结尾(如 __len__、__getitem__),不需要显式调用,而是在特定语法场景下由 Python 解释器自动触发。
3.2 字符串表示:__str__ 与 __repr__
python
class Book:
def __init__(self, title, author):
self.title = title
self.author = author
def __repr__(self):
"""面向开发者的表示,尽量无歧义"""
return f"Book('{self.title}', '{self.author}')"
def __str__(self):
"""面向用户的友好表示"""
return f"《{self.title}》 by {self.author}"
book = Book("三体", "刘慈欣")
print(repr(book)) # Book('三体', '刘慈欣')
print(str(book)) # 《三体》 by 刘慈欣
print(book) # 《三体》 by 刘慈欣(print 优先调用 __str__)
3.3 容器协议:__len__ 与 __getitem__
实现这两个方法,自定义类就能像列表一样支持 len() 和下标访问:
python
class Playlist:
def __init__(self, songs):
self._songs = songs
def __len__(self):
return len(self._songs)
def __getitem__(self, index):
return self._songs[index]
def __setitem__(self, index, song):
self._songs[index] = song
def __iter__(self):
return iter(self._songs)
pl = Playlist(["晴天", "七里香", "夜曲"])
print(len(pl)) # 3
print(pl[1]) # 七里香
pl[2] = "稻香"
for song in pl: # 有 __iter__ 后可迭代
print(song)
3.4 运算符重载
通过魔术方法,可以让自定义对象支持 +、==、< 等运算符:
python
class Vector:
def __init__(self, x, y):
self.x = x
self.y = y
def __add__(self, other):
"""实现 v1 + v2"""
return Vector(self.x + other.x, self.y + other.y)
def __eq__(self, other):
"""实现 v1 == v2"""
return self.x == other.x and self.y == other.y
def __lt__(self, other):
"""实现 v1 < v2(按模长比较)"""
return (self.x**2 + self.y**2) < (other.x**2 + other.y**2)
def __repr__(self):
return f"Vector({self.x}, {self.y})"
v1 = Vector(1, 2)
v2 = Vector(3, 4)
print(v1 + v2) # Vector(4, 6)
print(v1 == Vector(1, 2)) # True
print(v1 < v2) # True
3.5 可调用对象:__call__
给类加上 __call__ 方法后,实例可以像函数一样被调用:
python
class Counter:
def __init__(self):
self.count = 0
def __call__(self):
self.count += 1
return self.count
counter = Counter()
print(counter()) # 1
print(counter()) # 2
print(counter()) # 3
3.6 上下文管理器:__enter__ 与 __exit__
实现这两个方法,自定义类就能配合 with 语句使用:
python
class Timer:
def __enter__(self):
import time
self.start = time.time()
return self
def __exit__(self, exc_type, exc_val, exc_tb):
import time
elapsed = time.time() - self.start
print(f"耗时:{elapsed:.4f} 秒")
return False # False 表示不吞掉异常
with Timer():
total = sum(range(1_000_000))
3.7 常用魔术方法速查表
| 类别 | 方法 | 触发场景 |
|---|---|---|
| 创建与销毁 | __new__ / __init__ / __del__ |
实例化、销毁 |
| 字符串表示 | __str__ / __repr__ / __format__ |
print()、repr()、format() |
| 容器协议 | __len__ / __getitem__ / __setitem__ / __delitem__ / __contains__ / __iter__ |
len()、下标访问、in、迭代 |
| 属性访问 | __getattr__ / __setattr__ / __getattribute__ / __delattr__ |
属性读写删除 |
| 运算符 | __add__ / __sub__ / __mul__ / __eq__ / __lt__ 等 |
算术、比较运算 |
| 调用 | __call__ |
实例像函数一样调用 |
| 上下文管理 | __enter__ / __exit__ |
with 语句 |
| 哈希 | __hash__ |
hash()、作为字典键 |
4. 描述符:属性访问的底层协议
4.1 什么是描述符
描述符(Descriptor)是实现了 __get__、__set__、__delete__ 中至少一个方法的对象。它用于拦截属性访问,是 Python 属性机制的底层协议。
python
class Descriptor:
def __get__(self, instance, owner):
print("访问属性")
return "descriptor value"
def __set__(self, instance, value):
print(f"设置属性:{value}")
class MyClass:
attr = Descriptor()
obj = MyClass()
print(obj.attr) # 访问属性 → descriptor value
obj.attr = 100 # 设置属性:100
4.2 描述符的分类
根据实现的方法,描述符分为两类:
- 数据描述符(Data Descriptor) :同时实现
__get__和__set__,优先级最高。 - 非数据描述符(Non-data Descriptor) :只实现
__get__,优先级低于实例字典。
python
class DataDesc:
"""数据描述符:同时有 __get__ 和 __set__"""
def __get__(self, instance, owner):
return "data descriptor"
def __set__(self, instance, value):
print(f"set {value}")
class NonDataDesc:
"""非数据描述符:只有 __get__"""
def __get__(self, instance, owner):
return "non-data descriptor"
class Demo:
data = DataDesc()
non_data = NonDataDesc()
d = Demo()
d.data = 1 # 触发 __set__,不会写入实例字典
d.non_data = 1 # 非数据描述符不拦截 __set__,写入实例字典
print(d.__dict__) # {'non_data': 1}
print(d.non_data) # 实例字典优先 → 1(非数据描述符被遮蔽)
4.3 属性查找顺序
理解属性访问的优先级至关重要:
python
class Demo:
def __init__(self):
self.x = 100
@property
def x(self):
return self._x
@x.setter
def x(self, value):
self._x = value
属性查找顺序为:
- 数据描述符 (类属性中的
__get__+__set__) - 实例字典
__dict__ - 非数据描述符 (只有
__get__) - 类属性(非描述符)
__getattr__(兜底)
4.4 property 的实现原理
property 本质上就是一个数据描述符:
python
class MyProperty:
def __init__(self, fget=None, fset=None):
self.fget = fget
self.fset = fset
def __get__(self, instance, owner):
if instance is None:
return self
return self.fget(instance)
def __set__(self, instance, value):
if self.fset is None:
raise AttributeError("can't set attribute")
self.fset(instance, value)
def setter(self, fset):
self.fset = fset
return self
class Temperature:
def __init__(self, celsius):
self._celsius = celsius
@MyProperty
def celsius(self):
return self._celsius
@celsius.setter
def celsius(self, value):
self._celsius = value
t = Temperature(25)
print(t.celsius) # 25
t.celsius = 30
print(t.celsius) # 30
4.5 描述符实战:类型校验
描述符最常见的应用场景是属性类型校验:
python
class Typed:
def __init__(self, name, expected_type):
self.name = name
self.expected_type = expected_type
def __get__(self, instance, owner):
if instance is None:
return self
return instance.__dict__[self.name]
def __set__(self, instance, value):
if not isinstance(value, self.expected_type):
raise TypeError(f"{self.name} 必须是 {self.expected_type.__name__},收到 {type(value).__name__}")
instance.__dict__[self.name] = value
class Person:
name = Typed("name", str)
age = Typed("age", int)
def __init__(self, name, age):
self.name = name
self.age = age
p = Person("Alice", 30)
print(p.name, p.age) # Alice 30
# p.age = "三十" # TypeError: age 必须是 int,收到 str
5. 元类:类的「类」
5.1 什么是元类
元类(Metaclass)是创建类的类 。默认情况下,所有类都是由 type 创建的:
python
class Foo:
pass
print(type(Foo)) # <class 'type'>
元类的作用是在类创建时拦截并修改类的定义,常用于:
- 自动注册类
- 校验类定义
- 自动添加方法或属性
- 实现 ORM 中的模型映射
5.2 自定义元类
通过继承 type 并重写 __new__ 或 __init__ 来创建元类:
python
class Meta(type):
def __new__(mcs, name, bases, namespace):
print(f"创建类:{name}")
# 自动为所有类添加一个属性
namespace["created_by"] = "Meta"
return super().__new__(mcs, name, bases, namespace)
class MyClass(metaclass=Meta):
pass
print(MyClass.created_by) # Meta
5.3 元类实战:自动注册
元类最常见的应用是自动注册子类:
python
class PluginRegistry(type):
registry = {}
def __new__(mcs, name, bases, namespace):
cls = super().__new__(mcs, name, bases, namespace)
if name != "Plugin":
mcs.registry[name] = cls
return cls
class Plugin(metaclass=PluginRegistry):
"""所有插件的基类"""
class AudioPlugin(Plugin):
pass
class VideoPlugin(Plugin):
pass
print(PluginRegistry.registry)
# {'AudioPlugin': <class '__main__.AudioPlugin'>, 'VideoPlugin': <class '__main__.VideoPlugin'>}
5.4 元类实战:单例模式
利用元类实现单例,比装饰器方式更干净:
python
class SingletonMeta(type):
_instances = {}
def __call__(cls, *args, **kwargs):
if cls not in cls._instances:
cls._instances[cls] = super().__call__(*args, **kwargs)
return cls._instances[cls]
class Database(metaclass=SingletonMeta):
def __init__(self):
print("初始化数据库连接")
db1 = Database() # 初始化数据库连接
db2 = Database() # 不再打印,直接返回已有实例
print(db1 is db2) # True
5.5 元类实战:ORM 字段映射
元类在 ORM 框架中应用广泛,这里模拟一个简化版:
python
class Field:
def __init__(self, column_type):
self.column_type = column_type
class ModelMeta(type):
def __new__(mcs, name, bases, namespace):
if name == "Model":
return super().__new__(mcs, name, bases, namespace)
# 收集字段定义
fields = {}
for key, value in namespace.items():
if isinstance(value, Field):
fields[key] = value
# 从类命名空间中移除字段,避免被当作类属性
for key in fields:
del namespace[key]
namespace["_fields"] = fields
namespace["_table_name"] = name.lower()
return super().__new__(mcs, name, bases, namespace)
class Model(metaclass=ModelMeta):
def __init__(self, **kwargs):
for key, value in kwargs.items():
setattr(self, key, value)
def save(self):
fields = ", ".join(self._fields.keys())
print(f"INSERT INTO {self._table_name} ({fields}) VALUES ...")
class User(Model):
id = Field("INT")
name = Field("VARCHAR(100)")
user = User(id=1, name="Alice")
user.save() # INSERT INTO user (id, name) VALUES ...
5.6 元类与 __init_subclass__
Python 3.6+ 提供了 __init_subclass__,在多数场景下可以替代元类,更简洁:
python
class Plugin:
registry = {}
def __init_subclass__(cls, **kwargs):
super().__init_subclass__(**kwargs)
Plugin.registry[cls.__name__] = cls
class AudioPlugin(Plugin):
pass
print(Plugin.registry) # {'AudioPlugin': <class '__main__.AudioPlugin'>}
6. 三者的协同:一个综合示例
下面用一个「带校验的 ORM 模型」综合展示魔术方法、描述符与元类的协作:
python
# 1. 描述符:字段类型校验
class Field:
def __init__(self, name, column_type):
self.name = name
self.column_type = column_type
def __get__(self, instance, owner):
if instance is None:
return self