模型评估与过拟合——交叉验证、正则化与指标解读

个人主页: > for_ever_love__ <
其他栏目: > 我想学python了 <

其他栏目: > iOS项目总结大全 <

其他栏目: > iOS UI <

文章目录

模型评估与过拟合------交叉验证、正则化与指标解读

承上:前面几篇我们训练了各种模型,但还没系统回答「它到底好不好」。

本篇:本篇解决评估问题:交叉验证、正则化、指标解读,收官阶段 3。

启下:下一篇进入阶段 4《神经网络与反向传播原理》,正式进入深度学习。

学完这一节,你能动手做:

  1. 用学习曲线诊断过拟合与欠拟合,并知道该加数据还是加正则
  2. 用 K 折交叉验证得到稳定可靠的评估指标
  3. 在不平衡数据上正确选用 recall/F1/AUC 并按业务成本调阈值

到这里你已经会训练模型了。但一个更关键的问题来了:

你怎么知道模型是真的学会了,还是把训练答案背下来了?

这是机器学习的核心难题,也是新手最容易翻车的地方:训练集 99% 准确率,上线就崩。本篇系统解决三件事:

  1. 怎么可靠地评估(交叉验证)
  2. 怎么治过拟合(正则化)
  3. 用什么指标衡量(准确率远远不够)

这套东西后面会一直用到:大模型微调要防过拟合、RAG 系统要评估召回率、上线前要跑评测集------全是这一篇的内容。

一、过拟合与欠拟合:偏差-方差权衡

先用一个可复现的实验把概念钉死:

python 复制代码
import numpy as np
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline
from sklearn.metrics import mean_squared_error
import matplotlib.pyplot as plt

np.random.seed(0)
n = 30
X = np.sort(np.random.rand(n)) * 4 - 2                  # x ∈ [-2, 2]
y_true = 0.5 * X ** 3 - X                               # 真实规律(三次)
y = y_true + np.random.randn(n) * 1.5                   # 加噪声
X = X.reshape(-1, 1)

def poly_fit_eval(degree):
    model = make_pipeline(PolynomialFeatures(degree), LinearRegression())
    model.fit(X, y)
    return np.sqrt(mean_squared_error(y, model.predict(X))), model

for d in [1, 3, 9, 15]:
    rmse, _ = poly_fit_eval(d)
    print(f"degree={d:>2}  训练 RMSE = {rmse:.3f}")

degree 越大训练误差越小(15 次多项式几乎完美穿过每个点),但那是把噪声也当规律背下来了。

python 复制代码
# 看真实泛化:在密集的新点上比较
X_grid = np.linspace(-2, 2, 200).reshape(-1, 1)
y_grid_true = 0.5 * X_grid.ravel() ** 3 - X_grid.ravel()

fig, axes = plt.subplots(1, 3, figsize=(15, 4))
for ax, d in zip(axes, [1, 3, 15]):
    _, model = poly_fit_eval(d)
    ax.scatter(X, y, s=20, alpha=0.6, label="train data")
    ax.plot(X_grid, y_grid_true, "g--", lw=2, label="true function")
    ax.plot(X_grid, model.predict(X_grid), "r", lw=2, label=f"degree {d}")
    ax.set_title(f"degree={d}")
    ax.legend(); ax.set_ylim(-8, 8)
plt.tight_layout(); plt.show()
欠拟合 Underfit 刚好 过拟合 Overfit
训练误差 高 低 极低
验证误差 高 低 高
原因 模型太简单 --- 模型太复杂/数据太少
表现 学不到规律 ✅ 背答案,泛化差
解决 加大模型、加特征、减正则 --- 加数据、正则化、早停、简化模型

偏差-方差分解:误差 = 偏差²(模型假设错) + 方差(对数据敏感) + 噪声。欠拟合是偏差大,过拟合是方差大。

二、学习曲线:一眼诊断出问题

python 复制代码
from sklearn.model_selection import learning_curve
from sklearn.datasets import load_digits
from sklearn.ensemble import RandomForestClassifier

digits = load_digits()
Xd, yd = digits.data, digits.target

def plot_learning_curve(estimator, X, y, title, ax):
    sizes, train_scores, val_scores = learning_curve(
        estimator, X, y, train_sizes=np.linspace(0.1, 1.0, 8),
        cv=5, scoring="accuracy", n_jobs=-1)
    ax.plot(sizes, train_scores.mean(1), "o-", label="train")
    ax.fill_between(sizes, train_scores.mean(1) - train_scores.std(1),
                    train_scores.mean(1) + train_scores.std(1), alpha=0.2)
    ax.plot(sizes, val_scores.mean(1), "o-", label="val")
    ax.fill_between(sizes, val_scores.mean(1) - val_scores.std(1),
                    val_scores.mean(1) + val_scores.std(1), alpha=0.2)
    ax.set_title(title); ax.set_xlabel("training samples"); ax.set_ylabel("accuracy")
    ax.legend(loc="lower right"); ax.grid(alpha=0.3)

fig, axes = plt.subplots(1, 2, figsize=(13, 4))
plot_learning_curve(RandomForestClassifier(max_depth=2, n_estimators=50),
                    Xd, yd, "Underfit (shallow tree)", axes[0])
plot_learning_curve(RandomForestClassifier(max_depth=None, n_estimators=50),
                    Xd, yd, "Good fit (full tree)", axes[1])
plt.tight_layout(); plt.show()

怎么读学习曲线:

  • 两条曲线都低且靠拢 → 欠拟合,加大模型;
  • 训练高、验证低、差距大 → 过拟合,加数据/加正则;
  • 验证曲线还在上升 → 继续加数据有用;
  • 两条都高且靠拢 → 收敛了,加数据收益递减。

三、交叉验证:让评估更可靠

单次 train/test 划分有偶然性。K 折交叉验证:把数据分成 K 份,轮流用其中 1 份当验证集,跑 K 次取平均。

python 复制代码
from sklearn.model_selection import cross_val_score, KFold, StratifiedKFold
from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline

iris = load_iris()
Xi, yi = iris.data, iris.target

pipe = make_pipeline(StandardScaler(), LogisticRegression(max_iter=200))

# 普通 K 折
scores = cross_val_score(pipe, Xi, yi, cv=5, scoring="accuracy")
print("5-fold scores:", scores.round(3))
print(f"mean={scores.mean():.4f}  std={scores.std():.4f}")

# 分层 K 折(分类任务推荐:保证每折类别比例一致)
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores_skf = cross_val_score(pipe, Xi, yi, cv=skf, scoring="accuracy")
print("stratified 5-fold:", scores_skf.round(3), f"mean={scores_skf.mean():.4f}")

为什么分类要用 StratifiedKFold? 普通 K 折可能出现某一折里某个类一个样本都没有,评估结果失真。

时间序列不能随机切:

python 复制代码
from sklearn.model_selection import TimeSeriesSplit

tscv = TimeSeriesSplit(n_splits=5)
# 用法同 KFold,但保证"用过去预测未来",不会用未来数据训练

用未来的数据训练去预测过去 = 数据泄漏,回测收益虚高,实盘必亏。

四、正则化:给模型戴紧箍咒

4.1 原理

正则化在损失里加一项"惩罚",逼模型保持简单:

复制代码
L_total = L_data + λ · R(w)
  • L2(Ridge) :R(w) = Σw² → 让权重变小但不为 0,更平滑
  • L1(Lasso) :R(w) = Σ|w| → 让部分权重变成 0,自动特征选择
python 复制代码
from sklearn.linear_model import Ridge, Lasso
from sklearn.datasets import make_regression

Xr, yr = make_regression(n_samples=100, n_features=30, noise=15,
                         n_informative=5, random_state=0)   # 只有5个真特征

for alpha in [0.001, 0.1, 1.0, 10.0, 100.0]:
    l2 = Ridge(alpha=alpha).fit(Xr, yr)
    l1 = Lasso(alpha=alpha, max_iter=5000).fit(Xr, yr)
    print(f"alpha={alpha:>7}  L2非零权重={np.sum(l2.coef_!=0):>2}  "
          f"L1非零权重={np.sum(l1.coef_!=0):>2}  L1 系数和={np.sum(np.abs(l1.coef_)):.1f}")

L1 会把 30 个特征里没用的压成 0------这就是"自动特征选择",在特征很多时极其有用。

4.2 直观理解:为什么小权重 = 简单模型?

python 复制代码
# 用多项式看:系数小 → 曲线平滑
for degree, alpha in [(15, 0.0), (15, 1.0), (15, 50.0)]:
    m = make_pipeline(PolynomialFeatures(degree), Ridge(alpha=alpha)).fit(X, y)
    pred = m.predict(X_grid)
    print(f"degree=15 alpha={alpha:>5}  系数最大绝对值={np.abs(m[-1].coef_).max():>8.2f}  "
          f"曲线波动={np.std(np.diff(pred)):.3f}")

权重小 → 输入变化时输出变化小 → 不会因为个别噪声点就剧烈扭曲 → 泛化更好。

4.3 Dropout:神经网络专属正则

思路完全不同:训练时随机"关掉"一部分神经元,逼网络不能依赖任何单个神经元。

python 复制代码
def dropout(x, p=0.5, training=True):
    """Inverted dropout:训练时缩放,推理时直接用"""
    if not training:
        return x
    mask = (np.random.rand(*x.shape) > p).astype(float)
    return x * mask / (1 - p)     # 除以 (1-p) 保持期望不变

x = np.ones(10)
print("训练时:", dropout(x, 0.5, True))
print("推理时:", dropout(x, 0.5, False))

# 验证期望不变
samples = [dropout(x, 0.5, True).mean() for _ in range(5000)]
print("5000次平均:", round(np.mean(samples), 3), "(应接近 1.0)")

/ (1-p) 这步很关键(Inverted dropout):保证训练和推理时期望一致,推理时什么都不用改。

4.4 早停 Early Stopping

最简单也最实用的正则:验证集指标不再提升就停。

python 复制代码
def train_with_early_stopping(X_tr, y_tr, X_val, y_val, max_epochs=500, patience=20):
    best_val, best_epoch, wait = float("inf"), 0, 0
    history = {"train": [], "val": []}
    # 这里用一个简单模型示意;实际换成你的模型
    from sklearn.linear_model import SGDRegressor
    model = SGDRegressor(learning_rate="constant", eta0=0.01, warm_start=True)

    for ep in range(max_epochs):
        model.fit(X_tr, y_tr)                        # partial fit 式更新
        tr = mean_squared_error(y_tr, model.predict(X_tr))
        va = mean_squared_error(y_val, model.predict(X_val))
        history["train"].append(tr); history["val"].append(va)

        if va < best_val - 1e-6:
            best_val, best_epoch, wait = va, ep, 0
        else:
            wait += 1
            if wait >= patience:
                print(f"早停于 epoch {ep},最优在 epoch {best_epoch} (val={best_val:.4f})")
                break
    return model, history

# 说明:真实训练中应该在 best_val 时保存模型权重,结束后加载回来

早停一定要配合"保存最优权重"------停下的那一刻模型已经不是最优的了。

正则化手段全家福:

方法 做法 适用
L1/L2 正则 损失加权重惩罚 线性模型、神经网络(weight_decay)
Dropout 随机丢弃神经元 神经网络
早停 验证指标不涨就停 所有迭代训练
数据增强 造更多样本 图像/文本
简化模型 减层、减参数 所有
BatchNorm 稳定层间分布 神经网络
集成 多个模型投票 所有

五、分类指标:准确率是个骗子

5.1 为什么准确率不可靠

假设 1000 个样本里只有 10 个是欺诈交易。模型全部预测"正常",准确率 99% ------ 但它一条欺诈都没抓到,毫无价值。

python 复制代码
from sklearn.metrics import (accuracy_score, precision_score, recall_score,
                             f1_score, roc_auc_score, confusion_matrix,
                             classification_report, roc_curve, precision_recall_curve)
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression

# 构造不平衡数据(正类占 5%)
Ximb, yimb = make_classification(n_samples=3000, n_features=10, n_informative=5,
                                 weights=[0.95, 0.05], flip_y=0.02, random_state=42)
X_tr, X_te, y_tr, y_te = train_test_split(Ximb, yimb, test_size=0.3,
                                          random_state=42, stratify=yimb)
clf = LogisticRegression(max_iter=1000).fit(X_tr, y_tr)
pred = clf.predict(X_te)

print("正样本比例:", round(y_te.mean(), 3))
print("accuracy  :", round(accuracy_score(y_te, pred), 4))
print("precision :", round(precision_score(y_te, pred), 4))
print("recall    :", round(recall_score(y_te, pred), 4))
print("f1        :", round(f1_score(y_te, pred), 4))
print("roc_auc   :", round(roc_auc_score(y_te, clf.predict_proba(X_te)[:, 1]), 4))

accuracy 可能 0.95+,但 recall 只有 0.5 左右------一半的正类被漏掉了。

5.2 四个指标怎么记

混淆矩阵是原点:

复制代码
                 预测正    预测负
实际正    TP        FN
实际负    FP        TN
  • Precision 精确率 = TP/(TP+FP):我报的警里,多少是真的?(宁缺毋滥)
  • Recall 召回率 = TP/(TP+FN):真的里面,我抓到多少?(宁枉勿纵)
  • F1 = 二者的调和平均:2PR/(P+R)
  • AUC:随机抽一个正样本和一个负样本,模型给正样本打分更高的概率

选择指南:

场景 重点指标 理由
癌症筛查、欺诈检测 Recall 漏检代价太大
垃圾邮件、推荐推送 Precision 误杀正常邮件/推错内容惹人烦
一般分类 F1 / AUC 平衡
排序、召回(RAG) Recall@k / MRR 看前 k 个里有没有
python 复制代码
print("混淆矩阵:\n", confusion_matrix(y_te, pred))
print(classification_report(y_te, pred, target_names=["负类", "正类"]))

5.3 阈值调节:P-R 的取舍

模型输出的是概率,判定阈值默认是 0.5,但阈值可以调:

python 复制代码
proba = clf.predict_proba(X_te)[:, 1]
for thr in [0.2, 0.3, 0.5, 0.7]:
    p = (proba >= thr).astype(int)
    print(f"阈值={thr}  precision={precision_score(y_te, p, zero_division=0):.3f}  "
          f"recall={recall_score(y_te, p):.3f}  f1={f1_score(y_te, p):.3f}")

# PR 曲线与 ROC 曲线
prec, rec, _ = precision_recall_curve(y_te, proba)
fpr, tpr, _ = roc_curve(y_te, proba)

fig, axes = plt.subplots(1, 2, figsize=(12, 4))
axes[0].plot(rec, prec, lw=2); axes[0].set_xlabel("Recall"); axes[0].set_ylabel("Precision")
axes[0].set_title("PR Curve"); axes[0].grid(alpha=0.3)
axes[1].plot(fpr, tpr, lw=2); axes[1].plot([0,1],[0,1],"k--")
axes[1].set_xlabel("FPR"); axes[1].set_ylabel("TPR"); axes[1].set_title("ROC Curve")
axes[1].grid(alpha=0.3)
plt.tight_layout(); plt.show()

阈值低 → 召回高、精确低 (宁可错杀);阈值高 → 相反。业务上要根据成本选,不是一律 0.5。

不平衡数据怎么处理:

python 复制代码
# 方法1:类别权重
clf_bal = LogisticRegression(class_weight="balanced", max_iter=1000).fit(X_tr, y_tr)
pred_b = clf_bal.predict(X_te)
print("加 balanced 后 recall:", round(recall_score(y_te, pred_b), 4),
      " precision:", round(precision_score(y_te, pred_b), 4))

# 方法2:过/欠采样(需 imbalanced-learn)
# pip install imbalanced-learn
# from imblearn.over_sampling import SMOTE

六、回归指标

python 复制代码
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score

yt = np.array([100.0, 200.0, 300.0, 400.0])
yp = np.array([110.0, 190.0, 330.0, 380.0])

print("MAE :", round(mean_absolute_error(yt, yp), 2))                 # 平均差 20
print("MSE :", round(mean_squared_error(yt, yp), 2))
print("RMSE:", round(np.sqrt(mean_squared_error(yt, yp)), 2))
print("R²  :", round(r2_score(yt, yp), 4))                            # 越接近1越好
指标 特点
MAE 与原始单位一致,易解释,对异常值稳健
RMSE 单位一致,惩罚大误差,最常用
R² 相对基线(预测均值)提升了多少;1=完美,0=等于瞎猜,可负

七、超参数调优

python 复制代码
from sklearn.model_selection import GridSearchCV, RandomizedSearchCV
from sklearn.ensemble import RandomForestClassifier
from scipy.stats import randint

# 网格搜索:穷举(组合多时很慢)
param_grid = {"n_estimators": [50, 100], "max_depth": [3, 5, None]}
gs = GridSearchCV(RandomForestClassifier(random_state=42), param_grid,
                  cv=3, scoring="f1", n_jobs=-1)
gs.fit(X_tr, y_tr)
print("最优参数:", gs.best_params_, " 最优分数:", round(gs.best_score_, 4))

# 随机搜索:随机采样 N 组,性价比更高
param_dist = {"n_estimators": randint(50, 300), "max_depth": randint(2, 20)}
rs = RandomizedSearchCV(RandomForestClassifier(random_state=42), param_dist,
                        n_iter=15, cv=3, scoring="f1", random_state=42, n_jobs=-1)
rs.fit(X_tr, y_tr)
print("随机搜索最优:", rs.best_params_, round(rs.best_score_, 4))

随机搜索在超参数多时比网格搜索更高效------因为通常只有少数几个参数真正重要,网格搜索会在无关维度上浪费大量计算。

注意 :调参用验证集/交叉验证,测试集只在最后用一次。

八、常见坑与注意事项

坑 现象 解决
用测试集调参 上线效果远低于预期 严格三段划分
数据泄漏 指标虚高 scaler 只在 train fit;时序用 TimeSeriesSplit
只看 accuracy 不平衡数据下被骗 看 recall/F1/AUC
阈值一律 0.5 业务效果差 按成本调阈值
忘了 stratify 某折无某类别 StratifiedKFold
早停没存最优权重 拿到的不是最佳模型 保存 best checkpoint
交叉验证用在时序 未来预测过去 TimeSeriesSplit

九、本篇小结

  1. 过拟合 = 训练好、验证差 (方差大),欠拟合 = 都差(偏差大);学习曲线是最直接的诊断工具。
  2. K 折交叉验证 让评估更稳定;分类用 StratifiedKFold,时序用 TimeSeriesSplit。
  3. 正则化三件套:L1/L2 压权重、Dropout 随机失活、早停见好就收(记得存最优权重)。
  4. 准确率会骗人:不平衡数据必须看 precision/recall/F1/AUC,并按业务成本调阈值。
  5. 超参数调优用 GridSearchCV/RandomizedSearchCV,只在验证集上做。

阶段 3 机器学习基础到此完结。 你现在具备了评估一个模型好坏的完整方法论。

下一篇进入 阶段 4:深度学习与 PyTorch ,从神经网络与反向传播原理开始,然后正式上手 PyTorch------你会用 nn.Module 搭出第一个真正的神经网络,写出标准的训练循环。大模型的世界从这里真正开始。

本篇是《大模型开发从 0 到 1》专栏第 21 篇,阶段 3「机器学习基础」第 5 篇。专栏文章按「分类专栏」归类,顺序学习体验最佳。

相关推荐
2501_926978332 小时前
给自指系统接一个外部锚 —— 一个关于「用 AI 观察自己」的方法
人工智能·经验分享·笔记·机器学习·ai写作
for_ever_love__4 小时前
线性回归与梯度下降——从零手写一个模型
python·机器学习·线性回归·梯度下降
AOI小白新手上路6 小时前
anomalib 缺陷检测复现笔记:从跑通库到 EfficientAD 落地
人工智能·笔记·机器学习
I Am a robert girl6 小时前
稀有事件估计的迭代去对齐:从源码视角拆解重要性采样新范式
开发语言·python·机器学习·重要性采样·蒙特卡洛方法·稀有事件估计
ctlover9 小时前
Pandas进阶
人工智能·机器学习·pandas
袖清暮雨10 小时前
机器学习之K近邻算法(KNN)
人工智能·机器学习·近邻算法
Asa1213811 小时前
Nature Microbiology|不依赖系统发育的机器学习框架从基因组预测菌株水平噬菌体—宿主互作
人工智能·机器学习
for_ever_love__11 小时前
逻辑回归与 Softmax——分类问题的数学骨架
深度学习·机器学习·逻辑回归·分类算法
znx93912 小时前
手动交易 VS 量化交易:孰优孰劣?深度对比
人工智能·python·机器学习·期魔方