机器学习(六) — 评估模型

Evaluate model

1 test set

  1. split the training set into training set and a test set
  2. the test set is used to evaluate the model

1. linear regression

compute test error

J t e s t ( w ⃗ , b ) = 1 2 m t e s t ∑ i = 1 m t e s t ( f ( x t e s t ( i ) ) − y t e s t ( i ) ) 2 J_{test}(\vec w, b) = \frac{1}{2m_{test}}\sum_{i=1}^{m_{test}} \left (f(x_{test}\^{(i)}) - y_{test}\^{(i)})\^2 \\right Jtest(w ,b)=2mtest1i=1∑mtest(f(xtest(i))−ytest(i))2

2. classification regression

compute test error

J t e s t ( w ⃗ , b ) = − 1 m t e s t ∑ i = 1 m t e s t y t e s t ( i ) l o g ( f ( x t e s t ( i ) ) ) + ( 1 − y t e s t ( i ) ) l o g ( 1 − f ( x t e s t ( i ) ) J_{test}(\vec w, b) = -\frac{1}{m_{test}}\sum_{i=1}^{m_{test}} \left y_{test}\^{(i)}log(f(x_{test}\^{(i)})) + (1 - y_{test}\^{(i)})log(1 - f(x_{test}\^{(i)}) \\right Jtest(w ,b)=−mtest1i=1∑mtestytest(i)log(f(xtest(i)))+(1−ytest(i))log(1−f(xtest(i))

2 cross-validation set

  1. split the training set into training set, cross-validation set and test set
  2. the cross-validation set is used to automatically choose the better model, and the test set is used to evaluate the model that chosed

3 bias and variance

  1. high bias: J t r a i n J_{train} Jtrain and J c v J_{cv} Jcv is both high
  2. high variance: J t r a i n J_{train} Jtrain is low, but J c v J_{cv} Jcv is high
  1. if high bias: get more training set is helpless
  2. if high variance: get more training set is helpful

4 regularization

  1. if λ \lambda λ is too small, it will lead to overfitting(high variance)
  2. if λ \lambda λ is too large, it will lead to underfitting(high bias)

5 method

  1. fix high variance:
    • get more training set
    • try smaller set of features
    • reduce some of the higher-order terms
    • increase λ \lambda λ
  2. fix high bias:
    • get more addtional features
    • add polynomial features
    • decrease λ \lambda λ

6 neural network and bias variance

  1. a bigger network means a more complex model, so it will solve the high bias
  2. more data is helpful to solve high variance
  1. it turns out that a bigger(may be overfitting) and well regularized neural network is better than a small neural network
相关推荐
FL16238631292 分钟前
无人机视角海边沙滩垃圾检测数据集VOC+YOLO格式603张3类别
人工智能·yolo
Zguigo8 分钟前
【CUDA6】CUDA Stream 是什么,为什么 CUDA 是异步执行,如何正确测量 GPU 时间以及多个任务如何重叠执行
人工智能·pytorch·深度学习
C++ 老炮儿的技术栈25 分钟前
不一样的数值交换
数据结构·c++·人工智能·算法·c·csdn开发云
C++ 老炮儿的技术栈34 分钟前
main函数之后的调用
c语言·c++·人工智能·qt·mfc·c
zhizhizhuzhuxia39 分钟前
情感解说短视频的配音选型:从工程视角看情感音色与批量出片
人工智能·音视频
YOLO数据集集合41 分钟前
光伏组件异常检测数据集 | 光伏组件 异常检测 红外检测 热斑识别 光伏巡检9147期
人工智能·yolo·目标检测·语言模型·太阳能板·光伏组件
HIT_Weston1 小时前
242、【AI】【模型部署】基座模型研究:注意力机制深入
人工智能·模型部署
智碳能碳管理平台1 小时前
碳排放核算软件选型:历史台账断层补救与锁账回放实操指南
人工智能·数字孪生·能碳管理系统·智碳能碳管理平台·企业能碳管理系统·碳排放核算软件·绿色工厂申报saas
IT_陈寒1 小时前
SpringBoot自动配置坑了我一把,原来是这样绕过去的
前端·人工智能·后端
hahaha60161 小时前
Shades-of-Gray (SoG) 算法--白平衡算法
人工智能·嵌入式硬件·算法·计算机视觉