如何在Sklearn Pipeline中运行CatBoost

介绍

CatBoost的一大特点是可以很好的处理类别特征(Categorical Features)。当我们将其结合到Sklearn的Pipeline中时,会发生如下报错:

shell 复制代码
_catboost.CatBoostError: 'data' is numpy array of floating point numerical type, it means no categorical features, but 'cat_features' parameter specifies nonzero number of categorical features

因为CatBoost需要检查输入训练数据pandas.DataFrame中对应的cat_features。如果我们使用Pipeline后,输入给.fit()的数据是被修改过的,DataFrame中的columns的名字变为了数字。

解决方案

我们提前在数据上使用Pipeline,然后将原始数据转换为Pipeline处理后的数据,然后检索出其中包含的类别特征,将其传输给Catboost。

python 复制代码
# define your pipeline
pipeline = Pipeline(steps=[
    ('preprocessor', preprocessor),
    ('classifier', model),
])

preprocessor.fit(X_train)
transformed_X_train = pd.DataFrame(preprocessor.transform(X_train)).convert_dtypes()

new_cat_feature_idx = [transformed_X_train.columns.get_loc(col) for col in transformed_X_train.select_dtypes(include=['int64', 'bool']).columns]

pipeline.fit(X_train, y_train, classifier__cat_features=new_cat_feature_idx)
相关推荐
zhaowangji3 分钟前
mujoco仿真(机械臂推动正方体到指定位置)
人工智能·python
AI分享猿4 分钟前
百智云联网智能生图电商商品图场景
人工智能
数据管道工6 分钟前
解析一个老网站:GBK 编码、页面结构漂移与限流退避
python
小宋10217 分钟前
OpenTelemetry GenAI可观测性实战:串起模型、工具、Token与错误
java·人工智能·算法·贪心算法
只睡四小时8 分钟前
AI 生成 PPTX:11 页课件编译出 429 个形状
javascript·python·pptx·ai生成ppt·ooxml
leisoo809710 分钟前
股票筹码分布怎么用获利比例成本区间与集中度实战 IG50免费开源股票数据API接口
开发语言·jvm·数据库·python·开源
ControlM11 分钟前
从官方 CDN 里扒出 TRAE (TraeCode) 历史版本安装包
python·逆向·trae
tianyuanwo13 分钟前
Python 属性查找陷阱:从 `AttributeError: ‘X‘ object has no attribute ‘_children‘` 说起
python
yichengerp14 分钟前
国内中小电子工厂用哪个erp系统好?
大数据·运维·人工智能·云计算·制造
q275513004215 分钟前
微纳代理 WN8034F 国产降噪音频方案
人工智能·语音识别