transformers in tabular tiny survey 2024.4.8

推荐阅读

TabLLM

pmlr2023,

Few-shot Classification of Tabular Data with Large Language Models

方法

使用把tabular数据序列化成文字的方法进行classification。
使用的序列化方法有几个,有人工也有AI生成。

效果

做few shot learning的效果
看上去一般。

TransTab

Learning Transferable Tabular Transformers Across Tables

方法

属于transfer learning的方法。对category、binary和numeric值进行embedding后再进行transformers最后进行classification。

使用场景

原文:

  • S(1) Transfer learning . We collect data tables from multiple cancer trials for testing the efficacy

of the same drug on different patients. These tables were designed independently with overlapping

columns. How do we learn ML models for one trial by leveraging tables from all trials?

  • S(2) Incremental learning . Additional columns might be added over time. For example, additional

features are collected across different trial phases. How do we update the ML models using tables

from all trial phases?

  • S(3) Pretraining+Finetuning . The trial outcome label (e.g., mortality) might not be always available

from all table sources. Can we benefit pretraining on those tables without labels? How do we finetune

the model on the target table with labels?

  • S(4) Zero-shot inference . We model the drug efficacy based on our trial records. The next step is to

conduct inference with the model to find patients that can benefit from the drug. However, patient

tables do not share the same columns as trial tables so direct inference is not possible.

效果

具体看原文吧,与当时的baseline比有提升。

MET

Masked Encoding for Tabular Data

tabtransformer

2020年,arxiv,TabTransformer: Tabular Data Modeling Using Contextual Embeddings

方法

transformer无监督训练,mlp监督训练。

原文

we introduce a pre-training procedure to train the Transformer layers using unlabeled data . This is followed by fine-tuning of the pre-trained Transformer layers along with the top MLP layer using the labeled data

效果

跟mlp

跟其他模型

tabnet

2020, arxiv,Google Cloud AI,Attentive Interpretable Tabular Learning, 封装的非常好,都可以当工具包使用了。

方法

跟transformer没关系的。
feature selection用的是17年的某个选择模型,最后agg一下做predict。

相关推荐
冬奇Lab4 分钟前
开源项目第198期:Omarchy — DHH 打造的「意见鲜明」Linux 桌面系统
人工智能·开源·操作系统
zww89491118 分钟前
AI智能商城APP定制开发实战:从需求到上线完整指南
大数据·人工智能
wangjialelele11 分钟前
RAG 检索增强生成全面解析:从知识库构建到检索、重排序与生成优化
人工智能·ai·agent·rag·检索增强生成
从零开始学习人工智能25 分钟前
Labelme AI-Points 功能背后的ONNX模型:下载、配置与使用解析
人工智能
招财小梗38 分钟前
沈阳商贸公司AI企业服务优势几何?
大数据·人工智能·python
ZKKLLY1 小时前
生成式引擎优化赛道崛起,优质GEO服务商该如何筛选评估
大数据·人工智能
IT古董1 小时前
FDE(Forward Deployed Engineer)详解:AI时代正在崛起的新型工程师
大数据·人工智能·数据挖掘
ZISHU_9872 小时前
用 TLabel 给 SynTouch BioTac 数据做语义标注:从原始信号到结构化标注
开发语言·人工智能·python·数据·机器人触觉
hh9502 小时前
gent Plan x DeepSeek Harness:多 Agent 协作场景下的中央编排实践
人工智能·wpf·adg·火山引擎·agent plan·adg成都社区
cczixun2 小时前
2026 智能体平台落地实战:从 POC 验证到生产上线,避开选型与交付陷阱
人工智能·落地·选型·智能体·2026