多模态论文阅读之VLMo

VLMo泛读

Title

VLMo:Unified Vision_Langugae Pre-Training with Mixture-of-Modality-Experts

Motivation

  1. CLIP和ALIGN都采用dual-encoder 的方式分别编码图像和文本,模态之间的交互采用cosine similarity ,这种方法对retrieval tasks(检索任务)及其有效;但是如此shallow intersection between images and text is not enough to handle complex VL classfication tasks. In ViLT, find that CLIP gives a relatively low accuracy on visual resaoning(VR) task; 后来一系列的tasks,采用的fusion encoder 的方式,即一开始分来images and text 然后采用transformer的encoder 做cross-modal 的intersection,这样的architecture 弥补了dual encoder architecture的drawback,But it requires to jointly encode all possible image-text pairs to compute similarity scores for retrieval tasks. The quadratic time complexity leads to a much slower inference speed than the dual-encoder models models whos time complexity is linear. So, 有没**有一种融合上述两种架构的方法呢?**做检索任务的时候用 dual-encoder架构,做classfication的时候用fusion encoder,所以本文提出了Mixture-of-Modality-Experts
  2. VLMo的训练loss是image-text contrastive(ITC), image-text matching(ITM), masked Language modeling(MLM)和ALBEF是一样的。提出了一个stagewise的预训练方法分别vision 和NLP中的large-scale corpus:首先在vision上训练好,再预训练language experts on text-only data,最后将模型用于vision-language pre-training。

Contribution

  1. 模型上的改进:Mixture-of-Modality-Experts
  2. 训练方式上的改进:分阶段模型预训练

Model

  1. 模型中所有的multi-head self-Attention都是share weights的
  2. 模型inference的时候很灵活,要做那个任务,切换到那个架构上就行。
  3. 分阶段训练策略

Expertiments

  1. 比ALBEF性能好很多
  2. 在更大的数据集上训练,数据变得更好。

Summary

  1. 就是把transformer里的encoder中的FFN分为了几个FFN
相关推荐
Rocky Ding*11 小时前
深入浅出完整解析FLUX.2、Seedream(即梦)、Z-image、Qwen-Image、GLM-Image核心基础知识
论文阅读·人工智能·深度学习·机器学习·aigc·扩散模型·ai-native
m4Rk_13 小时前
【论文阅读】Agent 记忆机制(91):THEANINE——不删除过时记忆,让 Agent 记住事情是如何变化的
论文阅读·人工智能·学习·开源·github
Rocky Ding*19 小时前
深度解析LlamaGen核心基础知识
论文阅读·人工智能·深度学习·机器学习·aigc·ai-native·llamagen
a puppy slide 葱2 天前
ICLR | 2019 | DARTS:可微架构搜索
论文阅读·darts·神经网络架构搜索
m4Rk_2 天前
【论文阅读】Agent 记忆机制(90):HyperMem——用超图建模长期记忆中的高阶关联
论文阅读·人工智能·学习·开源·github
m4Rk_3 天前
【论文阅读】Agent 记忆机制(89):PersonaAgent——构建 Memory、Persona 与 Action 的持续反馈闭环
论文阅读·人工智能·学习·开源·github
Rocky Ding*3 天前
MaskGIT技术深度解析:图像生成如何从逐Token排队走向掩码并行预测
论文阅读·人工智能·深度学习·机器学习·aigc·ai-native·maskgit
Rocky Ding*3 天前
Muse技术深度解析:用掩码并行生成图像,速度、语义与编辑能力如何同时成立
论文阅读·人工智能·深度学习·机器学习·aigc·ai-native·muse
xx_xxxxx_3 天前
论文阅读-RoTTA
论文阅读·人工智能·深度学习·机器学习
大模型任我行4 天前
阿里:通义千问3.8全能版发布
人工智能·语言模型·自然语言处理·论文笔记