Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers | Zendy

Lisa Anne Hendricks | Zendy; John Mellor | Zendy; Rosalia Schneider | Zendy; Jean-Baptiste Alayrac | Zendy; Aida Nematzadeh | Zendy

AI Assistant Blog Pricing

Home ZAIA Blog

Open Access

Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers

Author(s) -

Lisa Anne Hendricks,

John Mellor,

Rosalia Schneider,

Jean-Baptiste Alayrac,

Aida Nematzadeh

Publication year - 2021

Publication title -

transactions of the association for computational linguistics

Language(s) - English

Resource type - Journals

ISSN - 2307-387X

DOI - 10.1162/tacl_a_00385

Subject(s) - computer science , transformer , artificial intelligence , machine learning , multimodal learning , physics , quantum mechanics , voltage

Recently, multimodal transformer models have gained popularity because their performance on downstream tasks suggests they learn rich visual-linguistic representations. Focusing on zero-shot image retrieval tasks, we study three important factors that can impact the quality of learned representations: pretraining data, the attention mechanism, and loss functions. By pretraining models on six datasets, we observe that dataset noise and language similarity to our downstream task are important indicators of model performance. Through architectural analysis, we learn that models with a multimodal attention mechanism can outperform deeper models with modality-specific attention mechanisms. Finally, we show that successful contrastive losses used in the self-supervised learning literature do not yield similar performance gains when used in multimodal transformers.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.

Having issues? You can contact us here

Accelerating Research