SOTAVerified

Multimodal Machine Translation

Multimodal machine translation is the task of doing machine translation with multiple data sources - for example, translating "a bird is flying over water" + an image of a bird over water to German text.

( Image credit: Findings of the Third Shared Task on Multimodal Machine Translation )

Papers

Showing 150 of 108 papers

TitleStatusHype
Seamless: Multilingual Expressive and Streaming Speech TranslationCode6
Attention Is All You NeedCode3
Distill the Image to Nowhere: Inversion Knowledge Distillation for Multimodal Machine TranslationCode1
MSCTD: A Multimodal Sentiment Chat Translation DatasetCode1
Dynamic Context-guided Capsule Network for Multimodal Machine TranslationCode1
On Vision Features in Multimodal Machine TranslationCode1
VALHALLA: Visual Hallucination for Machine TranslationCode1
Neural Machine Translation with Phrase-Level Universal Visual RepresentationsCode1
Self-Knowledge Distillation with Progressive Refinement of TargetsCode1
3AM: An Ambiguity-Aware Multi-Modal Machine Translation DatasetCode1
CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine TranslationCode1
BigVideo: A Large-scale Video Subtitle Translation Dataset for Multimodal Machine TranslationCode1
VISA: An Ambiguous Subtitles Dataset for Visual Scene-Aware Machine TranslationCode1
Cross-lingual Visual Pre-training for Multimodal Machine TranslationCode1
M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-trainingCode1
Multimodal Transformer for Multimodal Machine TranslationCode1
Scene Graph as Pivoting: Inference-time Image-free Unsupervised Multimodal Machine Translation with Visual Scene HallucinationCode1
Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive EvaluationCode1
BERTGEN: Multi-task Generation through BERTCode1
CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation0
A Survey of Vision-Language Pre-training from the Lens of Multimodal Machine Translation0
Iterative Adversarial Attack on Image-guided Story Ending Generation0
Input Combination Strategies for Multi-Source Transformer Decoder0
A Dataset and Reranking Method for Multimodal MT of User-Generated Image Captions0
A Shared Task on Multimodal Machine Translation and Crosslingual Image Description0
Doubly-Attentive Decoder for Multi-modal Neural Machine Translation0
Doubly Attentive Transformer Machine Translation0
Increasing Visual Awareness in Multimodal Neural Machine Translation from an Information Theoretic Perspective0
Adaptive Fusion Techniques for Multimodal Data0
Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps Grounding0
Investigating the Decoders of Maximum Likelihood Sequence Models: A Look-ahead Approach0
Low Resource Multimodal Neural Machine Translation of English-Hindi in News Domain0
Detecting Concrete Visual Tokens for Multimodal Machine Translation0
Generating Image Descriptions using Multilingual Data0
Debiasing Word Embeddings Improves Multimodal Machine Translation0
DCU-UvA Multimodal MT System Report0
CUNI System for WMT16 Automatic Post-Editing and Multimodal Translation Tasks0
Findings of the Second Shared Task on Multimodal Machine Translation and Multilingual Image Description0
A Visually-Grounded Parallel Corpus with Phrase-to-Region Linking0
Adversarial Evaluation of Multimodal Machine Translation0
Florenz: Scaling Laws for Systematic Generalization in Vision-Language Models0
Generalization algorithm of multimodal pre-training model based on graph-text self-supervised training0
Hindi Visual Genome: A Dataset for Multimodal English-to-Hindi Machine Translation0
Generative Imagination Elevates Machine Translation0
Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation0
Good for Misconceived Reasons: Revisiting Neural Multimodal Machine Translation0
Grounded Word Sense Translation0
Gumbel-Attention for Multi-modal Machine Translation0
Findings of the 2018 Conference on Machine Translation (WMT18)0
Findings of the 2017 Conference on Machine Translation (WMT17)0
Show:102550
← PrevPage 1 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1delMeteor (EN-FR)74.6Unverified
2ERNIE-UniX2BLEU (EN-DE)49.3Unverified
3IKD-MMTBLEU (EN-DE)41.28Unverified
4DCCNBLEU (EN-DE)39.7Unverified
5CaglayanBLEU (EN-DE)39.4Unverified
6Gumbel-Attention MMTBLEU (EN-DE)39.2Unverified
7Multimodal TransformerBLEU (EN-DE)38.7Unverified
8ImagiTBLEU (EN-DE)38.4Unverified
9del+objBLEU (EN-DE)38Unverified
10VMMTFBLEU (EN-DE)37.6Unverified
#ModelMetricClaimedVerifiedStatus
1ViTABLEU (EN-HI)51.6Unverified
#ModelMetricClaimedVerifiedStatus
1ViTABLEU (EN-HI)44.6Unverified