| On the Complementarity of Data Selection and Fine Tuning for Domain Adaptation | Oct 16, 2021 | Domain AdaptationDomain Generalization | —Unverified | 0 |
| Echo-Attention: Attend Once and Get N Attentions for Free | Oct 16, 2021 | Language ModelingLanguage Modelling | —Unverified | 0 |
| Semantic Search as Extractive Paraphrase Span Detection | Oct 16, 2021 | Extractive Question-AnsweringQuestion Answering | —Unverified | 0 |
| ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation | Oct 16, 2021 | Abstractive Text SummarizationMachine Translation | —Unverified | 0 |
| Leveraging Knowledge in Multilingual Commonsense Reasoning | Oct 16, 2021 | Language ModelingLanguage Modelling | —Unverified | 0 |
| Modeling Context With Linear Attention for Scalable Document-Level Translation | Oct 16, 2021 | Document TranslationMachine Translation | —Unverified | 0 |
| A Two-Stage Curriculum Training Framework for NMT | Oct 16, 2021 | Machine TranslationNMT | —Unverified | 0 |
| Rare but Severe Errors Induced by Minimal Deletions in English-Chinese Neural Machine Translation | Oct 16, 2021 | HallucinationMachine Translation | —Unverified | 0 |
| Alleviating the Inequality of Attention Heads for Neural Machine Translation | Oct 16, 2021 | Machine TranslationTranslation | —Unverified | 0 |
| Finding the Right Recipe for Low Resource Domain Adaptation in Neural Machine Translation | Oct 16, 2021 | 8kDomain Adaptation | —Unverified | 0 |