SOTAVerified

An Empirical Study of the Downstream Reliability of Pre-Trained Word Embeddings

2020-12-01COLING 2020Unverified0· sign in to hype

Anthony Rios, Brandon Lwowski

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

While pre-trained word embeddings have been shown to improve the performance of downstream tasks, many questions remain regarding their reliability: Do the same pre-trained word embeddings result in the best performance with slight changes to the training data? Do the same pre-trained embeddings perform well with multiple neural network architectures? Do imputation strategies for unknown words impact reliability? In this paper, we introduce two new metrics to understand the downstream reliability of word embeddings. We find that downstream reliability of word embeddings depends on multiple factors, including, the evaluation metric, the handling of out-of-vocabulary words, and whether the embeddings are fine-tuned.

Tasks

Reproductions