A dual-encoding system for dialect classification
2020-12-01VarDial (COLING) 2020Code Available0· sign in to hype
Petru Rebeja, Dan Cristea
Code Available — Be the first to reproduce this paper.
ReproduceCode
- github.com/repierre/vardial2020OfficialIn papertf★ 0
Abstract
In this paper we present the architecture, processing pipeline and results of the ensemble model developed for Romanian Dialect Identification task. The ensemble model consists of two TF-IDF encoders and a deep learning model aimed together at classifying input samples based on the writing patterns which are specific to each of the two dialects. Although the model performs well on the training set, its performance degrades heavily on the evaluation set. The drop in performance is due to the design decision which makes the model put too much weight on presence/lack of textual marks when determining the sample label.