SOTAVerified

Simple But Not Na\" : Fine-Grained Arabic Dialect Identification Using Only N-Grams

2019-08-01WS 2019Unverified0· sign in to hype

Sohaila Eltanbouly, May Bashendy, Tamer Elsayed

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

This paper presents the participation of Qatar University team in MADAR shared task, which addresses the problem of sentence-level fine-grained Arabic Dialect Identification over 25 different Arabic dialects in addition to the Modern Standard Arabic. Arabic Dialect Identification is not a trivial task since different dialects share some features, e.g., utilizing the same character set and some vocabularies. We opted to adopt a very simple approach in terms of extracted features and classification models; we only utilize word and character n-grams as features, and Na ̈ ve Bayes models as classifiers. Surprisingly, the simple approach achieved non-na ̈ ve performance. The official results, reported on a held-out testing set, show that the dialect of a given sentence can be identified at an accuracy of 64.58\% by our best submitted run.

Tasks

Reproductions