SOTAVerified

Language Identification

Language identification is the task of determining the language of a text.

Papers

Showing 151–200 of 794 papers

TitleStatusHype
Bilingual Streaming ASR with Grapheme units and Auxiliary Monolingual Loss—0
A Pre-trained Transformer and CNN Model with Joint Language ID and Part-of-Speech Tagging for Code-Mixed Social-Media Text—0
Albanian Language Identification in Text Documents—0
A deep-learning based native-language classification by using a latent semantic analysis for the NLI Shared Task 2017—0
BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition—0
BhamNLP at SemEval-2020 Task 12: An Ensemble of Different Word Embeddings and Emotion Transfer Learning for Arabic Offensive Language Identification in Social Media—0
BFCAI at ComMA@ICON 2021: Support Vector Machines for Multilingual Gender Biased and Communal Language Identification—0
A Portuguese Native Language Identification Dataset—0
SOLID: A Large-Scale Semi-Supervised Dataset for Offensive Language Identification—0
Beware Haters at ComMA@ICON: Sequence and Ensemble Classifiers for Aggression, Gender Bias and Communal Bias Identification in Indian Languages—0
A Perplexity-Based Method for Similar Languages Discrimination—0
BERT-based Multi-Task Model for Country and Province Level MSA and Dialectal Arabic Identification—0
BERT-based Multi-Task Model for Country and Province Level Modern Standard Arabic and Dialectal Arabic Identification—0
An Unsupervised Morphological Criterion for Discriminating Similar Languages—0
A language model based approach towards large scale and lightweight language identification systems—0
A Deep Generative Approach to Native Language Identification—0
A Code-Switching Corpus of Turkish-German Conversations—0
Beefmoves: Dissemination, Diversity, and Dynamics of English Borrowings in a German Hip Hop Forum—0
Babler - Data Collection from the Web to Support Speech Recognition and Keyword Search—0
An Overview of Indian Spoken Language Recognition from Machine Learning Perspective—0
Automatic Token and Turn Level Language Identification for Code-Switched Text Dialog: An Analysis Across Language Pairs and Corpora—0
Automatic Spoken Language Identification using a Time-Delay Neural Network—0
A Novel Learnable Dictionary Encoding Layer for End-to-End Language Identification—0
Automatic Spoken Language Identification Utilizing Acoustic and Phonetic Speech Information—0
Automatic language identity tagging on word and sentence-level in multilingual text sources: a case-study on Luxembourgish—0
AUTOMATIC LANGUAGE IDENTIFICATION USING DEEP NEURAL NETWORKS—0
Automatic Language Identification System for Hindi and Magahi—0
Annotation Efficient Language Identification from Weak Labels—0
Addition of Code Mixed Features to Enhance the Sentiment Prediction of Song Lyrics—0
Automatic Language Identification for Romance Languages using Stop Words and Diacritics—0
Automatic Language Identification for Celtic Texts—0
DCU-UVT: Word-Level Language Classification with Code-Mixed Data—0
Automatic language identification—0
Anlirika: An LSTM–CNN Flow Twister for Spoken Language Identification—0
Automatic Identification of Maghreb Dialects Using a Dictionary-Based Approach—0
CUSATNLP@DravidianLangTech-EACL2021:Language Agnostic Classification of Offensive Content in Tweets—0
Automatic Identification of Learners' Language Background Based on Their Writing in Czech—0
CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages—0
Automatic Identification of Closely-related Indian Languages: Resources and Experiments—0
cs@DravidianLangTech-EACL2021: Offensive Language Identification Based On Multilingual BERT Model—0
Cross-Linguistic Offensive Language Detection: BERT-Based Analysis of Bengali, Assamese, & Bodo Conversational Hateful Content from Social Media—0
Automatic discovery of Latin syntactic changes—0
Anglicized Words and Misspelled Cognates in Native Language Identification—0
Curriculum Design for Code-switching: Experiments with Language Identification and Language Modeling with Deep Neural Networks—0
A Federated Learning Approach to Privacy Preserving Offensive Language Identification—0
CUSATNLP@HASOC-Dravidian-CodeMix-FIRE2020:Identifying Offensive Language from ManglishTweets—0
A Dataset and Classifier for Recognizing Social Media English—0
Data Filtering using Cross-Lingual Word Embeddings—0
Accurate Pinyin-English Codeswitched Language Identification—0
Cross-lingual Inductive Transfer to Detect Offensive Language—0
Show:102550
← PrevPage 4 of 16Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1wav2vec 2.0 LV-60KError rate7.2—Unverified
2XLS-RError rate5.7—Unverified
#ModelMetricClaimedVerifiedStatus
1GlotLIDMacro F10.98—Unverified
#ModelMetricClaimedVerifiedStatus
1FastTextAccuracy0.97—Unverified
#ModelMetricClaimedVerifiedStatus
1Apple bi-LSTMAccuracy91.37—Unverified
#ModelMetricClaimedVerifiedStatus
1Apple bi-LSTMAccuracy86.93—Unverified
#ModelMetricClaimedVerifiedStatus
1ConformerG-PAccuracy99.8—Unverified