SOTAVerified

Language Identification

Language identification is the task of determining the language of a text.

Papers

Showing 151–200 of 794 papers

TitleStatusHype
LAMASSU: Streaming Language-Agnostic Multilingual Speech Recognition and Translation Using Neural Transducers—0
A Compact End-to-End Model with Local and Global Context for Spoken Language Identification—0
Italian Language and Dialect Identification and Regional French Variety Detection using Adaptive Naive BayesCode0
Neural Networks for Cross-domain Language Identification. Phlyers @Vardial 2022—0
OcWikiDisc: a Corpus of Wikipedia Talk Pages in Occitan—0
The Curious Case of Logistic Regression for Italian Languages and Dialects IdentificationCode0
Streaming End-to-End Multilingual Speech Recognition with Joint Language Identification—0
Evaluation of Off-the-Shelf Language Identification Tools on Bulgarian Social Media Posts—0
Unravelling Interlanguage Facts via Explainable Machine Learning—0
Extending RNN-T-based speech recognition systems with emotion and language classification—0
Distilled Non-Semantic Speech Embeddings with Binary Neural Networks for Low-Resource DevicesCode0
Huqariq: A Multilingual Speech Corpus of Native Languages of Peru for Speech Recognition—0
TechSSN at SemEval-2022 Task 6: Intended Sarcasm Detection using Transformer Models—0
Language Identification for Austronesian LanguagesCode0
HeLI-OTS, Off-the-shelf Language Identifier for Text—0
Huqariq: A Multilingual Speech Corpus of Native Languages of Peru forSpeech Recognition—0
Deep learning-based end-to-end spoken language identification system for domain-mismatched scenario—0
CoSwID, a Code Switching Identification Method Suitable for Under-Resourced Languages—0
GeezSwitch: Language Identification in Typologically Related Low-resourced East African LanguagesCode0
MHE: Code-Mixed Corpora for Similar Language Identification—0
Universal Dependencies Treebank for Tatar: Incorporating Intra-Word Code-Switching Information—0
Dialects Identification of Armenian Language—0
Adversarial synthesis based data-augmentation for code-switched spoken language identification—0
FLEURS: Few-shot Learning Evaluation of Universal Representations of SpeechCode0
Modernizing Open-Set Speech Language Identification—0
Automatic Spoken Language Identification using a Time-Delay Neural Network—0
Pretraining Approaches for Spoken Language Recognition: TalTech Submission to the OLR 2021 Challenge—0
Building Machine Translation Systems for the Next Thousand Languages—0
TuGeBiC: A Turkish German Bilingual Code-Switching Corpus—0
Unsupervised Preference-Aware Language IdentificationCode0
Findings of the Shared Task on Multi-task Learning in Dravidian Languages—0
Automated speech tools for helping communities process restricted-access corpora for language revival efforts—0
Transducer-based language embedding for spoken language identification—0
Partial Coupling of Optimal Transport for Spoken Language Identification—0
Improving Language Identification of Accented Speech—0
Code Switched and Code Mixed Speech Recognition for Indic languages—0
Geographic Adaptation of Pretrained Language ModelsCode0
Automatic Language Identification for Celtic Texts—0
Enhance Language Identification using Dual-mode Model with Knowledge DistillationCode0
Towards a Common Speech Analysis Engine—0
Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech—0
CALCS 2021 Shared Task: Machine Translation for Code-Switched Data—0
HaT5: Hate Language Identification using Text-to-Text Transfer Transformer—0
Translated Texts Under the Lens: From Machine Translation Detection to Source Language Identification—0
Cognitive Computing to Optimize IT Services—0
LUC at ComMA-2021 Shared Task: Multilingual Gender Biased and Communal Language Identification without using linguistic features—0
Integrating Knowledge in End-to-End Automatic Speech Recognition for Mandarin-English Code-Switching—0
Robust Speech Representation Learning via Flow-based Embedding Regularization—0
MUM at ComMA@ICON: Multilingual Gender Biased and Communal Language Identification Using Supervised Learning Approaches—0
BFCAI at ComMA@ICON 2021: Support Vector Machines for Multilingual Gender Biased and Communal Language Identification—0
Show:102550
← PrevPage 4 of 16Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1wav2vec 2.0 LV-60KError rate7.2—Unverified
2XLS-RError rate5.7—Unverified
#ModelMetricClaimedVerifiedStatus
1GlotLIDMacro F10.98—Unverified
#ModelMetricClaimedVerifiedStatus
1FastTextAccuracy0.97—Unverified
#ModelMetricClaimedVerifiedStatus
1Apple bi-LSTMAccuracy91.37—Unverified
#ModelMetricClaimedVerifiedStatus
1Apple bi-LSTMAccuracy86.93—Unverified
#ModelMetricClaimedVerifiedStatus
1ConformerG-PAccuracy99.8—Unverified