SOTAVerified

Speech-to-Text

Papers

Showing 201–250 of 403 papers

TitleStatusHype
Kencorpus: A Kenyan Language Corpus of Swahili, Dholuo and Luhya for Natural Language Processing Tasks—0
Improving Hypernasality Estimation with Automatic Speech Recognition in Cleft Palate Speech—0
Extending RNN-T-based speech recognition systems with emotion and language classification—0
RSD-GAN: Regularized Sobolev Defense GAN Against Speech-to-Text Adversarial Attacks—0
M-Adapter: Modality Adaptation for End-to-End Speech-to-Text TranslationCode0
Language Model Augmented Monotonic Attention for Simultaneous Translation—0
System Description on Automatic Simultaneous Translation Workshop—0
Findings of the Third Workshop on Automatic Simultaneous Translation—0
Swiss German Speech to Text system evaluation—0
Finstreder: Simple and fast Spoken Language Understanding with Finite State Transducers using modern Speech-to-Text modelsCode0
Developing a Speech Recognition System for Recognizing Tonal Speech Signals Using a Convolutional Neural Network—0
Revisiting End-to-End Speech-to-Text Translation From Scratch—0
The Nós Project: Opening routes for the Galician language in the field of language technologies—0
Towards Large Vocabulary Kazakh-Russian Sign Language Dataset: KRSL-OnlineSchool—0
A Semi-Automated Live Interlingual Communication Workflow Featuring Intralingual Respeaking: Evaluation and Benchmarking—0
Clinical Dialogue Transcription Error Correction using Seq2Seq Models—0
Semantic-preserved Communication System for Highly Efficient Speech Transmission—0
PaddleSpeech: An Easy-to-Use All-in-One Speech ToolkitCode6
SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation—0
Hearing voices at the National Library -- a speech corpus and acoustic model for the Swedish language—0
Cross-modal Contrastive Learning for Speech TranslationCode1
Design of a novel Korean learning application for efficient pronunciation correction—0
Wav2Seq: Pre-training Speech-to-Text Encoder-Decoder Models Using Pseudo LanguagesCode1
Learning Adaptive Segmentation Policy for End-to-End Simultaneous Translation—0
NAIST Simultaneous Speech-to-Text Translation System for IWSLT 2022—0
The HW-TSC’s Simultaneous Speech Translation System for IWSLT 2022 Evaluation—0
The AISP-SJTU Simultaneous Translation System for IWSLT 2022—0
LibriS2S: A German-English Speech-to-Speech Translation CorpusCode0
WaBERT: A Low-resource End-to-end Model for Spoken Language Understanding and Speech-to-BERT Alignment—0
Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation—0
A Study of Gender Impact in Self-supervised Models for Speech-to-Text Systems—0
Deep Speech Based End-to-End Automated Speech Recognition (ASR) for Indian-English Accents—0
The MIT Voice Name System—0
A Dataset for Speech Emotion Recognition in Greek Theatrical PlaysCode0
XTREME-S: Evaluating Cross-lingual Speech Representations—0
STEMM: Self-learning with Speech-text Manifold Mixup for Speech TranslationCode1
A^3T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and EditingCode1
A combined approach to the analysis of speech conversations in a contact center domain—0
Attacks as Defenses: Designing Robust Audio CAPTCHAs Using Attacks on Automatic Speech Recognition Systems—0
Which French speech recognition system for assistant robots?—0
Spanish and English Phoneme Recognition by Training on Simulated Classroom Audio Recordings of Collaborative Learning EnvironmentsCode0
Punctuation restoration in Swedish through fine-tuned KB-BERT—0
Semantic-aware Speech to Text Transmission with Redundancy Removal—0
Optimization of a Real-Time Wavelet-Based Algorithm for Improving Speech Intelligibility—0
CVSS Corpus and Massively Multilingual Speech-to-Speech TranslationCode2
A wearable sensor vest for social humanoid robots with GPGPU, IoT, and modular software architectureCode0
InstaIndoor and Multi-modal Deep Learning for Indoor Scene RecognitionCode0
Regularizing End-to-End Speech Translation with Triangular Decomposition AgreementCode1
Cross-modal Contrastive Learning for Speech Translation—0
X-Vector based voice activity detection for multi-genre broadcast speech-to-textCode1
Show:102550
← PrevPage 5 of 9Next →

No leaderboard results yet.