SOTAVerified

Speech-to-Text

Papers

Showing 151200 of 403 papers

TitleStatusHype
Online Hybrid CTC/Attention End-to-End Automatic Speech Recognition Architecture0
AudioPaLM: A Large Language Model That Can Speak and Listen0
Recent Advances in Direct Speech-to-text Translation0
Open Brain AI. Automatic Language Assessment0
Speech-to-Text Adapter and Speech-to-Entity Retriever Augmented LLMs for Speech Understanding0
Towards End-to-end Speech-to-text SummarizationCode0
Improved Cross-Lingual Transfer Learning For Automatic Speech Translation0
Strategies for improving low resource speech to text translation relying on pre-trained ASR models0
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions0
CIF-PT: Bridging Speech and Text Representations for Spoken Language Understanding via Continuous Integrate-and-Fire Pre-Training0
VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation0
ComSL: A Composite Speech-Language Model for End-to-End Speech-to-Text TranslationCode1
Improving Metrics for Speech Translation0
DUB: Discrete Unit Back-translation for Speech TranslationCode1
Application-Agnostic Language Modeling for On-Device ASR0
A Whisper transformer for audio captioning trained with synthetic captions and transfer learningCode1
Back Translation for Speech-to-text Translation Without TranscriptsCode1
Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks0
Improving Autoregressive NLP Tasks via Modular Linearized Attention0
ESPnet-ST-v2: Multipurpose Spoken Language Translation Toolkit0
Enhancing Speech-to-Speech Translation with Multiple TTS Targets0
Natural Language Robot Programming: NLP integrated with autonomous robotic grasping0
Improving the previous state-of-the-art Frisian ASR by fine-tuning XLS-R0
wav2vec and its current potential to Automatic Speech Recognition in German for the usage in Digital History: A comparative assessment of available ASR-technologies for the use in cultural heritage contexts0
Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages0
MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text TranslationCode2
Improving Medical Speech-to-Text Accuracy with Vision-Language Pre-training Model0
PATCorrect: Non-autoregressive Phoneme-augmented Transformer for ASR Error Correction0
Characterizing Financial Market Coverage using Artificial Intelligence0
PSST! Prosodic Speech Segmentation with TransformersCode1
Pre-training for Speech Translation: CTC Meets Optimal TransportCode1
Using External Off-Policy Speech-To-Text Mappings in Contextual End-To-End Automated Speech Recognition0
Pushing the performances of ASR models on English and Spanish accents0
WACO: Word-Aligned Contrastive Learning for Speech TranslationCode0
M3ST: Mix at Three Levels for Speech Translation0
MMSpeech: Multi-modal Multi-task Encoder-Decoder Pre-training for Speech Recognition0
Handling and extracting key entities from customer conversations using Speech recognition and Named Entity recognition0
Multilingual Speech Emotion Recognition With Multi-Gating Mechanism and Neural Architecture Search0
Phonemic Representation and Transcription for Speech to Text Applications for Under-resourced Indigenous African Languages: The Case of Kiswahili0
Efficient Speech Translation with Dynamic Latent PerceiversCode0
Don't Discard Fixed-Window Audio Segmentation in Speech-to-Text TranslationCode0
Information-Transport-based Policy for Simultaneous TranslationCode1
Named Entity Detection and Injection for Direct Speech Translation0
Improving Semi-supervised End-to-end Automatic Speech Recognition using CycleGAN and Inter-domain Losses0
Simple and Effective Unsupervised Speech Translation0
Anonymizing Speech with Generative Adversarial Networks to Preserve Speaker Privacy0
CTC Alignments Improve Autoregressive Translation0
SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingCode0
JoeyS2T: Minimalistic Speech-to-Text Modeling with JoeyNMTCode1
Speech-to-Text and Evaluation of Multiple Machine Translation Systems0
Show:102550
← PrevPage 4 of 9Next →

No leaderboard results yet.