SOTAVerified

Speech-to-Text

Papers

Showing 101150 of 403 papers

TitleStatusHype
Re-Translation Strategies For Long Form, Simultaneous, Spoken Language TranslationCode0
Optimizing Rare Word Accuracy in Direct Speech Translation with a Retrieval-and-Demonstration ApproachCode0
fairseq S2T: Fast Speech-to-Text Modeling with fairseqCode0
M-Adapter: Modality Adaptation for End-to-End Speech-to-Text TranslationCode0
Code-Switched Urdu ASR for Noisy Telephonic Environment using Data Centric Approach with Hybrid HMM and CNN-TDNNCode0
Measuring the Effect of Transcription Noise on Downstream Language Understanding TasksCode0
Pre-training on high-resource speech recognition improves low-resource speech-to-text translationCode0
LibriS2S: A German-English Speech-to-Speech Translation CorpusCode0
Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language ModelsCode0
Joint CTC-Attention based End-to-End Speech Recognition using Multi-task LearningCode0
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak SupervisionCode0
Kurdish (Sorani) Speech to Text: Presenting an Experimental DatasetCode0
Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text TranslationCode0
Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations GenerationCode0
Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language UnderstandingCode0
Finstreder: Simple and fast Spoken Language Understanding with Finite State Transducers using modern Speech-to-Text modelsCode0
FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the WildCode0
InstaIndoor and Multi-modal Deep Learning for Indoor Scene RecognitionCode0
Scribosermo: Fast Speech-to-Text models for German and other LanguagesCode0
Challenges and Opportunities of Speech Recognition for Bengali Language0
Efficient Monotonic Multihead Attention0
Effectively pretraining a speech translation decoder with Machine Translation data0
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?0
Application of Audio Fingerprinting Techniques for Real-Time Scalable Speech Retrieval and Speech Clusterization0
A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text Tasks0
BTS: Back TranScription for Speech-to-Text Post-Processor using Text-to-Speech-to-Text0
Direct Simultaneous Speech-to-Text Translation Assisted by Synchronized Streaming ASR0
Application-Agnostic Language Modeling for On-Device ASR0
Direct Punjabi to English speech translation using discrete units0
Digits micro-model for accurate and secure transactions0
Bridging the Modality Gap for Speech-to-Text Translation0
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum0
Development of Natural Language Processing Tools for Cook Islands M\=aori0
Bridging the gap between streaming and non-streaming ASR systems bydistilling ensembles of CTC and RNN-T models0
Anonymizing Speech with Generative Adversarial Networks to Preserve Speaker Privacy0
AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR0
A Comparative Study on End-to-end Speech to Text Translation0
Developing automatic verbatim transcripts for international multilingual meetings: an end-to-end solution0
Developing a Speech Recognition System for Recognizing Tonal Speech Signals Using a Convolutional Neural Network0
Design of a novel Korean learning application for efficient pronunciation correction0
Deep Speech Based End-to-End Automated Speech Recognition (ASR) for Indian-English Accents0
BCN2BRNO: ASR System Fusion for Albayzin 2020 Speech to Text Challenge0
Deep Learning Based Natural Language Processing for End to End Speech Translation0
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM0
An Experiment on Speech-to-Text Translation Systems for Manipuri to English on Low Resource Setting0
Adversarial Attacks against Neural Networks in Audio Domain: Exploiting Principal Components0
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models0
DeepCruiser: Automated Guided Testing for Stateful Deep Learning Systems0
Decision Attentive Regularization to Improve Simultaneous Speech Translation Systems0
Data Efficient Direct Speech-to-Text Translation with Modality Agnostic Meta-Learning0
Show:102550
← PrevPage 3 of 9Next →

No leaderboard results yet.