Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis Sep 20, 2024 Face Swapping Speech Synthesis
Code Code Available 0NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization Sep 19, 2024 Audio Compression Audio Generation
— Unverified 0Enhancing Multilingual Speech Generation and Recognition Abilities in LLMs with Constructed Code-switched Data Sep 17, 2024 Speech Synthesis
— Unverified 0Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation Sep 17, 2024 Knowledge Distillation Speech Synthesis
— Unverified 0StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion Sep 16, 2024 Speech Synthesis text-to-speech
— Unverified 0Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization Sep 16, 2024 Emotional Speech Synthesis In-Context Learning
— Unverified 0Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation Sep 14, 2024 Speech Synthesis text-to-speech
— Unverified 0LLM-Powered Grapheme-to-Phoneme Conversion: Benchmark and Case Study Sep 13, 2024 Benchmarking Grapheme-to-Phoneme Conversion
— Unverified 0Text-To-Speech Synthesis In The Wild Sep 13, 2024 Benchmarking Speaker Recognition
— Unverified 0Full-text Error Correction for Chinese Speech Recognition with Large Language Model Sep 12, 2024 Automatic Speech Recognition Automatic Speech Recognition (ASR)
— Unverified 0SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis Sep 11, 2024 Decoder Speech Synthesis
Code Code Available 2Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach Sep 10, 2024 Speech Synthesis text-to-speech
— Unverified 0What happens to diffusion model likelihood when your model is conditional? Sep 10, 2024 domain classification model
— Unverified 0AS-Speech: Adaptive Style For Speech Synthesis Sep 9, 2024 Rhythm Speech Synthesis
— Unverified 0Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP Sep 4, 2024 Audio Synthesis Computational Efficiency
— Unverified 0VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka Sep 3, 2024 Automatic Speech Recognition Automatic Speech Recognition (ASR)
— Unverified 0vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders Sep 3, 2024 Speech Synthesis Voice Conversion
— Unverified 0Sample-Efficient Diffusion for Text-To-Speech Synthesis Sep 1, 2024 Language Modeling Language Modelling
Code Code Available 2SelectTTS: Synthesizing Anyone's Voice via Discrete Unit-Based Frame Selection Aug 30, 2024 Self-Supervised Learning Speech Synthesis
— Unverified 0Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model Aug 30, 2024 Audio Compression Audio Generation
Code Code Available 3Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming Aug 29, 2024 Speech Synthesis
Code Code Available 7Literary and Colloquial Dialect Identification for Tamil using Acoustic Features Aug 27, 2024 Automatic Speech Recognition Dialect Identification
— Unverified 0SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description Aug 24, 2024 Descriptive Speech Synthesis
Code Code Available 2Which Prosodic Features Matter Most for Pragmatics? Aug 23, 2024 Speech Synthesis
— Unverified 0AI-Based IVR Aug 20, 2024 Speech Synthesis Speech-to-Text
— Unverified 0Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition Aug 17, 2024 Language Modeling Language Modelling
Code Code Available 0Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization Aug 15, 2024 Speech Synthesis
Code Code Available 3WavLM model ensemble for audio deepfake detection Aug 14, 2024 Audio Deepfake Detection Data Augmentation
Code Code Available 0PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation Aug 14, 2024 Speech Synthesis text-to-speech
Code Code Available 3SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis Aug 13, 2024 Speech Synthesis Spoken Dialogue Systems
Code Code Available 0PRESENT: Zero-Shot Text-to-Prosody Control Aug 13, 2024 Prosody Prediction Speech Synthesis
Code Code Available 1VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders Aug 13, 2024 Speech Synthesis
— Unverified 0Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation Aug 1, 2024 Representation Learning Speech Synthesis
— Unverified 0Generative Expressive Conversational Speech Synthesis Jul 31, 2024 Speech Synthesis
Code Code Available 1VoxSim: A perceptual voice similarity dataset Jul 26, 2024 Benchmarking Speaker Recognition
Code Code Available 1Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks Jul 26, 2024 Generative Adversarial Network Speech Enhancement
— Unverified 0Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models Jul 26, 2024 Speech Synthesis
— Unverified 0dMel: Speech Tokenization made Simple Jul 22, 2024 Decoder Language Modeling
Code Code Available 1Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning Jul 21, 2024 Representation Learning Self-Supervised Learning
— Unverified 0MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis Jul 19, 2024 Expressive Speech Synthesis Speech Synthesis
— Unverified 0Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings Jul 19, 2024 Expressive Speech Synthesis Speech Synthesis
Code Code Available 1Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models Jul 18, 2024 Language Modeling Language Modelling
— Unverified 0Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis Jul 13, 2024 Mamba speech-recognition
Code Code Available 2Toward accessible comics for blind and low vision readers Jul 11, 2024 Optical Character Recognition Prompt Engineering
— Unverified 0Autoregressive Speech Synthesis without Vector Quantization Jul 11, 2024 Audio Compression Diversity
— Unverified 0Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation Jul 8, 2024 Automatic Speech Recognition Emotion Recognition
— Unverified 0Fine-Grained and Interpretable Neural Speech Editing Jul 7, 2024 Data Augmentation Speech Synthesis
Code Code Available 1CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens Jul 7, 2024 Language Modelling Large Language Model
Code Code Available 11FA-GAN: Artifacts-free and Phase-aware High-fidelity GAN-based Vocoder Jul 5, 2024 Generative Adversarial Network Speech Synthesis
— Unverified 0We Need Variations in Speech Generation: Sub-center Modelling for Speaker Embeddings Jul 5, 2024 Speaker Recognition Speech Synthesis
— Unverified 0