SOTAVerified

Voice Cloning

Voice cloning is a highly desired feature for personalized speech interfaces. Neural voice cloning system learns to synthesize a person’s voice from only a few audio samples.

Papers

Showing 2650 of 112 papers

TitleStatusHype
Voice Adaptation for Swiss German0
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents0
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages0
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning0
Beyond Face Swapping: A Diffusion-Based Digital Human Benchmark for Multimodal Deepfake Detection0
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling0
VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning0
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder0
Voice Cloning: Comprehensive Survey0
ClonEval: An Open Voice Cloning BenchmarkCode0
"It's not a representation of me": Examining Accent Bias and Digital Exclusion in Synthetic AI Voice Services0
Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis0
SpeechDialogueFactory: Generating High-Quality Speech Dialogue Data to Accelerate Your Speech-LLM DevelopmentCode0
SoK: How Robust is Audio Watermarking in Generative AI models?0
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology0
Steganography Beyond Space-Time with Chain of Multimodal AI0
Deepfake Technology Unveiled: The Commoditization of AI and Its Impact on Digital Trust0
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement0
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model0
Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset0
Speech Watermarking with Discrete Intermediate Representations0
Parallel Stacked Aggregated Network for Voice Authentication in IoT-Enabled Smart Devices0
Hindi audio-video-Deepfake (HAV-DF): A Hindi language-based Audio-video Deepfake Dataset0
The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge: Tasks, Results and Findings0
DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis0
Show:102550
← PrevPage 2 of 5Next →

No leaderboard results yet.