Voice Conversion

I remember all the summer days Drinking wine in the sunshine I hope it never leaves And I remember all the summer nights Staring at you in the moonlight I hope you never leave 'cause baby You're so good to me You have all that all that I ever need It's easy to love you So easy to love you Ooh you know it's true The best part of being with you To know you're with me It's not so hard to say It's easy to love you I remember all those winter days frozen In the cold tryin' to get you home Should I be moving in, we can be together then Remember spending all those winter nights Stayin' inside by the warm fire Yeah you gotta know that I can never let you go You and I have the rest of our lives to say It's easy to love you So easy to love you Ooh you know it's true The best part of being with you To know you're with me It's not so hard to say It's easy to love you Can anybody else see it? Mm, can anybody else see what I do? Can anybody else feel it? Oh, can anybody else feel the way I do? But now I'm with you Hard to forget all the moments when We'd be sitting there hoping it would never end 'Cause this is meant to be So baby, will you marry me? It's easy to love you So easy to love you Ooh, you know it's true The best part of being with you To know you are with me It's not so hard to say It's easy to love you You and me will be together I know our love will last forever You and me will be together I know our love will last forever You know it's true The best part of being with you You're easy to love

Source: Joint training framework for text-to-speech and voice conversion using multi-source Tacotron and WaveNet

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 101–150 of 520 papers

Title	Date	Tasks	Status	Hype	Score
ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion	Mar 29, 2022	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	1	5
Seen and Unseen emotional style transfer for voice conversion with a new emotional speech dataset	Oct 28, 2020	DecoderEmotion Recognition	CodeCode Available	1	5
Spectrogram-channels u-net: a source separation model viewing each channel as the spectrogram of each source	Oct 26, 2018	Information RetrievalMusic Information Retrieval	CodeCode Available	1	5
Speak Like a Dog: Human to Non-human creature Voice Conversion	Jun 9, 2022	Generative Adversarial NetworkVoice Conversion	CodeCode Available	1	5
Voice Conversion Based on Cross-Domain Features Using Variational Auto Encoders	Aug 29, 2018	Voice Conversion	CodeCode Available	1	5
CSLP-AE: A Contrastive Split-Latent Permutation Autoencoder Framework for Zero-Shot Electroencephalography Signal Conversion	Nov 13, 2023	Contrastive LearningEEG	CodeCode Available	1	5
AutoVisual Fusion Suite: A Comprehensive Evaluation of Image Segmentation and Voice Conversion Tools on HuggingFace Platform	Dec 17, 2023	Image SegmentationSegmentation	CodeCode Available	1	5
Speech Representation Disentanglement with Adversarial Mutual Information Learning for One-shot Voice Conversion	Aug 18, 2022	DisentanglementRhythm	CodeCode Available	1	5
Evaluating Methods for Ground-Truth-Free Foreign Accent Conversion	Sep 5, 2023	Voice Conversion	CodeCode Available	1	5
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing	Oct 14, 2021	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	1	5
Rhythm Modeling for Voice Conversion	Jul 12, 2023	RhythmVoice Conversion	CodeCode Available	1	5
Emotional Voice Conversion: Theory, Databases and ESD	May 31, 2021	Voice Conversion	CodeCode Available	1	5
StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks	Jun 6, 2018	AttributeGenerative Adversarial Network	CodeCode Available	1	5
StarGAN-VC++: Towards Emotion Preserving Voice Conversion Using Deep Embeddings	Sep 14, 2023	Generative Adversarial NetworkVoice Conversion	CodeCode Available	1	5
Efficient Non-Autoregressive GAN Voice Conversion using VQWav2vec Features and Dynamic Convolution	Mar 31, 2022	Voice Conversion	CodeCode Available	1	5
StyleTTS-VC: One-Shot Voice Conversion by Knowledge Transfer from Style-Based TTS Models	Dec 29, 2022	Data Augmentationtext-to-speech	CodeCode Available	1	5
A unified one-shot prosody and speaker conversion system with self-supervised discrete speech units	Nov 12, 2022	RhythmVoice Conversion	CodeCode Available	1	5
Deep Learning Based Assessment of Synthetic Speech Naturalness	Apr 23, 2021	Deep LearningPrediction	CodeCode Available	1	5
Emo-StarGAN: A Semi-Supervised Any-to-Many Non-Parallel Emotion-Preserving Voice Conversion	Sep 14, 2023	Voice Conversion	CodeCode Available	1	5
Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data	Feb 1, 2020	Voice Conversion	CodeCode Available	1	5
F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder	Apr 15, 2020	Style TransferVoice Conversion	CodeCode Available	1	5
Training-Free Voice Conversion with Factorized Optimal Transport	Jun 11, 2025	Voice Conversion	CodeCode Available	1	5
Voice Conversion Using Speech-to-Speech Neuro-Style Transfer	Oct 25, 2020	Generative Adversarial NetworkStyle Transfer	CodeCode Available	1	5
Defending Your Voice: Adversarial Attack on Voice Conversion	May 18, 2020	Adversarial AttackVoice Conversion	CodeCode Available	1	5
Voice Spoofing Countermeasures: Taxonomy, State-of-the-art, experimental analysis of generalizability, open challenges, and the way forward	Oct 2, 2022	MisinformationSpeaker Verification	CodeCode Available	1	5
An Improved StarGAN for Emotional Voice Conversion: Enhancing Voice Quality and Data Augmentation	Jul 18, 2021	Data AugmentationEmotion Recognition	CodeCode Available	0	5
Defense for Black-box Attacks on Anti-spoofing Models by Self-Supervised Learning	Jun 5, 2020	Self-Supervised LearningSpeaker Verification	CodeCode Available	0	5
Universal Adaptor: Converting Mel-Spectrograms Between Different Configurations for Speech Synthesis	Apr 1, 2022	Speech SynthesisVoice Conversion	CodeCode Available	0	5
Deep Residual Neural Networks for Audio Spoofing Detection	Jun 30, 2019	Speaker VerificationSpeech Synthesis	CodeCode Available	0	5
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages	May 20, 2025	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	0	5
A Speech Representation Anonymization Framework via Selective Noise Perturbation	Mar 26, 2022	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	0	5
Unsupervised End-to-End Learning of Discrete Linguistic Units for Voice Conversion	May 28, 2019	DecoderVoice Conversion	CodeCode Available	0	5
SVSNet: An End-to-end Speaker Voice Similarity Assessment Model	Jul 20, 2021	Voice ConversionVoice Similarity	CodeCode Available	0	5
The Sequence-to-Sequence Baseline for the Voice Conversion Challenge 2020: Cascading ASR and TTS	Oct 6, 2020	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	0	5
Decoupling Speaker-Independent Emotions for Voice Conversion Via Source-Filter Networks	Oct 4, 2021	DecoderVoice Conversion	CodeCode Available	0	5
Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks	Sep 23, 2017	Speech Synthesistext-to-speech	CodeCode Available	0	5
AdaGAN: Adaptive GAN for Many-to-Many Non-Parallel Voice Conversion	Sep 25, 2019	Generative Adversarial NetworkStyle Transfer	CodeCode Available	0	5
STC Antispoofing Systems for the ASVspoof2019 Challenge	Apr 11, 2019	Speech SynthesisVoice Conversion	CodeCode Available	0	5
CycleGAN-VC2: Improved CycleGAN-based Non-parallel Voice Conversion	Apr 9, 2019	Voice Conversion	CodeCode Available	0	5
ASSERT: Anti-Spoofing with Squeeze-Excitation and Residual neTworks	Apr 1, 2019	Feature Engineeringtext-to-speech	CodeCode Available	0	5
Spoof detection using time-delay shallow neural network and feature switching	Apr 16, 2019	Speaker VerificationSpeech Synthesis	CodeCode Available	0	5
ACVAE-VC: Non-parallel many-to-many voice conversion with auxiliary classifier variational autoencoder	Aug 13, 2018	AttributeDecoder	CodeCode Available	0	5
SIG-VC: A Speaker Information Guided Zero-shot Voice Conversion System for Both Human Beings and Machines	Nov 6, 2021	DisentanglementSpeaker Verification	CodeCode Available	0	5
Scalable Factorized Hierarchical Variational Autoencoder Training	Apr 9, 2018	DisentanglementHyperparameter Optimization	CodeCode Available	0	5
StarGAN-VC2: Rethinking Conditional Methods for StarGAN-Based Voice Conversion	Jul 29, 2019	Voice Conversion	CodeCode Available	0	5
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech	Jun 2, 2025	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	0	5
Private kNN-VC: Interpretable Anonymization of Converted Speech	May 23, 2025	Speaker anonymizationSpeaker Recognition	CodeCode Available	0	5
Playing with Voices: Tabletop Role-Playing Game Recordings as a Diarization Challenge	Feb 18, 2025	Voice Conversion	CodeCode Available	0	5
Read the Room: Adapting a Robot's Voice to Ambient and Social Contexts	May 10, 2022	Speech SynthesisVoice Conversion	CodeCode Available	0	5
Parallel-Data-Free Voice Conversion Using Cycle-Consistent Adversarial Networks	Nov 30, 2017	Voice Conversion	CodeCode Available	0	5

Show:10 25 50

← PrevPage 3 of 11Next →

All datasets ZeroSpeech 2019 English LibriSpeech test-clean VCTK

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	VQ-CPC	Speaker Similarity	3.8	—	Unverified
2	VQ-VAE	Speaker Similarity	3.49	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	kNN-VC (prematched HiFiGAN)	Character Error Rate (CER)	2.96	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	DISSC	Total Length Error (TLE)	0.83	—	Unverified