Voice Conversion

I remember all the summer days Drinking wine in the sunshine I hope it never leaves And I remember all the summer nights Staring at you in the moonlight I hope you never leave 'cause baby You're so good to me You have all that all that I ever need It's easy to love you So easy to love you Ooh you know it's true The best part of being with you To know you're with me It's not so hard to say It's easy to love you I remember all those winter days frozen In the cold tryin' to get you home Should I be moving in, we can be together then Remember spending all those winter nights Stayin' inside by the warm fire Yeah you gotta know that I can never let you go You and I have the rest of our lives to say It's easy to love you So easy to love you Ooh you know it's true The best part of being with you To know you're with me It's not so hard to say It's easy to love you Can anybody else see it? Mm, can anybody else see what I do? Can anybody else feel it? Oh, can anybody else feel the way I do? But now I'm with you Hard to forget all the moments when We'd be sitting there hoping it would never end 'Cause this is meant to be So baby, will you marry me? It's easy to love you So easy to love you Ooh, you know it's true The best part of being with you To know you are with me It's not so hard to say It's easy to love you You and me will be together I know our love will last forever You and me will be together I know our love will last forever You know it's true The best part of being with you You're easy to love

Source: Joint training framework for text-to-speech and voice conversion using multi-source Tacotron and WaveNet

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 51–100 of 520 papers

Title	Date	Tasks	Status	Hype
Region-Based Optimization in Continual Learning for Audio Deepfake Detection	Dec 16, 2024	Audio Deepfake DetectionContinual Learning	CodeCode Available	1
Assem-VC: Realistic Voice Conversion by Assembling Modern Speech Synthesis Techniques	Apr 2, 2021	DecoderRhythm	CodeCode Available	1
Rhythm Modeling for Voice Conversion	Jul 12, 2023	RhythmVoice Conversion	CodeCode Available	1
Robust Disentangled Variational Speech Representation Learning for Zero-shot Voice Conversion	Mar 30, 2022	Data AugmentationDecoder	CodeCode Available	1
Investigation of F0 conditioning and Fully Convolutional Networks in Variational Autoencoder based Voice Conversion	May 2, 2019	DecoderDisentanglement	CodeCode Available	1
HM-Conformer: A Conformer-based audio deepfake detection system with hierarchical pooling and multi-level classification token aggregation methods	Sep 15, 2023	Audio Deepfake DetectionDeepFake Detection	CodeCode Available	1
Where are we in audio deepfake detection? A systematic analysis over generative and detection models	Oct 6, 2024	Audio Deepfake DetectionAudio Synthesis	CodeCode Available	1
Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised Units	Dec 19, 2022	RhythmVoice Conversion	CodeCode Available	1
GAN You Hear Me? Reclaiming Unconditional Speech Synthesis from Diffusion Models	Oct 11, 2022	DisentanglementGenerative Adversarial Network	CodeCode Available	1
SpeechLMScore: Evaluating speech generation using speech language model	Dec 8, 2022	Language ModelingLanguage Modelling	CodeCode Available	1
FSD: An Initial Chinese Dataset for Fake Song Detection	Sep 5, 2023	Audio Deepfake DetectionDeepFake Detection	CodeCode Available	1
Speech Resynthesis from Discrete Disentangled Self-Supervised Representations	Apr 1, 2021	DisentanglementRepresentation Learning	CodeCode Available	1
HiFi-VC: High Quality ASR-Based Voice Conversion	Mar 31, 2022	speech-recognitionSpeech Recognition	CodeCode Available	1
kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization	Apr 8, 2025	Voice Conversion	CodeCode Available	1
Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations	Oct 27, 2021	Voice Conversion	CodeCode Available	1
FastSVC: Fast Cross-Domain Singing Voice Conversion with Feature-wise Linear Modulation	Nov 11, 2020	Voice Conversion	CodeCode Available	1
Building Bilingual and Code-Switched Voice Conversion with Limited Training Data Using Embedding Consistency Loss	Apr 22, 2021	Voice CloningVoice Conversion	CodeCode Available	1
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions	May 19, 2022	Speech SynthesisStyle Transfer	CodeCode Available	1
Efficient Non-Autoregressive GAN Voice Conversion using VQWav2vec Features and Dynamic Convolution	Mar 31, 2022	Voice Conversion	CodeCode Available	1
Emotional Voice Conversion: Theory, Databases and ESD	May 31, 2021	Voice Conversion	CodeCode Available	1
BiSinger: Bilingual Singing Voice Synthesis	Sep 25, 2023	Singing Voice Synthesistext-to-speech	CodeCode Available	1
A Comparative Study of Self-supervised Speech Representation Based Voice Conversion	Jul 10, 2022	Voice Conversion	CodeCode Available	1
DuTa-VC: A Duration-aware Typical-to-atypical Voice Conversion Approach with Diffusion Probabilistic Model	Jun 18, 2023	Data AugmentationDecoder	CodeCode Available	1
Emo-StarGAN: A Semi-Supervised Any-to-Many Non-Parallel Emotion-Preserving Voice Conversion	Sep 14, 2023	Voice Conversion	CodeCode Available	1
Emotionless: Privacy-Preserving Speech Analysis for Voice Assistants	Aug 9, 2019	Emotion RecognitionPrivacy Preserving	CodeCode Available	1
Any-to-Many Voice Conversion with Location-Relative Sequence-to-Sequence Modeling	Sep 6, 2020	feature selectionspeech-recognition	CodeCode Available	1
Evaluating Methods for Ground-Truth-Free Foreign Accent Conversion	Sep 5, 2023	Voice Conversion	CodeCode Available	1
F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder	Apr 15, 2020	Style TransferVoice Conversion	CodeCode Available	1
FMFCC-A: A Challenging Mandarin Dataset for Synthetic Speech Detection	Oct 18, 2021	Speech SynthesisSynthetic Speech Detection	CodeCode Available	1
CinC-GAN for Effective F0 prediction for Whisper-to-Normal Speech Conversion	Aug 18, 2020	PredictionVoice Conversion	CodeCode Available	1
Baseline System of Voice Conversion Challenge 2020 with Cyclic Variational Autoencoder and Parallel WaveGAN	Oct 9, 2020	Generative Adversarial NetworkTask 2	CodeCode Available	1
Hiding speaker's sex in speech using zero-evidence speaker representation in an analysis/synthesis pipeline	Nov 29, 2022	Voice Conversion	CodeCode Available	1
Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme	Sep 28, 2021	Speech SynthesisVoice Conversion	CodeCode Available	1
Defending Your Voice: Adversarial Attack on Voice Conversion	May 18, 2020	Adversarial AttackVoice Conversion	CodeCode Available	1
Anonymizing Speech: Evaluating and Designing Speaker Anonymization Techniques	Aug 5, 2023	QuantizationSpeaker anonymization	CodeCode Available	1
Improving fairness for spoken language understanding in atypical speech with Text-to-Speech	Nov 16, 2023	Data AugmentationFairness	CodeCode Available	1
DeID-VC: Speaker De-identification via Zero-shot Pseudo Voice Conversion	Sep 9, 2022	De-identificationSpeaker Verification	CodeCode Available	1
LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech	Oct 18, 2021	Voice Conversion	CodeCode Available	1
Controllable and Interpretable Singing Voice Decomposition via Assem-VC	Oct 25, 2021	Voice Conversion	CodeCode Available	1
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed	Sep 23, 2022	Pitch controlSpeech Synthesis	CodeCode Available	1
Disentanglement in a GAN for Unconditional Speech Synthesis	Jul 4, 2023	DisentanglementGenerative Adversarial Network	CodeCode Available	1
AutoVisual Fusion Suite: A Comprehensive Evaluation of Image Segmentation and Voice Conversion Tools on HuggingFace Platform	Dec 17, 2023	Image SegmentationSegmentation	CodeCode Available	1
MaskCycleGAN-VC: Learning Non-parallel Voice Conversion with Filling in Frames	Feb 25, 2021	Voice Conversion	CodeCode Available	1
CycleTransGAN-EVC: A CycleGAN-based Emotional Voice Conversion Model with Transformer	Nov 30, 2021	Voice Conversion	CodeCode Available	1
Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep Learning	Jun 15, 2022	AttributeEmotion Classification	CodeCode Available	1
CSLP-AE: A Contrastive Split-Latent Permutation Autoencoder Framework for Zero-Shot Electroencephalography Signal Conversion	Nov 13, 2023	Contrastive LearningEEG	CodeCode Available	1
Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion	Jun 3, 2019	Audio GenerationVoice Conversion	CodeCode Available	1
crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder	Mar 4, 2021	Voice Conversion	CodeCode Available	1
A unified one-shot prosody and speaker conversion system with self-supervised discrete speech units	Nov 12, 2022	RhythmVoice Conversion	CodeCode Available	1
CycleGAN-VC3: Examining and Improving CycleGAN-VCs for Mel-spectrogram Conversion	Oct 22, 2020	Voice Conversion	CodeCode Available	1

Show:10 25 50

← PrevPage 2 of 11Next →

All datasets ZeroSpeech 2019 English LibriSpeech test-clean VCTK

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	VQ-CPC	Speaker Similarity	3.8	—	Unverified
2	VQ-VAE	Speaker Similarity	3.49	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	kNN-VC (prematched HiFiGAN)	Character Error Rate (CER)	2.96	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	DISSC	Total Length Error (TLE)	0.83	—	Unverified