SOTAVerified

Text to Speech

import gTTS import os def text_to_speech_kurdish(text, output_file="output.mp3"): # گۆڕینی نووسین بۆ دەنگ بە زمانی کوردی (هەڵبژاردنی زمانی "ku" بۆ کوردی) tts = gTTS(text=text, lang='ku', slow=False) tts.save(output_file) os.system(f"start {output_file}") # کردنەوەی فایلە دەنگییەکە (لە Windows) # نموونە: text_to_speech_kurdish("سڵاو، ئەمە دەنگی منە بە زمانی کوردی.")

Papers

Showing 701–750 of 1419 papers

TitleStatusHype
Large-Scale Automatic Audiobook Creation—0
GRASS: Unified Generation Model for Speech-to-Semantic Tasks—0
MuLanTTS: The Microsoft Speech Synthesis System for Blizzard Challenge 2023—0
PromptTTS 2: Describing and Generating Voices with Text Prompt—0
A Comparative Analysis of Pretrained Language Models for Text-to-Speech—0
The FruitShell French synthesis system at the Blizzard 2023 Challenge—0
Learning Speech Representation From Contrastive Token-Acoustic Pretraining—0
Improving Mandarin Prosodic Structure Prediction with Multi-level Contextual Information—0
Towards Spontaneous Style Modeling with Semi-supervised Pre-training for Conversational Text-to-Speech Synthesis—0
The DeepZen Speech Synthesis System for Blizzard Challenge 2023—0
Pruning Self-Attention for Zero-Shot Multi-Speaker Text-to-Speech—0
Rep2wav: Noise Robust text-to-speech Using self-supervised representations—0
Generalizable Zero-Shot Speaker Adaptive Speech Synthesis with Disentangled Representations—0
Multi-GradSpeech: Towards Diffusion-based Multi-Speaker Text-to-speech Using Consistent Diffusion Models—0
AffectEcho: Speaker Independent and Language-Agnostic Emotion and Affect Transfer for Speech Synthesis—0
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer—0
Text-to-Video: a Two-stage Framework for Zero-shot Identity-agnostic Talking-head GenerationCode0
Let's Give a Voice to Conversational Agents in Virtual RealityCode0
SALTTS: Leveraging Self-Supervised Speech Representations for improved Text-to-Speech Synthesis—0
Comparing normalizing flows and diffusion models for prosody and acoustic modelling in text-to-speech—0
Improving grapheme-to-phoneme conversion by learning pronunciations from speech recordings—0
Multilingual context-based pronunciation learning for Text-to-Speech—0
METTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer—0
Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic Coding—0
SLMGAN: Exploiting Speech Language Model Representations for Unsupervised Zero-Shot Voice Conversion in GANs—0
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis—0
Controllable Emphasis with zero data for text-to-speech—0
On the Use of Self-Supervised Speech Representations in Spontaneous Speech Synthesis—0
Artificial Eye for the Blind—0
ContextSpeech: Expressive and Efficient Text-to-Speech for Paragraph Reading—0
High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units—0
GenerTTS: Pronunciation Disentanglement for Timbre and Style Generalization in Cross-Lingual Text-to-Speech—0
DSE-TTS: Dual Speaker Embedding for Cross-Lingual Text-to-Speech—0
Voicebox: Text-Guided Multilingual Universal Speech Generation at ScaleCode0
Visual-Aware Text-to-Speech—0
Expressive Machine Dubbing Through Phrase-level Cross-lingual Prosody Transfer—0
Low-Resource Text-to-Speech Using Specific Data and Noise Augmentation—0
CML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource Languages—0
Improving Code-Switching and Named Entity Recognition in ASR with Speech Editing based Data Augmentation—0
PauseSpeech: Natural Speech Synthesis via Pre-trained Language Model and Pause-based Prosody Modeling—0
Learning Emotional Representations from Imbalanced Speech Data for Speech Emotion Recognition and Emotional Text-to-Speech—0
VIFS: An End-to-End Variational Inference for Foley Sound SynthesisCode0
Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias—0
Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar Synthesis—0
Cross-Lingual Transfer Learning for Phrase Break Prediction with Multilingual Language Model—0
Latent Optimal Paths by Gumbel Propagation for Variational Bayesian Dynamic ProgrammingCode0
Rhythm-controllable Attention with High Robustness for Long Sentence Speech Synthesis—0
Towards Robust FastSpeech 2 by Modelling Residual Multimodality—0
The Effects of Input Type and Pronunciation Dictionary Usage in Transfer Learning for Low-Resource Text-to-Speech—0
Text-to-Speech Pipeline for Swiss German -- A comparison—0
Show:102550
← PrevPage 15 of 29Next →

No leaderboard results yet.