SOTAVerified

Text to Speech

import gTTS import os def text_to_speech_kurdish(text, output_file="output.mp3"): # گۆڕینی نووسین بۆ دەنگ بە زمانی کوردی (هەڵبژاردنی زمانی "ku" بۆ کوردی) tts = gTTS(text=text, lang='ku', slow=False) tts.save(output_file) os.system(f"start {output_file}") # کردنەوەی فایلە دەنگییەکە (لە Windows) # نموونە: text_to_speech_kurdish("سڵاو، ئەمە دەنگی منە بە زمانی کوردی.")

Papers

Showing 526550 of 1419 papers

TitleStatusHype
E1 TTS: Simple and Fast Non-Autoregressive TTS0
Dynamic Prosody Generation for Speech Synthesis using Linguistics-Driven Acoustic Embedding Selection0
DurIAN-E: Duration Informed Attention Network For Expressive Text-to-Speech Synthesis0
Beyond Text-to-Text: An Overview of Multimodal and Generative Artificial Intelligence for Education Using Topic Modeling0
A Novel Approach to OCR using Image Recognition based Classification for Ancient Tamil Inscriptions in Temples0
DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis0
Duration-aware pause insertion using pre-trained language model for multi-speaker text-to-speech0
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing0
Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset0
Dual Supervised Learning0
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance0
BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model0
Dual Script E2E framework for Multilingual and Code-Switching ASR0
Dual Audio-Centric Modality Coupling for Talking Head Generation0
Anonymizing Speech with Generative Adversarial Networks to Preserve Speaker Privacy0
DTW-SiameseNet: Dynamic Time Warped Siamese Network for Mispronunciation Detection and Correction0
DSE-TTS: Dual Speaker Embedding for Cross-Lingual Text-to-Speech0
Benchmarking Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS20
DPP-TTS: Diversifying prosodic features of speech via determinantal point processes0
LAraBench: Benchmarking Arabic AI with Large Language Models0
An objective evaluation of the effects of recording conditions and speaker characteristics in multi-speaker deep neural speech synthesis0
Empowering Communication: Speech Technology for Indian and Western Accents through AI-powered Speech Synthesis0
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech0
Do Prosody Transfer Models Transfer Prosody?0
Does Audio Deepfake Detection Generalize?0
Show:102550
← PrevPage 22 of 57Next →

No leaderboard results yet.