SOTAVerified

Text to Speech

import gTTS import os def text_to_speech_kurdish(text, output_file="output.mp3"): # گۆڕینی نووسین بۆ دەنگ بە زمانی کوردی (هەڵبژاردنی زمانی "ku" بۆ کوردی) tts = gTTS(text=text, lang='ku', slow=False) tts.save(output_file) os.system(f"start {output_file}") # کردنەوەی فایلە دەنگییەکە (لە Windows) # نموونە: text_to_speech_kurdish("سڵاو، ئەمە دەنگی منە بە زمانی کوردی.")

Papers

Showing 601–650 of 1419 papers

TitleStatusHype
Building a Luganda Text-to-Speech Model From Crowdsourced Data—0
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text—0
Towards Evaluating the Robustness of Automatic Speech Recognition Systems via Audio Style Transfer—0
PolyGlotFake: A Novel Multilingual and Multimodal DeepFake DatasetCode0
Real-Time Pill Identification for the Visually Impaired Using Deep Learning—0
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech—0
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality—0
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations—0
Retrieval-Augmented Audio Deepfake Detection—0
Prior-agnostic Multi-scale Contrastive Text-Audio Pre-training for Parallelized TTS Frontend Modeling—0
Voice-Assisted Real-Time Traffic Sign Recognition System Using Convolutional Neural Network—0
The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge—0
Cross-Domain Audio Deepfake Detection: Dataset and Analysis—0
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis—0
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech—0
PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders—0
Humane Speech Synthesis through Zero-Shot Emotion and Disfluency GenerationCode0
A Review of Multi-Modal Large Language and Vision Models—0
Isometric Neural Machine Translation using Phoneme Count Ratio Reward-based Reinforcement Learning—0
Creating an African American-Sounding TTS: Guidelines, Technical Challenges,and Surprising Evaluations—0
EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech—0
Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation—0
AttentionStitch: How Attention Solves the Speech Editing Problem—0
Towards Accurate Lip-to-Speech Synthesis in-the-Wild—0
Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data—0
Efficient data selection employing Semantic Similarity-based Graph Structures for model training—0
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition—0
On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models—0
Bayesian Parameter-Efficient Fine-Tuning for Overcoming Catastrophic ForgettingCode0
Ain't Misbehavin' -- Using LLMs to Generate Expressive Robot Behavior in Conversations with the Tabletop Robot Haru—0
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech—0
Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like—0
BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data—0
A New Approach to Voice Authenticity—0
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations—0
Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech—0
MunTTS: A Text-to-Speech System for Mundari—0
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech—0
Maximizing Data Efficiency for Cross-Lingual TTS Adaptation by Self-Supervised Representation Mixing and Embedding Initialization—0
Adversarial speech for voice privacy protection from Personalized Speech generation—0
Empowering Communication: Speech Technology for Indian and Western Accents through AI-powered Speech Synthesis—0
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech—0
MCMChaos: Improvising Rap Music with MCMC Methods and Chaos Theory—0
ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering—0
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec2—0
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters—0
Evaluating and Personalizing User-Perceived Quality of Text-to-Speech Voices for Delivering Mindfulness Meditation with Different Physical Embodiments—0
Transfer the linguistic representations from TTS to accent conversion with non-parallel data—0
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction—0
Incremental FastPitch: Chunk-based High Quality Text to Speech—0
Show:102550
← PrevPage 13 of 29Next →

No leaderboard results yet.