SOTAVerified

Audio Generation

Audio generation (synthesis) is the task of generating raw audio such as speech.

( Image credit: MelNet )

Papers

Showing 211220 of 270 papers

TitleStatusHype
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation0
Modeling and Driving Human Body Soundfields through Acoustic Primitives0
Music Source Separation in the Waveform Domain0
Music Style Transfer With Diffusion Model0
NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization0
Neural Granular Sound Synthesis0
Nonparametric estimation of a factorizable density using diffusion models0
NU-GAN: High resolution neural upsampling with GAN0
On Target Representation in Continuous-output Neural Machine Translation0
On the Design of Diffusion-based Neural Speech Codecs0
Show:102550
← PrevPage 22 of 27Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1AudioGenFD_openl3185.53Unverified
2AudioLDM2-largeFD_openl3158.04Unverified
3Stable Audio 2.0FD_openl3110.62Unverified
4Stable AudioFD_openl3103.66Unverified
5ETTAFD_openl380.13Unverified
6TangoFlux-baseFD_openl379.7Unverified
7Stable Audio OpenFD_openl378.24Unverified
8TangoFluxFD_openl375.1Unverified
9ETTA-FT-AC-100kFD_openl361.79Unverified
10DiffsoundFAD7.75Unverified
#ModelMetricClaimedVerifiedStatus
1VAB-Encodec (Ours)Bits per byte40Unverified
2Sparse Transformer 152M (strided)Bits per byte1.97Unverified
#ModelMetricClaimedVerifiedStatus
1SymphonyNet Human listening average results3.5Unverified