SOTAVerified

Dialogue Evaluation

Papers

Showing 6170 of 97 papers

TitleStatusHype
xDial-Eval: A Multilingual Open-Domain Dialogue Evaluation BenchmarkCode0
Achieving Reliable Human Assessment of Open-Domain Dialogue SystemsCode0
A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsCode0
Adversarial Learning for Neural Dialogue GenerationCode0
A Human-machine Collaborative Framework for Evaluating Malevolence in DialoguesCode0
An Adversarially-Learned Turing Test for Dialog Generation ModelsCode0
Approximating Interactive Human Evaluation with Self-Play for Open-Domain Dialog SystemsCode0
BoK: Introducing Bag-of-Keywords Loss for Interpretable Dialogue Response GenerationCode0
C-PMI: Conditional Pointwise Mutual Information for Turn-level Dialogue EvaluationCode0
DEAM: Dialogue Coherence Evaluation using AMR-based Semantic ManipulationsCode0
Show:102550
← PrevPage 7 of 10Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MDD-EvalSpearman Correlation0.51Unverified
2Lin-Reg (all)Spearman Correlation0.49Unverified
3USRSpearman Correlation0.42Unverified
4USR - DR (x = c)Spearman Correlation0.32Unverified
5USR - MLMSpearman Correlation0.31Unverified
6USR - DR (x = f)Spearman Correlation0.14Unverified
#ModelMetricClaimedVerifiedStatus
1Lin-Reg (all)Spearman Correlation0.54Unverified
2USR - DR (x = c)Spearman Correlation0.48Unverified
3USRSpearman Correlation0.47Unverified
4USR - MLMSpearman Correlation0.08Unverified
5USR - DR (x = f)Spearman Correlation-0.05Unverified