SOTAVerified

Dialogue Evaluation

Papers

Showing 1120 of 97 papers

TitleStatusHype
Assessing Dialogue Systems with Distribution DistancesCode1
InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction TuningCode1
Automatic Evaluation and Moderation of Open-domain Dialogue SystemsCode1
PONE: A Novel Automatic Evaluation Metric for Open-Domain Generative Dialogue SystemsCode1
Improving Dialog Evaluation with a Multi-reference Adversarial Dataset and Large Scale PretrainingCode1
DEnsity: Open-domain Dialogue Evaluation Metric using Density EstimationCode1
Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue UtterancesCode1
DialogBench: Evaluating LLMs as Human-like Dialogue SystemsCode1
A Comprehensive Assessment of Dialog Evaluation MetricsCode1
RuNNE-2022 Shared Task: Recognizing Nested Named EntitiesCode1
Show:102550
← PrevPage 2 of 10Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MDD-EvalSpearman Correlation0.51Unverified
2Lin-Reg (all)Spearman Correlation0.49Unverified
3USRSpearman Correlation0.42Unverified
4USR - DR (x = c)Spearman Correlation0.32Unverified
5USR - MLMSpearman Correlation0.31Unverified
6USR - DR (x = f)Spearman Correlation0.14Unverified
#ModelMetricClaimedVerifiedStatus
1Lin-Reg (all)Spearman Correlation0.54Unverified
2USR - DR (x = c)Spearman Correlation0.48Unverified
3USRSpearman Correlation0.47Unverified
4USR - MLMSpearman Correlation0.08Unverified
5USR - DR (x = f)Spearman Correlation-0.05Unverified