SOTAVerified

Dialogue Evaluation

Papers

Showing 21–30 of 97 papers

TitleStatusHype
Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationCode1
DEnsity: Open-domain Dialogue Evaluation Metric using Density EstimationCode1
DialogBench: Evaluating LLMs as Human-like Dialogue SystemsCode1
Assessing Dialogue Systems with Distribution DistancesCode1
Generating Negative Samples by Manipulating Golden Responses for Unsupervised Learning of a Response Evaluation ModelCode0
GCDF1: A Goal- and Context- Driven F-Score for Evaluating User ModelsCode0
Achieving Reliable Human Assessment of Open-Domain Dialogue SystemsCode0
Exploring the Impact of Human Evaluator Group on Chat-Oriented Dialogue EvaluationCode0
ECoh: Turn-level Coherence Evaluation for Multilingual DialoguesCode0
Deconstruct to Reconstruct a Configurable Evaluation Metric for Open-Domain Dialogue SystemsCode0
Show:102550
← PrevPage 3 of 10Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MDD-EvalSpearman Correlation0.51—Unverified
2Lin-Reg (all)Spearman Correlation0.49—Unverified
3USRSpearman Correlation0.42—Unverified
4USR - DR (x = c)Spearman Correlation0.32—Unverified
5USR - MLMSpearman Correlation0.31—Unverified
6USR - DR (x = f)Spearman Correlation0.14—Unverified
#ModelMetricClaimedVerifiedStatus
1Lin-Reg (all)Spearman Correlation0.54—Unverified
2USR - DR (x = c)Spearman Correlation0.48—Unverified
3USRSpearman Correlation0.47—Unverified
4USR - MLMSpearman Correlation0.08—Unverified
5USR - DR (x = f)Spearman Correlation-0.05—Unverified