SOTAVerified

Dialogue Evaluation

Papers

Showing 41–50 of 97 papers

TitleStatusHype
PoE: a Panel of Experts for Generalized Automatic Dialogue Assessment—0
Pragmatically Appropriate Diversity for Dialogue Evaluation—0
Predicting Ratings of Real Dialogue Participants from Artificial Data and Ratings of Human Dialogue Observers—0
Dialogue Evaluation with Offline Reinforcement Learning—0
RADE: Reference-Assisted Dialogue Evaluation for Open-Domain Dialogue—0
Re-evaluating ADEM: A Deeper Look at Scoring Dialogue Responses—0
Report from the NSF Future Directions Workshop on Automatic Evaluation of Dialog: Research Directions and Challenges—0
Dialogue You Can Trust: Human and AI Perspectives on Generated Conversations—0
DRE: An Effective Dual-Refined Method for Integrating Small and Large Language Models in Open-Domain Dialogue Evaluation—0
Enhancing the Open-Domain Dialogue Evaluation in Latent Space—0
Show:102550
← PrevPage 5 of 10Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MDD-EvalSpearman Correlation0.51—Unverified
2Lin-Reg (all)Spearman Correlation0.49—Unverified
3USRSpearman Correlation0.42—Unverified
4USR - DR (x = c)Spearman Correlation0.32—Unverified
5USR - MLMSpearman Correlation0.31—Unverified
6USR - DR (x = f)Spearman Correlation0.14—Unverified
#ModelMetricClaimedVerifiedStatus
1Lin-Reg (all)Spearman Correlation0.54—Unverified
2USR - DR (x = c)Spearman Correlation0.48—Unverified
3USRSpearman Correlation0.47—Unverified
4USR - MLMSpearman Correlation0.08—Unverified
5USR - DR (x = f)Spearman Correlation-0.05—Unverified