SOTAVerified

Dialogue Evaluation

Papers

Showing 9197 of 97 papers

TitleStatusHype
Emphasising Structured Information: Integrating Abstract Meaning Representation into LLMs for Enhanced Open-Domain Dialogue EvaluationCode0
Synthesizing Adversarial Negative Responses for Robust Response Ranking and EvaluationCode0
Towards an Automatic Turing Test: Learning to Evaluate Dialogue ResponsesCode0
Towards Multilingual Automatic Dialogue EvaluationCode0
Transformers for Headline Selection for Russian News ClustersCode0
What is wrong with you?: Leveraging User Sentiment for Automatic Dialog EvaluationCode0
Towards Best Experiment Design for Evaluating Dialogue System OutputCode0
Show:102550
← PrevPage 10 of 10Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MDD-EvalSpearman Correlation0.51Unverified
2Lin-Reg (all)Spearman Correlation0.49Unverified
3USRSpearman Correlation0.42Unverified
4USR - DR (x = c)Spearman Correlation0.32Unverified
5USR - MLMSpearman Correlation0.31Unverified
6USR - DR (x = f)Spearman Correlation0.14Unverified
#ModelMetricClaimedVerifiedStatus
1Lin-Reg (all)Spearman Correlation0.54Unverified
2USR - DR (x = c)Spearman Correlation0.48Unverified
3USRSpearman Correlation0.47Unverified
4USR - MLMSpearman Correlation0.08Unverified
5USR - DR (x = f)Spearman Correlation-0.05Unverified