SOTAVerified

Dialogue Evaluation

Papers

Showing 51–60 of 97 papers

TitleStatusHype
CodingTeachLLM: Empowering LLM's Coding Ability via AST Prior Knowledge—0
Explaining Dialogue Evaluation Metrics using Adversarial Behavioral Analysis—0
Treating Dialogue Quality Evaluation as an Anomaly Detection Problem—0
U-NEED: A Fine-grained Dataset for User Needs-Centric E-commerce Conversational Recommendation—0
User Response and Sentiment Prediction for Automatic Dialogue Evaluation—0
WeChat AI & ICT's Submission for DSTC9 Interactive Dialogue Evaluation Track—0
FlowEval: A Consensus-Based Dialogue Evaluation Framework Using Segment Act Flows—0
Better Automatic Evaluation of Open-Domain Dialogue Systems with Contextualized Embeddings—0
How to Choose How to Choose Your Chatbot: A Massively Multi-System MultiReference Data Set for Dialog Metric Evaluation—0
How to Evaluate the Next System: Automatic Dialogue Evaluation from the Perspective of Continual Learning—0
Show:102550
← PrevPage 6 of 10Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MDD-EvalSpearman Correlation0.51—Unverified
2Lin-Reg (all)Spearman Correlation0.49—Unverified
3USRSpearman Correlation0.42—Unverified
4USR - DR (x = c)Spearman Correlation0.32—Unverified
5USR - MLMSpearman Correlation0.31—Unverified
6USR - DR (x = f)Spearman Correlation0.14—Unverified
#ModelMetricClaimedVerifiedStatus
1Lin-Reg (all)Spearman Correlation0.54—Unverified
2USR - DR (x = c)Spearman Correlation0.48—Unverified
3USRSpearman Correlation0.47—Unverified
4USR - MLMSpearman Correlation0.08—Unverified
5USR - DR (x = f)Spearman Correlation-0.05—Unverified