SOTAVerified

Dialogue Evaluation

Papers

Showing 51–97 of 97 papers

TitleStatusHype
SelF-Eval: Self-supervised Fine-grained Dialogue EvaluationCode0
Simple LLM Prompting is State-of-the-Art for Robust and Multilingual Dialogue EvaluationCode0
SLIDE: A Framework Integrating Small and Large Language Models for Open-Domain Dialogues EvaluationCode0
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMsCode0
Emphasising Structured Information: Integrating Abstract Meaning Representation into LLMs for Enhanced Open-Domain Dialogue EvaluationCode0
Synthesizing Adversarial Negative Responses for Robust Response Ranking and EvaluationCode0
Towards an Automatic Turing Test: Learning to Evaluate Dialogue ResponsesCode0
Towards Multilingual Automatic Dialogue EvaluationCode0
Transformers for Headline Selection for Russian News ClustersCode0
What is wrong with you?: Leveraging User Sentiment for Automatic Dialog EvaluationCode0
Towards Best Experiment Design for Evaluating Dialogue System OutputCode0
MME-CRS: Multi-Metric Evaluation Based on Correlation Re-Scaling for Evaluating Open-Domain Dialogue—0
One "Ruler" for All Languages: Multi-Lingual Dialogue Evaluation with Adversarial Multi-Task Learning—0
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation—0
U-NEED: A Fine-grained Dataset for User Needs-Centric E-commerce Conversational Recommendation—0
PoE: a Panel of Experts for Generalized Automatic Dialogue Assessment—0
Dialogue You Can Trust: Human and AI Perspectives on Generated Conversations—0
Pragmatically Appropriate Diversity for Dialogue Evaluation—0
Predicting Ratings of Real Dialogue Participants from Artificial Data and Ratings of Human Dialogue Observers—0
ACUTE-EVAL: Improved Dialogue Evaluation with Optimized Questions and Multi-turn Comparisons—0
User Response and Sentiment Prediction for Automatic Dialogue Evaluation—0
Dialogue Evaluation with Offline Reinforcement Learning—0
RADE: Reference-Assisted Dialogue Evaluation for Open-Domain Dialogue—0
Re-evaluating ADEM: A Deeper Look at Scoring Dialogue Responses—0
Report from the NSF Future Directions Workshop on Automatic Evaluation of Dialog: Research Directions and Challenges—0
DCH-2: A Parallel Customer-Helpdesk Dialogue Corpus with Distributions of Annotators' Labels—0
FlowEval: A Consensus-Based Dialogue Evaluation Framework Using Segment Act Flows—0
How to Choose How to Choose Your Chatbot: A Massively Multi-System MultiReference Data Set for Dialog Metric Evaluation—0
How to Evaluate the Next System: Automatic Dialogue Evaluation from the Perspective of Continual Learning—0
Human Evaluation of Conversations is an Open Problem: comparing the sensitivity of various methods for evaluating dialogue agents—0
Better Automatic Evaluation of Open-Domain Dialogue Systems with Contextualized Embeddings—0
Explaining Dialogue Evaluation Metrics using Adversarial Behavioral Analysis—0
Improving Open-Domain Dialogue Evaluation with a Causal Inference Model—0
Enhancing the Open-Domain Dialogue Evaluation in Latent Space—0
CodingTeachLLM: Empowering LLM's Coding Ability via AST Prior Knowledge—0
Investigating the Impact of Pre-trained Language Models on Dialog Evaluation—0
Joint Goal Segmentation and Goal Success Prediction on Multi-Domain Conversations—0
DRE: An Effective Dual-Refined Method for Integrating Small and Large Language Models in Open-Domain Dialogue Evaluation—0
Learning the Human Judgment for the Automatic Evaluation of Chatbot—0
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation—0
Leveraging LLMs for Dialogue Quality Measurement—0
LLM as a Scorer: The Impact of Output Order on Dialogue Evaluation—0
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation—0
Achieving Reliable Human Assessment of Open-Domain Dialogue Systems—0
AdaCoach: A Virtual Coach for Training Customer Service Agents—0
WeChat AI & ICT's Submission for DSTC9 Interactive Dialogue Evaluation Track—0
Treating Dialogue Quality Evaluation as an Anomaly Detection Problem—0
Show:102550
← PrevPage 2 of 2Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MDD-EvalSpearman Correlation0.51—Unverified
2Lin-Reg (all)Spearman Correlation0.49—Unverified
3USRSpearman Correlation0.42—Unverified
4USR - DR (x = c)Spearman Correlation0.32—Unverified
5USR - MLMSpearman Correlation0.31—Unverified
6USR - DR (x = f)Spearman Correlation0.14—Unverified
#ModelMetricClaimedVerifiedStatus
1Lin-Reg (all)Spearman Correlation0.54—Unverified
2USR - DR (x = c)Spearman Correlation0.48—Unverified
3USRSpearman Correlation0.47—Unverified
4USR - MLMSpearman Correlation0.08—Unverified
5USR - DR (x = f)Spearman Correlation-0.05—Unverified