SOTAVerified

Dialogue Evaluation

Papers

Showing 1–50 of 97 papers

TitleStatusHype
Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language ModelsCode2
Assessing Dialogue Systems with Distribution DistancesCode1
GLM-Dialog: Noise-tolerant Pre-training for Knowledge-grounded Dialogue GenerationCode1
GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue SystemsCode1
DialogBench: Evaluating LLMs as Human-like Dialogue SystemsCode1
DEnsity: Open-domain Dialogue Evaluation Metric using Density EstimationCode1
RUBER: An Unsupervised Method for Automatic Evaluation of Open-Domain Dialog SystemsCode1
Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationCode1
Improving Dialog Evaluation with a Multi-reference Adversarial Dataset and Large Scale PretrainingCode1
Automatic Evaluation and Moderation of Open-domain Dialogue SystemsCode1
InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction TuningCode1
Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue UtterancesCode1
Towards Quantifiable Dialogue Coherence EvaluationCode1
Findings of the The RuATD Shared Task 2022 on Artificial Text Detection in RussianCode1
Learning an Unreferenced Metric for Online Dialogue EvaluationCode1
RuNNE-2022 Shared Task: Recognizing Nested Named EntitiesCode1
Q^2: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringCode1
Unsupervised Evaluation of Interactive Dialog with DialoGPTCode1
Don't Forget Your ABC's: Evaluating the State-of-the-Art in Chat-Oriented Dialogue SystemsCode1
USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationCode1
A Comprehensive Assessment of Dialog Evaluation MetricsCode1
PONE: A Novel Automatic Evaluation Metric for Open-Domain Generative Dialogue SystemsCode1
DynaEval: Unifying Turn and Dialogue Level EvaluationCode1
FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationCode1
Human Evaluation of Conversations is an Open Problem: comparing the sensitivity of various methods for evaluating dialogue agents—0
Achieving Reliable Human Assessment of Open-Domain Dialogue Systems—0
Improving Open-Domain Dialogue Evaluation with a Causal Inference Model—0
Investigating the Impact of Pre-trained Language Models on Dialog Evaluation—0
Joint Goal Segmentation and Goal Success Prediction on Multi-Domain Conversations—0
Learning the Human Judgment for the Automatic Evaluation of Chatbot—0
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation—0
Leveraging LLMs for Dialogue Quality Measurement—0
LLM as a Scorer: The Impact of Output Order on Dialogue Evaluation—0
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation—0
DCH-2: A Parallel Customer-Helpdesk Dialogue Corpus with Distributions of Annotators' Labels—0
AdaCoach: A Virtual Coach for Training Customer Service Agents—0
ACUTE-EVAL: Improved Dialogue Evaluation with Optimized Questions and Multi-turn Comparisons—0
MME-CRS: Multi-Metric Evaluation Based on Correlation Re-Scaling for Evaluating Open-Domain Dialogue—0
One "Ruler" for All Languages: Multi-Lingual Dialogue Evaluation with Adversarial Multi-Task Learning—0
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation—0
PoE: a Panel of Experts for Generalized Automatic Dialogue Assessment—0
Pragmatically Appropriate Diversity for Dialogue Evaluation—0
Predicting Ratings of Real Dialogue Participants from Artificial Data and Ratings of Human Dialogue Observers—0
Dialogue Evaluation with Offline Reinforcement Learning—0
RADE: Reference-Assisted Dialogue Evaluation for Open-Domain Dialogue—0
Re-evaluating ADEM: A Deeper Look at Scoring Dialogue Responses—0
Report from the NSF Future Directions Workshop on Automatic Evaluation of Dialog: Research Directions and Challenges—0
Dialogue You Can Trust: Human and AI Perspectives on Generated Conversations—0
DRE: An Effective Dual-Refined Method for Integrating Small and Large Language Models in Open-Domain Dialogue Evaluation—0
Enhancing the Open-Domain Dialogue Evaluation in Latent Space—0
Show:102550
← PrevPage 1 of 2Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MDD-EvalSpearman Correlation0.51—Unverified
2Lin-Reg (all)Spearman Correlation0.49—Unverified
3USRSpearman Correlation0.42—Unverified
4USR - DR (x = c)Spearman Correlation0.32—Unverified
5USR - MLMSpearman Correlation0.31—Unverified
6USR - DR (x = f)Spearman Correlation0.14—Unverified
#ModelMetricClaimedVerifiedStatus
1Lin-Reg (all)Spearman Correlation0.54—Unverified
2USR - DR (x = c)Spearman Correlation0.48—Unverified
3USRSpearman Correlation0.47—Unverified
4USR - MLMSpearman Correlation0.08—Unverified
5USR - DR (x = f)Spearman Correlation-0.05—Unverified