| Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models | Apr 7, 2025 | Dialogue EvaluationFairness | CodeCode Available | 2 |
| DialogBench: Evaluating LLMs as Human-like Dialogue Systems | Nov 3, 2023 | Dialogue Evaluation | CodeCode Available | 1 |
| DEnsity: Open-domain Dialogue Evaluation Metric using Density Estimation | May 8, 2023 | Contrastive LearningDensity Estimation | CodeCode Available | 1 |
| GLM-Dialog: Noise-tolerant Pre-training for Knowledge-grounded Dialogue Generation | Feb 28, 2023 | Dialogue EvaluationDialogue Generation | CodeCode Available | 1 |
| Don't Forget Your ABC's: Evaluating the State-of-the-Art in Chat-Oriented Dialogue Systems | Dec 18, 2022 | ChatbotDialogue Evaluation | CodeCode Available | 1 |
| FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation | Oct 25, 2022 | Dialogue Evaluation | CodeCode Available | 1 |
| Findings of the The RuATD Shared Task 2022 on Artificial Text Detection in Russian | Jun 3, 2022 | Binary ClassificationDialogue Evaluation | CodeCode Available | 1 |
| InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning | May 25, 2022 | Dialogue EvaluationDialogue Generation | CodeCode Available | 1 |
| RuNNE-2022 Shared Task: Recognizing Nested Named Entities | May 23, 2022 | Dialogue Evaluationnamed-entity-recognition | CodeCode Available | 1 |
| Automatic Evaluation and Moderation of Open-domain Dialogue Systems | Nov 3, 2021 | ChatbotDialogue Evaluation | CodeCode Available | 1 |