| MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Chatbots and Dialogue Evaluators | May 28, 2025 | BenchmarkingChatbot | CodeCode Available | 0 |
| Methods for Recognizing Nested Terms | Apr 22, 2025 | Dialogue Evaluationnamed-entity-recognition | CodeCode Available | 0 |
| PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison | Apr 1, 2024 | Dialogue Evaluation | CodeCode Available | 0 |
| Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems | Nov 4, 2019 | Dialogue Evaluation | CodeCode Available | 0 |
| Proxy Indicators for the Quality of Open-domain Dialogues | Nov 1, 2021 | Dialogue Evaluation | CodeCode Available | 0 |
| RuOpinionNE-2024: Extraction of Opinion Tuples from Russian News Texts | Apr 9, 2025 | Dialogue EvaluationLanguage Modeling | CodeCode Available | 0 |
| SelF-Eval: Self-supervised Fine-grained Dialogue Evaluation | Aug 17, 2022 | Contrastive LearningDialogue Evaluation | CodeCode Available | 0 |
| Simple LLM Prompting is State-of-the-Art for Robust and Multilingual Dialogue Evaluation | Aug 31, 2023 | Dialogue Evaluation | CodeCode Available | 0 |
| SLIDE: A Framework Integrating Small and Large Language Models for Open-Domain Dialogues Evaluation | May 24, 2024 | Contrastive LearningDialogue Evaluation | CodeCode Available | 0 |
| Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs | Aug 20, 2024 | Dialogue Evaluation | CodeCode Available | 0 |