SOTAVerified

nlg evaluation

Evaluate the generated text by NLG (Natural Language Generation) systems, like large language models

Papers

Showing 61–70 of 71 papers

TitleStatusHype
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects—0
Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications—0
DeepSeek vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?—0
SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text—0
Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation—0
DHP Benchmark: Are LLMs Good NLG Evaluators?—0
Dialect-robust Evaluation of Generated Text—0
Dolphin: A Challenging and Diverse Benchmark for Arabic NLG—0
Ev2R: Evaluating Evidence Retrieval in Automated Fact-Checking—0
A Tutorial on Evaluation Metrics used in Natural Language Generation—0
Show:102550
← PrevPage 7 of 8Next →

No leaderboard results yet.