SOTAVerified

Language Model Evaluation

The task of using LLMs as evaluators of large language and vision language models.

Papers

Showing 125 of 69 papers

TitleStatusHype
Enterprise Large Language Model Evaluation Benchmark0
Finance Language Model Evaluation (FLaME)0
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models0
FABLE: A Novel Data-Flow Analysis Benchmark on Procedural Text for Large Language Model EvaluationCode0
Role-Playing Evaluation for Large Language ModelsCode1
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation0
Confidence in Large Language Model Evaluation: A Bayesian Approach to Limited-Sample Challenges0
CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation0
UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation0
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation0
Predicting Liquidity-Aware Bond Yields using Causal GANs and Deep Reinforcement Learning with LLM Evaluation0
M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment AnalysisCode1
Environmental large language model Evaluation (ELLE) dataset: A Benchmark for Evaluating Generative AI applications in Eco-environment DomainCode0
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model EvaluationCode1
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation0
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests0
Template Matters: Understanding the Role of Instruction Templates in Multimodal Language Model Evaluation and TrainingCode1
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNACode1
C^2LEVA: Toward Comprehensive and Contamination-Free Language Model EvaluationCode2
Benchmarking Harmonized Tariff Schedule Classification Models0
Large Language Model Evaluation via Matrix Nuclear-NormCode0
Enterprise Benchmarks for Large Language Model EvaluationCode0
ViDAS: Vision-based Danger Assessment and Scoring0
Mitigating the Bias of Large Language Model EvaluationCode0
Salmon: A Suite for Acoustic Language Model EvaluationCode1
Show:102550
← PrevPage 1 of 3Next →

No leaderboard results yet.