SOTAVerified

Multiple-choice

Papers

Showing 171180 of 1107 papers

TitleStatusHype
Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language ModelsCode1
LogiDynamics: Unraveling the Dynamics of Logical Inference in Large Language Model Reasoning0
VisCon-100K: Leveraging Contextual Web Data for Fine-tuning Vision Language Models0
Objective quantification of mood states using large language models0
Truth Knows No Language: Evaluating Truthfulness Beyond EnglishCode0
SB-Bench: Stereotype Bias Benchmark for Large Multimodal Models0
A Semantic Parsing Algorithm to Solve Linear Ordering Problems0
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs0
PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian0
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark0
Show:102550
← PrevPage 18 of 111Next →

No leaderboard results yet.