SOTAVerified

Multiple-choice

Papers

Showing 651675 of 1107 papers

TitleStatusHype
DeSIQ: Towards an Unbiased, Challenging Benchmark for Social Intelligence Understanding0
POE: Process of Elimination for Multiple Choice ReasoningCode0
Dataset Bias Mitigation in Multiple-Choice Visual Question Answering and Beyond0
StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical UnderstandingCode0
Field-testing items using artificial intelligence: Natural language processing with transformers0
Investigating Uncertainty Calibration of Aligned Language Models under the Multiple-Choice Setting0
Evaluating the Symbol Binding Ability of Large Language Models for Multiple-Choice Questions in Vietnamese General Education0
JMedLoRA:Medical Domain Adaptation on Japanese Large Language Models using Instruction-tuningCode1
KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language ModelsCode0
Mitigating Bias for Question Answering Models by Tracking Bias Influence0
OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language ModelsCode1
BRAINTEASER: Lateral Thinking Puzzles for Large Language ModelsCode1
Analyzing Zero-Shot Abilities of Vision-Language Models on Video Understanding Tasks0
LLM-Coordination: Evaluating and Analyzing Multi-agent Coordination Abilities in Large Language ModelsCode1
On the Performance of Multimodal Language Models0
AutoCast++: Enhancing World Event Prediction with Zero-shot Ranking-based Context RetrievalCode0
Can Large Language Models Provide Security & Privacy Advice? Measuring the Ability of LLMs to Refute MisconceptionsCode0
Language Models as Knowledge Bases for Visual Word Sense DisambiguationCode0
Fusing Models with Complementary ExpertiseCode0
Fool Your (Vision and) Language Model With Embarrassingly Simple PermutationsCode1
Automating question generation from educational text0
HANS, are you clever? Clever Hans Effect Analysis of Neural Systems0
Exploring Iterative Enhancement for Improving Learnersourced Multiple-Choice Question Explanations with Large Language ModelsCode0
Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model EvaluationCode1
Benchmarks for Pirá 2.0, a Reading Comprehension Dataset about the Ocean, the Brazilian Coast, and Climate Change0
Show:102550
← PrevPage 27 of 45Next →

No leaderboard results yet.