SOTAVerified

Sentence Completion

Papers

Showing 41–50 of 91 papers

TitleStatusHype
Investigating Subtler Biases in LLMs: Ageism, Beauty, Institutional, and Nationality Bias in Generative ModelsCode0
Exploiting Language Models as a Source of Knowledge for Cognitive Agents—0
I-WAS: a Data Augmentation Method with GPT-2 for Simile Detection—0
Stay on topic with Classifier-Free Guidance—0
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context LearningCode0
PaLM 2 Technical Report—0
BloombergGPT: A Large Language Model for Finance—0
Numeracy from Literacy: Data Science as an Emergent Skill from Large Language Models—0
POIBERT: A Transformer-based Model for the Tour Recommendation Problem—0
Implicit causality in GPT-2: a case study—0
Show:102550
← PrevPage 5 of 10Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1CompassMTL 567M with TailorAccuracy96.1—Unverified
2CompassMTL 567MAccuracy95.6—Unverified
3DeBERTa-Large 304M (classification-based)Accuracy95.6—Unverified
4GPT-4 (10-shot)Accuracy95.3—Unverified
5LLaMA3+MoSLoRAAccuracy95—Unverified
6LLaMA-2 13B + MixLoRAAccuracy94.7—Unverified
7DeBERTa-Large 304MAccuracy94.7—Unverified
8Unicorn 11B (fine-tuned)Accuracy93.9—Unverified
9LLaMA-3 8B + MixLoRAAccuracy93.3—Unverified
10LLaMA-2 7B + MixLoRAAccuracy93.1—Unverified