SOTAVerified

General Knowledge

This task aims to evaluate the ability of a model to answer general-knowledge questions.

Source: BIG-bench

Papers

Showing 351–375 of 399 papers

TitleStatusHype
Domain Generalization via Model-Agnostic Learning of Semantic FeaturesCode0
Model Compression with Two-stage Multi-teacher Knowledge Distillation for Web Question Answering System—0
Spoken Conversational Search for General Knowledge—0
A Human-Centered Data-Driven Planner-Actor-Critic Architecture via Logic Programming—0
QuaRTz: An Open-Domain Dataset of Qualitative Relationship Questions—0
GeoSQA: A Benchmark for Scenario-based Question Answering in the Geography Domain at High School Level—0
Joey NMT: A Minimalist NMT Toolkit for NovicesCode0
T-Norms Driven Loss Functions for Machine Learning—0
Integration of Imitation Learning using GAIL and Reinforcement Learning using Task-achievement Rewards via Probabilistic Graphical Model—0
A Joint Planning and Learning Framework for Human-Aided Decision-Making—0
Generating Question Relevant Captions to Aid Visual Question Answering—0
The World in My Mind: Visual Dialog with Adversarial Multi-modal Feature Encoding—0
Integrating Semantic Knowledge to Tackle Zero-shot Text ClassificationCode0
Transferable Natural Language Interface to Structured Queries aided by Adversarial Generation—0
Specifying Conceptual Models Using Restricted Natural Language—0
Learning to Specialize with Knowledge Distillation for Visual Question Answering—0
Visual Question Answering as Reading Comprehension—0
Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering—0
Explicit Utilization of General Knowledge in Machine Reading Comprehension—0
Straight to the Facts: Learning Knowledge Base Retrieval for Factual Visual Question Answering—0
Controversy Rules - Discovering Regions Where Classifiers (Dis-)Agree Exceptionally—0
Knowledge Representation and Extraction at Scale—0
Luminoso at SemEval-2018 Task 10: Distinguishing Attributes Using Text Corpora and Relational KnowledgeCode0
Utilisation d'une base de connaissances de sp\'ecialit\'e et de sens commun pour la simplification de comptes-rendus radiologiques (Radiological text simplification using a general knowledge base)—0
Context and Humor: Understanding Amul advertisements of India—0
Show:102550
← PrevPage 15 of 16Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Chinchilla-70B (few-shot, k=5)Accuracy94.3—Unverified
2Gopher-280B (few-shot, k=5)Accuracy93.9—Unverified
3Chinchilla-70B (few-shot, k=5)Accuracy 85.7—Unverified
4Gopher-280B (few-shot, k=5)Accuracy 84.8—Unverified
5Gopher-280B (few-shot, k=5)Accuracy84.2—Unverified
6Gopher-280B (few-shot, k=5)Accuracy 84.1—Unverified
7Gopher-280B (few-shot, k=5)Accuracy 83.9—Unverified
8Gopher-280B (few-shot, k=5)Accuracy83.3—Unverified
9Gopher-280B (few-shot, k=5)Accuracy 81.8—Unverified
10Gopher-280B (few-shot, k=5)Accuracy 81—Unverified