SOTAVerified

General Knowledge

This task aims to evaluate the ability of a model to answer general-knowledge questions.

Source: BIG-bench

Papers

Showing 51100 of 399 papers

TitleStatusHype
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding0
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context AccurayCode3
PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models0
FlexiCrackNet: A Flexible Pipeline for Enhanced Crack Segmentation with General Features Transfered from SAM0
Enabling Autonomic Microservice Management through Self-Learning Agents0
CALM: Unleashing the Cross-Lingual Self-Aligning Ability of Language Model Question Answering0
Sample-Efficient Behavior Cloning Using General Domain Knowledge0
DAGPrompT: Pushing the Limits of Graph Prompting with a Distribution-aware Graph Prompt Tuning ApproachCode0
Pilot: Building the Federated Multimodal Instruction Tuning Framework0
The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling CapabilitiesCode3
How to Complete Domain Tuning while Keeping General Ability in LLM: Adaptive Layer-wise and Element-wise Regularization0
Generating Diverse Q&A Benchmarks for RAG Evaluation with DataMorgana0
LLM4WM: Adapting LLM for Wireless Multi-Tasking0
Comparative Insights from 12 Machine Learning Models in Extracting Economic Ideology from Political Text0
Super-class guided Transformer for Zero-Shot Attribute ClassificationCode1
Collective inference of the truth of propositions from crowd probability judgments0
Advancing Retrieval-Augmented Generation for Persian: Development of Language Models, Comprehensive Benchmarks, and Best Practices for Optimization0
KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration0
Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric PerspectivesCode5
The Scaling Law for LoRA Base on Mutual Information Upper Bound0
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning0
KnowRA: Knowledge Retrieval Augmented Method for Document-level Relation Extraction with Comprehensive Reasoning Abilities0
RAG with Differential PrivacyCode1
scReader: Prompting Large Language Models to Interpret scRNA-seq Data0
Survey on Abstractive Text Summarization: Dataset, Models, and MetricsCode0
Extending TWIG: Zero-Shot Predictive Hyperparameter Selection for KGEs based on Graph Structure0
Are Longer Prompts Always Better? Prompt Selection in Large Language Models for Recommendation Systems0
LLM-RG4: Flexible and Factual Radiology Report Generation across Diverse Input ContextsCode2
What Makes Cryptic Crosswords Challenging for LLMs?Code0
MoSLD: An Extremely Parameter-Efficient Mixture-of-Shared LoRAs for Multi-Task Learning0
TRIM: Token Reduction and Inference Modeling for Cost-Effective Language Generation0
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of ExpertsCode1
Adapter-based Approaches to Knowledge-enhanced Language Models -- A Survey0
GOT4Rec: Graph of Thoughts for Sequential Recommendation0
GRL-Prompt: Towards Knowledge Graph based Prompt Optimization via Reinforcement Learning0
Efficient Transfer Learning for Video-language Foundation ModelsCode0
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMsCode0
Exploring Zero-Shot Anomaly Detection with CLIP in Medical Imaging: Are We There Yet?0
SHARP: Unlocking Interactive Hallucination via Stance Transfer in Role-Playing Agents0
Extracting Unlearned Information from LLMs with Activation Steering0
SAFE: Slow and Fast Parameter-Efficient Tuning for Continual Learning with Pre-Trained ModelsCode1
Evaluating Company-specific Biases in Financial Sentiment Analysis using Large Language Models0
A Comparison of Prompt Engineering Techniques for Task Planning and Execution in Service RoboticsCode0
AdaptGCD: Multi-Expert Adapter Tuning for Generalized Category Discovery0
Point-PRC: A Prompt Learning Based Regulation Framework for Generalizable Point Cloud AnalysisCode1
Bridge-Coder: Unlocking LLMs' Potential to Overcome Language Gaps in Low-Resource Code0
Should We Really Edit Language Models? On the Evaluation of Edited Language ModelsCode0
Fast constrained sampling in pre-trained diffusion models0
VoiceBench: Benchmarking LLM-Based Voice AssistantsCode3
Boosting LLM Translation Skills without General Ability Loss via Rationale Distillation0
Show:102550
← PrevPage 2 of 8Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Chinchilla-70B (few-shot, k=5)Accuracy94.3Unverified
2Gopher-280B (few-shot, k=5)Accuracy93.9Unverified
3Chinchilla-70B (few-shot, k=5)Accuracy 85.7Unverified
4Gopher-280B (few-shot, k=5)Accuracy 84.8Unverified
5Gopher-280B (few-shot, k=5)Accuracy84.2Unverified
6Gopher-280B (few-shot, k=5)Accuracy 84.1Unverified
7Gopher-280B (few-shot, k=5)Accuracy 83.9Unverified
8Gopher-280B (few-shot, k=5)Accuracy83.3Unverified
9Gopher-280B (few-shot, k=5)Accuracy 81.8Unverified
10Gopher-280B (few-shot, k=5)Accuracy 81Unverified