General Knowledge

This task aims to evaluate the ability of a model to answer general-knowledge questions.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 51–100 of 399 papers

Title	Date	Tasks	Status	Hype
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding	Feb 14, 2025	General KnowledgeQuestion Answering	—Unverified	0
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuray	Feb 7, 2025	4kGeneral Knowledge	CodeCode Available	3
PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models	Feb 3, 2025	General Knowledge	—Unverified	0
FlexiCrackNet: A Flexible Pipeline for Enhanced Crack Segmentation with General Features Transfered from SAM	Jan 31, 2025	Computational EfficiencyCrack Segmentation	—Unverified	0
Enabling Autonomic Microservice Management through Self-Learning Agents	Jan 31, 2025	General KnowledgeManagement	—Unverified	0
CALM: Unleashing the Cross-Lingual Self-Aligning Ability of Language Model Question Answering	Jan 30, 2025	General KnowledgeLanguage Modeling	—Unverified	0
Sample-Efficient Behavior Cloning Using General Domain Knowledge	Jan 27, 2025	Car RacingFeature Engineering	—Unverified	0
DAGPrompT: Pushing the Limits of Graph Prompting with a Distribution-aware Graph Prompt Tuning Approach	Jan 25, 2025	General KnowledgeGraph Classification	CodeCode Available	0
Pilot: Building the Federated Multimodal Instruction Tuning Framework	Jan 23, 2025	General Knowledge	—Unverified	0
The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities	Jan 23, 2025	General KnowledgeInstruction Following	CodeCode Available	3
How to Complete Domain Tuning while Keeping General Ability in LLM: Adaptive Layer-wise and Element-wise Regularization	Jan 23, 2025	General Knowledge	—Unverified	0
Generating Diverse Q&A Benchmarks for RAG Evaluation with DataMorgana	Jan 22, 2025	General KnowledgeRAG	—Unverified	0
LLM4WM: Adapting LLM for Wireless Multi-Tasking	Jan 22, 2025	General KnowledgeLanguage Modeling	—Unverified	0
Comparative Insights from 12 Machine Learning Models in Extracting Economic Ideology from Political Text	Jan 16, 2025	General Knowledge	—Unverified	0
Super-class guided Transformer for Zero-Shot Attribute Classification	Jan 10, 2025	AttributeClassification	CodeCode Available	1
Collective inference of the truth of propositions from crowd probability judgments	Jan 9, 2025	General Knowledge	—Unverified	0
Advancing Retrieval-Augmented Generation for Persian: Development of Language Models, Comprehensive Benchmarks, and Best Practices for Optimization	Jan 8, 2025	BenchmarkingGeneral Knowledge	—Unverified	0
KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration	Jan 7, 2025	Anomaly DetectionAnomaly Segmentation	—Unverified	0
Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives	Jan 7, 2025	Autonomous DrivingGeneral Knowledge	CodeCode Available	5
The Scaling Law for LoRA Base on Mutual Information Upper Bound	Jan 6, 2025	General Knowledge	—Unverified	0
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning	Jan 3, 2025	DiagnosticGeneral Knowledge	—Unverified	0
KnowRA: Knowledge Retrieval Augmented Method for Document-level Relation Extraction with Comprehensive Reasoning Abilities	Dec 31, 2024	Common Sense ReasoningDocument-level Relation Extraction	—Unverified	0
RAG with Differential Privacy	Dec 26, 2024	General KnowledgeRAG	CodeCode Available	1
scReader: Prompting Large Language Models to Interpret scRNA-seq Data	Dec 24, 2024	General Knowledge	—Unverified	0
Survey on Abstractive Text Summarization: Dataset, Models, and Metrics	Dec 22, 2024	Abstractive Text SummarizationGeneral Knowledge	CodeCode Available	0
Extending TWIG: Zero-Shot Predictive Hyperparameter Selection for KGEs based on Graph Structure	Dec 19, 2024	General KnowledgeKnowledge Graph Embeddings	—Unverified	0
Are Longer Prompts Always Better? Prompt Selection in Large Language Models for Recommendation Systems	Dec 19, 2024	General KnowledgeRecommendation Systems	—Unverified	0
LLM-RG4: Flexible and Factual Radiology Report Generation across Diverse Input Contexts	Dec 16, 2024	General KnowledgeInstruction Following	CodeCode Available	2
What Makes Cryptic Crosswords Challenging for LLMs?	Dec 12, 2024	General Knowledge	CodeCode Available	0
MoSLD: An Extremely Parameter-Efficient Mixture-of-Shared LoRAs for Multi-Task Learning	Dec 12, 2024	Domain GeneralizationGeneral Knowledge	—Unverified	0
TRIM: Token Reduction and Inference Modeling for Cost-Effective Language Generation	Dec 10, 2024	General KnowledgeText Generation	—Unverified	0
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts	Dec 7, 2024	General KnowledgeMixture-of-Experts	CodeCode Available	1
Adapter-based Approaches to Knowledge-enhanced Language Models -- A Survey	Nov 25, 2024	General KnowledgeKnowledge Graphs	—Unverified	0
GOT4Rec: Graph of Thoughts for Sequential Recommendation	Nov 22, 2024	General KnowledgeSequential Recommendation	—Unverified	0
GRL-Prompt: Towards Knowledge Graph based Prompt Optimization via Reinforcement Learning	Nov 19, 2024	General KnowledgePrompt Engineering	—Unverified	0
Efficient Transfer Learning for Video-language Foundation Models	Nov 18, 2024	Action RecognitionFew-Shot Learning	CodeCode Available	0
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs	Nov 14, 2024	General KnowledgeMath	CodeCode Available	0
Exploring Zero-Shot Anomaly Detection with CLIP in Medical Imaging: Are We There Yet?	Nov 14, 2024	Anomaly DetectionGeneral Knowledge	—Unverified	0
SHARP: Unlocking Interactive Hallucination via Stance Transfer in Role-Playing Agents	Nov 12, 2024	General KnowledgeHallucination	—Unverified	0
Extracting Unlearned Information from LLMs with Activation Steering	Nov 4, 2024	General KnowledgeInformation Retrieval	—Unverified	0
SAFE: Slow and Fast Parameter-Efficient Tuning for Continual Learning with Pre-Trained Models	Nov 4, 2024	Continual LearningGeneral Knowledge	CodeCode Available	1
Evaluating Company-specific Biases in Financial Sentiment Analysis using Large Language Models	Nov 1, 2024	General KnowledgeSentiment Analysis	—Unverified	0
A Comparison of Prompt Engineering Techniques for Task Planning and Execution in Service Robotics	Oct 30, 2024	General KnowledgePrompt Engineering	CodeCode Available	0
AdaptGCD: Multi-Expert Adapter Tuning for Generalized Category Discovery	Oct 29, 2024	General KnowledgePrompt Learning	—Unverified	0
Point-PRC: A Prompt Learning Based Regulation Framework for Generalizable Point Cloud Analysis	Oct 27, 2024	Domain GeneralizationGeneral Knowledge	CodeCode Available	1
Bridge-Coder: Unlocking LLMs' Potential to Overcome Language Gaps in Low-Resource Code	Oct 24, 2024	General KnowledgeIn-Context Learning	—Unverified	0
Should We Really Edit Language Models? On the Evaluation of Edited Language Models	Oct 24, 2024	General KnowledgeModel Editing	CodeCode Available	0
Fast constrained sampling in pre-trained diffusion models	Oct 24, 2024	General Knowledge	—Unverified	0
VoiceBench: Benchmarking LLM-Based Voice Assistants	Oct 22, 2024	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	3
Boosting LLM Translation Skills without General Ability Loss via Rationale Distillation	Oct 17, 2024	General KnowledgeInstruction Following	—Unverified	0

Show:10 25 50

← PrevPage 2 of 8Next →

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	Chinchilla-70B (few-shot, k=5)	Accuracy	94.3	—	Unverified
2	Gopher-280B (few-shot, k=5)	Accuracy	93.9	—	Unverified
3	Chinchilla-70B (few-shot, k=5)	Accuracy	85.7	—	Unverified
4	Gopher-280B (few-shot, k=5)	Accuracy	84.8	—	Unverified
5	Gopher-280B (few-shot, k=5)	Accuracy	84.2	—	Unverified
6	Gopher-280B (few-shot, k=5)	Accuracy	84.1	—	Unverified
7	Gopher-280B (few-shot, k=5)	Accuracy	83.9	—	Unverified
8	Gopher-280B (few-shot, k=5)	Accuracy	83.3	—	Unverified
9	Gopher-280B (few-shot, k=5)	Accuracy	81.8	—	Unverified
10	Gopher-280B (few-shot, k=5)	Accuracy	81	—	Unverified