Instruction Following

Instruction following is the basic task of the model. This task is dedicated to evaluating the ability of the large model to follow human instructions. It is hoped that the model can generate controllable and safe answers.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 101–125 of 1135 papers

Title	Date	Tasks	Status	Hype
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs	Mar 10, 2025	Code GenerationInstruction Following	CodeCode Available	2
Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model	Mar 10, 2025	Image DescriptionImage Generation	CodeCode Available	2
RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs	Mar 8, 2025	Instruction FollowingMathematical Reasoning	CodeCode Available	2
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems	Feb 26, 2025	Instruction Following	CodeCode Available	2
Rank1: Test-Time Compute for Reranking in Information Retrieval	Feb 25, 2025	Information RetrievalInstruction Following	CodeCode Available	2
TESS 2: A Large-Scale Generalist Diffusion Language Model	Feb 19, 2025	Instruction FollowingLanguage Modeling	CodeCode Available	2
mFollowIR: a Multilingual Benchmark for Instruction Following in Retrieval	Jan 31, 2025	Instruction FollowingRetrieval	CodeCode Available	2
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs	Jan 29, 2025	AllInstruction Following	CodeCode Available	2
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate	Jan 29, 2025	Instruction FollowingMath	CodeCode Available	2
Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback	Jan 22, 2025	Instruction Following	CodeCode Available	2
LLM-RG4: Flexible and Factual Radiology Report Generation across Diverse Input Contexts	Dec 16, 2024	General KnowledgeInstruction Following	CodeCode Available	2
GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding	Nov 16, 2024	Instruction FollowingLanguage Modeling	CodeCode Available	2
LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation	Nov 14, 2024	Earth ObservationInstruction Following	CodeCode Available	2
Open6DOR: Benchmarking Open-instruction 6-DoF Object Rearrangement and A VLM-based Approach	Oct 24, 2024	BenchmarkingInstruction Following	CodeCode Available	2
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models	Oct 23, 2024	Instruction FollowingLanguage Modelling	CodeCode Available	2
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following	Oct 21, 2024	BenchmarkingInstruction Following	CodeCode Available	2
Toward General Instruction-Following Alignment for Retrieval-Augmented Generation	Oct 12, 2024	Instruction FollowingRAG	CodeCode Available	2
TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data	Oct 8, 2024	Change DetectionEarth Observation	CodeCode Available	2
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data	Sep 30, 2024	Instruction FollowingLanguage Modeling	CodeCode Available	2
Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning	Sep 30, 2024	Instruction FollowingLanguage Modeling	CodeCode Available	2
OmniBench: Towards The Future of Universal Omni-Language Models	Sep 23, 2024	Instruction Following	CodeCode Available	2
Archon: An Architecture Search Framework for Inference-Time Techniques	Sep 23, 2024	Hyperparameter OptimizationInstruction Following	CodeCode Available	2
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding	Aug 28, 2024	Instruction Followingscientific discovery	CodeCode Available	2
Autonomous Improvement of Instruction Following Skills via Foundation Models	Jul 30, 2024	Image GenerationInstruction Following	CodeCode Available	2
SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages	Jul 29, 2024	DiversityInstruction Following	CodeCode Available	2

Show:10 25 50

← PrevPage 5 of 46Next →

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	AutoIF (Llama3 70B)	Inst-level loose-accuracy	90.4	—	Unverified
2	AutoIF (Qwen2 72B)	Inst-level loose-accuracy	88	—	Unverified
3	GPT-4	Inst-level loose-accuracy	85.37	—	Unverified
4	PaLM 2 S	Inst-level loose-accuracy	59.11	—	Unverified