Instruction Following

Instruction following is the basic task of the model. This task is dedicated to evaluating the ability of the large model to follow human instructions. It is hoped that the model can generate controllable and safe answers.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 101–110 of 1135 papers

Title	Date	Tasks	Status	Hype
Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model	Mar 10, 2025	Image DescriptionImage Generation	CodeCode Available	2
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs	Mar 10, 2025	Code GenerationInstruction Following	CodeCode Available	2
RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs	Mar 8, 2025	Instruction FollowingMathematical Reasoning	CodeCode Available	2
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems	Feb 26, 2025	Instruction Following	CodeCode Available	2
Rank1: Test-Time Compute for Reranking in Information Retrieval	Feb 25, 2025	Information RetrievalInstruction Following	CodeCode Available	2
TESS 2: A Large-Scale Generalist Diffusion Language Model	Feb 19, 2025	Instruction FollowingLanguage Modeling	CodeCode Available	2
mFollowIR: a Multilingual Benchmark for Instruction Following in Retrieval	Jan 31, 2025	Instruction FollowingRetrieval	CodeCode Available	2
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs	Jan 29, 2025	AllInstruction Following	CodeCode Available	2
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate	Jan 29, 2025	Instruction FollowingMath	CodeCode Available	2
Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback	Jan 22, 2025	Instruction Following	CodeCode Available	2

Show:10 25 50

← PrevPage 11 of 114Next →

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	AutoIF (Llama3 70B)	Inst-level loose-accuracy	90.4	—	Unverified
2	AutoIF (Qwen2 72B)	Inst-level loose-accuracy	88	—	Unverified
3	GPT-4	Inst-level loose-accuracy	85.37	—	Unverified
4	PaLM 2 S	Inst-level loose-accuracy	59.11	—	Unverified