Instruction Following

Instruction following is the basic task of the model. This task is dedicated to evaluating the ability of the large model to follow human instructions. It is hoped that the model can generate controllable and safe answers.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 201–210 of 1135 papers

Title	Date	Tasks	Status	Hype
Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models	Feb 3, 2024	Instruction FollowingSafety Alignment	CodeCode Available	2
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding	Aug 28, 2024	Instruction Followingscientific discovery	CodeCode Available	2
BLSP-Emo: Towards Empathetic Large Speech-Language Models	Jun 6, 2024	Emotion RecognitionInstruction Following	CodeCode Available	2
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning	Feb 18, 2024	HallucinationInstruction Following	CodeCode Available	2
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts	May 24, 2023	In-Context LearningInstruction Following	CodeCode Available	2
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate	Jan 29, 2025	Instruction FollowingMath	CodeCode Available	2
EditWorld: Simulating World Dynamics for Instruction-Following Image Editing	May 23, 2024	Instruction Following	CodeCode Available	2
Lion: Adversarial Distillation of Proprietary Large Language Models	May 22, 2023	Instruction FollowingKnowledge Distillation	CodeCode Available	2
MAPLM: A Real-World Large-Scale Vision-Language Benchmark for Map and Traffic Scene Understanding	Jan 1, 2024	Autonomous DrivingInstruction Following	CodeCode Available	2
Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning	Oct 1, 2023	In-Context LearningInstruction Following	CodeCode Available	1

Show:10 25 50

← PrevPage 21 of 114Next →

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	AutoIF (Llama3 70B)	Inst-level loose-accuracy	90.4	—	Unverified
2	AutoIF (Qwen2 72B)	Inst-level loose-accuracy	88	—	Unverified
3	GPT-4	Inst-level loose-accuracy	85.37	—	Unverified
4	PaLM 2 S	Inst-level loose-accuracy	59.11	—	Unverified