Instruction Following

Instruction following is the basic task of the model. This task is dedicated to evaluating the ability of the large model to follow human instructions. It is hoped that the model can generate controllable and safe answers.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 961–970 of 1135 papers

Title	Date	Tasks	Status
Taking Action Towards Graceful Interaction: The Effects of Performing Actions on Modelling Policies for Instruction Clarification Requests	Jan 30, 2024	Instruction Following	CodeCode Available
KAUCUS: Knowledge Augmented User Simulators for Training Language Model Assistants	Jan 29, 2024	DiversityInstruction Following	—Unverified
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents	Jan 23, 2024	Instruction FollowingScene Understanding	—Unverified
COCO is "ALL'' You Need for Visual Instruction Fine-tuning	Jan 17, 2024	AllImage Captioning	—Unverified
PUB: A Pragmatics Understanding Benchmark for Assessing LLMs' Pragmatics Capabilities	Jan 13, 2024	Instruction FollowingMultiple-choice	—Unverified
Human-Instruction-Free LLM Self-Alignment with Limited Samples	Jan 6, 2024	In-Context LearningInstruction Following	—Unverified
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models	Jan 6, 2024	Instruction FollowingMixture-of-Experts	—Unverified
Multilingual Instruction Tuning With Just a Pinch of Multilinguality	Jan 3, 2024	Cross-Lingual TransferInstruction Following	—Unverified
SSP: A Simple and Safe automatic Prompt engineering method towards realistic image synthesis on LVM	Jan 2, 2024	Image GenerationInstruction Following	—Unverified
Generate Subgoal Images before Act: Unlocking the Chain-of-Thought Reasoning in Diffusion Model for Robot Manipulation with Multimodal Prompts	Jan 1, 2024	Image GenerationInstruction Following	—Unverified

Show:10 25 50

← PrevPage 97 of 114Next →

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	AutoIF (Llama3 70B)	Inst-level loose-accuracy	90.4	—	Unverified
2	AutoIF (Qwen2 72B)	Inst-level loose-accuracy	88	—	Unverified
3	GPT-4	Inst-level loose-accuracy	85.37	—	Unverified
4	PaLM 2 S	Inst-level loose-accuracy	59.11	—	Unverified