Instruction Following

Instruction following is the basic task of the model. This task is dedicated to evaluating the ability of the large model to follow human instructions. It is hoped that the model can generate controllable and safe answers.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 951–975 of 1135 papers

Title	Date	Tasks	Status
Multi-Query Focused Disaster Summarization via Instruction-Based Prompting	Feb 14, 2024	Instruction FollowingLanguage Modeling	—Unverified
Policy Improvement using Language Feedback Models	Feb 12, 2024	Behavioural cloningImitation Learning	CodeCode Available
Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation	Feb 12, 2024	Instruction FollowingText to SQL	—Unverified
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs	Feb 12, 2024	Instruction FollowingLogical Reasoning	—Unverified
Nevermind: Instruction Override and Moderation in Large Language Models	Feb 5, 2024	Instruction FollowingLanguage Modeling	—Unverified
Vision-Language Models Provide Promptable Representations for Reinforcement Learning	Feb 5, 2024	Common Sense ReasoningInstruction Following	—Unverified
Diversity Measurement and Subset Selection for Instruction Tuning Datasets	Feb 4, 2024	DiversityInstruction Following	—Unverified
IndiVec: An Exploration of Leveraging Large Language Models for Media Bias Detection with Fine-Grained Bias Indicators	Feb 1, 2024	Bias DetectionInstruction Following	CodeCode Available
Instruction Makes a Difference	Feb 1, 2024	HallucinationInstruction Following	CodeCode Available
Mitigating the Influence of Distractor Tasks in LMs with Prior-Aware Decoding	Jan 31, 2024	Instruction Following	—Unverified
Taking Action Towards Graceful Interaction: The Effects of Performing Actions on Modelling Policies for Instruction Clarification Requests	Jan 30, 2024	Instruction Following	CodeCode Available
KAUCUS: Knowledge Augmented User Simulators for Training Language Model Assistants	Jan 29, 2024	DiversityInstruction Following	—Unverified
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents	Jan 23, 2024	Instruction FollowingScene Understanding	—Unverified
COCO is "ALL'' You Need for Visual Instruction Fine-tuning	Jan 17, 2024	AllImage Captioning	—Unverified
PUB: A Pragmatics Understanding Benchmark for Assessing LLMs' Pragmatics Capabilities	Jan 13, 2024	Instruction FollowingMultiple-choice	—Unverified
Human-Instruction-Free LLM Self-Alignment with Limited Samples	Jan 6, 2024	In-Context LearningInstruction Following	—Unverified
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models	Jan 6, 2024	Instruction FollowingMixture-of-Experts	—Unverified
Multilingual Instruction Tuning With Just a Pinch of Multilinguality	Jan 3, 2024	Cross-Lingual TransferInstruction Following	—Unverified
SSP: A Simple and Safe automatic Prompt engineering method towards realistic image synthesis on LVM	Jan 2, 2024	Image GenerationInstruction Following	—Unverified
Generate Subgoal Images before Act: Unlocking the Chain-of-Thought Reasoning in Diffusion Model for Robot Manipulation with Multimodal Prompts	Jan 1, 2024	Image GenerationInstruction Following	—Unverified
Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision Language Audio and Action	Jan 1, 2024	Image GenerationInstruction Following	—Unverified
Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey	Dec 27, 2023	Instruction FollowingSurvey	—Unverified
LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding	Dec 21, 2023	Instruction FollowingLanguage Modeling	—Unverified
Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning	Dec 19, 2023	DiversityInstruction Following	—Unverified
Rethinking the Instruction Quality: LIFT is What You Need	Dec 12, 2023	Code GenerationInstruction Following	—Unverified

Show:10 25 50

← PrevPage 39 of 46Next →

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	AutoIF (Llama3 70B)	Inst-level loose-accuracy	90.4	—	Unverified
2	AutoIF (Qwen2 72B)	Inst-level loose-accuracy	88	—	Unverified
3	GPT-4	Inst-level loose-accuracy	85.37	—	Unverified
4	PaLM 2 S	Inst-level loose-accuracy	59.11	—	Unverified