SOTAVerified

Instruction Following

Instruction following is the basic task of the model. This task is dedicated to evaluating the ability of the large model to follow human instructions. It is hoped that the model can generate controllable and safe answers.

Papers

Showing 451–500 of 1135 papers

TitleStatusHype
Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing—0
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling—0
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding—0
AC/DC: LLM-based Audio Comprehension via Dialogue Continuation—0
Learning Human Perception Dynamics for Informative Robot Communication—0
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models—0
Better Instruction-Following Through Minimum Bayes Risk—0
InstructBooth: Instruction-following Personalized Text-to-Image Generation—0
Efficient Prompt Optimization Through the Lens of Best Arm Identification—0
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective—0
InsightEdit: Towards Better Instruction Following for Image Editing—0
Large Language Models for Autonomous Driving (LLM4AD): Concept, Benchmark, Experiments, and Challenges—0
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training—0
Differential Information: An Information-Theoretic Perspective on Preference Optimization—0
Inference-Time Language Model Alignment via Integrated Value Guidance—0
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models—0
DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation—0
A Monte Carlo Language Model Pipeline for Zero-Shot Sociopolitical Event Extraction—0
Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge—0
In-Context Watermarks for Large Language Models—0
In-context Learning vs. Instruction Tuning: The Case of Small and Multilingual Language Models—0
Incentivizing Inclusive Contributions in Model Sharing Markets—0
Improving the Robustness to Variations of Objects and Instructions with a Neuro-Symbolic Approach for Interactive Instruction Following—0
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment—0
Improving the Robustness of Large Language Models via Consistency Alignment—0
Improving Reward Models with Synthetic Critiques—0
Improving Open Information Extraction with Large Language Models: A Study on Demonstration Uncertainty—0
Leveraging LLMs for Influence Path Planning in Proactive Recommendation—0
Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets—0
Improving Instruct Models for Free: A Study on Partial Adaptation—0
Improving Instruction-Following in Language Models through Activation Steering—0
DEM: Distribution Edited Model for Training with Mixed Data Distributions—0
Improving Coherence and Consistency in Neural Sequence Models with Dual-System, Neuro-Symbolic Reasoning—0
Benchmarking and Improving Generator-Validator Consistency of Language Models—0
Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models—0
Adaptive Decoding via Latent Preference Optimization—0
Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course—0
Becoming self-instruct: introducing early stopping criteria for minimal instruct tuning—0
If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs—0
Language Models Benefit from Preparation with Elicited Knowledge—0
DecIF: Improving Instruction-Following through Meta-Decomposition—0
IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval—0
Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning—0
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy—0
Space-LLaVA: a Vision-Language Model Adapted to Extraterrestrial Applications—0
InstructionCP: A fast approach to transfer Large Language Models into target language—0
DataMan: Data Manager for Pre-training Large Language Models—0
Instruction Following by Boosting Attention of Large Language Models—0
AdaGrad under Anisotropic Smoothness—0
ICCO: Learning an Instruction-conditioned Coordinator for Language-guided Task-aligned Multi-robot Control—0
Show:102550
← PrevPage 10 of 23Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1AutoIF (Llama3 70B)Inst-level loose-accuracy90.4—Unverified
2AutoIF (Qwen2 72B)Inst-level loose-accuracy88—Unverified
3GPT-4Inst-level loose-accuracy85.37—Unverified
4PaLM 2 SInst-level loose-accuracy59.11—Unverified