SOTAVerified

Instruction Following

Instruction following is the basic task of the model. This task is dedicated to evaluating the ability of the large model to follow human instructions. It is hoped that the model can generate controllable and safe answers.

Papers

Showing 751–800 of 1135 papers

TitleStatusHype
Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models—0
Recommendation as Instruction Following: A Large Language Model Empowered Recommendation Approach—0
Audio-Aware Large Language Models as Judges for Speaking Styles—0
Unveiling the Misuse Potential of Base Large Language Models via In-Context Learning—0
ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Format Restriction, and Column Exploration—0
A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens—0
ATEB: Evaluating and Improving Advanced NLP Tasks for Text Embedding Models—0
UrduLLaMA 1.0: Dataset Curation, Preprocessing, and Evaluation in Low-Resource Settings—0
Releasing the CRaQAn (Coreference Resolution in Question-Answering): An open-source dataset and dataset creation methodology using instruction-following models—0
RELIC: Evaluating Compositional Instruction Following via Language Recognition—0
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags—0
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models—0
Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers—0
Retrieval Augmented Chest X-Ray Report Generation using OpenAI GPT models—0
URO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models—0
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References—0
Revisiting the Superficial Alignment Hypothesis—0
Shuttle Between the Instructions and the Parameters of Large Language Models—0
A Systematic Examination of Preference Learning through the Lens of Instruction-Following—0
A Survey of Reinforcement Learning Informed by Natural Language—0
RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being—0
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation—0
Direct Preference Optimization for LLM-Enhanced Recommendation Systems—0
3D-Properties: Identifying Challenges in DPO and Charting a Path Forward—0
RNR: Teaching Large Language Models to Follow Roles and Rules—0
Assessing Robustness to Spurious Correlations in Post-Training Language Models—0
Robust Anti-Backdoor Instruction Tuning in LVLMs—0
Robust Instruction-Following in a Situated Agent via Transfer-Learning from Text—0
Robust Learning of Diverse Code Edits—0
Role-Play Zero-Shot Prompting with Large Language Models for Open-Domain Human-Machine Conversation—0
3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow—0
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization—0
Rethinking the Instruction Quality: LIFT is What You Need—0
VeRA: Vector-based Random Matrix Adaptation—0
Verifiable Format Control for Large Language Model Generations—0
AC/DC: LLM-based Audio Comprehension via Dialogue Continuation—0
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding—0
S2S-Arena, Evaluating Speech2Speech Protocols on Instruction Following with Paralinguistic Information—0
SAG: Style-Aligned Article Generation via Model Collaboration—0
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models—0
SAIL: Search-Augmented Instruction Learning—0
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation—0
SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal Domain—0
Scalable Ensembling For Mitigating Reward Overoptimisation—0
Scalable Vision Language Model Training via High Quality Data Curation—0
ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting—0
Video Instruction Tuning With Synthetic Data—0
Video Unlearning via Low-Rank Refusal Vector—0
Argument Quality Assessment in the Age of Instruction-Following Large Language Models—0
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding—0
Show:102550
← PrevPage 16 of 23Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1AutoIF (Llama3 70B)Inst-level loose-accuracy90.4—Unverified
2AutoIF (Qwen2 72B)Inst-level loose-accuracy88—Unverified
3GPT-4Inst-level loose-accuracy85.37—Unverified
4PaLM 2 SInst-level loose-accuracy59.11—Unverified