SOTAVerified

Instruction Following

Instruction following is the basic task of the model. This task is dedicated to evaluating the ability of the large model to follow human instructions. It is hoped that the model can generate controllable and safe answers.

Papers

Showing 651–700 of 1135 papers

TitleStatusHype
Modular Networks for Compositional Instruction Following—0
Can Large Language Models Understand Symbolic Graphics Programs?—0
Traffic Sign Interpretation in Real Road Scene—0
CamelEval: Advancing Culturally Aligned Arabic Language Models and Benchmarks—0
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons—0
CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation—0
MrSteve: Instruction-Following Agents in Minecraft with What-Where-When Memory—0
MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following—0
CachePrune: Neural-Based Attribution Defense Against Indirect Prompt Injection Attacks—0
Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks—0
Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard Newspaper—0
Multi-lingual Functional Evaluation for Large Language Models—0
Multilingual Instruction Tuning With Just a Pinch of Multilinguality—0
Multilingual Multimodal Software Developer for Code Generation—0
Bridging Offline and Online Reinforcement Learning for LLMs—0
Multi-Modal Instruction-Tuning Small-Scale Language-and-Vision Assistant for Semiconductor Electron Micrograph Analysis—0
Transformer-based Causal Language Models Perform Clustering—0
Multimodal Sequential Generative Models for Semi-Supervised Language Instruction Following—0
Multimodal Situational Safety—0
Multimodal Web Navigation with Instruction-Finetuned Foundation Models—0
Multi-Query Focused Disaster Summarization via Instruction-Based Prompting—0
Multi-Reward as Condition for Instruction-based Image Editing—0
Multi-Task Instruction Tuning of LLaMa for Specific Scenarios: A Preliminary Study on Writing Assistance—0
Treasure Hunt: Real-time Targeting of the Long Tail using Training-Time Markers—0
Nadine: An LLM-driven Intelligent Social Robot with Affective Capabilities and Human-like Memory—0
Natural Language-conditioned Reinforcement Learning with Inside-out Task Language Development and Translation—0
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs—0
Boosting LLM Translation Skills without General Ability Loss via Rationale Distillation—0
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation—0
Navigating the Alpha Jungle: An LLM-Powered MCTS Framework for Formulaic Factor Mining—0
Neural Semantic Parsing—0
Nevermind: Instruction Override and Moderation in Large Language Models—0
Zero-shot Object-Centric Instruction Following: Integrating Foundation Models with Traditional Navigation—0
Triple Phase Transitions: Understanding the Learning Dynamics of Large Language Models from a Neuroscience Perspective—0
TS-Align: A Teacher-Student Collaborative Framework for Scalable Iterative Finetuning of Large Language Models—0
Nudging: Inference-time Alignment of LLMs via Guided Decoding—0
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation—0
TuneShield: Mitigating Toxicity in Conversational AI while Fine-tuning on Untrusted Data—0
OmniGeo: Towards a Multimodal Large Language Models for Geospatial Artificial Intelligence—0
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents—0
On Instruction-Finetuning Neural Machine Translation Models—0
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training—0
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models—0
UAV-VLN: End-to-End Vision Language guided Navigation for UAVs—0
UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function—0
On the Mechanism of Reasoning Pattern Selection in Reinforcement Learning for Language Models—0
UGIF: UI Grounded Instruction Following—0
BioMistral-NLU: Towards More Generalizable Medical Language Understanding through Instruction Tuning—0
XIFBench: Evaluating Large Language Models on Multilingual Instruction Following—0
OpenSearch-SQL: Enhancing Text-to-SQL with Dynamic Few-shot and Consistency Alignment—0
Show:102550
← PrevPage 14 of 23Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1AutoIF (Llama3 70B)Inst-level loose-accuracy90.4—Unverified
2AutoIF (Qwen2 72B)Inst-level loose-accuracy88—Unverified
3GPT-4Inst-level loose-accuracy85.37—Unverified
4PaLM 2 SInst-level loose-accuracy59.11—Unverified