SOTAVerified

StrategyQA

StrategyQA aims to measure the ability of models to answer questions that require multi-step implicit reasoning.

Source: BIG-bench

Papers

Showing 1–40 of 40 papers

TitleStatusHype
Fusing Bidirectional Chains of Thought and Reward Mechanisms A Method for Enhancing Question-Answering Capabilities of Large Language Models for Chinese Intangible Cultural Heritage—0
Rule-Guided Feedback: Enhancing Reasoning by Enforcing Rule Adherence in Large Language Models—0
DeLTa: A Decoding Strategy based on Logit Trajectory Prediction Improves Factuality and Reasoning AbilityCode0
Voting or Consensus? Decision-Making in Multi-Agent DebateCode0
Unraveling Indirect In-Context Learning Using Influence Functions—0
AutoReason: Automatic Few-Shot Reasoning DecompositionCode1
Dialectical Behavior Therapy Approach to LLM Prompting—0
Rationale-Aware Answer Verification by Pairwise Self-EvaluationCode0
A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions—0
Proof of Thought : Neurosymbolic Program Synthesis allows Robust and Interpretable Reasoning—0
Mutual Reasoning Makes Smaller LLMs Stronger Problem-SolversCode4
Meta-prompting Optimized Retrieval-augmented Generation—0
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks—0
Advancing Process Verification for Large Language Models via Tree-Based Preference Learning—0
Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-ContrastCode1
Improving Attributed Text Generation of Large Language Models via Preference Learning—0
CR-LT-KGQA: A Knowledge Graph Question Answering Dataset Requiring Commonsense Reasoning and Long-Tail KnowledgeCode1
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and DistillationCode1
Towards Uncertainty-Aware Language Agent—0
Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step ReasoningCode1
IAG: Induction-Augmented Generation Framework for Answering Reasoning Questions—0
The ART of LLM Refinement: Ask, Refine, and Trust—0
Tailoring Self-Rationalizers with Multi-Reward DistillationCode0
Improving Planning with Large Language Models: A Modular Agentic ArchitectureCode1
Large Language Models Are Also Good Prototypical Commonsense Reasoners—0
Answering Unseen Questions With Smaller Language Models Using Rationale Generation and Dense Retrieval—0
Teaching Smaller Language Models To Generalise To Unseen Compositional QuestionsCode0
Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive TasksCode1
Deduction under Perturbed Evidence: Probing Student Simulation Capabilities of Large Language Models—0
Hint of Thought prompting: an explainable and zero-shot approach to reasoning tasks with LLMs—0
Self-Evaluation Guided Beam Search for Reasoning—0
Visconde: Multi-document QA with GPT-3 and Neural RerankingCode1
Distilling Reasoning Capabilities into Smaller Language ModelsCode0
Learning to Decompose: Hypothetical Question Decomposition Based on Comparable Texts—0
Better Retrieval May Not Lead to Better Question Answering—0
PaLM: Scaling Language Modeling with PathwaysCode2
Training Compute-Optimal Large Language ModelsCode6
Self-Consistency Improves Chain of Thought Reasoning in Language ModelsCode1
Scaling Language Models: Methods, Analysis & Insights from Training GopherCode2
Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning StrategiesCode1
Show:102550

No leaderboard results yet.