SOTAVerified

mbpp

Papers

Showing 76100 of 129 papers

TitleStatusHype
Multiple-Choice Questions are Efficient and Robust LLM EvaluatorsCode1
MHPP: Exploring the Capabilities and Limitations of Language Models Beyond Basic Code GenerationCode1
MapCoder: Multi-Agent Code Generation for Competitive Problem SolvingCode2
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User PromptsCode2
Better & Faster Large Language Models via Multi-token PredictionCode1
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-ExpertsCode1
NExT: Teaching Large Language Models to Reason about Code Execution0
Comments as Natural Logic Pivots: Improve Code Generation via Comment PerspectiveCode0
CYCLE: Learning to Self-Refine the Code GenerationCode1
SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents0
Software Vulnerability and Functionality Assessment using LLMs0
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code0
InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language ModelsCode1
Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-stepCode4
OpenCodeInterpreter: Integrating Code Generation with Execution and RefinementCode5
Test-Driven Development for Code Generation0
DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction TuningCode1
Unsupervised Evaluation of Code LLMs with Round-Trip CorrectnessCode1
Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision0
Getting the most out of your tokenizer for pre-training and domain adaptationCode1
OOP: Object-Oriented Programming Evaluation Benchmark for Large Language ModelsCode1
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs0
Instruction Fusion: Advancing Prompt Evolution through HybridizationCode0
AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and OptimisationCode2
ComplexityNet: Increasing LLM Inference Efficiency by Learning Task Complexity0
Show:102550
← PrevPage 4 of 6Next →

No leaderboard results yet.