SOTAVerified

HumanEval

Papers

Showing 101110 of 264 papers

TitleStatusHype
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM WatermarkingCode1
RES-Q: Evaluating Code-Editing Large Language Model Systems at the Repository ScaleCode1
Concept Distillation from Strong to Weak Models via Hypotheses-to-Theories Prompting0
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Enhancement Protocol0
Addressing Data Leakage in HumanEval Using Combinatorial Test Design0
BASS: Batched Attention-optimized Speculative Sampling0
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models0
AutoTest: Evolutionary Code Solution Selection with Test Cases0
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks0
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models0
Show:102550
← PrevPage 11 of 27Next →

No leaderboard results yet.