| CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models | Feb 23, 2025 | Code GenerationHumanEval | CodeCode Available | 1 | 5 |
| CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules | Oct 13, 2023 | Code GenerationHumanEval | CodeCode Available | 1 | 5 |
| Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models | Feb 24, 2024 | HumanEvalMemorization | CodeCode Available | 1 | 5 |
| ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation | Aug 3, 2023 | Class-level Code GenerationCode Generation | CodeCode Available | 1 | 5 |
| Fault-Aware Neural Code Rankers | Jun 4, 2022 | Code GenerationHumanEval | CodeCode Available | 1 | 5 |
| ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models | Aug 2, 2024 | Code GenerationHumanEval | CodeCode Available | 1 | 5 |
| Generation Meets Verification: Accelerating Large Language Model Inference with Smart Parallel Auto-Correct Decoding | Feb 19, 2024 | HumanEvalLanguage Modeling | CodeCode Available | 1 | 5 |
| Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet' | Oct 29, 2024 | Code CompletionCode Generation | CodeCode Available | 1 | 5 |
| Learning to Generate Unit Tests for Automated Debugging | Feb 3, 2025 | HumanEvalLarge Language Model | CodeCode Available | 1 | 5 |
| DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning | Feb 14, 2024 | Code GenerationHumanEval | CodeCode Available | 1 | 5 |