| ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models | Aug 2, 2024 | Code GenerationHumanEval | CodeCode Available | 1 |
| HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation | Dec 30, 2024 | Code GenerationHumanEval | CodeCode Available | 1 |
| HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization | Feb 26, 2024 | Code GenerationHumanEval | CodeCode Available | 1 |
| Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking | May 20, 2025 | HumanEvalmbpp | CodeCode Available | 1 |
| HALO: Hierarchical Autonomous Logic-Oriented Orchestration for Multi-Agent LLM Systems | May 17, 2025 | Arithmetic ReasoningCode Generation | CodeCode Available | 1 |
| Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet' | Oct 29, 2024 | Code CompletionCode Generation | CodeCode Available | 1 |
| How Do Your Code LLMs Perform? Empowering Code Instruction Tuning with High-Quality Data | Sep 5, 2024 | Code GenerationDiversity | CodeCode Available | 1 |
| Generation Meets Verification: Accelerating Large Language Model Inference with Smart Parallel Auto-Correct Decoding | Feb 19, 2024 | HumanEvalLanguage Modeling | CodeCode Available | 1 |
| Getting the most out of your tokenizer for pre-training and domain adaptation | Feb 1, 2024 | Code GenerationDomain Adaptation | CodeCode Available | 1 |
| How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark | Jun 10, 2024 | HumanEvalProgram Synthesis | CodeCode Available | 1 |