SOTAVerified

Code Search

The goal of Code Search is to retrieve code fragments from a large code corpus that most closely match a developer’s intent, which is expressed in natural language.

Source: When Deep Learning Met Code Search

Papers

Showing 1–50 of 125 papers

TitleStatusHype
MGS3: A Multi-Granularity Self-Supervised Code Search Framework—0
DeepRTL2: A Versatile Model for RTL-Related Tasks—0
LEANCODE: Understanding Models Better for Code Simplification of Pre-trained Large Language Models—0
Knowledge Graph Based Repository-Level Code Generation—0
Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks—0
Towards Leveraging Large Language Model Summaries for Topic Modeling in Source Code—0
A Study on Mixup-Inspired Augmentation Methods for Software Vulnerability Detection—0
Zero-Shot Cross-Domain Code Search without Fine-TuningCode1
OASIS: Order-Augmented Strategy for Improved Code Search—0
LoRACode: LoRA Adapters for Code Embeddings—0
MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings—0
Beyond Natural Language Perplexity: Detecting Dead Code Poisoning in Code Generation Datasets—0
URECA: The Chain of Two Minimum Set Cover Problems exists behind Adaptation to Shifts in Semantic Code Search—0
Repository-level Code Search with Neural Retrieval MethodsCode0
OrcaLoca: An LLM Agent Framework for Software Issue Localization—0
On the Compression of Language Models for Code: An Empirical Study on CodeBERT—0
Isotropy Matters: Soft-ZCA Whitening of Embeddings for Semantic Code SearchCode0
CodeSAM: Source Code Representation Learning by Infusing Self-Attention with Multi-Code-View GraphsCode2
In-the-loop Hyper-Parameter Optimization for LLM-Based Automated Design of Heuristics—0
Deep Code Search with Naming-Agnostic Contrastive Multi-View Learning—0
ViC: Virtual Compiler Is All You Need For Assembly Code SearchCode1
Natural Language Outlines for Code: Literate Programming in the LLM Era—0
LLM Agents Improve Semantic Code Search—0
SpecRover: Code Intent Extraction via LLMs—0
CoIR: A Comprehensive Benchmark for Code Information Retrieval ModelsCode2
Toward Exploring the Code Understanding Capabilities of Pre-trained Code Generation Models—0
CoSQA+: Pioneering the Multi-Choice Code Search Benchmark with Test-Driven AgentsCode0
RepoQA: Evaluating Long Context Code UnderstandingCode2
Advanced Detection of Source Code Clones via an Ensemble of Unsupervised Similarity MeasuresCode1
AutoCodeRover: Autonomous Program ImprovementCode7
ProCQA: A Large-scale Community-based Programming Question Answering Dataset for Code SearchCode0
Source Code Clone Detection Using Unsupervised Similarity MeasuresCode1
Rewriting the Code: A Simple Method for Large Language Model Augmented Code SearchCode1
Code Search Debiasing:Improve Search Results beyond Overall Ranking Performance—0
GenCodeSearchNet: A Benchmark Test Suite for Evaluating Generalization in Programming Language UnderstandingCode0
TransformCode: A Contrastive Learning Framework for Code Embedding via Subtree TransformationCode0
Noisy Pair Corrector for Dense Retrieval—0
ACES: Generating Diverse Programming Puzzles with with Autotelic Generative Models—0
Language Models are Universal EmbeddersCode1
Rethinking Negative Pairs in Code SearchCode1
Contrastive Prompt Learning-based Code Search based on Interaction Matrix—0
Code Representation Pre-training with Complements from Program Executions—0
Laminar: A New Serverless Stream-based Framework with Semantic Code Search and Code Completion—0
MELT: Mining Effective Lightweight Transformations from Pull RequestsCode0
Evaluating and Optimizing the Effectiveness of Neural Machine Translation in Supporting Code Retrieval Models: A Study on the CAT Benchmark—0
Constructing Multilingual Code Search Dataset Using Neural Machine TranslationCode0
Structure-Aware Language Model Pretraining Improves Dense Retrieval on Structured DataCode1
Backdooring Neural Code SearchCode1
CCT-Code: Cross-Consistency Training for Multilingual Clone Detection and Code Search—0
Searching by Code: a New SearchBySnippet Dataset and SnippeR Retrieval Model for Searching by Code Snippets—0
Show:102550
← PrevPage 1 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1cpt-code MOverall93.5—Unverified
2cpt-code SOverall93.4—Unverified
3CodeT5+ 770MOverall77.4—Unverified
4GraphCodeBERTOverall77.4—Unverified
5CodeT5+ 220MOverall77.1—Unverified
6CodeBERTOverall76—Unverified
#ModelMetricClaimedVerifiedStatus
1Self-attentionTest MRR0.84—Unverified
2NBOWTest MRR0.81—Unverified
3RNNTest MRR0.77—Unverified
#ModelMetricClaimedVerifiedStatus
1CodeT5+ 770MMRR44.7—Unverified
2CodeT5+ 220MMRR43.3—Unverified
3CodeBERTMRR27.19—Unverified
#ModelMetricClaimedVerifiedStatus
1Uni-SBTMRR0.36—Unverified
#ModelMetricClaimedVerifiedStatus
1CodeBERTAccuracy47.8—Unverified
#ModelMetricClaimedVerifiedStatus
1Voyage-code-002nDCG@1056.26—Unverified