| DQI: Measuring Data Quality in NLP | May 2, 2020 | Active LearningBenchmarking | CodeCode Available | 0 |
| ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks | Jan 29, 2024 | BenchmarkingCross-Lingual Transfer | CodeCode Available | 0 |
| Domain-Expanded ASTE: Rethinking Generalization in Aspect Sentiment Triplet Extraction | May 23, 2023 | Aspect-Based Sentiment AnalysisAspect-Based Sentiment Analysis (ABSA) | CodeCode Available | 0 |
| WebSuite: Systematically Evaluating Why Web Agents Fail | Jun 1, 2024 | BenchmarkingDiagnostic | CodeCode Available | 0 |
| Domain2Vec: Domain Embedding for Unsupervised Domain Adaptation | Jul 17, 2020 | BenchmarkingDisentanglement | CodeCode Available | 0 |
| Benchmarking Machine Learning Robustness in Covid-19 Genome Sequence Classification | Jul 18, 2022 | BenchmarkingBIG-bench Machine Learning | CodeCode Available | 0 |
| Do Localization Methods Actually Localize Memorized Data in LLMs? A Tale of Two Benchmarks | Nov 15, 2023 | BenchmarkingNetwork Pruning | CodeCode Available | 0 |
| Do LLMs Memorize Recommendation Datasets? A Preliminary Study on MovieLens-1M | May 15, 2025 | BenchmarkingMemorization | CodeCode Available | 0 |
| Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs | Apr 7, 2025 | BenchmarkingFairness | CodeCode Available | 0 |
| A Review of Testing Object-Based Environment Perception for Safe Automated Driving | Feb 16, 2021 | BenchmarkingSensor Modeling | CodeCode Available | 0 |