| Single and Multi-Hop Question-Answering Datasets for Reticular Chemistry with GPT-4-Turbo | May 3, 2024 | BenchmarkingMulti-hop Question Answering | CodeCode Available | 0 |
| Benchmarking machine learning for bowel sound pattern classification from tabular features to pretrained models | Feb 21, 2025 | BenchmarkingDiagnostic | CodeCode Available | 0 |
| On dataset transferability in medical image classification | Dec 28, 2024 | BenchmarkingClassification | CodeCode Available | 0 |
| Are Synthetic Corruptions A Reliable Proxy For Real-World Corruptions? | May 7, 2025 | BenchmarkingSemantic Segmentation | CodeCode Available | 0 |
| Do LLM Evaluators Prefer Themselves for a Reason? | Apr 4, 2025 | BenchmarkingCode Generation | CodeCode Available | 0 |
| YOLOBench: Benchmarking Efficient Object Detectors on Embedded Systems | Jul 26, 2023 | BenchmarkingCPU | CodeCode Available | 0 |
| Benchmarking Long-tail Generalization with Likelihood Splits | Oct 13, 2022 | BenchmarkingLanguage Modeling | CodeCode Available | 0 |
| UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking | May 21, 2025 | BenchmarkingClaim Verification | CodeCode Available | 0 |
| On Empirical Comparisons of Optimizers for Deep Learning | Oct 11, 2019 | BenchmarkingDeep Learning | CodeCode Available | 0 |
| Benchmarking LLMs' Judgments with No Gold Standard | Nov 11, 2024 | BenchmarkingMachine Translation | CodeCode Available | 0 |