| Model Performance-Guided Evaluation Data Selection for Effective Prompt Optimization | May 15, 2025 | BenchmarkingClustering | —Unverified | 0 |
| GNN-Suite: a Graph Neural Network Benchmarking Framework for Biomedical Informatics | May 15, 2025 | BenchmarkingGraph Neural Network | CodeCode Available | 0 |
| JointDistill: Adaptive Multi-Task Distillation for Joint Depth Estimation and Scene Segmentation | May 15, 2025 | BenchmarkingDepth Estimation | —Unverified | 0 |
| What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs | May 15, 2025 | AllBenchmarking | —Unverified | 0 |
| On the Evaluation of Engineering Artificial General Intelligence | May 15, 2025 | Benchmarking | —Unverified | 0 |
| DIF: A Framework for Benchmarking and Verifying Implicit Bias in LLMs | May 15, 2025 | BenchmarkingFairness | —Unverified | 0 |
| Do LLMs Memorize Recommendation Datasets? A Preliminary Study on MovieLens-1M | May 15, 2025 | BenchmarkingMemorization | CodeCode Available | 0 |
| PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language | May 15, 2025 | BenchmarkingOptical Character Recognition | CodeCode Available | 0 |
| Real-World fNIRS-Based Brain-Computer Interfaces: Benchmarking Deep Learning and Classical Models in Interactive Gaming | May 15, 2025 | BenchmarkingData Augmentation | —Unverified | 0 |
| RobustSpring: Benchmarking Robustness to Image Corruptions for Optical Flow, Scene Flow and Stereo | May 14, 2025 | BenchmarkingOptical Flow Estimation | —Unverified | 0 |