| QCPINN: Quantum-Classical Physics-Informed Neural Networks for Solving PDEs | Mar 20, 2025 | BenchmarkingPhysics-informed machine learning | CodeCode Available | 1 |
| A Statistical Analysis for Per-Instance Evaluation of Stochastic Optimizers: How Many Repeats Are Enough? | Mar 20, 2025 | Benchmarking | —Unverified | 0 |
| Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models | Mar 20, 2025 | BenchmarkingReinforcement Learning (RL) | CodeCode Available | 4 |
| ECKGBench: Benchmarking Large Language Models in E-commerce Leveraging Knowledge Graph | Mar 20, 2025 | BenchmarkingHallucination | —Unverified | 0 |
| The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination | Mar 20, 2025 | BenchmarkingLarge Language Model | CodeCode Available | 1 |
| DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs | Mar 20, 2025 | BenchmarkingHallucination | —Unverified | 0 |
| Empirical Analysis of Privacy-Fairness-Accuracy Trade-offs in Federated Learning: A Step Towards Responsible AI | Mar 20, 2025 | BenchmarkingFairness | —Unverified | 0 |
| FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding | Mar 19, 2025 | BenchmarkingMultiple-choice | —Unverified | 0 |
| Benchmarking Open-Source Large Language Models on Healthcare Text Classification Tasks | Mar 19, 2025 | BenchmarkingDomain Adaptation | —Unverified | 0 |
| Language-based Image Colorization: A Benchmark and Beyond | Mar 19, 2025 | BenchmarkingColorization | CodeCode Available | 0 |
| Kolmogorov-Arnold Network for Transistor Compact Modeling | Mar 19, 2025 | Benchmarking | —Unverified | 0 |
| Benchmarking Large Language Models for Handwritten Text Recognition | Mar 19, 2025 | BenchmarkingHandwritten Text Recognition | —Unverified | 0 |
| VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning | Mar 19, 2025 | BenchmarkingLanguage Modeling | CodeCode Available | 2 |
| SUM Parts: Benchmarking Part-Level Semantic Segmentation of Urban Meshes | Mar 19, 2025 | 3D Semantic SegmentationBenchmarking | —Unverified | 0 |
| ImputeGAP: A Comprehensive Library for Time Series Imputation | Mar 19, 2025 | BenchmarkingImputation | —Unverified | 0 |
| Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack | Mar 18, 2025 | 8kBenchmarking | —Unverified | 0 |
| COPA: Comparing the Incomparable to Explore the Pareto Front | Mar 18, 2025 | AutoMLBenchmarking | —Unverified | 0 |
| ConSCompF: Consistency-focused Similarity Comparison Framework for Generative Large Language Models | Mar 18, 2025 | BenchmarkingChatbot | —Unverified | 0 |
| JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System | Mar 18, 2025 | BenchmarkingIn-Context Learning | CodeCode Available | 1 |
| Benchmarking Failures in Tool-Augmented Language Models | Mar 18, 2025 | BenchmarkingText Generation | CodeCode Available | 0 |
| HA-VLN: A Benchmark for Human-Aware Navigation in Discrete-Continuous Environments with Dynamic Multi-Human Interactions, Real-World Validation, and an Open Leaderboard | Mar 18, 2025 | BenchmarkingHuman Dynamics | —Unverified | 0 |
| Stable Virtual Camera: Generative View Synthesis with Diffusion Models | Mar 18, 2025 | Benchmarking | —Unverified | 0 |
| Benchmarking community drug response prediction models: datasets, models, tools, and metrics for cross-dataset generalization analysis | Mar 18, 2025 | BenchmarkingDrug Response Prediction | CodeCode Available | 0 |
| Organ-aware Multi-scale Medical Image Segmentation Using Text Prompt Engineering | Mar 18, 2025 | BenchmarkingDescriptive | —Unverified | 0 |
| CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models | Mar 18, 2025 | BenchmarkingSpatial Reasoning | CodeCode Available | 0 |