| MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research | Mar 17, 2025 | ArticlesBenchmarking | CodeCode Available | 1 |
| Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos | Mar 17, 2025 | BenchmarkingQuestion Answering | CodeCode Available | 1 |
| VeriContaminated: Assessing LLM-Driven Verilog Coding for Data Contamination | Mar 17, 2025 | BenchmarkingCode Generation | —Unverified | 0 |
| CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era | Mar 16, 2025 | BenchmarkingImage Captioning | —Unverified | 0 |
| Advancing Human-Machine Teaming: Concepts, Challenges, and Applications | Mar 16, 2025 | BenchmarkingDecision Making | —Unverified | 0 |
| Genicious: Contextual Few-shot Prompting for Insights Discovery | Mar 15, 2025 | BenchmarkingDecision Making | —Unverified | 0 |
| Language Models for Automated Classification of Brain MRI Reports and Growth Chart Generation | Mar 15, 2025 | Benchmarking | —Unverified | 0 |
| Dataset Properties Shape the Success of Neuroimaging-Based Patient Stratification: A Benchmarking Analysis Across Clustering Algorithms | Mar 15, 2025 | BenchmarkingBrain Morphometry | —Unverified | 0 |
| Challenges and Advancements in Modeling Shock Fronts with Physics-Informed Neural Networks: A Review and Benchmarking Study | Mar 14, 2025 | Benchmarking | —Unverified | 0 |
| LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama | Mar 14, 2025 | BenchmarkingMMLU | —Unverified | 0 |
| Heterogeneous graph neural networks for species distribution modeling | Mar 14, 2025 | Benchmarking | —Unverified | 0 |
| RESPONSE: Benchmarking the Ability of Language Models to Undertake Commonsense Reasoning in Crisis Situation | Mar 14, 2025 | Benchmarking | —Unverified | 0 |
| InverseBench: Benchmarking Plug-and-Play Diffusion Priors for Inverse Problems in Physical Sciences | Mar 14, 2025 | BenchmarkingImage Restoration | —Unverified | 0 |
| VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity | Mar 14, 2025 | BenchmarkingDecision Making | —Unverified | 0 |
| V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning | Mar 14, 2025 | BenchmarkingRelational Reasoning | —Unverified | 0 |
| A Benchmarking Study of Vision-based Robotic Grasping Algorithms | Mar 14, 2025 | BenchmarkingRobotic Grasping | CodeCode Available | 0 |
| GNNs as Predictors of Agentic Workflow Performances | Mar 14, 2025 | BenchmarkingPosition | CodeCode Available | 1 |
| Enhancing Hand Palm Motion Gesture Recognition by Eliminating Reference Frame Bias via Frame-Invariant Similarity Measures | Mar 14, 2025 | BenchmarkingGesture Recognition | —Unverified | 0 |
| Dynamic Obstacle Avoidance with Bounded Rationality Adversarial Reinforcement Learning | Mar 14, 2025 | BenchmarkingNavigate | —Unverified | 0 |
| DarkBench: Benchmarking Dark Patterns in Large Language Models | Mar 13, 2025 | Benchmarking | —Unverified | 0 |
| VisTai: Benchmarking Vision-Language Models for Traditional Chinese in Taiwan | Mar 13, 2025 | BenchmarkingDialogue Generation | CodeCode Available | 1 |
| TIME: Temporal-sensitive Multi-dimensional Instruction Tuning and Benchmarking for Video-LLMs | Mar 13, 2025 | BenchmarkingQuestion Answering | —Unverified | 0 |
| ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content | Mar 13, 2025 | BenchmarkingImage Generation | —Unverified | 0 |
| CULEMO: Cultural Lenses on Emotion -- Benchmarking LLMs for Cross-Cultural Emotion Understanding | Mar 12, 2025 | BenchmarkingEmotion Recognition | —Unverified | 0 |
| SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models | Mar 12, 2025 | BenchmarkingFairness | —Unverified | 0 |