| SimBank: from Simulation to Solution in Prescriptive Process Monitoring | Mar 28, 2025 | Benchmarking | —Unverified | 0 |
| Generalization Bias in Large Language Model Summarization of Scientific Research | Mar 28, 2025 | BenchmarkingLanguage Modeling | —Unverified | 0 |
| EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos | Mar 28, 2025 | BenchmarkingQuestion Answering | CodeCode Available | 1 |
| Why Stop at One Error? Benchmarking LLMs as Data Science Code Debuggers for Multi-Hop and Multi-Bug Errors | Mar 28, 2025 | BenchmarkingCode Generation | CodeCode Available | 0 |
| Benchmarking Ultra-Low-Power μNPUs | Mar 28, 2025 | Benchmarking | —Unverified | 0 |
| An Advanced Ensemble Deep Learning Framework for Stock Price Prediction Using VAE, Transformer, and LSTM Model | Mar 28, 2025 | Algorithmic TradingBenchmarking | —Unverified | 0 |
| LIM: Large Interpolator Model for Dynamic Reconstruction | Mar 28, 2025 | 4D reconstructionBenchmarking | —Unverified | 0 |
| Assessing Foundation Models for Sea Ice Type Segmentation in Sentinel-1 SAR Imagery | Mar 28, 2025 | BenchmarkingSegmentation | —Unverified | 0 |
| Benchmarking Deep Learning-Based Methods for Irradiance Nowcasting with Sky Images | Mar 27, 2025 | Benchmarking | —Unverified | 0 |
| CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers? | Mar 27, 2025 | BenchmarkingSpecificity | CodeCode Available | 0 |
| Evaluating Text-to-Image Synthesis with a Conditional Fréchet Distance | Mar 27, 2025 | BenchmarkingImage Generation | —Unverified | 0 |
| ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition | Mar 27, 2025 | Benchmarkingscientific discovery | —Unverified | 0 |
| GateLens: A Reasoning-Enhanced LLM Agent for Automotive Software Release Analytics | Mar 27, 2025 | BenchmarkingNatural Language Queries | —Unverified | 0 |
| FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs | Mar 27, 2025 | AttributeBenchmarking | CodeCode Available | 1 |
| A Comprehensive Benchmark for RNA 3D Structure-Function Modeling | Mar 27, 2025 | BenchmarkingDeep Learning | CodeCode Available | 1 |
| CSPO: Cross-Market Synergistic Stock Price Movement Forecasting with Pseudo-volatility Optimization | Mar 26, 2025 | Benchmarking | —Unverified | 0 |
| Can geometric combinatorics improve RNA branching predictions? | Mar 26, 2025 | Benchmarking | CodeCode Available | 0 |
| RxRx3-core: Benchmarking drug-target interactions in High-Content Microscopy | Mar 26, 2025 | BenchmarkingRepresentation Learning | —Unverified | 0 |
| StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs | Mar 26, 2025 | Benchmarking | CodeCode Available | 3 |
| Benchmarking and optimizing organism wide single-cell RNA alignment methods | Mar 26, 2025 | BenchmarkingDecoder | CodeCode Available | 0 |
| TerraTorch: The Geospatial Foundation Models Toolkit | Mar 26, 2025 | BenchmarkingDecoder | CodeCode Available | 4 |
| Benchmarking Machine Learning Methods for Distributed Acoustic Sensing | Mar 26, 2025 | BenchmarkingData Augmentation | —Unverified | 0 |
| Reservoir Computing with a Single Oscillating Gas Bubble: Emphasizing the Chaotic Regime | Mar 25, 2025 | BenchmarkingLearning Theory | —Unverified | 0 |
| Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy | Mar 25, 2025 | Benchmarkingspeech-recognition | —Unverified | 0 |
| Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models | Mar 25, 2025 | BenchmarkingImage Captioning | CodeCode Available | 1 |