| Benchmarking General-Purpose In-Context Learning | May 27, 2024 | BenchmarkingDecision Making | —Unverified | 0 |
| GeneAgent: Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases | May 25, 2024 | BenchmarkingHallucination | —Unverified | 0 |
| BOLD: Boolean Logic Deep Learning | May 25, 2024 | BenchmarkingDeep Learning | —Unverified | 0 |
| NuwaTS: a Foundation Model Mending Every Incomplete Time Series | May 24, 2024 | BenchmarkingContrastive Learning | —Unverified | 0 |
| MCDFN: Supply Chain Demand Forecasting via an Explainable Multi-Channel Data Fusion Network Model | May 24, 2024 | BenchmarkingDemand Forecasting | —Unverified | 0 |
| Application based Evaluation of an Efficient Spike-Encoder, "Spiketrum" | May 24, 2024 | BenchmarkingClassification | —Unverified | 0 |
| Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks | May 24, 2024 | BenchmarkingDecoder | —Unverified | 0 |
| Harnessing Large Language Models for Software Vulnerability Detection: A Comprehensive Benchmarking Study | May 24, 2024 | BenchmarkingVulnerability Detection | —Unverified | 0 |
| Full-stack evaluation of Machine Learning inference workloads for RISC-V systems | May 24, 2024 | BenchmarkingDeep Learning | —Unverified | 0 |
| Benchmarking Hierarchical Image Pyramid Transformer for the classification of colon biopsies and polyps in histopathology images | May 24, 2024 | BenchmarkingClassification | —Unverified | 0 |
| Free Performance Gain from Mixing Multiple Partially Labeled Samples in Multi-label Image Classification | May 24, 2024 | BenchmarkingData Augmentation | —Unverified | 0 |
| A Gap in Time: The Challenge of Processing Heterogeneous IoT Data in Digitalized Buildings | May 23, 2024 | BenchmarkingData Integration | —Unverified | 0 |
| An Empirical Study of Training State-of-the-Art LiDAR Segmentation Models | May 23, 2024 | Autonomous DrivingBenchmarking | —Unverified | 0 |
| CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models | May 22, 2024 | BenchmarkingHallucination | —Unverified | 0 |
| EXACT: Towards a platform for empirically benchmarking Machine Learning model explanation methods | May 20, 2024 | BenchmarkingExplainable artificial intelligence | —Unverified | 0 |
| CT-Eval: Benchmarking Chinese Text-to-Table Performance in Large Language Models | May 20, 2024 | BenchmarkingDiversity | —Unverified | 0 |
| DispaRisk: Auditing Fairness Through Usable Information | May 20, 2024 | BenchmarkingBias Detection | CodeCode Available | 0 |
| EnviroExam: Benchmarking Environmental Science Knowledge of Large Language Models | May 18, 2024 | BenchmarkingSpecificity | —Unverified | 0 |
| From Generalist to Specialist: Improving Large Language Models for Medical Physics Using ARCoT | May 17, 2024 | BenchmarkingMultiple-choice | —Unverified | 0 |
| SMP Challenge: An Overview and Analysis of Social Media Prediction Challenge | May 17, 2024 | BenchmarkingSocial Media Popularity Prediction | —Unverified | 0 |
| BraTS-Path Challenge: Assessing Heterogeneous Histopathologic Brain Tumor Sub-regions | May 17, 2024 | BenchmarkingPrognosis | —Unverified | 0 |
| An Integrated Framework for Multi-Granular Explanation of Video Summarization | May 16, 2024 | BenchmarkingPanoptic Segmentation | CodeCode Available | 0 |
| Simulation-Based Benchmarking of Reinforcement Learning Agents for Personalized Retail Promotions | May 16, 2024 | BenchmarkingReinforcement Learning (RL) | CodeCode Available | 0 |
| A Robust Autoencoder Ensemble-Based Approach for Anomaly Detection in Text | May 16, 2024 | Anomaly DetectionBenchmarking | —Unverified | 0 |
| SpeechVerse: A Large-scale Generalizable Audio Language Model | May 14, 2024 | Automatic Speech RecognitionBenchmarking | —Unverified | 0 |