| Benchmarking Active Learning Strategies for Materials Optimization and Discovery | Apr 12, 2022 | Active LearningBenchmarking | —Unverified | 0 | 0 |
| A critical analysis of metrics used for measuring progress in artificial intelligence | Aug 6, 2020 | Benchmarking | —Unverified | 0 | 0 |
| True Online TD-Replan(lambda) Achieving Planning through Replaying | Jan 31, 2025 | Benchmarking | —Unverified | 0 | 0 |
| Benchmarking Active Learning for NILM | Nov 24, 2024 | Active LearningBenchmarking | —Unverified | 0 | 0 |
| Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles | Jan 13, 2025 | ArticlesBenchmarking | —Unverified | 0 | 0 |
| Parsing Any Domain English text to CoNLL dependencies | May 1, 2012 | BenchmarkingDependency Parsing | —Unverified | 0 | 0 |
| Trust but Verify: Programmatic VLM Evaluation in the Wild | Oct 17, 2024 | BenchmarkingLanguage Modelling | —Unverified | 0 | 0 |
| Participatory Personalization in Classification | Feb 8, 2023 | BenchmarkingClassification | —Unverified | 0 | 0 |
| 'Part'ly first among equals: Semantic part-based benchmarking for state-of-the-art object recognition systems | Nov 23, 2016 | BenchmarkingObject | —Unverified | 0 | 0 |
| When Safety Detectors Aren't Enough: A Stealthy and Effective Jailbreak Attack on LLMs via Steganographic Techniques | May 22, 2025 | Benchmarking | —Unverified | 0 | 0 |