| Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation | Sep 23, 2024 | Multiple-choiceQuestion Answering | —Unverified | 0 | 0 |
| Developing A Framework to Support Human Evaluation of Bias in Generated Free Response Text | May 5, 2025 | Multiple-choice | —Unverified | 0 | 0 |
| Development and Evaluation of a Personalized Computer-aided Question Generation for English Learners to Improve Proficiency and Correct Mistakes | Aug 29, 2018 | Multiple-choiceQuestion Generation | —Unverified | 0 | 0 |
| DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response | May 26, 2025 | Multiple-choice | —Unverified | 0 | 0 |
| D-GEN: Automatic Distractor Generation and Evaluation for Reliable Assessment of Generative Model | Apr 18, 2025 | Distractor GenerationMultiple-choice | —Unverified | 0 | 0 |
| DGRC: An Effective Fine-tuning Framework for Distractor Generation in Chinese Multi-choice Reading Comprehension | May 29, 2024 | Distractor GenerationMultiple-choice | —Unverified | 0 | 0 |
| Instructions and Guide for Diagnostic Questions: The NeurIPS 2020 Education Challenge | Jul 23, 2020 | DiagnosticMisconceptions | —Unverified | 0 | 0 |
| Dialogue-Based Simulation For Cultural Awareness Training | Feb 1, 2020 | Automatic Speech RecognitionAutomatic Speech Recognition (ASR) | —Unverified | 0 | 0 |
| Dienstplanerstellung in Krankenhaeusern mittels genetischer Algorithmen | May 30, 2013 | Multiple-choice | —Unverified | 0 | 0 |
| Differentiable Open-Ended Commonsense Reasoning | Oct 24, 2020 | Multiple-choice | —Unverified | 0 | 0 |
| Plug-in, Trainable Gate for Streamlining Arbitrary Neural Networks | Apr 24, 2019 | Efficient Neural Networkimage-classification | —Unverified | 0 | 0 |
| Different Questions, Different Models: Fine-Grained Evaluation of Uncertainty and Calibration in Clinical QA with LLMs | Jun 12, 2025 | Multiple-choiceQuestion Answering | —Unverified | 0 | 0 |
| Digital Comprehensibility Assessment of Simplified Texts among Persons with Intellectual Disabilities | Feb 20, 2024 | Multiple-choiceText Simplification | —Unverified | 0 | 0 |
| Disaggregating Hops: Can We Guide a Multi-Hop Reasoning Language Model to Incrementally Learn at each Hop? | Jan 16, 2022 | Language ModelingLanguage Modelling | —Unverified | 0 | 0 |
| DISTO: Evaluating Textual Distractors for Multi-Choice Questions using Negative Sampling based Approach | Apr 10, 2023 | Distractor GenerationMachine Translation | —Unverified | 0 | 0 |
| Distractor Analysis and Selection for Multiple-Choice Cloze Questions for Second-Language Learners | Jul 1, 2020 | Multiple-choice | —Unverified | 0 | 0 |
| Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and Evaluation | Feb 2, 2024 | Distractor GenerationMultiple-choice | —Unverified | 0 | 0 |
| Distributional semantics beyond words: Supervised learning of analogy and paraphrase | Oct 18, 2013 | Multiple-choiceTask 2 | —Unverified | 0 | 0 |
| DiverseNet: When One Right Answer is not Enough | Aug 24, 2020 | Multiple-choiceStructured Prediction | —Unverified | 0 | 0 |
| DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain | Apr 18, 2025 | Multiple-choice | —Unverified | 0 | 0 |
| Document-level Event Factuality Identification via Machine Reading Comprehension Frameworks with Transfer Learning | Oct 1, 2022 | Data AugmentationMachine Reading Comprehension | —Unverified | 0 | 0 |
| Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla | Jul 18, 2023 | Multiple-choiceQuestion Answering | —Unverified | 0 | 0 |
| Do Fine-tuned Commonsense Language Models Really Generalize? | Nov 18, 2020 | Multiple-choiceQuestion Answering | —Unverified | 0 | 0 |
| Do Large Language Models Know Folktales? A Case Study of Yokai in Japanese Folktales | Jun 4, 2025 | Multiple-choice | —Unverified | 0 | 0 |
| Do LLMs Act as Repositories of Causal Knowledge? | Dec 14, 2024 | Causal InferenceMultiple-choice | —Unverified | 0 | 0 |