| Ontology-based knowledge representation for bone disease diagnosis: a foundation for safe and sustainable medical artificial intelligence systems | Jun 5, 2025 | DiagnosticMultimodal Deep Learning | —Unverified | 0 | 0 |
| Assistive Image Annotation Systems with Deep Learning and Natural Language Capabilities: A Review | Jun 28, 2024 | Active LearningImage Captioning | —Unverified | 0 | 0 |
| Towards Models that Can See and Read | Jan 18, 2023 | DecoderImage Captioning | —Unverified | 0 | 0 |
| Active Data Curation Effectively Distills Large-Scale Multimodal Models | Nov 27, 2024 | DecoderImage Captioning | —Unverified | 0 | 0 |
| Assisting Scene Graph Generation with Self-Supervision | Aug 8, 2020 | Graph GenerationImage Captioning | —Unverified | 0 | 0 |
| Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method | May 20, 2025 | HallucinationObject Localization | —Unverified | 0 | 0 |
| Towards Reasoning-Aware Explainable VQA | Nov 9, 2022 | DecoderExplanation Generation | —Unverified | 0 | 0 |
| Assessing the Robustness of Visual Question Answering Models | Nov 30, 2019 | Question AnsweringVisual Question Answering | —Unverified | 0 | 0 |
| Towards Semantic Equivalence of Tokenization in Multimodal LLM | Jun 7, 2024 | Visual Question Answering | —Unverified | 0 | 0 |
| Towards Top-Down Reasoning: An Explainable Multi-Agent Approach for Visual Question Answering | Nov 29, 2023 | Common Sense ReasoningQuestion Answering | —Unverified | 0 | 0 |
| Towards Transparent AI Systems: Interpreting Visual Question Answering Models | Aug 31, 2016 | Question AnsweringVisual Question Answering | —Unverified | 0 | 0 |
| Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers | Jan 3, 2024 | Question AnsweringVisual Grounding | —Unverified | 0 | 0 |
| Assessing Image Quality Issues for Real-World Problems | Mar 27, 2020 | Image CaptioningQuestion Answering | —Unverified | 0 | 0 |
| Towards Unsupervised Visual Reasoning: Do Off-The-Shelf Features Know How to Reason? | Dec 20, 2022 | Question AnsweringRepresentation Learning | —Unverified | 0 | 0 |
| Ask Me Anything: Free-form Visual Question Answering Based on Knowledge from External Sources | Nov 22, 2015 | FormGeneral Knowledge | —Unverified | 0 | 0 |
| Asking questions on handwritten document collections | Oct 2, 2021 | Optical Character Recognition (OCR)Question Answering | —Unverified | 0 | 0 |
| Towards Visual Dialog for Radiology | Jul 1, 2020 | Question AnsweringVisual Dialog | —Unverified | 0 | 0 |
| A Corpus for Visual Question Answering Annotated with Frame Semantic Information | May 1, 2020 | Question AnsweringVisual Question Answering | —Unverified | 0 | 0 |
| Toward Unsupervised Realistic Visual Question Answering | Mar 9, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 | 0 |
| Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images | Feb 23, 2025 | Adversarial AttackQuestion Answering | —Unverified | 0 | 0 |
| VSA4VQA: Scaling a Vector Symbolic Architecture to Visual Question Answering on Natural Images | May 6, 2024 | AttributeLanguage Modeling | —Unverified | 0 | 0 |
| Training Recurrent Answering Units with Joint Loss Minimization for VQA | Jun 12, 2016 | Question AnsweringVisual Question Answering | —Unverified | 0 | 0 |
| Transferable Adversarial Attacks on Black-Box Vision-Language Models | May 2, 2025 | Image CaptioningObject Recognition | —Unverified | 0 | 0 |
| A Confidence-Based Interface for Neuro-Symbolic Visual Question Answering | Nov 21, 2021 | Question AnsweringTranslation | —Unverified | 0 | 0 |
| MMPKUBase: A Comprehensive and High-quality Chinese Multi-modal Knowledge Graph | Aug 3, 2024 | AttributeContrastive Learning | —Unverified | 0 | 0 |
| Transformers in Vision: A Survey | Jan 4, 2021 | Action RecognitionActivity Recognition | —Unverified | 0 | 0 |
| Transform-Retrieve-Generate: Natural Language-Centric Outside-Knowledge Visual Question Answering | Jan 1, 2022 | Generative Question AnsweringImage to text | —Unverified | 0 | 0 |
| Translation Deserves Better: Analyzing Translation Artifacts in Cross-lingual Visual Question Answering | Jun 4, 2024 | Data AugmentationMachine Translation | —Unverified | 0 | 0 |
| TransMamba: Fast Universal Architecture Adaption from Transformers to Mamba | Feb 21, 2025 | image-classificationImage Classification | —Unverified | 0 | 0 |
| WangLab at MEDIQA-M3G 2024: Multimodal Medical Answer Generation using Large Language Models | Apr 22, 2024 | Answer Generationimage-classification | —Unverified | 0 | 0 |
| Asking More Informative Questions for Grounded Retrieval | Nov 14, 2023 | Question AnsweringQuestion Selection | —Unverified | 0 | 0 |
| Yin and Yang: Balancing and Answering Binary Visual Questions | Nov 16, 2015 | Question AnsweringVisual Question Answering | —Unverified | 0 | 0 |
| TraveLLaMA: Facilitating Multi-modal Large Language Models to Understand Urban Scenes and Provide Travel Assistance | Apr 23, 2025 | Question AnsweringScene Understanding | —Unverified | 0 | 0 |
| A Concept-Centric Approach to Multi-Modality Learning | Dec 18, 2024 | Image-text matchingQuestion Answering | —Unverified | 0 | 0 |
| Tree Memory Networks for Modelling Long-term Temporal Dependencies | Mar 12, 2017 | Machine TranslationPart-Of-Speech Tagging | —Unverified | 0 | 0 |
| Triplet-Aware Scene Graph Embeddings | Sep 19, 2019 | Data AugmentationGraph Embedding | —Unverified | 0 | 0 |
| Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis | Jun 21, 2024 | AttributeMedical Visual Question Answering | —Unverified | 0 | 0 |
| TrojVLM: Backdoor Attack Against Vision Language Models | Sep 28, 2024 | Backdoor AttackImage Captioning | —Unverified | 0 | 0 |
| As Firm As Their Foundations: Can open-sourced foundation models be used to create adversarial examples for downstream tasks? | Mar 19, 2024 | Adversarial AttackImage Captioning | —Unverified | 0 | 0 |
| TRRNet: Tiered Relation Reasoning for Compositional Visual Question Answering | Aug 1, 2020 | ObjectQuestion Answering | —Unverified | 0 | 0 |
| TruthLens:A Training-Free Paradigm for DeepFake Detection | Mar 19, 2025 | Binary ClassificationDeepFake Detection | —Unverified | 0 | 0 |
| Prompting Medical Large Vision-Language Models to Diagnose Pathologies by Visual Question Answering | Jul 31, 2024 | DiagnosticHallucination | —Unverified | 0 | 0 |
| A scoping review on multimodal deep learning in biomedical images and texts | Jul 14, 2023 | Cross-Modal RetrievalDecision Making | —Unverified | 0 | 0 |
| Two can play this Game: Visual Dialog with Discriminative Question Generation and Answering | Mar 29, 2018 | Image CaptioningQuestion Answering | —Unverified | 0 | 0 |
| TxT: Crossmodal End-to-End Learning with Transformers | Sep 9, 2021 | Multimodal ReasoningQuestion Answering | —Unverified | 0 | 0 |
| UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training | Apr 1, 2021 | Image-text matchingImage-text Retrieval | —Unverified | 0 | 0 |
| U-CAM: Visual Explanation using Uncertainty based Class Activation Maps | Aug 17, 2019 | Deep LearningProbabilistic Deep Learning | —Unverified | 0 | 0 |
| SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge | May 23, 2024 | Question AnsweringRAG | —Unverified | 0 | 0 |
| UFO: A UniFied TransfOrmer for Vision-Language Representation Learning | Nov 19, 2021 | Image CaptioningImage-text matching | —Unverified | 0 | 0 |
| UIT-Saviors at MEDVQA-GI 2023: Improving Multimodal Learning with Image Enhancement for Gastrointestinal Visual Question Answering | Jul 6, 2023 | DiagnosticImage Enhancement | —Unverified | 0 | 0 |