| Towards Top-Down Reasoning: An Explainable Multi-Agent Approach for Visual Question Answering | Nov 29, 2023 | Common Sense ReasoningQuestion Answering | —Unverified | 0 |
| Towards Transparent AI Systems: Interpreting Visual Question Answering Models | Aug 31, 2016 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers | Jan 3, 2024 | Question AnsweringVisual Grounding | —Unverified | 0 |
| Towards Unsupervised Visual Reasoning: Do Off-The-Shelf Features Know How to Reason? | Dec 20, 2022 | Question AnsweringRepresentation Learning | —Unverified | 0 |
| Towards Visual Dialog for Radiology | Jul 1, 2020 | Question AnsweringVisual Dialog | —Unverified | 0 |
| Toward Unsupervised Realistic Visual Question Answering | Mar 9, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images | Feb 23, 2025 | Adversarial AttackQuestion Answering | —Unverified | 0 |
| Training Recurrent Answering Units with Joint Loss Minimization for VQA | Jun 12, 2016 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Transferable Adversarial Attacks on Black-Box Vision-Language Models | May 2, 2025 | Image CaptioningObject Recognition | —Unverified | 0 |
| Transformers in Vision: A Survey | Jan 4, 2021 | Action RecognitionActivity Recognition | —Unverified | 0 |
| Transform-Retrieve-Generate: Natural Language-Centric Outside-Knowledge Visual Question Answering | Jan 1, 2022 | Generative Question AnsweringImage to text | —Unverified | 0 |
| Translation Deserves Better: Analyzing Translation Artifacts in Cross-lingual Visual Question Answering | Jun 4, 2024 | Data AugmentationMachine Translation | —Unverified | 0 |
| TransMamba: Fast Universal Architecture Adaption from Transformers to Mamba | Feb 21, 2025 | image-classificationImage Classification | —Unverified | 0 |
| TraveLLaMA: Facilitating Multi-modal Large Language Models to Understand Urban Scenes and Provide Travel Assistance | Apr 23, 2025 | Question AnsweringScene Understanding | —Unverified | 0 |
| Tree Memory Networks for Modelling Long-term Temporal Dependencies | Mar 12, 2017 | Machine TranslationPart-Of-Speech Tagging | —Unverified | 0 |
| Triplet-Aware Scene Graph Embeddings | Sep 19, 2019 | Data AugmentationGraph Embedding | —Unverified | 0 |
| Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis | Jun 21, 2024 | AttributeMedical Visual Question Answering | —Unverified | 0 |
| TrojVLM: Backdoor Attack Against Vision Language Models | Sep 28, 2024 | Backdoor AttackImage Captioning | —Unverified | 0 |
| TRRNet: Tiered Relation Reasoning for Compositional Visual Question Answering | Aug 1, 2020 | ObjectQuestion Answering | —Unverified | 0 |
| TruthLens:A Training-Free Paradigm for DeepFake Detection | Mar 19, 2025 | Binary ClassificationDeepFake Detection | —Unverified | 0 |
| Two can play this Game: Visual Dialog with Discriminative Question Generation and Answering | Mar 29, 2018 | Image CaptioningQuestion Answering | —Unverified | 0 |
| TxT: Crossmodal End-to-End Learning with Transformers | Sep 9, 2021 | Multimodal ReasoningQuestion Answering | —Unverified | 0 |
| UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training | Apr 1, 2021 | Image-text matchingImage-text Retrieval | —Unverified | 0 |
| U-CAM: Visual Explanation using Uncertainty based Class Activation Maps | Aug 17, 2019 | Deep LearningProbabilistic Deep Learning | —Unverified | 0 |
| SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge | May 23, 2024 | Question AnsweringRAG | —Unverified | 0 |