SOTAVerified

multimodal interaction

Papers

Showing 51–75 of 106 papers

TitleStatusHype
RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba—0
Dissecting Dissonance: Benchmarking Large Multimodal Models Against Self-Contradictory InstructionsCode0
A Unified Understanding of Adversarial Vulnerability Regarding Unimodal Models and Vision-Language Pre-training Models—0
Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic—0
Empathic Grounding: Explorations using Multimodal Interaction and Large Language Models with Conversational AgentsCode0
HGNET: A Hierarchical Feature Guided Network for Occupancy Flow Field Prediction—0
A look under the hood of the Interactive Deep Learning Enterprise (No-IDLE)—0
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents—0
EMMI -- Empathic Multimodal Motivational Interviews Dataset: Analyses and Annotations—0
Revisiting Multimodal Emotion Recognition in Conversation from the Perspective of Graph Spectrum—0
BlendScape: Enabling End-User Customization of Video-Conferencing Environments through Generative AI—0
Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction—0
On the Arrow of Inference—0
Memory-Inspired Temporal Prompt Interaction for Text-Image Classification—0
Dynamic Hand Gesture-Featured Human Motor Adaptation in Tool Delivery using Voice Recognition—0
Adaptive User-centered Neuro-symbolic Learning for Multimodal Interaction with Autonomous Systems—0
Expanding the Role of Affective Phenomena in Multimodal Interaction Research—0
A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPTCode0
HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object Segmentation—0
InterMulti:Multi-view Multimodal Interactions with Text-dominated Hierarchical High-order Fusion for Emotion Analysis—0
A novel multimodal dynamic fusion network for disfluency detection in spoken utterances—0
Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness PredictionsCode0
Adaptive User-Centered Multimodal Interaction towards Reliable and Trusted Automotive Interfaces—0
Semantics-Consistent Cross-domain Summarization via Optimal Transport Alignment—0
MAMO: Masked Multimodal Modeling for Fine-Grained Vision-Language Representation Learning—0
Show:102550
← PrevPage 3 of 5Next →

No leaderboard results yet.