SOTAVerified

Human-Object Interaction Detection

Human-Object Interaction (HOI) detection is a task of identifying "a set of interactions" in an image, which involves the i) localization of the subject (i.e., humans) and target (i.e., objects) of interaction, and ii) the classification of the interaction labels.

Papers

Showing 51–75 of 449 papers

TitleStatusHype
Eye Gaze as a Signal for Conveying User Attention in Contextual AI Systems—0
Dynamic Scene Understanding from Vision-Language Representations—0
DAViD: Modeling Dynamic Affordance of 3D Objects using Pre-trained Video Diffusion Models—0
From My View to Yours: Ego-Augmented Learning in Large Vision Language Models for Understanding Exocentric Daily Living ActivitiesCode1
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction GenerationCode0
Vision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D Scenes—0
PersonaHOI: Effortlessly Improving Face Personalization in Human-Object Interaction GenerationCode0
Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance Grounding—0
HORP: Human-Object Relation Priors Guided HOI Detection—0
InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation—0
ChatHuman: Chatting about 3D Humans with Tools—0
ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation—0
PICO: Reconstructing 3D People In Contact with Objects—0
Diffgrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model—0
SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis—0
Interacted Object Grounding in Spatio-Temporal Human-Object InteractionsCode1
ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation—0
Human-Humanoid Robots Cross-Embodiment Behavior-Skill Transfer Using Decomposed Adversarial Learning from DemonstrationCode4
Descriptive Caption Enhancement with Visual Specialists for Multimodal PerceptionCode0
ContextHOI: Spatial Context Learning for Human-Object Interaction Detection—0
Orchestrating the Symphony of Prompt Distribution Learning for Human-Object Interaction Detection—0
TriDi: Trilateral Diffusion of 3D Humans, Objects, and Interactions—0
Lifting Motion to the 3D World via 2D Diffusion—0
OOD-HOI: Text-Driven 3D Whole-Body Human-Object Interactions Generation Beyond Training Domains—0
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis—0
Show:102550
← PrevPage 3 of 18Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Ours (PViC+)mAP46.49—Unverified
2RLIPv2 (Swin-L)mAP45.09—Unverified
3PViC-SwinLmAP44.32—Unverified
4SOV-STG (Swin-L)mAP43.35—Unverified
5DiffHOImAP41.5—Unverified
6ViPLOmAP37.22—Unverified
7FGAHOImAP37.18—Unverified
8ERNetmAP36.89—Unverified
9CQL+GEN-VLKT-LmAP36.03—Unverified
10QAHOI (Swin-L)mAP35.78—Unverified
#ModelMetricClaimedVerifiedStatus
1RLIPv2AP(S1)72.1—Unverified
2MURENAP(S1)68.8—Unverified
3STIPAP(S1)66—Unverified
4DiffHOIAP(S1)65.7—Unverified
5OCN (ResNet101)AP(S1)65.3—Unverified
6OCN (ResNet50)AP(S1)64.2—Unverified
7CDN (ResNet101)AP(S1)63.91—Unverified
8HOICLIPAP(S1)63.5—Unverified
9QPIC + CPCMAP63.1—Unverified
10Body Part InteractivenessAP(S1)63—Unverified
#ModelMetricClaimedVerifiedStatus
1DEFRmAP65.6—Unverified
2HAKEmAP47.1—Unverified
3PaStaNetmAP46.3—Unverified
4RelViTmAP43.98—Unverified
5Pairwise-PartmAP39.9—Unverified
6Mallya & LazebnikmAP36.1—Unverified
7Girdhar & RamananmAP34.6—Unverified
8R*CNNmAP28.5—Unverified
#ModelMetricClaimedVerifiedStatus
1HOI4ABOTDetection: Full ([email protected])11.12—Unverified
2ST-GAZEDetection: Full ([email protected])10.4—Unverified
3STTRANDetection: Full ([email protected])7.61—Unverified
#ModelMetricClaimedVerifiedStatus
1DJ-RNmAP10.37—Unverified
2iCANmAP8.14—Unverified
#ModelMetricClaimedVerifiedStatus
1SlowFast + FasterRCNN[email protected] role25.93—Unverified