SOTAVerified

Human-Object Interaction Detection

Human-Object Interaction (HOI) detection is a task of identifying "a set of interactions" in an image, which involves the i) localization of the subject (i.e., humans) and target (i.e., objects) of interaction, and ii) the classification of the interaction labels.

Papers

Showing 76–100 of 449 papers

TitleStatusHype
AnchorCrafter: Animate CyberAnchors Saling Your Products via Human-Object Interacting Video Generation—0
VioPose: Violin Performance 4D Pose Estimation by Hierarchical Audiovisual InferenceCode0
GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping—0
EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI DetectionCode1
Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models—0
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models—0
GraspDiffusion: Synthesizing Realistic Whole-body Hand-Object Interaction—0
Visual-Geometric Collaborative Guidance for Affordance LearningCode0
3DArticCyclists: Generating Synthetic Articulated 8D Pose-Controllable Cyclist Data for Computer Vision Applications—0
Understanding Spatio-Temporal Relations in Human-Object Interaction using Pyramid Graph Convolutional Network—0
AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation—0
Autonomous Character-Scene Interaction Synthesis from Text Instruction—0
Space-time 2D Gaussian Splatting for Accurate Surface Reconstruction under Complex Dynamic ScenesCode2
A Comprehensive Methodological Survey of Human Activity Recognition Across Divers Data Modalities—0
DreamHOI: Subject-Driven Generation of 3D Human-Object Interactions with Diffusion Priors—0
InterTrack: Tracking Human Object Interaction without Object Templates—0
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding—0
A Review of Human-Object Interaction Detection—0
Towards Flexible Visual Relationship Segmentation—0
UAHOI: Uncertainty-aware Robust Interaction Learning for HOI Detection—0
Efficient Human-Object-Interaction (EHOI) Detection via Interaction Label Coding and Conditional Decision—0
SkillMimic: Learning Basketball Interaction Skills from DemonstrationsCode3
Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI DetectionCode1
An analysis of HOI: using a training-free method with multimodal visual foundation models when only the test set is available, without the training set—0
Exploring Conditional Multi-Modal Prompts for Zero-shot HOI DetectionCode1
Show:102550
← PrevPage 4 of 18Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Ours (PViC+)mAP46.49—Unverified
2RLIPv2 (Swin-L)mAP45.09—Unverified
3PViC-SwinLmAP44.32—Unverified
4SOV-STG (Swin-L)mAP43.35—Unverified
5DiffHOImAP41.5—Unverified
6ViPLOmAP37.22—Unverified
7FGAHOImAP37.18—Unverified
8ERNetmAP36.89—Unverified
9CQL+GEN-VLKT-LmAP36.03—Unverified
10QAHOI (Swin-L)mAP35.78—Unverified
#ModelMetricClaimedVerifiedStatus
1RLIPv2AP(S1)72.1—Unverified
2MURENAP(S1)68.8—Unverified
3STIPAP(S1)66—Unverified
4DiffHOIAP(S1)65.7—Unverified
5OCN (ResNet101)AP(S1)65.3—Unverified
6OCN (ResNet50)AP(S1)64.2—Unverified
7CDN (ResNet101)AP(S1)63.91—Unverified
8HOICLIPAP(S1)63.5—Unverified
9QPIC + CPCMAP63.1—Unverified
10Body Part InteractivenessAP(S1)63—Unverified
#ModelMetricClaimedVerifiedStatus
1DEFRmAP65.6—Unverified
2HAKEmAP47.1—Unverified
3PaStaNetmAP46.3—Unverified
4RelViTmAP43.98—Unverified
5Pairwise-PartmAP39.9—Unverified
6Mallya & LazebnikmAP36.1—Unverified
7Girdhar & RamananmAP34.6—Unverified
8R*CNNmAP28.5—Unverified
#ModelMetricClaimedVerifiedStatus
1HOI4ABOTDetection: Full ([email protected])11.12—Unverified
2ST-GAZEDetection: Full ([email protected])10.4—Unverified
3STTRANDetection: Full ([email protected])7.61—Unverified
#ModelMetricClaimedVerifiedStatus
1DJ-RNmAP10.37—Unverified
2iCANmAP8.14—Unverified
#ModelMetricClaimedVerifiedStatus
1SlowFast + FasterRCNN[email protected] role25.93—Unverified