SOTAVerified

Temporal Localization

Papers

Showing 76–100 of 153 papers

TitleStatusHype
Exploring State Change Capture of Heterogeneous Backbones @ Ego4D Hands and Objects Challenge 2022—0
Exploring Temporal Preservation Networks for Precise Temporal Action Localization—0
Few-Shot Transformation of Common Actions into Time and Space—0
Fine-Tuning Large Audio-Language Models with LoRA for Precise Temporal Localization of Prolonged Exposure Therapy Elements—0
Fusion of Millimeter-wave Radar and Pulse Oximeter Data for Low-burden Diagnosis of Obstructive Sleep Apnea-Hypopnea Syndrome—0
Identity-aware Graph Memory Network for Action Detection—0
Impact of Noisy Labels on Sound Event Detection: Deletion Errors Are More Detrimental Than Insertion Errors—0
Impact of temporal resolution on convolutional recurrent networks for audio tagging and sound event detection—0
Inceptive Event Time-Surfaces for Object Classification Using Neuromorphic Cameras—0
Joint Visual-Temporal Embedding for Unsupervised Learning of Actions in Untrimmed Sequences—0
Learning to track for spatio-temporal action localization—0
Measure Twice, Cut Once: Grasping Video Structures and Event Semantics with LLMs for Video Temporal Localization—0
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval—0
Modality Shifting Attention Network for Multi-modal Video Question Answering—0
Modeling Spatio-Temporal Human Track Structure for Action Localization—0
Objects2action: Classifying and localizing actions without any video example—0
OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog—0
Optimizing Temporal Resolution Of Convolutional Recurrent Neural Networks For Sound Event Detection—0
OWL (Observe, Watch, Listen): Audiovisual Temporal Context for Localizing Actions in Egocentric Videos—0
PcmNet: Position-Sensitive Context Modeling Network for Temporal Action Localization—0
Pointly-Supervised Action Localization—0
Poselet Key-Framing: A Model for Human Activity Recognition—0
Spatio-Temporal Attention Models for Grounded Video Captioning—0
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding—0
Spectro-Temporal RF Identification using Deep Learning—0
Show:102550
← PrevPage 4 of 7Next →

No leaderboard results yet.