| Technical Report of the Video Event Reconstruction and Analysis (VERA) System -- Shooter Localization, Models, Interface, and Beyond | May 26, 2019 | Gunshot DetectionShooter Localization | CodeCode Available | 0 | 5 |
| Video Anomaly Detection for Smart Surveillance | Apr 1, 2020 | Anomaly DetectionTemporal Localization | —Unverified | 0 | 0 |
| Exploring State Change Capture of Heterogeneous Backbones @ Ego4D Hands and Objects Challenge 2022 | Nov 16, 2022 | Human-Object Interaction DetectionObject | —Unverified | 0 | 0 |
| Efficient Action Localization with Approximately Normalized Fisher Vectors | Jun 1, 2014 | Action LocalizationAction Recognition | —Unverified | 0 | 0 |
| Optimizing Temporal Resolution Of Convolutional Recurrent Neural Networks For Sound Event Detection | Oct 18, 2022 | Event DetectionSound Event Detection | —Unverified | 0 | 0 |
| OWL (Observe, Watch, Listen): Audiovisual Temporal Context for Localizing Actions in Egocentric Videos | Feb 10, 2022 | Action LocalizationTemporal Action Localization | —Unverified | 0 | 0 |
| PcmNet: Position-Sensitive Context Modeling Network for Temporal Action Localization | Mar 9, 2021 | Action LocalizationBoundary Detection | —Unverified | 0 | 0 |
| Pointly-Supervised Action Localization | May 29, 2018 | Action LocalizationMultiple Instance Learning | —Unverified | 0 | 0 |
| Poselet Key-Framing: A Model for Human Activity Recognition | Jun 1, 2013 | Activity RecognitionHuman Activity Recognition | —Unverified | 0 | 0 |
| Practitioner-Centric Approach for Early Incident Detection Using Crowdsourced Data for Emergency Services | Dec 3, 2021 | Event DetectionManagement | —Unverified | 0 | 0 |
| Pseudo Strong Labels from Frame-Level Predictions for Weakly Supervised Sound Event Detection | Jan 7, 2025 | Event DetectionSound Event Detection | —Unverified | 0 | 0 |
| ReActNet: Temporal Localization of Repetitive Activities in Real-World Videos | Oct 14, 2019 | Temporal Localization | —Unverified | 0 | 0 |
| A benchmark of state-of-the-art sound event detection systems evaluated on synthetic soundscapes | Feb 3, 2022 | Data AugmentationEvent Detection | —Unverified | 0 | 0 |
| Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos | Sep 18, 2020 | cross-modal alignmentreinforcement-learning | —Unverified | 0 | 0 |
| Scalable Temporal Localization of Sensitive Activities in Movies and TV Episodes | Jun 16, 2022 | Temporal Localization | —Unverified | 0 | 0 |
| Efficient Action Detection in Untrimmed Videos via Multi-Task Learning | Dec 22, 2016 | Action DetectionAction Localization | —Unverified | 0 | 0 |
| AdapNet: Adaptability Decomposing Encoder-Decoder Network for Weakly Supervised Action Recognition and Localization | Nov 27, 2019 | Action ClassificationAction Recognition | —Unverified | 0 | 0 |
| Sequential End-to-End Intent and Slot Label Classification and Localization | Jun 8, 2021 | Automatic Speech RecognitionAutomatic Speech Recognition (ASR) | —Unverified | 0 | 0 |
| ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries | Dec 17, 2024 | Human Detectionimage-classification | —Unverified | 0 | 0 |
| Single-Stage Visual Query Localization in Egocentric Videos | Jun 15, 2023 | object-detectionObject Detection | —Unverified | 0 | 0 |
| Activity Recognition on a Large Scale in Short Videos - Moments in Time Dataset | Sep 1, 2018 | Action RecognitionActivity Recognition | —Unverified | 0 | 0 |
| SocialGesture: Delving into Multi-person Gesture Understanding | Apr 3, 2025 | Gesture RecognitionQuestion Answering | —Unverified | 0 | 0 |
| Action Shuffling for Weakly Supervised Temporal Localization | May 10, 2021 | Action LocalizationTemporal Localization | —Unverified | 0 | 0 |
| Spatio-Temporal Attention Models for Grounded Video Captioning | Oct 17, 2016 | image-classificationImage Classification | —Unverified | 0 | 0 |
| Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding | Mar 28, 2023 | Action LocalizationAction Recognition | —Unverified | 0 | 0 |