Towards Holistic Surgical Scene Understanding

2022-12-08Code Available1· sign in to hype

Natalia Valderrama, Paola Ruiz Puentes, Isabela Hernández, Nicolás Ayobi, Mathilde Verlyk, Jessica Santander, Juan Caicedo, Nicolás Fernández, Pablo Arbeláez

arXiv PDF

Code Available — Be the first to reproduce this paper.

Reproduce

Code

github.com/bcv-uniandes/tapir
OfficialIn paperpytorch★ 37

Abstract

Most benchmarks for studying surgical interventions focus on a specific challenge instead of leveraging the intrinsic complementarity among different tasks. In this work, we present a new experimental framework towards holistic surgical scene understanding. First, we introduce the Phase, Step, Instrument, and Atomic Visual Action recognition (PSI-AVA) Dataset. PSI-AVA includes annotations for both long-term (Phase and Step recognition) and short-term reasoning (Instrument detection and novel Atomic Action recognition) in robot-assisted radical prostatectomy videos. Second, we present Transformers for Action, Phase, Instrument, and steps Recognition (TAPIR) as a strong baseline for surgical scene understanding. TAPIR leverages our dataset's multi-level annotations as it benefits from the learned representation on the instrument detection task to improve its classification capacity. Our experimental results in both PSI-AVA and other publicly available databases demonstrate the adequacy of our framework to spur future research on holistic surgical scene understanding.

Tasks

Action Recognition Atomic action recognition Scene Understanding Surgical phase recognition

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
MISAW	TAPIR	mAP	94.24	—	Unverified

Towards Holistic Surgical Scene Understanding

Code

Abstract

Tasks

Benchmark Results

Reproductions