Scene Understanding

Scene understanding involves interpreting the visual information of a scene, including objects, their spatial relationships, and the overall layout. It goes beyond simple object recognition by considering the context and how objects relate to each other and the environment.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 451–500 of 1723 papers

Title	Date	Tasks	Status	Hype
Image Segmentation Using Deep Learning: A Survey	Jan 15, 2020	DecoderDeep Learning	CodeCode Available	1
NODIS: Neural Ordinary Differential Scene Understanding	Jan 14, 2020	AllGraph Generation	CodeCode Available	1
Visual-Semantic Graph Attention Networks for Human-Object Interaction Detection	Jan 7, 2020	Graph AttentionHuman-Object Interaction Detection	CodeCode Available	1
IRS: A Large Naturalistic Indoor Robotics Stereo Dataset to Train Deep Models for Disparity and Surface Normal Estimation	Dec 20, 2019	Disparity EstimationScene Understanding	CodeCode Available	1
AeroRIT: A New Scene for Hyperspectral Image Analysis	Dec 17, 2019	Hyperspectral image analysisImage Super-Resolution	CodeCode Available	1
TextSLAM: Visual SLAM with Planar Text Features	Nov 26, 2019	Object SLAMScene Understanding	CodeCode Available	1
Towards Ghost-free Shadow Removal via Dual Hierarchical Aggregation Network and Shadow Matting GAN	Nov 20, 2019	2kGenerative Adversarial Network	CodeCode Available	1
DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames	Nov 1, 2019	Autonomous NavigationGPU	CodeCode Available	1
Underwater Image Super-Resolution using Deep Residual Multipliers	Sep 20, 2019	Image Super-ResolutionScene Understanding	CodeCode Available	1
Global Aggregation then Local Distribution in Fully Convolutional Networks	Sep 16, 2019	Instance Segmentationobject-detection	CodeCode Available	1
Dynamic Graph Message Passing Networks	Aug 19, 2019	Image Classificationobject-detection	CodeCode Available	1
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering	Aug 14, 2019	Embodied Question AnsweringQuestion Answering	CodeCode Available	1
M3D-RPN: Monocular 3D Region Proposal Network for Object Detection	Jul 13, 2019	3D Object Detection3D Object Detection From Monocular Images	CodeCode Available	1
From Points to Parts: 3D Object Detection from Point Cloud with Part-aware and Part-aggregation Network	Jul 8, 2019	3D Object DetectionObject	CodeCode Available	1
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge	May 31, 2019	object-detectionObject Detection	CodeCode Available	1
GFF: Gated Fully Fusion for Semantic Segmentation	Apr 3, 2019	Scene ParsingScene Understanding	CodeCode Available	1
Curriculum Model Adaptation with Synthetic and Real Data for Semantic Foggy Scene Understanding	Jan 5, 2019	Domain AdaptationScene Understanding	CodeCode Available	1
Unified Perceptual Parsing for Scene Understanding	Jul 26, 2018	2D Semantic SegmentationScene Understanding	CodeCode Available	1
Visual Graphs from Motion (VGfM): Scene understanding with object geometry reasoning	Jul 16, 2018	3d scene graph generationGraph Generation	CodeCode Available	1
Digging Into Self-Supervised Monocular Depth Estimation	Jun 4, 2018	Camera Pose EstimationDepth Estimation	CodeCode Available	1
DeepScores -- A Dataset for Segmentation, Detection and Classification of Tiny Objects	Mar 27, 2018	General ClassificationObject	CodeCode Available	1
Semantic Line Detection and Its Applications	Oct 1, 2017	ClassificationGeneral Classification	CodeCode Available	1
LinkNet: Exploiting Encoder Representations for Efficient Semantic Segmentation	Jun 14, 2017	GPUScene Understanding	CodeCode Available	1
ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes	Feb 14, 2017	3D Object ClassificationGeneral Classification	CodeCode Available	1
Joint 2D-3D-Semantic Data for Indoor Scene Understanding	Feb 3, 2017	Scene Understanding	CodeCode Available	1
The Cityscapes Dataset for Semantic Urban Scene Understanding	Apr 6, 2016	object-detectionObject Detection	CodeCode Available	1
Bayesian SegNet: Model Uncertainty in Deep Convolutional Encoder-Decoder Architectures for Scene Understanding	Nov 9, 2015	Decision MakingDecoder	CodeCode Available	1
SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation	Nov 2, 2015	Crowd CountingDecoder	CodeCode Available	1
Microsoft COCO: Common Objects in Context	May 1, 2014	Instance SegmentationObject	CodeCode Available	1
Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models	Jul 17, 2025	3D Point Cloud ReconstructionPoint cloud reconstruction	—Unverified	0
Advancing Complex Wide-Area Scene Understanding with Hierarchical Coresets Selection	Jul 17, 2025	Scene Understanding	—Unverified	0
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning	Jul 17, 2025	Question AnsweringScene Understanding	—Unverified	0
Tactical Decision for Multi-UGV Confrontation with a Vision-Language Model-Based Commander	Jul 15, 2025	Language ModelingLanguage Modelling	—Unverified	0
Seeing the Signs: A Survey of Edge-Deployable OCR Models for Billboard Visibility Analysis	Jul 15, 2025	MarketingOptical Character Recognition	—Unverified	0
EmbRACE-3K: Embodied Reasoning and Action in Complex Environments	Jul 14, 2025	Scene UnderstandingSpatial Reasoning	—Unverified	0
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding	Jul 10, 2025	Scene UnderstandingSpatial Reasoning	CodeCode Available	0
MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation	Jul 10, 2025	NeRFObject	—Unverified	0
What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies	Jul 9, 2025	Scene UnderstandingSurvey	—Unverified	0
VoteSplat: Hough Voting Gaussian Splatting for 3D Scene Understanding	Jun 28, 2025	3DGSInstance Segmentation	—Unverified	0
CoPa-SG: Dense Scene Graphs with Parametric and Proto-Relations	Jun 26, 2025	Graph GenerationRelation	—Unverified	0
Case-based Reasoning Augmented Large Language Model Framework for Decision Making in Realistic Safety-Critical Driving Scenarios	Jun 25, 2025	Autonomous DrivingDecision Making	—Unverified	0
DreamAnywhere: Object-Centric Panoramic 3D Scene Generation	Jun 25, 2025	Novel View SynthesisObject	—Unverified	0
IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals	Jun 25, 2025	Scene Understanding	—Unverified	0
HOIverse: A Synthetic Scene Graph Dataset With Human Object Interactions	Jun 24, 2025	Graph GenerationHuman-Object Interaction Detection	—Unverified	0
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations	Jun 21, 2025	Question AnsweringScene Understanding	—Unverified	0
SceneAware: Scene-Constrained Pedestrian Trajectory Prediction with LLM-Guided Walkability	Jun 17, 2025	Pedestrian Trajectory PredictionScene Understanding	CodeCode Available	0
Leader360V: The Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment	Jun 17, 2025	Autonomous DrivingInstance Segmentation	—Unverified	0
Unified Representation Space for 3D Visual Grounding	Jun 17, 2025	3D visual groundingContrastive Learning	—Unverified	0
Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems	Jun 17, 2025	Autonomous DrivingImage Segmentation	—Unverified	0
FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding	Jun 16, 2025	FormGraph Generation	—Unverified	0

Show:10 25 50

← PrevPage 10 of 35Next →

All datasets Semantic Scene Understanding Challenge (passive actuation & ground-truth localisation)ADE20K val Semantic Scene Understanding Challenge (active actuation & ground-truth localisation)

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	ACRV Baseline	OMQ	0.44	—	Unverified
2	Team VGAI (TCS Research)	OMQ	0.37	—	Unverified
3	Demo_semantic_SLAM	OMQ	0.11	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	CPN(ResNet-101)	Mean IoU	46.3	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	ACRV Baseline	OMQ	0.35	—	Unverified