SOTAVerified

Depth Estimation

Depth Estimation is the task of measuring the distance of each pixel relative to the camera. Depth is extracted from either monocular (single) or stereo (multiple views of a scene) images. Traditional methods use multi-view geometry to find the relationship between the images. Newer methods can directly estimate depth by minimizing the regression loss, or by learning to generate a novel view from a sequence. The most popular benchmarks are KITTI and NYUv2. Models are typically evaluated according to a RMS metric.

Source: DIODE: A Dense Indoor and Outdoor DEpth Dataset

Papers

Showing 251–300 of 2454 papers

TitleStatusHype
RSGaussian:3D Gaussian Splatting with LiDAR for Aerial Remote Sensing Novel View Synthesis—0
LiRCDepth: Lightweight Radar-Camera Depth Estimation via Knowledge Distillation and Uncertainty GuidanceCode1
Scaling 4D Representations—0
Flowing from Words to Pixels: A Framework for Cross-Modality Evolution—0
Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion—0
Foundation Models Meet Low-Cost Sensors: Test-Time Adaptation for Rescaling Disparity for Zero-Shot Metric Depth Estimation—0
Prompting Depth Anything for 4K Resolution Accurate Metric Depth EstimationCode5
PromptDet: A Lightweight 3D Object Detection Framework with LiDAR Prompts—0
V-MIND: Building Versatile Monocular Indoor 3D Detector with Diverse 2D Annotations—0
Depth-Centric Dehazing and Depth-Estimation from Real-World Hazy Driving Video—0
ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy PredictionCode1
MAL: Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance—0
Cross-View Completion Models are Zero-shot Correspondence Estimators—0
Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos—0
T-SVG: Text-Driven Stereoscopic Video Generation—0
BLADE: Single-view Body Mesh Learning through Accurate Depth Estimation—0
Utilizing Multi-step Loss for Single Image Reflection RemovalCode0
Dense Depth from Event Focal Stack—0
Balancing Shared and Task-Specific Representations: A Hybrid Approach to Depth-Aware Video Panoptic Segmentation—0
SphereUFormer: A U-Shaped Transformer for Spherical 360 Perception—0
On-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only Events—0
Driv3R: Learning Dense 4D Reconstruction for Autonomous DrivingCode2
Event fields: Capturing light fields at high speed, resolution, and dynamic range—0
Omni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstruction—0
GVDepth: Zero-Shot Monocular Depth Estimation for Ground Vehicles based on Probabilistic Cue Fusion—0
TACO: Learning Multi-modal Action Models with Synthetic Chains-of-Thought-and-ActionCode2
PanoDreamer: Optimization-Based Single Image to 360 3D Scene With DiffusionCode2
SimC3D: A Simple Contrastive 3D Pretraining Framework Using RGB ImagesCode0
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic VideosCode5
DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction—0
MT3DNet: Multi-Task learning Network for 3D Surgical Scene Reconstruction—0
LAA-Net: A Physical-prior-knowledge Based Network for Robust Nighttime Depth Estimation—0
MultiGO: Towards Multi-level Geometry Learning for Monocular 3D Textured Human Reconstruction—0
Align3R: Aligned Monocular Depth Estimation for Dynamic Videos—0
Dense Scene Reconstruction from Light-Field Images Affected by Rolling ShutterCode0
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models—0
GSGTrack: Gaussian Splatting-Guided Object Pose Tracking from RGB Videos—0
Amodal Depth Anything: Amodal Depth Estimation in the Wild—0
Dual Exposure Stereo for Extended Dynamic Range 3D Imaging—0
Single-Shot Metric Depth from Focused Plenoptic Cameras—0
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving—0
AVS-Net: Audio-Visual Scale Net for Self-supervised Monocular Metric Depth Estimation—0
STATIC : Surface Temporal Affine for TIme Consistency in Video Monocular Depth Estimation—0
Mutli-View 3D Reconstruction using Knowledge DistillationCode0
FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth Estimation—0
SpaRC: Sparse Radar-Camera Fusion for 3D Object DetectionCode0
Gaussian Splashing: Direct Volumetric Rendering Underwater—0
MonoPP: Metric-Scaled Self-Supervised Monocular Depth Estimation by Planar-Parallax Geometry in Automotive Applications—0
Video Depth without Video Models—0
360Recon: An Accurate Reconstruction Method Based on Depth Fusion from 360 Images—0
Show:102550
← PrevPage 6 of 50Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1OmniDepthRMSE0.62—Unverified
2SphereDepthRMSE0.45—Unverified
3Jin et al.RMSE0.42—Unverified
4BiFuse with fusionRMSE0.41—Unverified
5HoHoNet (ResNet-101)RMSE0.38—Unverified
6PanoDepthRMSE0.37—Unverified
7BiFuse++RMSE0.37—Unverified
8UniFuse with fusionRMSE0.37—Unverified
9DisConvRMSE0.37—Unverified
10SliceNetRMSE0.37—Unverified
#ModelMetricClaimedVerifiedStatus
1A2JmAP8.61—Unverified
2PAD-NetRMS0.79—Unverified
3MS-CRFRMS0.59—Unverified
4DORNRMS0.51—Unverified
5FreeformRMS0.43—Unverified
6Optimized, freeformRMS0.43—Unverified
7VNLRMS0.42—Unverified
8BTSRMS0.41—Unverified
9TransDepth (AGD+ ViT)RMS0.37—Unverified
10AdaBinsRMS0.36—Unverified
#ModelMetricClaimedVerifiedStatus
1T2NetAbs Rel0.35—Unverified
2MIDASAbs Rel0.31—Unverified
3Bhattacharjee et al.Abs Rel0.25—Unverified
#ModelMetricClaimedVerifiedStatus
1T2NetAbs Rel0.49—Unverified
2MIDASAbs Rel0.42—Unverified
3Bhattacharjee et al.Abs Rel0.38—Unverified
#ModelMetricClaimedVerifiedStatus
1LeReSabsolute relative error0.1—Unverified
2DELTASabsolute relative error0.09—Unverified
3Distill Any Depthabsolute relative error0.04—Unverified
#ModelMetricClaimedVerifiedStatus
1SDC-DepthRMSE6.92—Unverified
2SwinMTLRMSE6.35—Unverified
#ModelMetricClaimedVerifiedStatus
1AIP-BrownDelta < 1.250.36—Unverified
2LeResDelta < 1.250.23—Unverified
#ModelMetricClaimedVerifiedStatus
1H-Net (Ours)Absolute relative error (AbsRel)0.09—Unverified
2H-Net (Ours) Full EigenAbsolute relative error (AbsRel)0.08—Unverified
#ModelMetricClaimedVerifiedStatus
1GLPDepthDelta < 1.250.43—Unverified
2SRDINET (Model A)Delta < 1.250.4—Unverified
#ModelMetricClaimedVerifiedStatus
1Atlas (finetuned)RMSE0.17—Unverified
2Atlas (plain)RMSE0.17—Unverified
#ModelMetricClaimedVerifiedStatus
1LFattNetBadPix(0.01)17.23—Unverified
#ModelMetricClaimedVerifiedStatus
1LightDepthNumber of parameters (M)42.6—Unverified
#ModelMetricClaimedVerifiedStatus
1UniFuseAbs Rel0.11—Unverified
#ModelMetricClaimedVerifiedStatus
1X-TC (Cross-Task Consistency)L1 error1.63—Unverified