SOTAVerified

Visual Navigation

Visual Navigation is the problem of navigating an agent, e.g. a mobile robot, in an environment using camera input only. The agent is given a target image (an image it will see from the target position), and its goal is to move from its current position to the target by applying a sequence of actions, based on the camera observations only.

Source: Vision-based Navigation Using Deep Reinforcement Learning

Papers

Showing 1–50 of 316 papers

TitleStatusHype
LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction—0
Enhancing Safety of Foundation Models for Visual Navigation through Collision Avoidance via Repulsive Estimation—0
Visual Planning: Let's Think Only with ImagesCode3
Learning to Drive Anywhere with Model-Based Reannotation—0
Task-Oriented Communications for Visual Navigation with Edge-Aerial Collaboration in Low Altitude EconomyCode1
Unreal Robotics Lab: A High-Fidelity Robotics Simulator with Advanced Physics and Rendering—0
Decision-based AI Visual Navigation for Cardiac Ultrasounds—0
Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge ModelsCode2
NaviDiffusor: Cost-Guided Diffusion Model for Visual NavigationCode2
The Composite Visual-Laser Navigation Method Applied in Indoor Poultry Farming Environments—0
UAS Visual Navigation in Large and Unseen Environments via a Meta Agent—0
Good Actions Succeed, Bad Actions Generalize: A Case Study on Why RL Generalizes Better—0
ViVa-SAFELAND: a New Freeware for Safe Validation of Vision-based Navigation in Aerial Vehicles—0
Reasoning in visual navigation of end-to-end trained agents: a dynamical systems approach—0
A Map-free Deep Learning-based Framework for Gate-to-Gate Monocular Visual Navigation aboard Miniaturized Aerial Vehicles—0
EndoMamba: An Efficient Foundation Model for Endoscopic Videos via Hierarchical Pre-trainingCode1
High-precision visual navigation device calibration method based on collimator—0
Improving Collision-Free Success Rate For Object Goal Visual Navigation Via Two-Stage Training With Collision Prediction—0
RAPID: Robust and Agile Planner Using Inverse Reinforcement Learning for Vision-Based Drone Navigation—0
Enhancing Feature Tracking Reliability for Visual Navigation using Real-Time Safety Filter—0
VR-Robo: A Real-to-Sim-to-Real Framework for Visual Robot Navigation and Locomotion—0
Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation—0
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI—0
FloNa: Floor Plan Guided Embodied Visual Navigation—0
Multi-View Pedestrian Occupancy Prediction with a Novel Synthetic Dataset—0
Agent Journey Beyond RGB: Unveiling Hybrid Semantic-Spatial Environmental Representations for Vision-and-Language NavigationCode1
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of ExpertsCode1
Navigation World ModelsCode4
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual PreferencesCode2
CityWalker: Learning Embodied Urban Navigation from Web-Scale VideosCode3
MetaCropFollow: Few-Shot Adaptation with Meta-Learning for Under-Canopy Navigation—0
Memory Proxy Maps for Visual Navigation—0
Grounding Video Models to Actions through Goal Conditioned Exploration—0
Personalized Instance-based Navigation Toward User-Specific Objects in Realistic EnvironmentsCode1
Visual Navigation of Digital Libraries: Retrieval and Classification of Images in the National Library of Norway's Digitised Book CollectionCode0
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features—0
RNR-Nav: A Real-World Visual Navigation System Using Renderable Neural Radiance Maps—0
Fast Object Detection with a Machine Learning Edge Device—0
Initialization of Monocular Visual Navigation for Autonomous Agents Using Modified Structure from Small Motion—0
HM3D-OVON: A Dataset and Benchmark for Open-Vocabulary Object Goal Navigation—0
Causality-Aware Transformer Networks for Robotic Navigation—0
Addressing the challenges of loop detection in agricultural environmentsCode0
NOLO: Navigate Only Look Once—0
IN-Sight: Interactive Navigation through Sight—0
Visuospatial navigation without distance, prediction, integration, or maps—0
CAMON: Cooperative Agents for Multi-Object Navigation with LLM-based Conversations—0
Solving Vision Tasks with Simple Photoreceptors Instead of Cameras—0
SPIN: Spacecraft Imagery for NavigationCode1
RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation—0
Sim2Real Transfer for Audio-Visual Navigation with Frequency-Adaptive Acoustic Field Prediction—0
Show:102550
← PrevPage 1 of 7Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1NaviLLMdist_to_end_reduction7.9—Unverified
2VLN-PETLdist_to_end_reduction6.13—Unverified
3early to beddist_to_end_reduction6.03—Unverified
4HAMTdist_to_end_reduction5.58—Unverified
5s-agent (NDH-Full)dist_to_end_reduction5.27—Unverified
6BabyWalk (r2r-pretrain)dist_to_end_reduction4.46—Unverified
7Environment-agnostic Multitask Learningdist_to_end_reduction3.91—Unverified
8BabyWalkdist_to_end_reduction3.65—Unverified
9Test2-NDHdist_to_end_reduction3.44—Unverified
10SCoAdist_to_end_reduction3.37—Unverified
#ModelMetricClaimedVerifiedStatus
1SUSAspl0.64—Unverified
2Meta-Explorespl0.61—Unverified
3NaviLLMspl0.6—Unverified
4BEV-BERTspl0.6—Unverified
5HOPspl0.59—Unverified
6DUETspl0.58—Unverified
7VLN-PETLspl0.58—Unverified
8VLN-BERTspl0.57—Unverified
9Prevalentspl0.51—Unverified
10RCM+SIL(no early exploration)spl0.38—Unverified
#ModelMetricClaimedVerifiedStatus
1AutoVLNNav-SPL27.83—Unverified
2NaviLLMNav-SPL26.26—Unverified
3Meta-ExploreNav-SPL25.8—Unverified
4SUSANav-SPL25.47—Unverified
5DUETNav-SPL21.42—Unverified
6GBENav-SPL13.3—Unverified
#ModelMetricClaimedVerifiedStatus
1MVV-INSPL (All)17.27—Unverified
2SAVNSPL (All)16.15—Unverified
#ModelMetricClaimedVerifiedStatus
1PopArt-IMPALAMedium Human-Normalized Score72.8—Unverified
#ModelMetricClaimedVerifiedStatus
1Prevalentspl28.72—Unverified