SOTAVerified

Object Recognition

Object recognition is a computer vision technique for detecting + classifying objects in images or videos. Since this is a combined task of object detection plus image classification, the state-of-the-art tables are recorded for each component task here and here.

( Image credit: Tensorflow Object Detection API )

Papers

Showing 2650 of 2042 papers

TitleStatusHype
Topology-Guided Knowledge Distillation for Efficient Point Cloud ProcessingCode0
Visually Interpretable Subtask Reasoning for Visual Question AnsweringCode0
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding0
Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models0
Transferable Adversarial Attacks on Black-Box Vision-Language Models0
Zoomer: Adaptive Image Focus Optimization for Black-box MLLM0
LM-MCVT: A Lightweight Multi-modal Multi-view Convolutional-Vision Transformer Approach for 3D Object Recognition0
Benchmarking Multimodal Mathematical Reasoning with Explicit Visual DependencyCode1
Disaggregated Deep Learning via In-Physics Computing at Radio Frequency0
V^2R-Bench: Holistically Evaluating LVLM Robustness to Fundamental Visual Variations0
Naturally Computed Scale Invariance in the Residual Stream of ResNet18Code0
Quantum Doubly Stochastic Transformers0
Taccel: Scaling Up Vision-based Tactile Robotics via High-performance GPU SimulationCode2
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment0
Visual Language Models show widespread visual deficits on neuropsychological tests0
MASSeg : 2nd Technical Report for 4th PVUW MOSE TrackCode0
Hardware, Algorithms, and Applications of the Neuromorphic Vision Sensor: a Review0
P2Object: Single Point Supervised Object Detection and Instance SegmentationCode2
D-Feat Occlusions: Diffusion Features for Robustness to Partial Visual Occlusions in Object Recognition0
Advancing Egocentric Video Question Answering with Multimodal Large Language Models0
ForcePose: A Deep Learning Approach for Force Calculation Based on Action Recognition Using MediaPipe Pose Estimation Combined with Object Detection0
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users0
Foveated Instance SegmentationCode0
DuckSegmentation: A segmentation model based on the AnYue Hemp Duck Dataset0
Leveraging 3D Geometric Priors in 2D Rotation Symmetry Detection0
Show:102550
← PrevPage 2 of 82Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Imagenshape bias98.7Unverified
2Stable Diffusionshape bias92.7Unverified
3Partishape bias91.7Unverified
4ViT-22B-384shape bias86.4Unverified
5ViT-22B-560shape bias83.8Unverified
6CLIP (ViT-B)shape bias79.9Unverified
7ViT-22B-224shape bias78Unverified
8ResNet-50 (L2 eps 5.0 adv trained)shape bias69.5Unverified
9ResNet-50 (with strong augmentations)shape bias62.2Unverified
10SWSL (ResNeXt-101)shape bias49.8Unverified
#ModelMetricClaimedVerifiedStatus
1Spike-VGG11Accuracy (% )85.55Unverified
2SSNNAccuracy (% )78.57Unverified
#ModelMetricClaimedVerifiedStatus
1Spike-VGG11Accuracy (% )85.62Unverified
2SSNNAccuracy (% )79.25Unverified
#ModelMetricClaimedVerifiedStatus
1ObjectNet-BaselineTop 5 Accuracy18.75Unverified
2yunTop 5 Accuracy14.75Unverified
#ModelMetricClaimedVerifiedStatus
1ObjectNet-BaselineTop 5 Accuracy52.24Unverified
2DYTop 5 Accuracy0.08Unverified
#ModelMetricClaimedVerifiedStatus
1ObjectNet-BaselineTop 5 Accuracy52.24Unverified
2AJ2021Top 5 Accuracy27.68Unverified
#ModelMetricClaimedVerifiedStatus
1SSNNAccuracy (% )94.91Unverified
#ModelMetricClaimedVerifiedStatus
1Faster-RCNNmAP30.39Unverified
#ModelMetricClaimedVerifiedStatus
1Spike-VGG11Accuracy (% )96Unverified