SOTAVerified

Image Captioning

Image Captioning is the task of describing the content of an image in words. This task lies at the intersection of computer vision and natural language processing. Most image captioning systems use an encoder-decoder framework, where an input image is encoded into an intermediate representation of the information in the image, and then decoded into a descriptive text sequence. The most popular benchmarks are nocaps and COCO, and models are typically evaluated according to a BLEU or CIDER metric.

( Image credit: Reflective Decoding Network for Image Captioning, ICCV'19)

Papers

Showing 1–10 of 1878 papers

Show:102550
← PrevPage 1 of 188Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GIT2CIDEr124.18—Unverified
2GITCIDEr122.4—Unverified
3VLAF2CIDEr106.36—Unverified
4Microsoft Cognitive Services teamCIDEr100.62—Unverified
5test_cbs2CIDEr90.73—Unverified
6icp2ssi1_coco_si_0.02_5_testCIDEr82.86—Unverified
7HumanCIDEr80.61—Unverified
8UpDown + ELMo + CBSCIDEr76.02—Unverified
9UpDownCIDEr74.27—Unverified
10Neural Baby Talk + CBSCIDEr62.96—Unverified