SOTAVerified

Human Judgment Correlation

A task where an algorithm should generate the judgment scores correlating with human judgments.

Papers

Showing 1–5 of 5 papers

TitleStatusHype
CLIPScore: A Reference-free Evaluation Metric for Image CaptioningCode1
FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph ParsingCode1
Mutual Information Divergence: A Unified Metric for Multimodal Generative ModelsCode1
Improving Image Captioning Evaluation by Considering Inter References Variance—0
PerSEval: Assessing Personalization in Text Summarizers—0
Show:102550

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MIDKendall's Tau-c54.9—Unverified
2SoftSPICEKendall's Tau-c54.2—Unverified
3RefCLIP-SKendall's Tau-c53—Unverified
4CLIP-SKendall's Tau-c51.2—Unverified
#ModelMetricClaimedVerifiedStatus
1MIDKendall's Tau-b37.3—Unverified
2RefCLIP-SKendall's Tau-b36.4—Unverified
3CLIP-SKendall's Tau-b34.4—Unverified