SOTAVerified

Benchmarking

Papers

Showing 50815090 of 5548 papers

TitleStatusHype
DQI: Measuring Data Quality in NLPCode0
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling TasksCode0
Domain-Expanded ASTE: Rethinking Generalization in Aspect Sentiment Triplet ExtractionCode0
WebSuite: Systematically Evaluating Why Web Agents FailCode0
Domain2Vec: Domain Embedding for Unsupervised Domain AdaptationCode0
Benchmarking Machine Learning Robustness in Covid-19 Genome Sequence ClassificationCode0
Do Localization Methods Actually Localize Memorized Data in LLMs? A Tale of Two BenchmarksCode0
Do LLMs Memorize Recommendation Datasets? A Preliminary Study on MovieLens-1MCode0
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIsCode0
A Review of Testing Object-Based Environment Perception for Safe Automated DrivingCode0
Show:102550
← PrevPage 509 of 555Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified