SOTAVerified

Long-Context Understanding

Papers

Showing 1–10 of 81 papers

TitleStatusHype
Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language ModelsCode0
Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?Code1
PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding—0
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference AccelerationCode1
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training—0
ATLAS: Learning to Optimally Memorize the Context at Test Time—0
SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long SequencesCode0
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM CompressionCode1
MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language ModelsCode1
Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning—0
Show:102550
← PrevPage 1 of 9Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4-Turbo-11062k18.5—Unverified
2GPT-4-Turbo-01252k15.5—Unverified
3Vicuna-13b-v1.5-16k2k5.4—Unverified
4Vicuna-7b-v1.5-16k2k5.3—Unverified
5LongChat-7b-v1.5-32k2k5.3—Unverified
6InternLM2-7b2k5.1—Unverified
7Claude-22k5—Unverified
8GPT-3.5-Turbo-11062k4—Unverified
9ChatGLM3-6b-32k2k2.3—Unverified
10ChatGLM2-6b-32k2k0.9—Unverified