SOTAVerified

Chunking

Chunking, also known as shallow parsing, identifies continuous spans of tokens that form syntactic units such as noun phrases or verb phrases.

Example:

| Vinken | , | 61 | years | old | | --- | ---| --- | --- | --- | | B-NLP| I-NP | I-NP | I-NP | I-NP |

Papers

Showing 101–150 of 447 papers

TitleStatusHype
Advanced System Integration: Analyzing OpenAPI Chunking for Retrieval-Augmented Generation—0
Performance Evaluation of Geospatial Images based on Zarr and Tiff—0
Unlocking Legal Knowledge with Multi-Layered Embedding-Based Retrieval—0
LLM-Ref: Enhancing Reference Handling in Technical Writing with Large Language Models—0
ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems—0
ProveRAG: Provenance-Driven Vulnerability Analysis with Automated Retrieval-Augmented LLMsCode0
EPIC: Efficient Position-Independent Caching for Serving Large Language Models—0
Action abstractions for amortized sampling—0
Is Semantic Chunking Worth the Computational Cost?—0
SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented GenerationCode0
ChuLo: Chunk-Level Key Information Representation for Long Document ProcessingCode0
SciGisPy: a Novel Metric for Biomedical Text Simplification via Gist Inference Score—0
Integrating Supertag Features into Neural Discontinuous Constituent ParsingCode0
UncertaintyRAG: Span-Level Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation—0
Medha: Efficiently Serving Multi-Million Context Length LLM Inference Requests Without Approximations—0
J2N -- Nominal Adjective Identification and its ApplicationCode0
InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation—0
3D Gaussian Splatting for Large-scale Surface Reconstruction from Aerial Images—0
TalkLoRA: Low-Rank Adaptation for Speech-Driven Animation—0
Meta Knowledge for Retrieval Augmented Large Language Models—0
Hierarchical Working Memory and a New Magic Number—0
BLAZE: Cross-Language and Cross-Project Bug Localization via Dynamic Chunking and Hard Example Learning—0
From Imitation to Refinement -- Residual RL for Precise Assembly—0
Two eyes, Two views, and finally, One summary! Towards Multi-modal Multi-tasking Knowledge-Infused Medical Dialogue SummarizationCode0
CUSIDE-array: A Streaming Multi-Channel End-to-End Speech Recognition System with Realistic Evaluations—0
Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic TransformationsCode0
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs—0
Leveraging Large Language Models for Web Scraping—0
Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented GenerationCode0
PathOCL: Path-Based Prompt Augmentation for OCL Generation with GPT-4—0
Equipping Transformer with Random-Access Reading for Long-Context Understanding—0
ExACT: An End-to-End Autonomous Excavator System Using Action Chunking With Transformers—0
Sequence Compression Speeds Up Credit Assignment in Reinforcement LearningCode0
Multi-view Content-aware Indexing for Long Document Retrieval—0
Improving Retrieval for RAG based Question Answering Models on Financial Documents—0
Opening the black box of language acquisitionCode0
BGE Landmark Embedding: A Chunking-Free Embedding Method For Retrieval Augmented Long-Context Large Language Models—0
Grounding Language Model with Chunking-Free In-Context Retrieval—0
Punctuation Restoration Improves Structure Understanding Without SupervisionCode0
Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT—0
Financial Report Chunking for Effective Retrieval Augmented GenerationCode0
Def2Vec: Extensible Word Embeddings from Dictionary DefinitionsCode0
Releasing the CRaQAn (Coreference Resolution in Question-Answering): An open-source dataset and dataset creation methodology using instruction-following models—0
A recurrent connectionist model of melody perception : An exploration using TRACX2—0
Breaking the Token Barrier: Chunking and Convolution for Efficient Long Text Classification with BERT—0
Symmetrical SyncMap for Imbalanced General Chunking Problems—0
Abstractive Summarization of Large Document Collections Using GPT—0
Chunking: Continual Learning is not just about Distribution ShiftCode0
Fine-tuned vs. Prompt-tuned Supervised Representations: Which Better Account for Brain Language Representations?—0
Exploring RWKV for Memory Efficient and Low Latency Streaming ASR—0
Show:102550
← PrevPage 3 of 9Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ACEExact Span F197.3—Unverified
2BERT-CRF (Replicated in AdaSeq)Exact Span F197.18—Unverified
3ELMo + MAT + Multi-TaskExact Span F197.04—Unverified
4CVT+Multi-Task+LargeExact Span F196.98—Unverified
5ELMo + Multi-TaskExact Span F196.83—Unverified
6FlairExact Span F196.72—Unverified
7SeqVATExact Span F195.45—Unverified
8Adversarial TrainingExact Span F195.25—Unverified
9BiLSTM-CRFExact Span F195.18—Unverified
#ModelMetricClaimedVerifiedStatus
1ACEF1 score97.3—Unverified
2Flair embeddingsF1 score96.72—Unverified
3JMTF1 score95.77—Unverified
4Low supervisionF1 score95.57—Unverified
5IntNet + BiLSTM-CRFF1 score95.29—Unverified
6Suzuki and IsozakiF1 score95.15—Unverified
7NCRF++F1 score95.06—Unverified
8BI-LSTM-CRF (Senna) (ours)F1 score94.46—Unverified
#ModelMetricClaimedVerifiedStatus
1ACEF195—Unverified
2Wang et al., 2020F194.4—Unverified
3AINF194.04—Unverified
#ModelMetricClaimedVerifiedStatus
1Wang et al., 2020F192—Unverified
2AINF191.71—Unverified
#ModelMetricClaimedVerifiedStatus
1Def2VecAUC93.07—Unverified