| Segment Any Text: A Universal Approach for Robust, Efficient and Adaptable Sentence Segmentation | Jun 24, 2024 | parameter-efficient fine-tuningSentence | CodeCode Available | 7 | 5 |
| Where's the Point? Self-Supervised Multilingual Punctuation-Agnostic Sentence Segmentation | May 30, 2023 | Machine TranslationSegmentation | CodeCode Available | 3 | 5 |
| Abstractive Summarization of Spoken andWritten Instructions with BERT | Aug 21, 2020 | Abstractive Text SummarizationArticles | CodeCode Available | 2 | 5 |
| A unified approach to sentence segmentation of punctuated text in many languages | Aug 1, 2021 | SentenceSentence segmentation | CodeCode Available | 1 | 5 |
| KG-GPT: A General Framework for Reasoning on Knowledge Graphs Using Large Language Models | Oct 17, 2023 | Fact VerificationKnowledge Graphs | CodeCode Available | 1 | 5 |
| Trankit: A Light-Weight Transformer-based Toolkit for Multilingual Natural Language Processing | Jan 9, 2021 | Dependency ParsingLanguage Modeling | CodeCode Available | 1 | 5 |
| Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine Translation | Sep 20, 2020 | Machine TranslationSentence | CodeCode Available | 1 | 5 |
| Opera Graeca Adnotata: Building a 34M+ Token Multilayer Corpus for Ancient Greek | Mar 31, 2024 | LemmatizationSentence | CodeCode Available | 1 | 5 |
| Lexical Semantic Recognition | Apr 30, 2020 | Natural Language UnderstandingSentence | CodeCode Available | 1 | 5 |
| Ascle: A Python Natural Language Processing Toolkit for Medical Text Generation | Nov 28, 2023 | Machine TranslationQuestion Answering | CodeCode Available | 1 | 5 |
| Mukayese: Turkish NLP Strikes Back | Mar 2, 2022 | BenchmarkingLanguage Modeling | CodeCode Available | 1 | 5 |
| Abstractive Summarization of Spoken and Written Instructions with BERT | Aug 21, 2020 | Abstractive Text SummarizationArticles | CodeCode Available | 1 | 5 |
| Using Punkt for Sentence Segmentation in non-Latin Scripts: Experiments on Kurdish (Sorani) Texts | Apr 9, 2020 | SentenceSentence segmentation | CodeCode Available | 0 | 5 |
| Creating a Universal Dependencies Treebank of Spoken Frisian-Dutch Code-switched Data | Feb 22, 2021 | SentenceSentence segmentation | CodeCode Available | 0 | 5 |
| Towards JointUD: Part-of-speech Tagging and Lemmatization using Recurrent Neural Networks | Sep 10, 2018 | Dependency ParsingLemmatization | CodeCode Available | 0 | 5 |
| Universal Dependency Parsing from Scratch | Jan 29, 2019 | AllDependency Parsing | CodeCode Available | 0 | 5 |
| Prosodic features improve sentence segmentation and parsing | Feb 23, 2023 | SentenceSentence segmentation | CodeCode Available | 0 | 5 |
| Fine-Grained Argument Unit Recognition and Classification | Apr 22, 2019 | Argument MiningArgument Retrieval | CodeCode Available | 0 | 5 |
| Evaluating Sentence Segmentation and Word Tokenization Systems on Estonian Web Texts | Nov 16, 2020 | SegmentationSentence | CodeCode Available | 0 | 5 |
| LeConTra: A Learner Corpus of English-to-Dutch News Translation | Jun 1, 2022 | SentenceSentence segmentation | CodeCode Available | 0 | 5 |
| Human Genome Book: Words, Sentences and Paragraphs | Jan 23, 2025 | Protein Structure PredictionSentence segmentation | CodeCode Available | 0 | 5 |
| SLATE: A Sequence Labeling Approach for Task Extraction from Free-form Inked Content | Nov 8, 2022 | FormSegmentation | CodeCode Available | 0 | 5 |
| CUNI Systems in WMT21: Revisiting Backtranslation Techniques for English-Czech NMT | Nov 1, 2021 | NMTSegmentation | —Unverified | 0 | 0 |
| Creating Training Corpora for NLG Micro-Planners | Jul 1, 2017 | Data-to-Text GenerationReferring Expression | —Unverified | 0 | 0 |
| A Statistical, Grammar-Based Approach to Microplanning | Apr 1, 2017 | SentenceSentence segmentation | —Unverified | 0 | 0 |