| Capability-Based Scaling Laws for LLM Red-Teaming | May 26, 2025 | MMLUPrompt Engineering | CodeCode Available | 0 | 5 |
| CHAIR -- Classifier of Hallucination as Improver | Jan 5, 2025 | HallucinationMMLU | CodeCode Available | 0 | 5 |
| Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations | Jul 7, 2025 | AttributeMMLU | CodeCode Available | 0 | 5 |
| NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning | Mar 30, 2024 | Language ModelingLanguage Modelling | —Unverified | 0 | 0 |
| Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models | Feb 20, 2025 | HellaSwagMemorization | —Unverified | 0 | 0 |
| Octopus v4: Graph of language models | Apr 30, 2024 | MMLU | —Unverified | 0 | 0 |
| OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning | Feb 16, 2025 | MedQAMMLU | —Unverified | 0 | 0 |
| On the Reasoning Capacity of AI Models and How to Quantify It | Jan 23, 2025 | MemorizationMMLU | —Unverified | 0 | 0 |
| OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models | Feb 29, 2024 | Medical Question AnsweringMedQA | —Unverified | 0 | 0 |
| Optimised Grouped-Query Attention Mechanism for Transformers | Jun 21, 2024 | MMLU | —Unverified | 0 | 0 |
| Order Independence With Finetuning | Mar 30, 2025 | ARCLanguage Modeling | —Unverified | 0 | 0 |
| ORI: O Routing Intelligence | Feb 14, 2025 | ARCMMLU | —Unverified | 0 | 0 |
| Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone | Apr 22, 2024 | Language ModelingLanguage Modelling | —Unverified | 0 | 0 |
| Pistis-RAG: Enhancing Retrieval-Augmented Generation with Human Feedback | Jun 21, 2024 | Information RetrievalLearning-To-Rank | —Unverified | 0 | 0 |
| PolyPrompt: Automating Knowledge Extraction from Multilingual Language Models with Dynamic Prompt Generation | Feb 27, 2025 | MMLU | —Unverified | 0 | 0 |
| Predicting Emergent Capabilities by Finetuning | Nov 25, 2024 | CoLAGSM8K | —Unverified | 0 | 0 |
| BOTS-LM: Training Large Language Models for Setswana | Aug 5, 2024 | Computational EfficiencyLanguage Modeling | —Unverified | 0 | 0 |
| Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs | Sep 30, 2024 | ARCDiversity | —Unverified | 0 | 0 |
| Project MPG: towards a generalized performance benchmark for LLM capabilities | Oct 28, 2024 | BenchmarkingChatbot | —Unverified | 0 | 0 |
| Pruning Large Language Models via Accuracy Predictor | Sep 18, 2023 | MMLUModel Compression | —Unverified | 0 | 0 |
| ConceptPsy:A Benchmark Suite with Conceptual Comprehensiveness in Psychology | Nov 16, 2023 | MMLUMultiple-choice | —Unverified | 0 | 0 |
| Quantifying Variance in Evaluation Benchmarks | Jun 14, 2024 | MMLU | —Unverified | 0 | 0 |
| ALLaM: Large Language Models for Arabic and English | Jul 22, 2024 | DecoderLanguage Acquisition | —Unverified | 0 | 0 |
| AgentInstruct: Toward Generative Teaching with Agentic Flows | Jul 3, 2024 | GSM8KMMLU | —Unverified | 0 | 0 |
| Reactor Mk.1 performances: MMLU, HumanEval and BBH test results | Jun 15, 2024 | BenchmarkingHumanEval | —Unverified | 0 | 0 |