SOTAVerified

Causal Language Modeling

Papers

Showing 1–50 of 52 papers

TitleStatusHype
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization—0
GRITHopper: Decomposition-Free Multi-Hop Dense RetrievalCode1
Trojan Detection Through Pattern Recognition for Large Language Models—0
Towards the Anonymization of the Language Modeling—0
Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language ModelsCode0
AntLM: Bridging Causal and Masked Language Models—0
Enhancing Trust in Large Language Models with Uncertainty-Aware Fine-Tuning—0
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation—0
Interpretable Language Modeling via Induction-head Ngram ModelsCode1
GPT or BERT: why not both?Code2
A Simple Baseline for Predicting Events with Auto-Regressive Tabular TransformersCode0
QuAILoRA: Quantization-Aware Initialization for LoRA—0
Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic PuzzlesCode1
Generating Synthetic Free-text Medical Records with Low Re-identification Risk using Masked Language ModelingCode0
N-gram Prediction and Word Difference Representations for Language Modeling—0
Masked Mixers for Language Generation and RetrievalCode0
Novel-WD: Exploring acquisition of Novel World Knowledge in LLMs Using Prefix-Tuning—0
Predictability and Causality in Spanish and English Natural Language Generation—0
Conditional Language Learning with ContextCode0
Understanding Token Probability Encoding in Output Embeddings—0
Transformer based neural networks for emotion recognition in conversationsCode0
NIFTY Financial News Headlines Dataset—0
SERPENT-VLM : Self-Refining Radiology Report Generation Using Vision Language Models—0
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation—0
Looking Right is Sometimes Right: Investigating the Capabilities of Decoder-only LLMs for Sequence Labeling—0
Linear Attention via Orthogonal Memory—0
Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image CaptioningCode1
DavIR: Data Selection via Implicit Reward for Large Language Models—0
Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning AbilityCode0
A Meta-Learning Perspective on Transformers for Causal Language Modeling—0
What's the Magic Word? A Control Theory of LLM PromptingCode1
AstroLLaMA: Towards Specialized Foundation Models in Astronomy—0
CodeGen2: Lessons for Training LLMs on Programming and Natural LanguagesCode5
ProtFIM: Fill-in-Middle Protein Sequence Design via Protein Language Models—0
Video Pre-trained Transformer: A Multimodal Mixture of Pre-trained ExpertsCode1
Cross-lingual Similarity of Multilingual Representations RevisitedCode0
Suffix Retrieval-Augmented Language ModelingCode0
A Simple, Yet Effective Approach to Finding Biases in Code Generation—0
A Closer Look at Parameter Contributions When Training Neural Language and Translation Models—0
AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq ModelCode2
Learning from flowsheets: A generative transformer model for autocompletion of flowsheets—0
Self-Supervised Learning of Brain Dynamics from Broad Neuroimaging DataCode1
Language Models are General-Purpose Interfaces—0
Transcormer: Transformer for Sentence Scoring with Sliding Language ModelingCode1
Multitask Finetuning for Improving Neural Machine Translation in Indian Languages—0
Prix-LM: Pretraining for Multilingual Knowledge Base Construction—0
Prix-LM: Pretraining for Multilingual Knowledge Base ConstructionCode0
Multi-Task Learning for Situated Multi-Domain End-to-End Dialogue Systems—0
IntenT5: Search Result Diversification using Causal Language Models—0
Large Product Key Memory for Pretrained Language ModelsCode0
Show:102550
← PrevPage 1 of 2Next →

No leaderboard results yet.