SOTAVerified

Mixture-of-Experts

Papers

Showing 401425 of 1312 papers

TitleStatusHype
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought0
Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert ModelsCode0
CoLA: Collaborative Low-Rank AdaptationCode0
MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual DecodingCode0
Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks0
FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation0
Towards Rehearsal-Free Continual Relation Extraction: Capturing Within-Task Variance with Adaptive PromptingCode0
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach0
Multimodal Cultural Safety: Evaluation Frameworks and Alignment StrategiesCode0
THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation0
StPR: Spatiotemporal Preservation and Routing for Exemplar-Free Video Class-Incremental Learning0
Multimodal Mixture of Low-Rank Experts for Sentiment Analysis and Emotion Recognition0
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training0
EfficientLLM: Efficiency in Large Language Models0
Balanced and Elastic End-to-end Training of Dynamic LLMs0
True Zero-Shot Inference of Dynamical Systems Preserving Long-Term Statistics0
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via CompetitionCode0
Model Selection for Gaussian-gated Gaussian Mixture of Experts Using Dendrograms of Mixing Measures0
Seeing the Unseen: How EMoE Unveils Bias in Text-to-Image Diffusion Models0
Multi-modal Collaborative Optimization and Expansion Network for Event-assisted Single-eye Expression RecognitionCode0
MINGLE: Mixtures of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging0
Model Merging in Pre-training of Large Language Models0
Improving Coverage in Combined Prediction Sets with Weighted p-values0
A Fast Kernel-based Conditional Independence test with Application to Causal Discovery0
On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating0
Show:102550
← PrevPage 17 of 53Next →

No leaderboard results yet.