SOTAVerified

Mixture-of-Experts

Papers

Showing 601–650 of 1312 papers

TitleStatusHype
FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation—0
FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion—0
Continual Traffic Forecasting via Mixture of Experts—0
Improving Transformer Performance for French Clinical Notes Classification Using Mixture of Experts on a Limited Dataset—0
Functional mixture-of-experts for classification—0
Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs—0
Continual Pre-training of MoEs: How robust is your router?—0
Full-Precision Free Binary Graph Neural Networks—0
Continual Learning Using Task Conditional Neural Networks—0
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts—0
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models—0
ContextWIN: Whittle Index Based Mixture-of-Experts Neural Model For Restless Bandits Via Deep RL—0
From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape—0
Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning—0
Contextual Policy Transfer in Reinforcement Learning Domains via Deep Mixtures-of-Experts—0
A Simple Architecture for Enterprise Large Language Model Applications based on Role based security and Clearance Levels using Retrieval-Augmented Generation or Mixture of Experts—0
Contextual Mixture of Experts: Integrating Knowledge into Predictive Modeling—0
FreqMoE: Dynamic Frequency Enhancement for Neural PDE Solvers—0
Free Agent in Agent-Based Mixture-of-Experts Generative AI Framework—0
ConstitutionalExperts: Training a Mixture of Principle-based Prompts—0
A similarity-based Bayesian mixture-of-experts model—0
A Generalist Cross-Domain Molecular Learning Framework for Structure-Based Drug Discovery—0
Adapted-MoE: Mixture of Experts with Test-Time Adaption for Anomaly Detection—0
ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation—0
FMT:A Multimodal Pneumonia Detection Model Based on Stacking MOE Framework—0
Connector-S: A Survey of Connectors in Multi-modal Large Language Models—0
fMoE: Fine-Grained Expert Offloading for Large Mixture-of-Experts Serving—0
FloE: On-the-Fly MoE Inference on Memory-constrained GPU—0
Configurable Foundation Models: Building LLMs from a Modular Perspective—0
FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement—0
Mixture-of-Experts Meets Instruction Tuning:A Winning Combination for Large Language Models—0
Conditional computation in neural networks: principles and research trends—0
Fixing MoE Over-Fitting on Low-Resource Languages in Multilingual Machine Translation—0
FinTeamExperts: Role Specialized MOEs For Financial Analysis—0
On the Adaptation to Concept Drift for CTR Prediction—0
A Review of Sparse Expert Models in Deep Learning—0
FineQuant: Unlocking Efficiency with Fine-Grained Weight-Only Quantization for LLMs—0
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations—0
Filtered not Mixed: Stochastic Filtering-Based Online Gating for Mixture of Large Language Models—0
Complexity Experts are Task-Discriminative Learners for Any Image Restoration—0
A Review of DeepSeek Models' Key Innovative Techniques—0
AdaMV-MoE: Adaptive Multi-Task Vision Mixture-of-Experts—0
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings—0
FedMoE: Personalized Federated Learning via Heterogeneous Mixture of Experts—0
FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation—0
FedMerge: Federated Personalization via Model Merging—0
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts—0
Affect in Tweets Using Experts Model—0
Federated Mixture of Experts—0
Federated learning using mixture of experts—0
Show:102550
← PrevPage 13 of 27Next →

No leaderboard results yet.