SOTAVerified

Mixture-of-Experts

Papers

Showing 351–400 of 1312 papers

TitleStatusHype
AT-MoE: Adaptive Task-planning Mixture of Experts via LoRA Approach—0
HDformer: A Higher Dimensional Transformer for Diabetes Detection Utilizing Long Range Vascular Signals—0
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs—0
DADNN: Multi-Scene CTR Prediction via Domain-Aware Deep Neural Network—0
D^2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving—0
A Theoretical View on Sparsely Activated Networks—0
A Large-scale Medical Visual Task Adaptation Benchmark—0
HMDN: Hierarchical Multi-Distribution Network for Click-Through Rate Prediction—0
CSAOT: Cooperative Multi-Agent System for Active Object Tracking—0
Cross-Topic Rumor Detection using Topic-Mixtures—0
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning—0
AIREX: Neural Network-based Approach for Air Quality Inference in Unmonitored Cities—0
PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning—0
Hard Mixtures of Experts for Large Scale Weakly Supervised Vision—0
Heuristic-Informed Mixture of Experts for Link Prediction in Multilayer Networks—0
Agent4Ranking: Semantic Robust Ranking via Personalized Query Rewriting Using Multi-agent LLM—0
Adapted-MoE: Mixture of Experts with Test-Time Adaption for Anomaly Detection—0
CoSMoEs: Compact Sparse Mixture of Experts—0
Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt Tuning—0
GRIN: GRadient-INformed MoE—0
Core-Periphery Principle Guided State Space Model for Functional Connectome Classification—0
Coordination with Humans via Strategy Matching—0
A Survey on Dynamic Neural Networks for Natural Language Processing—0
Convolutional Neural Networks and Mixture of Experts for Intrusion Detection in 5G Networks and beyond—0
Convergence Rates for Softmax Gating Mixture of Experts—0
Astrea: A MOE-based Visual Understanding Model with Progressive Alignment—0
HAECcity: Open-Vocabulary Scene Understanding of City-Scale Point Clouds with Superpoint Graph Clustering—0
Continual Traffic Forecasting via Mixture of Experts—0
Improving Transformer Performance for French Clinical Notes Classification Using Mixture of Experts on a Limited Dataset—0
Continual Pre-training of MoEs: How robust is your router?—0
Continual Learning Using Task Conditional Neural Networks—0
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts—0
A Generalist Cross-Domain Molecular Learning Framework for Structure-Based Drug Discovery—0
ContextWIN: Whittle Index Based Mixture-of-Experts Neural Model For Restless Bandits Via Deep RL—0
Contextual Policy Transfer in Reinforcement Learning Domains via Deep Mixtures-of-Experts—0
A Simple Architecture for Enterprise Large Language Model Applications based on Role based security and Clearance Levels using Retrieval-Augmented Generation or Mixture of Experts—0
Contextual Mixture of Experts: Integrating Knowledge into Predictive Modeling—0
ConstitutionalExperts: Training a Mixture of Principle-based Prompts—0
A similarity-based Bayesian mixture-of-experts model—0
Half-Space Feature Learning in Neural Networks—0
Connector-S: A Survey of Connectors in Multi-modal Large Language Models—0
Configurable Foundation Models: Building LLMs from a Modular Perspective—0
3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow—0
Conditional computation in neural networks: principles and research trends—0
On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating—0
On the Adaptation to Concept Drift for CTR Prediction—0
A Review of Sparse Expert Models in Deep Learning—0
Complexity Experts are Task-Discriminative Learners for Any Image Restoration—0
Filtered not Mixed: Stochastic Filtering-Based Online Gating for Mixture of Large Language Models—0
A Review of DeepSeek Models' Key Innovative Techniques—0
Show:102550
← PrevPage 8 of 27Next →

No leaderboard results yet.