SOTAVerified

Mixture-of-Experts

Papers

Showing 1101–1150 of 1312 papers

TitleStatusHype
Sparse Mixers: Combining MoE and Mixing to build a more efficient BERT—0
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services—0
Pluralistic Image Completion with Probabilistic Mixture-of-Experts—0
Unified Modeling of Multi-Domain Multi-Device ASR Systems—0
ST-ExpertNet: A Deep Expert Framework for Traffic Prediction—0
Optimizing Mixture of Experts using Dynamic Recompilations—0
How Can Cross-lingual Knowledge Contribute Better to Fine-Grained Entity Typing?—0
On the Representation Collapse of Sparse Mixture of Experts—0
Residual Mixture of Experts—0
Towards Efficient Single Image Dehazing and Desnowing—0
Table-based Fact Verification with Self-adaptive Mixture of ExpertsCode0
Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners—0
Mixture of Experts for Biomedical Question Answering—0
Mixture-of-experts VAEs can disregard variation in surjective multimodal data—0
Learning to Adapt Clinical Sequences with Residual Mixture of ExpertsCode0
Combining Spectral and Self-Supervised Features for Low Resource Speech Recognition and Translation—0
On the Adaptation to Concept Drift for CTR Prediction—0
Efficient Reflectance Capture with a Deep Gated Mixture-of-Experts—0
Build a Robust QA System with Transformer-based Mixture of ExpertsCode0
Efficient Language Modeling with Sparse all-MLP—0
SkillNet-NLU: A Sparsely Activated Model for General-Purpose Natural Language Understanding—0
Functional mixture-of-experts for classification—0
Mixture-of-Experts with Expert Choice Routing—0
A Survey on Dynamic Neural Networks for Natural Language Processing—0
Physics-Guided Problem Decomposition for Scaling Deep Learning of High-dimensional Eigen-Solvers: The Case of Schrödinger's Equation—0
One Student Knows All Experts Know: From Sparse to Dense—0
MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation—0
Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners—0
DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleCode0
Towards Lightweight Neural Animation : Exploration of Neural Network Pruning in Mixture of Experts-based Animation Models—0
Combinations of Adaptive Filters—0
Efficient Large Scale Language Modeling with Mixtures of Experts—0
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts—0
Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition—0
Specializing Versatile Skill Libraries using Local Mixture of ExpertsCode0
Anchoring to Exemplars for Training Mixture-of-Expert Cell Embeddings—0
A Mixture of Expert Based Deep Neural Network for Improved ASR—0
TAL: Two-stream Adaptive Learning for Generalizable Person Re-identification—0
Expert Aggregation for Financial Forecasting—0
SpeechMoE2: Mixture-of-Experts Model with Improved Routing—0
Table-based Fact Verification with Self-adaptive Mixture of Experts—0
MoEfication: Conditional Computation of Transformer Models for Efficient Inference—0
StableMoE: Stable Routing Strategy for Mixture of Experts—0
M6-T: Exploring Sparse Expert Models and Beyond—0
SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization—0
Elucidating Robust Learning with Uncertainty-Aware Corruption Pattern EstimationCode0
RTM Super Learner Results at Quality Estimation Task—0
Polynomial-Spline Neural Networks with Exact Integrals—0
P-Adapters: Robustly Extracting Factual Information from Language Models with Diverse Prompts—0
Simple or Complex? Complexity-Controllable Question Generation with Soft Templates and Deep Mixture of Experts Model—0
Show:102550
← PrevPage 23 of 27Next →

No leaderboard results yet.