SOTAVerified

Mixture-of-Experts

Papers

Showing 1051–1100 of 1312 papers

TitleStatusHype
Generalizing Multimodal Variational Methods to Sets—0
Mod-Squad: Designing Mixture of Experts As Modular Multi-Task Learners—0
Fixing MoE Over-Fitting on Low-Resource Languages in Multilingual Machine Translation—0
SMILE: Scaling Mixture-of-Experts with Efficient Bi-level Routing—0
Incorporating Polar Field Data for Improved Solar Flare Prediction—0
Named Entity and Relation Extraction with Multi-Modal Retrieval—0
Automatically Extracting Information in Medical Dialogue: Expert System And Attention for Labelling—0
Double Deep Q-Learning in Opponent Modeling—0
Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production—0
A Bird's-eye View of Reranking: from List Level to Page LevelCode0
HMOE: Hypernetwork-based Mixture of Experts for Domain Generalization—0
Handling Trade-Offs in Speech Separation with Sparsely-Gated Mixture of Experts—0
SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations—0
Using Deep Mixture-of-Experts to Detect Word Meaning Shift for TempoWiC—0
Safe Real-World Autonomous Driving by Learning to Predict and Plan with a Mixture of Experts—0
Contextual Mixture of Experts: Integrating Knowledge into Predictive Modeling—0
Prediction Sets for High-Dimensional Mixture of Experts Models—0
Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models—0
Coordination with Humans via Strategy Matching—0
On the Adversarial Robustness of Mixture of Experts—0
Tiny-Attention Adapter: Contexts Are More Important Than the Number of Parameters—0
FEAMOE: Fair, Explainable and Adaptive Mixture of Experts—0
Deep Learning Mixture-of-Experts Approach for Cytotoxic Edema Assessment in Infants and Children—0
Probabilistic partition of unity networks for high-dimensional regression problems—0
Parameter-varying neural ordinary differential equations with partition-of-unity networks—0
Table-based Fact Verification with Self-labeled Keypoint Alignment—0
Sparsity-Constrained Optimal Transport—0
Mixture of experts models for multilevel data: modelling framework and approximation theory—0
Tuning of Mixture-of-Experts Mixed-Precision Neural Networks—0
Diversified Dynamic Routing for Vision Tasks—0
Parameter-Efficient Conformers via Sharing Sparsely-Gated Experts for End-to-End Speech Recognition—0
Sparse Video Representation Using Steered Mixture-of-Experts With Global Motion Compensation—0
A Review of Sparse Expert Models in Deep Learning—0
ADMoE: Anomaly Detection with Mixture-of-Experts from Noisy Labels—0
Context-aware Mixture-of-Experts for Unbiased Scene Graph Generation—0
A Theoretical View on Sparsely Activated Networks—0
Edge-Aware Autoencoder Design for Real-Time Mixture-of-Experts Image Compression—0
Adaptive Mixture of Experts Learning for Generalizable Face Anti-Spoofing—0
MoEC: Mixture of Expert Clusters—0
Learning Large-scale Universal User Representation with Sparse Mixture of Experts—0
RoME: Role-aware Mixture-of-Expert Transformer for Text-to-Video RetrievalCode0
Scalable Neural Data Server: A Data Recommender for Transfer Learning—0
Adaptive Expert Models for Personalization in Federated LearningCode0
Quantitative Stock Investment by Routing Uncertainty-Aware Trading Experts: A Multi-Task Learning Approach—0
Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts—0
Interpretable Mixture of Experts—0
Task-Specific Expert Pruning for Sparse Mixture-of-Experts—0
Gating Dropout: Communication-efficient Regularization for Sparsely Activated Transformers—0
Automatic Expert Selection for Multi-Scenario and Multi-Task Search—0
Eliciting and Understanding Cross-Task Skills with Task-Level Mixture-of-ExpertsCode0
Show:102550
← PrevPage 22 of 27Next →

No leaderboard results yet.