SOTAVerified

Mixture-of-Experts

Papers

Showing 426450 of 1312 papers

TitleStatusHype
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production0
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems0
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures0
PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt TuningCode0
AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale0
PWC-MoE: Privacy-Aware Wireless Collaborative Mixture of Experts0
UMoE: Unifying Attention and FFN with Shared Experts0
FreqMoE: Dynamic Frequency Enhancement for Neural PDE Solvers0
Seed1.5-VL Technical Report0
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts0
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration0
FloE: On-the-Fly MoE Inference on Memory-constrained GPU0
Divide-and-Conquer: Cold-Start Bundle Recommendation via Mixture of Diffusion Experts0
SToLa: Self-Adaptive Touch-Language Framework with Tactile Commonsense Reasoning in Open-Ended Scenarios0
LLM-e Guess: Can LLMs Capabilities Advance Without Hardware Progress?Code0
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs0
Faster MoE LLM Inference for Extremely Large Models0
3D Gaussian Splatting Data Compression with Mixture of Priors0
STAR-Rec: Making Peace with Length Variance and Pattern Diversity in Sequential Recommendation0
Towards Smart Point-and-Shoot Photography0
Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques0
Finger Pose Estimation for Under-screen Fingerprint SensorCode0
Multimodal Deep Learning-Empowered Beam Prediction in Future THz ISAC Systems0
Perception-Informed Neural Networks: Beyond Physics-Informed Neural Networks0
CoCoAFusE: Beyond Mixtures of Experts via Model Fusion0
Show:102550
← PrevPage 18 of 53Next →

No leaderboard results yet.