SOTAVerified

Routing for Large ML Models

2025-03-07Code Available0· sign in to hype

Ofir Cohen, Jose Yallouz Michael Schapira, Shahar Belkar, Tal Mizrahi

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

Training large language models (LLMs), and other large machine learning models, involves repeated communication of large volumes of data across a data center network. The communication patterns induced by these training process exhibit high regularity and persistence, giving rise to significant opportunities for optimizing the manner in which flows are routed across the network. We present an algorithmic framework for quantifying network-wide efficiency in the context of training LLMs (and other large-scale ML models), and for periodically optimizing routing with respect to this global metric.

Reproductions