During AI training, GPUs need to frequently synchronize parameters. Network bandwidth and stability directly determine whether the cluster scale can be expanded or whether it can idle and wait. Networking based on the traditional computing cluster approach often fails to meet demand. This article summarizes the key design points of training cluster networks.

Understanding the two currents of AI traffic

The first stream is the collective communication traffic between computing nodes: bursty, highly synchronized, and extremely sensitive to delay and packet loss. Any slow link will drag down the entire group; the second stream is the read traffic of data sets flowing from storage to computing nodes: continuous, large throughput, and relatively tolerant to delays. The two streams of traffic should be physically or logically isolated, otherwise the long stream of data reading will crowd out the short burst of parameter synchronization.

Access layer design

Convergence and No Damage Guarantee

Spine-Leaf is used for the aggregation of the Parameter Network to provide non-blocking forwarding, and ECN and PFC lossless mechanisms are enabled end-to-end; when the cluster is expanded, the entire Parameter Network is maintained at the same rate to avoid shortcomings caused by the mixing of old and new equipment. The storage network can be independently networked and added according to the 100G module.LC-LC single mode dual core jumperPlan long-distance links.

The core of AI network design is global balance rather than single-point stacking. EB-LINK can supply high-speed modules and network cards in batches and write codes according to the equipment brand. For cluster link selection, please call 0755-83179002.