[1] Verbraeken J, Wolting M, Katzy J, et al. A survey on distributed machine learning[J]. Acm computing surveys (csur), 2020, 53(2): 1-33.
[2] 王恩东, 闫瑞栋, 郭振华, 等. 分布式训练系统及其优化算法综述[J]. 计算机学报, 2024, 47(1): 1-28.
WANG En-Dong, YAN Rui-Dong, Guo Zhen-Hua, et al. A Survey of Distributed Traing System and Its Optimization Algorithms[J]. Chinese Journal of Computers, 2024, 47(1): 1-28.
[3] Navarro C A, Hitschfeld-Kahler N, Mateu L. A survey on parallel computing and its applications in data-parallel problems using GPU architectures[J]. Communications in computational physics, 2014, 15(2): 285-329.
[4] 聂飞, 江波. 基于软件定义的并行计算架构设计与实现[J]. 计算机工程, 2026, 52(8): 238-246.
NIE Fei, JIANG Bo. Design and Implementation of Parallel Computing Architecture Based on Software Definition[J]. Computer Engineering, 2026, 52(8): 238-246.
[5] Keuper J, Preundt F J. Distributed training of deep neural networks: Theoretical and practical limits of parallel scalability[C]//2016 2nd Workshop on Machine Learning in HPC Environments (MLHPC). IEEE, 2016: 19-26.
[6] 朱泓睿, 元国军, 姚成吉, 等. 分布式深度学习训练网络综述[J]. 计算机研究与发展, 2021, 58(1): 98-115.
Zhu Hong-Rui, Yuan Guo-Jun, Yao Cheng-Ji, et al. Survey on Network of Distributed Deep Learning Training[J]. Journal of Computer Research and Development, 2021, 58(1): 98-115.
[7] 董德尊, 欧阳硕. 分布式深度学习系统网络通信优化技术[J]. 中兴通讯技术, 2020.
Dong De-Zun, Ou Yang. Optimization Techniques of Network Communication in Distributed Deep Learning Systems[J]. ZTE Technology Journal, 2020.
[8] Al-Fares M, Loukissas A, Vahdat A. A scalable, commodity data center network architecture[C]//Proceedings of ACM SIGCOMM. New York: ACM, 2008: 63-74.
[9] Agarwal S, Rajakrishnan S, Narayan A, et al. Sincronia: Near-optimal network design for coflows[C]//Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication. 2018: 16-29.
[10] Chowdhury M, Stoica I. Coflow: A networking abstraction for cluster applications[C]//Proceedings of the 11th ACM Workshop on Hot Topics in Networks. 2012: 31-36.
[11] Chowdhury M, Zhong Y, Stoica I. Efficient coflow scheduling with varys[C]//Proceedings of the 2014 ACM conference on SIGCOMM. 2014: 443-454.
[12] NVIDIA. NVIDIA Collective Communications Library (NCCL) Technical Report[R]. USA: NVIDIA Corporation, 2024.
[13] Shah A, Chidambaram V, Cowan M, et al. {TACCL}: Guiding collective algorithm synthesis using communication sketches[C]//20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). 2023: 593-612.
[14] Liu T, Hei C, Li F, et al. ResCCL: Resource-efficient scheduling for collective communication[C]//Proceedings of the ACM SIGCOMM 2025 Conference. 2025: 55-70.
[15] Cao J, Shi S, Gao J, et al. Syccl: Exploiting symmetry for efficient collective communication scheduling[C]//Proceedings of the ACM SIGCOMM 2025 Conference. 2025: 645-662.
[16] Mahajan K, Chu C H, Sridharan S, et al. Better together: Jointly optimizing {ML} collective scheduling and execution planning using {SYNDICATE}[C]//20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). 2023: 809-824.
[17] Jiang Y, Zhu Y, Lan C, et al. A unified architecture for accelerating distributed {DNN} training in heterogeneous {GPU/CPU} clusters[C]//14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 2020: 463-479.
[18] Pan R, Lei Y, Li J, et al. Efficient flow scheduling in distributed deep learning training with echelon formation[C]//Proceedings of the 21st ACM Workshop on Hot Topics in Networks. 2022: 93-100.
[19] Wang W, Das S, Wu X C, et al. Mxdag: A hybrid abstraction for emerging applications[C]//Proceedings of the 20th ACM Workshop on Hot Topics in Networks. 2021: 221-228.
[20] Li J, Jiang Y, Zhu Y, et al. Accelerating distributed {MoE} training and inference with lina[C]//2023 USENIX Annual Technical Conference (USENIX ATC 23). 2023: 945-959.
[21] Rajasekaran S, Ghobadi M, Akella A. {CASSINI}:{Network-Aware} job scheduling in machine learning clusters[C]//21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). 2024: 1403-1420.
[22] Cao J, Guan Y, Qian K, et al. Crux: Gpu-efficient communication scheduling for deep learning training[C]//Proceedings of the ACM SIGCOMM 2024 Conference. 2024: 1-15.
[23] Xu H, Du C, Yu Z, et al. Gsched: Coordinated flow-control and priority scheduling for dnn training in ai cluster[J]. IEEE Transactions on Network Science and Engineering, 2026.
[24] Li M, Andersen D G, Smola A J, et al. Communication efficient distributed machine learning with the parameter server[J]. Advances in Neural Information Processing Systems, 2014, 27.
[25] Patarasuk P, Yuan X. Bandwidth optimal all-reduce algorithms for clusters of workstations[J]. Journal of Parallel and Distributed Computing, 2009, 69(2): 117-124.
[26] Chaskar H M, Madhow U. Fair scheduling with tunable latency: a round-robin approach[J]. IEEE/ACM Transactions on networking, 2003, 11(4): 592-601.
[27] Tabatabaee S M, Le Boudec J Y, Boyer M. Interleaved weighted round-robin: A network calculus analysis[J]. IEICE Transactions on Communications, 2021, 104(12): 1479-1493.
[28] Ozfatura E, Ulukus S, Gündüz D. Straggler-aware distributed learning: Communication–computation latency trade-off[J]. Entropy, 2020, 22(5): 544.
[29] Qian K, Xi Y, Cao J, et al. Alibaba hpn: A data center network for large language model training[C]//Proceedings of the ACM SIGCOMM 2024 Conference. 2024: 691-706.
[30] Open Networking Foundation. OpenFlow Switch Specification Version 1.3.5 [S]. Menlo Park, CA: Open Networking Foundation, 2015.
|