[1] 李柯泉,陈燕,刘佳晨,等. 基于深度学习的目标检测算法综述[J]. 计算机工程, 2022, 48(07):1-12.
LI K Q, CHEN Y, LIU J C, et al. Survey of object detection algorithms based on deep learning[J]. Computer Engineering, 2022, 48(07): 1-12. (in Chinese)
[2] Otter D W, Medina J R, Kalita J K. A survey of the usages of deep learning for natural language processing[J]. IEEE transactions on neural networks and learning systems, 2020, 32(2): 604-624.
[3] He K, Zhang X, Ren S, et al. Deep residual learning for image recognition[C]// Proceedings of the IEEE conference on computer vision and pattern recognition. Washington, DC, USA. IEEE Press, 2016: 770-778.
[4] Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale[J]. arXiv preprint arXiv, 2020, 1-22.
[5] Devlin J, Chang M W, Lee K, et al. Bert: Pre-training of deep bidirectional transformers for language understanding[C]// Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies. Minneapolis, Minnesota, ACL Press, 2019: 4171-4186.
[6] Brown T, Mann B, Ryder N, et al. Language models are few-shot learners[J]. Advances in neural information processing systems, 2020, 33: 1877-1901.
[7] LeCun Y, Bottou L, Bengio Y, et al. Gradient-based learning applied to document recognition[J]. Proceedings of the IEEE, 2002, 86(11): 2278-2324.
[8] Floridi L, Chiriatti M. GPT-3: Its nature, scope, limits, and consequences[J]. Minds and machines, 2020, 30(4): 681-694.
[9] Chen Y, Xie Y, Song L, et al. A survey of accelerator architectures for deep neural networks[J]. Engineering, 2020, 6(3): 264-274.
[10] 章晋睿,龙婷婷,张德宇,等. 端智能推理加速技术综述[J]. 电子学报, 2025, 53(04): 1063-1102.
ZHANG J R, LONG T T, ZHANG D Y, et al. Survey of inference acceleration technologies for edge intelligence[J]. Acta Electronica Sinica, 2025, 53(04): 1063-1102. (in Chinese)
[11] 尹经纬,李志强,刘裕彤.大语言模型混合量化压缩与加速推理技术[J]. 计算机工程与设计, 2026, 47(01): 187-194.
YIN J W, LI Z Q, LIU Y T. Hybrid quantization compression and accelerated inference techniques for large language models[J]. Computer Engineering and Design, 2026, 47(01): 187-194. (in Chinese)
[12] Ma L, Xie Z, Yang Z, et al. Rammer: Enabling holistic deep learning compiler optimizations with rTasks[C]// 14th USENIX Symposium on Operating Systems Design and Implementation. Berkeley, CA, USA. USENIX Association, 2020: 881-897.
[13] Zhang C, Ma L, Xue J, et al. Cocktailer: Analyzing and optimizing dynamic control flow in deep learning[C]// 17th USENIX Symposium on Operating Systems Design and Implementation. Berkeley, CA, USA. USENIX Association, 2023: 681-699.
[14] Liu L, Deng J. Dynamic deep neural networks: Optimizing accuracy-efficiency trade-offs by selective execution[C]// Proceedings of the AAAI conference on artificial intelligence. Washington, DC, USA. AAAI Press, 2018, 32(1).
[15] Kwon W, Yu G I, Jeong E, et al. Nimble: Lightweight and parallel gpu task scheduling for deep learning[J]. Advances in Neural Information Processing Systems, 2020, 33: 8343-8354.
[16] Chen A, Xu F, Han L, et al. Opara: Exploiting operator parallelism for expediting dnn inference on gpus[J]. IEEE Transactions on Computers, 2024, 74(1): 325-333.
[17] Chen T, Moreau T, Jiang Z, et al. {TVM}: An automated {End-to-End} optimizing compiler for deep learning[C]// 13th USENIX symposium on operating systems design and implementation. Berkeley, CA, USA. USENIX Association, 2018: 578-594.
[18] OpenXLA Project. XLA[EB/OL]. [2024-12-03]. https://openxla.org/xla.
[19] Niu W, Guan J, Wang Y, et al. Dnnfusion: accelerating deep neural networks execution with advanced operator fusion[C]// Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation. New York, USA: ACM Press, 2021: 883-898.
[20] Ansel J, Yang E, He H, et al. PyTorch 2: Faster machine learning through dynamic Python bytecode transformation and graph compilation[C]// Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. New York: ACM, 2024: 929-947.
[21] NVIDIA Corporation. NVIDIA TensorRT [EB/OL]. [2026-03-26]. https://docs.nvidia.com/deeplearning/tensorrt.
[22] Microsoft. ONNX Runtime[EB/OL]. [2026-03-26]. https://onnxruntime.ai/docs/.
[23] Shi Y, Yang Z, Xue J, et al. Welder: Scheduling deep learning memory access via tile-graph[C]// 17th USENIX Symposium on Operating Systems Design and Implementation. Berkeley, CA, USA. USENIX Association, 2023: 701-718.
[24] Recasens P G, Agullo F, Zhu Y, et al. Mind the memory gap: Unveiling gpu bottlenecks in large-batch llm inference[C]// 2025 IEEE 18th International Conference on Cloud Computing. Washington, DC, USA. IEEE Press, 2025: 277-287.
[25] Chen F, Cheng Y, Wang L, et al. MetaAttention: A unified and performant attention framework across hardware backends[C]// Proceedings of the 31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming. Sydney, Australia: ACM Press, 2026: 635-647.
[26] Zhu K, Gao Y, Zhao Y, et al. NanoFlow: Towards optimal large language model serving throughput[C]// 19th USENIX Symposium on Operating Systems Design and Implementation. Boston, MA, USA: USENIX Association, 2025: 749-765.
[27] Ding Y, Zhu L, Jia Z, et al. Ios: Inter-operator scheduler for cnn acceleration[J]. Proceedings of Machine Learning and Systems, 2021, 3: 167-180.
[28] Zhao Y, Sun Q, He Z, Bai Y, Yu B. AutoGraph: Optimizing DNN Computation Graph for Parallel GPU Kernel Execution[C]// Proceedings of the AAAI Conference on Artificial Intelligence. Washington, DC, USA: AAAI Press, 2023, 37(9): 11354-11362.
|