| 1 |
ZHAO G , SUN N W , SHEN S , et al. GPU-accelerated target strength prediction based on multiresolution shooting and bouncing ray method. Applied Sciences, 2022, 12 (12): 6119.
doi: 10.3390/app12126119
|
| 2 |
GOLOSIO B , VILLAMAR J , TIDDIA G , et al. Runtime construction of large-scale spiking neuronal network models on GPU devices. Applied Sciences, 2023, 13 (17): 9598.
doi: 10.3390/app13179598
|
| 3 |
HU Y X, LIU Y H, LIU Z Y. A survey on convolutional neural network accelerators: GPU, FPGA and ASIC[C]//Proceedings of the 14th International Conference on Computer Research and Development (ICCRD). Shenzhen, China: IEEE Press, 2022: 100-107.
|
| 4 |
CHEN Y P, DAI X Y, LIU M C, et al. Dynamic convolution: attention over convolution kernels[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA: IEEE Press, 2020: 11027-11036.
|
| 5 |
KWON H , CHATARASI P , SARKAR V , et al. Maestro: a data-centric approach to understand reuse, performance, and hardware cost of dnn mappings. IEEE Micro, 2020, 40 (3): 20- 29.
doi: 10.1109/MM.2020.2985963
|
| 6 |
XIE X R , LIN J , WANG Z F , et al. An efficient and flexible accelerator design for sparse convolutional neural networks. IEEE Transactions on Circuits and Systems I: Regular Papers, 2021, 68 (7): 2936- 2949.
doi: 10.1109/TCSI.2021.3074300
|
| 7 |
JIA Z, MAGGIONI M, STAIGER B, et al. Dissecting the NVIDIA volta GPU architecture via microbenchmarking[EB/OL]. [2024-10-05]. http://arxiv.org/abs/1804.06826.
|
| 8 |
ALWAN E H , KETRAN R M , HUSSEIN I A . A comprehensive survey on loop unrolling technique in code optimization. Journal of University of Babylon for Pure and Applied Sciences, 2024, 32 (1): 108- 117.
|
| 9 |
DAGHAGHI S, MEISBURGER N, ZHAO M N, et al. Accelerating slide deep learning on modern CPUs: vectorization, quantizations, memory optimizations, and more[EB/OL]. [2024-10-05]. http://arxiv.org/abs/2103.10891.
|
| 10 |
庞文豪, 王嘉伦, 翁楚良. GPGPU和CUDA统一内存研究现状综述. 计算机工程, 2024, 50 (12): 1- 15.
doi: 10.19678/j.issn.1000-3428.0068694
|
|
PANG W H , WANG J L , WENG C L . Survey on GPGPU and CUDA unified memory research status. Computer Engineering, 2024, 50 (12): 1- 15.
doi: 10.19678/j.issn.1000-3428.0068694
|
| 11 |
曹义魁, 陆忠华, 张鉴, 等. 面向国产加速器的CFD核心算法并行优化. 数据与计算发展前沿, 2021, 3 (4): 93- 103.
|
|
CAO Y K , LU Z H , ZHANG J , et al. Parallel optimization of CFD core algorithms based on domestic processor. Frontiers of Data & Computing, 2021, 3 (4): 93- 103.
|
| 12 |
|
| 13 |
SUN W , LI A , GENG T , et al. Dissecting tensor cores via microbenchmarks: latency, throughput and numeric behaviors. IEEE Transactions on Parallel and Distributed Systems, 2023, 34 (1): 246- 261.
doi: 10.1109/TPDS.2022.3217824
|
| 14 |
SUN W , LI A , STUIJK S , et al. How much can we gain from tensor kernel fusion on GPUs?. IEEE Access, 2024, 12, 126135- 126144.
doi: 10.1109/ACCESS.2024.3411473
|
| 15 |
JIANG Z H, ZHENG L M, YAN E, et al. TVM: an automated optimizing compiler for deep learning[C]//Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). [S. l. ]: USENIX, 2018: 578-594.
|
| 16 |
KATEL N, KHANDELWAL V, BONDHUGULA U. MLIR-based code generation for GPU tensor cores[C]//Proceedings of the 31st ACM SIGPLAN International Conference on Compiler Construction. New York, USA: ACM Press, 2022: 117-128.
|
| 17 |
TRIPATHI M . Analysis of convolutional neural network based image classification techniques. Journal of Innovative Image Processing, 2021, 3 (2): 100- 117.
doi: 10.36548/jiip.2021.2.003
|
| 18 |
ZHANG Z Y, ZHANG P F, XU Z P, et al. Im2col-winograd: an efficient and flexible fused-Winograd convolution for NHWC format on GPUs[C]//Proceedings of the 53rd International Conference on Parallel Processing. New York, USA: ACM Press, 2024: 1072-1081.
|
| 19 |
HIGHAM N J , MARY T . Mixed precision algorithms in numerical linear algebra. Acta Numerica, 2022, 31, 347- 414.
doi: 10.1017/S0962492922000022
|
| 20 |
XU R, MA S, GUO Y. Performance analysis of different convolution algorithms in GPU environment[C]//Proceedings of the IEEE International Conference on Networking, Architecture and Storage (NAS). Chongqing, China: IEEE Press, 2018: 1-10.
|
| 21 |
SHEVGUNOV T , EFIMOV E , GUSCHINA O . Estimation of a spectral correlation function using a time-smoothing cyclic periodogram and FFT interpolation—2N-FFT algorithm. Sensors, 2022, 23 (1): 215.
doi: 10.3390/s23010215
|
| 22 |
童敢, 黄立波. Winograd快速卷积相关研究综述. 计算机科学与探索, 2022, 16 (5): 959- 971.
|
|
TONG G , HUANG L B . Review of Winograd fast convolution technique research. Journal of Frontiers of Computer Science and Technology, 2022, 16 (5): 959- 971.
|
| 23 |
NAKASATO N . A fast GEMM implementation on the cypress GPU. ACM SIGMETRICS Performance Evaluation Review, 2011, 38 (4): 50- 55.
doi: 10.1145/1964218.1964227
|
| 24 |
李茂文, 曲国远, 魏大洲, 等. 面向GPU计算平台的神经网络卷积性能优化. 计算机研究与发展, 2022, 59 (6): 1181- 1191.
|
|
LI M W , QU G Y , WEI D Z , et al. Performance optimization of neural network convolution based on GPU platform. Journal of Computer Research and Development, 2022, 59 (6): 1181- 1191.
|
| 25 |
KORCH M, RAITHEL P, WERNER T. Implementation and optimization of a 1D2V PIC method for nonlinear kinetic models on GPUs[C]//Proceedings of the 28th Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP). Västerås, Sweden: IEEE Press, 2020: 30-37.
|
| 26 |
ZACHARIADIS O , SATPUTE N , GÓMEZ-LUNA J , et al. Accelerating sparse matrix-matrix multiplication with GPU Tensor Cores. Computers & Electrical Engineering, 2020, 88, 106848.
|
| 27 |
MARKIDIS S, DER CHIEN S W, LAURE E, et al. NVIDIA tensor core programmability, performance & precision[C]// Proceedings of the IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). Vancouver, Canada: IEEE Press, 2018: 522-531.
|
| 28 |
NEMATOLLAHI N , SADROSADATI M , FALAHATI H , et al. Efficient nearest-neighbor data sharing in GPUs. ACM Transactions on Architecture and Code Optimization, 2021, 18 (1): 1- 26.
|
| 29 |
BASSO P M , DOS SANTOS F F , RECH P . Impact of tensor cores and mixed precision on the reliability of matrix multiplication in GPUs. IEEE Transactions on Nuclear Science, 2020, 67 (7): 1560- 1565.
doi: 10.1109/TNS.2020.2977583
|
| 30 |
WILLEMSEN F J , SCHOONHOVEN R , FILIPOVIČ J , et al. A methodology for comparing optimization algorithms for auto-tuning. Future Generation Computer Systems, 2024, 159, 489- 504.
doi: 10.1016/j.future.2024.05.021
|
| 31 |
刘仲, 李程, 田希, 等. MVSim: 面向VLIW多核向量处理器的快速、可扩展和精确的体系结构模拟器. 计算机工程与科学, 2024, 46 (2): 191- 199.
|
|
LIU Z , LI C , TIAN X , et al. MVSim: a fast, scalable and accurate architecture simulator for VLIW multi-core vector processors. Computer Engineering & Science, 2024, 46 (2): 191- 199.
|
| 32 |
ITO Y , NAKANO K . A GPU implementation of dynamic programming for the optimal polygon triangulation. IEICE Transactions on Information and Systems, 2013, E96.D (12): 2596- 2603.
doi: 10.1587/transinf.E96.D.2596
|