| 1 |
庞文豪, 王嘉伦, 翁楚良. GPGPU和CUDA统一内存研究现状综述. 计算机工程, 2024, 50 (12): 1- 15.
doi: 10.19678/j.issn.1000-3428.0068694
|
|
PANG W H , WANG J L , WENG C L . Survey on GPGPU and CUDA unified memory research status. Computer Engineering, 2024, 50 (12): 1- 15.
doi: 10.19678/j.issn.1000-3428.0068694
|
| 2 |
LINDHOLM E , NICKOLLS J , OBERMAN S , et al. NVIDIA Tesla: a unified graphics and computing architecture. IEEE Micro, 2008, 28 (2): 39- 55.
doi: 10.1109/MM.2008.31
|
| 3 |
STEFFEN M, ZAMBRENO J. Improving SIMT efficiency of global rendering algorithms with architectural support for dynamic micro-kernels[C]//Proceedings of the 43rd Annual IEEE/ACM International Symposium on Microarchitecture. Washington D.C., USA: IEEE Press, 2011: 237-248.
|
| 4 |
NUGTEREN C, VAN DEN BRAAK G J, CORPORAAL H. Future of GPGPU micro-architectural parameters[C]//Proceedings of the Design, Automation & Test in Europe Conference & Exhibition (DATE). Washington D.C., USA: IEEE Press, 2013: 392-395.
|
| 5 |
KHORASANI F, GUPTA R, BHUYAN L N. Efficient warp execution in presence of divergence with collaborative context collection[C]//Proceedings of the 48th International Symposium on Microarchitecture. New York, USA: ACM Press, 2015: 204-215.
|
| 6 |
RHU M, EREZ M. CAPRI: prediction of compaction-adequacy for handling control-divergence in GPGPU architectures[C]//Proceedings of the 39th Annual International Symposium on Computer Architecture (ISCA). Washington D.C., USA: IEEE Press, 2012: 61-71.
|
| 7 |
蒋奎. 面向嵌入式系统优化安全加固的程序控制流静态分析关键问题研究[D]. 西安: 西安电子科技大学, 2023.
|
|
JIANG K. Research on key problems of static analysis of program control flow for embedded system optimization security hardening[D]. Xi'an: XIDIAN UNIVERSITY, 2023. (in Chinese)
|
| 8 |
CHEN W K, LI B G, GUPTA R. Code compaction of matching single-entry multiple-exit regions[M]//Static Analysis. Berlin, Germany: Springer, 2003: 401-417.
|
| 9 |
COUTINHO B, SAMPAIO D, PEREIRA F M Q, et al. Divergence analysis and optimizations[C]//Proceedings of the International Conference on Parallel Architectures and Compilation Techniques. Washington D.C., USA: IEEE Press, 2011: 320-329.
|
| 10 |
SAUMYA C, SUNDARARAJAH K, KULKARNI M. DARM: control-flow melding for SIMT thread divergence reduction[C]//Proceedings of the IEEE/ACM International Symposium on Code Generation and Optimization (CGO). Washington D.C., USA: IEEE Press, 2022: 1-13.
|
| 11 |
SMITH T F , WATERMAN M S . Identification of common molecular subsequences. Journal of Molecular Biology, 1981, 147 (1): 195- 197.
doi: 10.1016/0022-2836(81)90087-5
|
| 12 |
LATTNER C, ADVE V. LLVM: a compilation framework for lifelong program analysis & transformation[C]//Proceedings of the International Symposium on Code Generation and Optimization. Washington D.C., USA: IEEE Press, 2004: 75-86.
|
| 13 |
LLVM Foundation. The LLVM compiler infrastructure[EB/OL]. [2024-09-20]. https://llvm.org/.
|
| 14 |
|
| 15 |
LOZANO R C, CARLSSON M, DREJHAMMAR F, et al. Constraint-based register allocation and instruction scheduling[M]//Principles and Practice of Constraint Programming. Berlin, Germany: Springer, 2012: 750-766.
|
| 16 |
杨太龙, 赵红朋, 张磊. 基于国产异构平台的奇异值分解法. 计算机工程, 2024, 50 (9): 216- 225.
doi: 10.19678/j.issn.1000-3428.0068183
|
|
YANG T L , ZHAO H P , ZHANG L . Singular value decomposition method based on domestic heterogeneous platforms. Computer Engineering, 2024, 50 (9): 216- 225.
doi: 10.19678/j.issn.1000-3428.0068183
|
| 17 |
LIU J , WU Z H , FU D L , et al. HeterPS: distributed deep learning with reinforcement learning based scheduling in heterogeneous environments. Future Generation Computer Systems, 2023 (c): 106- 117.
|
| 18 |
张军, 魏继桢, 沈凡凡, 等. 基于GPGPU-sim的多kernel场景下GPGPU性能优化实验方法. 实验技术与管理, 2024, 41 (7): 87- 93.
|
|
ZHANG J , WEI J Z , SHEN F F , et al. Experimental method for optimizing GPGPU performance in a multiple-kernel environment based on GPGPU-sim. Experimental Technology and Management, 2024, 41 (7): 87- 93.
|
| 19 |
|
| 20 |
CYTRON R , FERRANTE J , ROSEN B K , et al. Efficiently computing static single assignment form and the control dependence graph. ACM Transactions on Programming Languages and Systems, 1991, 13 (4): 451- 490.
doi: 10.1145/115372.115320
|
| 21 |
|
| 22 |
HUANG J C, LENG T. Generalized loop-unrolling: a method for program speedup[C]//Proceedings of IEEE Symposium on Application-Specific Systems and Software Engineering and Technology. Washington D.C., USA: IEEE Press, 1999: 51-60.
|
| 23 |
RODRIGUEZ-CANCIO M, COMBEMALE B, BAUDRY B. Automatic microbenchmark generation to prevent dead code elimination and constant folding[C]//Proceedings of the 31st IEEE/ACM International Conference on Automated Software Engineering. New York, USA: ACM Press, 2016: 132-143.
|
| 24 |
JIN Z M, VETTER J S. A benchmark suite for improving performance portability of the SYCL programming model[C]//Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). Washington D.C., USA: IEEE Press, 2023: 325-327.
|
| 25 |
|