Location via proxy:   [ UP ]  
[Report a bug]   [Manage cookies]                
skip to main content
research-article

Characterizing and Exploiting Soft Error Vulnerability Phase Behavior in GPU Applications

Published: 01 January 2022 Publication History

Abstract

System reliability has become a first-class design constraint. As the use of Graphics Processing Units (GPU) continues to increase in compute applications, including High Performance Computing (HPC) and safety-critical applications, so do the number of transient faults in GPUs. However, our understanding of the potential impact of transient fault propagation in GPU applications remains limited. This study shows that the resilience characteristics of GPU programs change significantly during program execution and these characteristics show repetitive, time-varying behavior. Interestingly, these repetitive, time-varying, resilience characteristics of GPU programs do not align or correlate well with the performance phases of GPU programs. Furthermore, this work discovers and validates that temporal changes in the vulnerability behavior during a kernel execution tends to coincide with changes in basic block execution paths. Finally, we demonstrate how these observations can be exploited to accelerate the fault injection campaigns for reliability assessment of GPU programs by an order of magnitude and open opportunities for designing other effective resilience mitigation strategies.

Cited By

View all
  • (2024)ApproxDup: Developing an Approximate Instruction Duplication Mechanism for Efficient SDC Detection in GPGPUsIEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems10.1109/TCAD.2023.333082143:4(1051-1064)Online publication date: 1-Apr-2024
  • (2023)Photon: A Fine-grained Sampled Simulation Methodology for GPU WorkloadsProceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture10.1145/3613424.3623773(1227-1241)Online publication date: 28-Oct-2023
  • (2022)GPU Devices for Safety-Critical Systems: A SurveyACM Computing Surveys10.1145/354952655:7(1-37)Online publication date: 15-Dec-2022
  • Show More Cited By

Recommendations

Comments

Information & Contributors

Information

Published In

cover image IEEE Transactions on Dependable and Secure Computing
IEEE Transactions on Dependable and Secure Computing  Volume 19, Issue 1
Jan.-Feb. 2022
716 pages

Publisher

IEEE Computer Society Press

Washington, DC, United States

Publication History

Published: 01 January 2022

Qualifiers

  • Research-article

Contributors

Other Metrics

Bibliometrics & Citations

Bibliometrics

Article Metrics

  • Downloads (Last 12 months)0
  • Downloads (Last 6 weeks)0
Reflects downloads up to 16 Feb 2025

Other Metrics

Citations

Cited By

View all
  • (2024)ApproxDup: Developing an Approximate Instruction Duplication Mechanism for Efficient SDC Detection in GPGPUsIEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems10.1109/TCAD.2023.333082143:4(1051-1064)Online publication date: 1-Apr-2024
  • (2023)Photon: A Fine-grained Sampled Simulation Methodology for GPU WorkloadsProceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture10.1145/3613424.3623773(1227-1241)Online publication date: 28-Oct-2023
  • (2022)GPU Devices for Safety-Critical Systems: A SurveyACM Computing Surveys10.1145/354952655:7(1-37)Online publication date: 15-Dec-2022
  • (2022)Studying error propagation on application data structure and hardwareThe Journal of Supercomputing10.1007/s11227-022-04625-x78:17(18691-18724)Online publication date: 1-Nov-2022
  • (2021)G-SEPMProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis10.1145/3458817.3476170(1-15)Online publication date: 14-Nov-2021

View Options

View options

Figures

Tables

Media

Share

Share

Share this Publication link

Share on social media