research-article

Whippletree: task-based scheduling of dynamic workloads on the GPU

Authors:

Markus Steinberger,

Michael Kenzel,

Bernhard Kerbl,

Dieter SchmalstiegAuthors Info & Claims

ACM Transactions on Graphics (TOG), Volume 33, Issue 6

Article No.: 228, Pages 1 - 11

https://doi.org/10.1145/2661229.2661250

Published: 19 November 2014 Publication History

Abstract

In this paper, we present Whippletree, a novel approach to scheduling dynamic, irregular workloads on the GPU. We introduce a new programming model which offers the simplicity and expressiveness of task-based parallelism while retaining all aspects of the multi-level execution hierarchy essential to unlocking the full potential of a modern GPU. At the same time, our programming model lends itself to efficient implementation on the SIMD-based architecture typical of a current GPU. We demonstrate the practical utility of our model by providing a reference implementation on top of current CUDA hardware. Furthermore, we show that our model compares favorably to traditional approaches in terms of both performance as well as the range of applications that can be covered. We demonstrate the benefits of our model for recursive Reyes rendering, procedural geometry generation and volume rendering with concurrent irradiance caching.

Supplementary Material

ZIP File (a228.zip)

Supplemental material.

Download
45.11 MB

References

[1]

Aila, T., and Laine, S. 2009. Understanding the efficiency of ray traversal on GPUs. In Proc. HPG, 145--149.

Digital Library

[2]

Breitbart, J. 2011. Static GPU threads and an improved scan algorithm. In Proc. Euro-Par 2010, 373--380.

Digital Library

[3]

Cederman, D., and Tsigas, P. 2008. On dynamic load balancing on graphics processors. In Proc. Graphics Hardware, 57--64.

Digital Library

[4]

Chatterjee, S., Grossman, M., Sbirlea, A., and Sarkar, V. 2011. Dynamic task parallelism with a GPU work-stealing runtime system. In Proc. Languages and Compilers for Parallel Computing.

[5]

Chen, L., Villa, O., Krishnamoorthy, S., and Gao, G. 2010. Dynamic load balancing on single- and multi-gpu systems. In IEEE Parallel Distributed Processing.

[6]

Cook, R. L., Carpenter, L., and Catmull, E. 1987. The reyes image rendering architecture. SIGGRAPH Comput. Graph. 21, 4 (Aug.), 95--102.

Digital Library

[7]

Hargreaves, S. 2005. Generating shaders from HLSL fragments. ShaderX3: Advanced rendering with DirectX and OpenGL.

[8]

Hoberock, J., Lu, V., Jia, Y., and Hart, J. C. 2009. Stream compaction for deferred shading. In Proc. HPG, 173--180.

Digital Library

[9]

Kroes, T., Post, F. H., and Botha, C. P. 2012. Exposure render: An interactive photo-realistic volume rendering framework. PLoS ONE 7, 7 (07).

[10]

Laine, S., Karras, T., and Aila, T. 2013. Megakernels considered harmful: Wavefront path tracing on GPUs. In Proc. HPG.

Digital Library

[11]

Liu, F., Huang, M.-C., Liu, X.-H., and Wu, E.-H. 2010. Freepipe: A programmable parallel rendering architecture for efficient multi-fragment effects. In Proc. I3D, 75--82.

Digital Library

[12]

Luo, L., Wong, M., and Hwu, W.-M. 2010. An effective GPU implementation of breadth-first search. In Proc. Design Automation Conference, ACM, 52--55.

Digital Library

[13]

NVIDIA. 2012. CUDA Dynamic Parallelism Programming Guide.

[14]

Parker, S. G., Bigler, J., Dietrich, A., Friedrich, H., Hoberock, J., Luebke, D., McAllister, D., McGuire, M., Morley, K., Robison, A., and Stich, M. 2010. Optix: a general purpose ray tracing engine. ACM TOG 29, 4(66).

Digital Library

[15]

Patney, A., and Owens, J. D. 2008. Real-time Reyes-style adaptive surface subdivision. ACM TOG 27, 5(143).

Digital Library

[16]

Satish, N., Harris, M., and Garland, M. 2009. Designing efficient sorting algorithms for manycore GPUs. In Proc. IEEE Parallel&Distributed Processing.

Digital Library

[17]

Steinberger, M., Kainz, B., Kerbl, B., Hauswiesner, S., Kenzel, M., and Schmalstieg, D. 2012. Softshell: Dynamic scheduling on GPUs. ACM TOG 31, 6(161).

Digital Library

[18]

Steinberger, M., Kenzel, M., Kainz, B., Müller, J., Wonka, P., and Schmalstieg, D. 2014. Parallel generation of architecture on the GPU. In Computer Graphics Forum, vol. 33, 73--82.

Digital Library

[19]

Stuart, J. A., and Owens, J. D. 2009. Message passing on data-parallel architectures. In Proc. Parallel&Distributed Processing, IEEE.

Digital Library

[20]

Sugerman, J., Fatahalian, K., Boulos, S., Akeley, K., and Hanrahan, P. 2009. GRAMPS: A programming model for graphics pipelines. ACM TOG 28, 1, 4:1--4:11.

Digital Library

[21]

Tzeng, S., Patney, A., and Owens, J. D. 2010. Task management for irregular-parallel workloads on the GPU. In Proc. HPG, 29--37.

Digital Library

[22]

Xiao, S., and Feng, W. 2010. Inter-block GPU communication via fast barrier synchronization. In IEEE Parallel Distributed Processing.

[23]

Yan, S., Long, G., and Zhang, Y. 2013. Streamscan: fast scan algorithms for GPUs without global barrier synchronization. In ACM Principles and Practice of Parallel Programming, 229--238.

Digital Library

[24]

Zhou, K., Hou, Q., Ren, Z., Gong, M., Sun, X., and Guo, B. 2009. RenderAnts: interactive Reyes rendering on GPUs. ACM TOG 28, 5 (Dec.), 155:1--155:11.

Digital Library

Cited By

Kuth BOberberger MFaber CBaumeister DChajdas MMeyer Q(2024)Real-Time Procedural Generation with GPU Work GraphsProceedings of the ACM on Computer Graphics and Interactive Techniques10.1145/36753767:3(1-16)Online publication date: 9-Aug-2024
https://dl.acm.org/doi/10.1145/3675376
Durvasula SZhao AKiguru RGuan YChen ZVijaykumar N(2024)ACE: Efficient GPU Kernel Concurrency for Input-Dependent Irregular Computational GraphsProceedings of the 2024 International Conference on Parallel Architectures and Compilation Techniques10.1145/3656019.3676897(258-270)Online publication date: 14-Oct-2024
https://dl.acm.org/doi/10.1145/3656019.3676897
Turimbetov ISasongko MUnat D(2024)GPU-Initiated Resource Allocation for Irregular WorkloadsProceedings of the 3rd International Workshop on Extreme Heterogeneity Solutions10.1145/3642961.3643799(1-8)Online publication date: 2-Mar-2024
https://dl.acm.org/doi/10.1145/3642961.3643799
Show More Cited By

Index Terms

Whippletree: task-based scheduling of dynamic workloads on the GPU
1. Computing methodologies
  1. Computer graphics
    1. Graphics systems and interfaces
  2. Parallel computing methodologies
2. Human-centered computing
  1. Human computer interaction (HCI)
    1. Interaction techniques

Recommendations

Kokkos

The manycore revolution can be characterized by increasing thread counts, decreasing memory per thread, and diversity of continually evolving manycore architectures. High performance computing (HPC) applications and libraries must exploit increasingly ...
Improving performance of GPU code using novel features of the NVIDIA kepler architecture

Graphics processing unit GPU computing is a popular approach to simulating complex models and performing massive calculations. GPUs have attracted a great deal of interest because they offer both high performance and energy efficiency. Efficient General-...
SIMD Monte-Carlo Numerical Simulations Accelerated on GPU and Xeon Phi

The efficiency of a pleasingly parallel application is studied for several computing platforms. A real world problem, i.e., Monte-Carlo numerical simulations of stratospheric balloon envelope drift descent is considered. We detail the optimization of ...

Comments

Information & Contributors

Information

Published In

cover image ACM Transactions on Graphics

ACM Transactions on Graphics Volume 33, Issue 6

November 2014

704 pages

ISSN:0730-0301

EISSN:1557-7368

DOI:10.1145/2661229

Issue’s Table of Contents

Copyright © 2014 ACM.

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]

Publisher

Association for Computing Machinery

New York, NY, United States

Publication History

Published: 19 November 2014

Published in TOG Volume 33, Issue 6

Permissions

Request permissions for this article.

Request Permissions

Check for updates

Author Tags

Qualifiers

Research-article

Funding Sources

Austrian Science Fund

Contributors

Other Metrics

View Article Metrics

Bibliometrics & Citations

Bibliometrics

Article Metrics

65
Total Citations
View Citations
833
Total Downloads

Downloads (Last 12 months)66
Downloads (Last 6 weeks)3

Reflects downloads up to 15 Oct 2024

Other Metrics

View Author Metrics

Citations

Cited By

Kuth BOberberger MFaber CBaumeister DChajdas MMeyer Q(2024)Real-Time Procedural Generation with GPU Work GraphsProceedings of the ACM on Computer Graphics and Interactive Techniques10.1145/36753767:3(1-16)Online publication date: 9-Aug-2024
https://dl.acm.org/doi/10.1145/3675376
Durvasula SZhao AKiguru RGuan YChen ZVijaykumar N(2024)ACE: Efficient GPU Kernel Concurrency for Input-Dependent Irregular Computational GraphsProceedings of the 2024 International Conference on Parallel Architectures and Compilation Techniques10.1145/3656019.3676897(258-270)Online publication date: 14-Oct-2024
https://dl.acm.org/doi/10.1145/3656019.3676897
Turimbetov ISasongko MUnat D(2024)GPU-Initiated Resource Allocation for Irregular WorkloadsProceedings of the 3rd International Workshop on Extreme Heterogeneity Solutions10.1145/3642961.3643799(1-8)Online publication date: 2-Mar-2024
https://dl.acm.org/doi/10.1145/3642961.3643799
Park SHong JSong JKim HKim YLee JLee IChabbi MSteuwer M(2024)AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read MappingProceedings of the 29th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming10.1145/3627535.3638474(431-444)Online publication date: 2-Mar-2024
https://dl.acm.org/doi/10.1145/3627535.3638474
Ismayilov IBaydamirli JSağbili DWahib MUnat DGallivan KNikolopoulos DBeivide RGallopoulos E(2023)Multi-GPU Communication Schemes for Iterative Solvers: When CPUs are Not in ChargeProceedings of the 37th International Conference on Supercomputing10.1145/3577193.3593713(192-202)Online publication date: 21-Jun-2023
https://dl.acm.org/doi/10.1145/3577193.3593713
Nourazar MBooth BGoossens B(2023)A GPU optimization workflow for real-time execution of ultra-high frame rate computer vision applicationsJournal of Real-Time Image Processing10.1007/s11554-023-01384-721:1Online publication date: 26-Nov-2023
https://dl.acm.org/doi/10.1007/s11554-023-01384-7
Chen YBrock BPorumbescu SBuluc AYelick KOwens J(2022)Atos: A Task-Parallel GPU Scheduler for Graph AnalyticsProceedings of the 51st International Conference on Parallel Processing10.1145/3545008.3545056(1-11)Online publication date: 29-Aug-2022
https://dl.acm.org/doi/10.1145/3545008.3545056
Huang TLin DLin CLin Y(2022)Taskflow: A Lightweight Parallel and Heterogeneous Task Graph Computing SystemIEEE Transactions on Parallel and Distributed Systems10.1109/TPDS.2021.310425533:6(1303-1320)Online publication date: 1-Jun-2022
https://doi.org/10.1109/TPDS.2021.3104255
Wang QRen BChen JEdwards R(2022)MICCO: An Enhanced Multi-GPU Scheduling Framework for Many-Body Correlation Functions2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS)10.1109/IPDPS53621.2022.00022(135-145)Online publication date: May-2022
https://doi.org/10.1109/IPDPS53621.2022.00022
Stavrinides GKaratza H(2021)Orchestrating Bag-of-Tasks Applications with Dynamically Spawned Tasks in a Distributed Environment2021 International Symposium on Performance Evaluation of Computer and Telecommunication Systems (SPECTS)10.23919/SPECTS52716.2021.9639275(1-8)Online publication date: 19-Jul-2021
https://doi.org/10.23919/SPECTS52716.2021.9639275
Show More Cited By

View Options

Get Access

Login options

Check if you have access through your login credentials or your institution to get full access on this article.

Full Access

Get this Article

View options

PDF

View or Download as a PDF file.

eReader

View online with eReader.

Media

Figures

Other

Tables

View Issue’s Table of Contents