Logo image
Augmented Block Cimmino Distributed Algorithm for solving tridiagonal systems on GPU
Book chapter

Augmented Block Cimmino Distributed Algorithm for solving tridiagonal systems on GPU

Y.-C. Chen and C.-R. Lee
Advances in GPU Research and Practice, pp.233-246
09/2016

Abstract

GPU Numerical linear algebra Parallel algorithm Performance optimization Tridiagonal solver Computer Science (all)
Tridiagonal systems appear in many scientific and engineering problems, such as Alternating Direction Implicit methods, fluid simulation, and Poisson equation. This chapter presents the parallelization of the Augmented Block Cimmino Distributed method for solving tridiagonal systems on graphics processing units (GPUs). Because of the special structure of tridiagonal matrices, we investigate the boundary padding technique to eliminate the execution branches on GPUs. Various performance optimization techniques, such as memory coalescing, are also incorporated to further enhance the performance. We evaluate the performance of our GPU implementation and analyze the effectiveness of each optimization technique. Over 24 times speedups can be obtained on the GPU as compared to speedups on the CPU version.

Metrics

1 Record Views

Details

Logo image