Abstract
In this thesis, the implementation of GPU with CUDA architecture on lattice Boltzmann model is presented. Simulations of 3D lid-driven cavity flow is conducted as a test case with multi-GPU. The optimization of multi-GPU with two-dimensional domain decomposition is also discussed here. The numerical results are validated with benchmark solutions and the performance of the GPU implementation is also discussed. In the present work, we can achieve 17664.23 MLUPS for 384X384X384 grids with 96 nVIDIAr Tesla M2070 GPU cards.