z-logo
open-access-imgOpen Access
CUDA Memory Optimizations for Large Data-Structures in the Gravit Simulator
Author(s) -
Jakob Siegel,
Juergen Ributzka,
Xiaoming Li
Publication year - 2011
Publication title -
journal of algorithms and computational technology
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.234
H-Index - 13
eISSN - 1748-3026
pISSN - 1748-3018
DOI - 10.1260/1748-3018.5.2.341
Subject(s) - speedup , computer science , cuda , parallel computing , memory hierarchy , programmer , set (abstract data type) , operating system , programming language , cache
Modern GPUs open a completely new field to optimize embarrassingly parallel algorithms. Implementing an algorithm on a GPU confronts the programmer with a new set of challenges for program optimization. Especially tuning the program for the GPU memory hierarchy whose organization and performance implications are radically different from those of general purpose CPUs; and optimizing programs at the instruction-level for the GPU. In this paper we analyze different approaches for optimizing the memory usage and access patterns for GPUs and propose a class of memory layout optimizations that can take full advantage of the unique memory hierarchy of NVIDIA CUDA. Furthermore, we analyze some classical optimization techniques and how they effect the performance on a GPU. We used the Gravit gravity simulator to demonstrate these optimizations. The final optimized GPU version achieves a 87× speedup compared to the original CPU version. Almost 30% of this speedup are direct results of the optimizations discussed in this paper.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom