CUDA Memory Management

Overview

On GPUs the bottleneck is rarely arithmetic — it is getting the data there. A single transfer across PCIe can easily cost more time than the kernel that consumes it. CUDA offers several strategies for this — device, managed, pinned and constant memory — and picking the wrong one costs more performance than kernel tuning can win back afterwards. The underlying trade-offs (explicit copies versus automatic migration, page-locked transfers, keeping read-only data close to the compute units) are not vendor-specific and transfer to other accelerators.

NVIDIA GPUs and host CPUs have separate memory spaces, and choosing the right strategy for moving and placing data between them is one of the most important decisions in CUDA programming. This module walks through the main CUDA memory management strategies — device memory, unified memory, pinned host memory, and constant memory — and the synchronisation rules that go with them, so that learners can pick the right tool for their use case.

After completing the module, learners can allocate data in each of these memory types, move it between host and device explicitly or let the CUDA runtime migrate it, and place the synchronisation calls that make data written by an asynchronously running kernel safe to read on the host. The material consists of a software-setup episode, four lesson episodes — unified memory, explicit (manual) memory management, pinned and constant memory, and a synchronisation recap with a summary of all memory kinds — and a quiz. All code is shown in CUDA C/C++ and CUDA Fortran side by side, and the two longer episodes contain hands-on exercises that are compiled and run by the learners.

Prerequisites

  • The Introduction to CUDA module, or equivalent knowledge of CUDA kernels, execution configuration and error handling

  • Access to an NVIDIA GPU with CUDA Toolkit 12.x or newer installed (12.x if the GPU has compute capability 7.0, see the software-setup episode), either on your own machine or on a cluster, to compile and run the exercises

  • The ability to edit and compile a program on the machine that has the GPU — from a Linux shell (for example this tutorial) or from an IDE that can call nvcc/nvfortran: the exercises are downloaded, edited and compiled by you

This module is aimed at researchers, engineers and students who can already write sequential programs in C, C++ or Fortran and want to start programming NVIDIA GPUs. It assumes the content of the Introduction to CUDA module.

Software setup

Learning outcomes

After completing this module, learners will be able to:

  • Choose between device memory, managed memory, pinned memory and constant memory for a given workload

  • Allocate and free managed memory using cudaMallocManaged in C/C++ (and the managed attribute in CUDA Fortran), and explain how data migration works

  • Allocate device memory with cudaMalloc or device in CUDA Fortran and copy data between host and device

  • Allocate pinned host memory with cudaMallocHost (or pinned in CUDA Fortran) and explain when pinned memory is required (e.g. for asynchronous copies)

  • Use constant memory for small read-only data accessed by many threads, and stay within its 64 KB limit

  • Insert cudaDeviceSynchronize() correctly before reading GPU-modified data on the host, and recognise the race conditions that occur otherwise

Estimated commitment: about 2.5 hours (150 minutes of lecture and exercises; provisional estimate, to be re-measured at the first delivery).

EVITA skill tree: SD1.2.6 GPU Programming with CUDA — the module covers the refinement SD1.2.6.2 CUDA Memory Management (proposed).

Related topics not covered here:

  • Streams and asynchronous copies

  • Shared memory

  • Atomics

  • Profiling

  • The vendor-neutral view of heterogeneous memory management, covered by the EVITA PP.HS1 modules (e.g. PP.HS1-K1.5.4 Memory management in heterogeneous systems)

See also

Credit

This module is based on the CUDA course developed at HLRS, the High-Performance Computing Center Stuttgart (University of Stuttgart). Authors: Tobias Haas and Jasper Seehofer, HLRS.

If you want to reuse the material beyond the terms of the licence below, or if you find errors, please contact the authors through the module repository.

Related EVITA module: Introduction to CUDA

License