Reference for learners¶
Glossary¶
Foundational CUDA terms (host, device, kernel, thread, grid, …) are defined in the Introduction to CUDA glossary and not repeated here. This module adds the memory-specific terms below.
- managed memory¶
Memory allocated with
cudaMallocManaged(C/C++) or themanagedattribute (Fortran). The CUDA runtime automatically migrates the data between host memory and device memory as needed (see page migration). Builds on UVA. Managed memory is the explicitly allocated form of unified memory.- UVA¶
Unified Virtual Addressing — a single address space shared by the CPU and all GPUs (available since CUDA 4). Managed memory builds on it to provide automatic data migration.
- explicit (manual) memory management¶
The classical approach of explicitly allocating device memory with
cudaMalloc(or thedeviceattribute in CUDA Fortran) and copying data between host and device withcudaMemcpy(or array assignment). NVIDIA’s Programming Guide calls this explicit memory management; episode 2 uses the shorter manual memory management.- pinned memory¶
Page-locked host memory that is guaranteed to stay in physical memory. It can be transferred directly over PCIe without a staging buffer (higher bandwidth) and is required for asynchronous copies. Allocated with
cudaMallocHost.- constant memory¶
A read-only memory space that resides in global memory but is served by a dedicated constant cache on each streaming multiprocessor. Suited to small (max 64 KB), read-only data read by all threads. Kernel arguments are also passed through constant memory (a separate bank, up to 32 KB in total on Volta and newer with CUDA 12.1 or later).
- unified memory¶
NVIDIA’s umbrella term for memory that both host and device can access through the same address; the CUDA runtime keeps the data where it is used. On most Linux systems it is obtained with an explicit allocation (managed memory); systems with full unified-memory support (HMM, ATS, Grace Hopper) extend it to every host allocation.
- page migration¶
The automatic movement of managed memory pages between host and device performed by the CUDA runtime, triggered on access.
Further reading¶
Additional resources for each episode, beyond the sources cited in the text. All links were checked on 2026-09-30.
Software setup¶
CUDA Installation Guide for Linux — installing the CUDA Toolkit and driver
NVIDIA CUDA Compiler Driver NVCC — all
nvccoptions, including the optimisation, debug and-lineinfoflags used in this moduleNVIDIA HPC SDK Documentation —
nvfortranand the CUDA Fortran toolchainCUDA Fortran Programming Guide — the language reference for the Fortran tabs of this module
cuda-samples — NVIDIA’s official example programs, useful to check that a fresh installation compiles and runs CUDA code, and as reference code for specific features
Episode 1: Unified memory¶
CUDA Programming Guide — Unified Memory — the in-depth chapter on managed memory, support levels, prefetching and memory advice
CUDA Fortran Programming Guide — Managed data — the
managedattribute and its restrictionsAn Even Easier Introduction to CUDA (NVIDIA Developer Blog) — a first CUDA program built on
cudaMallocManagedUnified Memory for CUDA Beginners (NVIDIA Developer Blog) — what happens on page faults and why prefetching helps
Maximizing Unified Memory Performance in CUDA (NVIDIA Developer Blog) —
cudaMemPrefetchAsyncandcudaMemAdvisein practice
Episode 2: Manual memory management¶
CUDA C++ Best Practices Guide — Data Transfer Between Host and Device — why transfers dominate and how to minimise them
CUDA Runtime API — Memory Management — reference pages for
cudaMalloc,cudaMemcpy,cudaMemset,cudaFreeand the other functions used in this moduleHow to Optimize Data Transfers in CUDA C/C++ (NVIDIA Developer Blog) — measuring transfer bandwidth
How to Optimize Data Transfers in CUDA Fortran (NVIDIA Developer Blog) — the same material for CUDA Fortran
CUDA Pro Tip: Write Flexible Kernels with Grid-Stride Loops (NVIDIA Developer Blog) — execution configurations that work for any vector size, relevant to the vector-addition exercise
Episode 3: Pinned and constant memory¶
CUDA Programming Guide — Page-Locked Host Memory — pinned memory, mapped memory and their restrictions
CUDA C++ Best Practices Guide — Pinned Memory — when pinning pays off and why not to pin everything
CUDA Programming Guide — Constant Memory — the constant cache and
__constant__variablesCUDA Fortran Programming Guide — Pinned arrays and Constant data — the
pinnedandconstantattributesCUDA Runtime API — Memory Management (see Episode 2 above) — also documents
cudaMallocHost,cudaHostAlloc,cudaMemcpyToSymbolandcudaMemcpyFromSymbol
Episode 4: Synchronization and summary¶
CUDA Programming Guide — Synchronizing CPU and GPU — why kernel launches are asynchronous and what
cudaDeviceSynchronizeguaranteesCUDA Programming Guide — Error Checking in CUDA — catching errors from asynchronous operations
CUDA Programming Guide — Asynchronous Execution — streams and stream synchronisation, the next step after this module
CUDA Runtime API — Device Management — reference page for
cudaDeviceSynchronizeCUDA C++ Best Practices Guide — Memory Optimizations — the memory-related performance advice that builds on this module’s summary table
How to Access Global Memory Efficiently in CUDA C/C++ Kernels (NVIDIA Developer Blog) — coalesced access to device memory
Textbooks¶
Kirk, D. B., Hwu, W. W., El Hajj, I.: Programming Massively Parallel Processors: A Hands-on Approach, 4th edition, Morgan Kaufmann, 2022 — the standard textbook; its memory chapters cover the trade-offs discussed here in depth
Cheng, J., Grossman, M., McKercher, T.: Professional CUDA C Programming, Wiley, 2014 — a practical treatment of the CUDA memory model with many measured examples
Wilt, N.: The CUDA Handbook: A Comprehensive Guide to GPU Programming, Addison-Wesley, 2013 — the author maintains a free online edition and the sample code on the linked site; its chapters on address spaces and host memory cover every host-memory variant of this module
Ruetsch, G., Fatica, M.: CUDA Fortran for Scientists and Engineers, 2nd edition, Morgan Kaufmann, 2024 — the reference for the CUDA Fortran tabs, including data transfers and memory types
Reading materials for further learning¶
The external sources cited throughout this module are collected here. Each is defined once in references.bib and cited with {cite} where relevant in the episodes.
NVIDIA Corporation. CUDA Programming Guide — Unified and System Memory. URL: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/understanding-memory.html#unified-and-system-memory (visited on 2026-09-05).
NVIDIA Corporation. CUDA Programming Guide — Unified Memory. URL: https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/unified-memory.html#unified-memory (visited on 2026-09-05).
NVIDIA Corporation. CUDA Programming Guide — Explicit Memory Management. URL: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/intro-to-cuda-cpp.html#explicit-memory-management (visited on 2026-09-05).
NVIDIA Corporation. CUDA Programming Guide — Page-Locked Host Memory. URL: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/understanding-memory.html#page-locked-host-memory (visited on 2026-09-05).
NVIDIA Corporation. CUDA Programming Guide — Constant Memory. URL: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/writing-cuda-kernels.html#writing-cuda-kernels-constant-memory (visited on 2026-09-05).