Reference for learners

Glossary

Foundational CUDA terms (host, device, kernel, thread, grid, …) are defined in the Introduction to CUDA glossary and not repeated here. This module adds the memory-specific terms below.

managed memory

Memory allocated with cudaMallocManaged (C/C++) or the managed attribute (Fortran). The CUDA runtime automatically migrates the data between host memory and device memory as needed (see page migration). Builds on UVA. Managed memory is the explicitly allocated form of unified memory.

UVA

Unified Virtual Addressing — a single address space shared by the CPU and all GPUs (available since CUDA 4). Managed memory builds on it to provide automatic data migration.

explicit (manual) memory management

The classical approach of explicitly allocating device memory with cudaMalloc (or the device attribute in CUDA Fortran) and copying data between host and device with cudaMemcpy (or array assignment). NVIDIA’s Programming Guide calls this explicit memory management; episode 2 uses the shorter manual memory management.

pinned memory

Page-locked host memory that is guaranteed to stay in physical memory. It can be transferred directly over PCIe without a staging buffer (higher bandwidth) and is required for asynchronous copies. Allocated with cudaMallocHost.

constant memory

A read-only memory space that resides in global memory but is served by a dedicated constant cache on each streaming multiprocessor. Suited to small (max 64 KB), read-only data read by all threads. Kernel arguments are also passed through constant memory (a separate bank, up to 32 KB in total on Volta and newer with CUDA 12.1 or later).

unified memory

NVIDIA’s umbrella term for memory that both host and device can access through the same address; the CUDA runtime keeps the data where it is used. On most Linux systems it is obtained with an explicit allocation (managed memory); systems with full unified-memory support (HMM, ATS, Grace Hopper) extend it to every host allocation.

page migration

The automatic movement of managed memory pages between host and device performed by the CUDA runtime, triggered on access.

Further reading

Additional resources for each episode, beyond the sources cited in the text. All links were checked on 2026-09-30.

Software setup

Episode 1: Unified memory

Episode 2: Manual memory management

Episode 3: Pinned and constant memory

Episode 4: Synchronization and summary

Textbooks

Reading materials for further learning

The external sources cited throughout this module are collected here. Each is defined once in references.bib and cited with {cite} where relevant in the episodes.

[1]

NVIDIA Corporation. CUDA Programming Guide — Unified and System Memory. URL: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/understanding-memory.html#unified-and-system-memory (visited on 2026-09-05).

[2]

NVIDIA Corporation. CUDA Programming Guide — Unified Memory. URL: https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/unified-memory.html#unified-memory (visited on 2026-09-05).

[3]

NVIDIA Corporation. CUDA Programming Guide — Explicit Memory Management. URL: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/intro-to-cuda-cpp.html#explicit-memory-management (visited on 2026-09-05).

[4]

NVIDIA Corporation. CUDA Programming Guide — Page-Locked Host Memory. URL: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/understanding-memory.html#page-locked-host-memory (visited on 2026-09-05).

[5]

NVIDIA Corporation. CUDA Programming Guide — Constant Memory. URL: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/writing-cuda-kernels.html#writing-cuda-kernels-constant-memory (visited on 2026-09-05).