Software Setup

This episode describes how to compile and run the CUDA exercises. To follow the exercises you need a machine with an NVIDIA GPU — either your own workstation or a node of a cluster — with the CUDA Toolkit (nvcc) or, for CUDA Fortran, the NVIDIA HPC SDK (nvfortran) installed. This episode shows how to compile CUDA C/C++ and CUDA Fortran programs, which compiler flags to use for optimisation, debugging and profiling, and where to get the header file that the C/C++ exercise templates need. Connecting to a specific cluster — logging in, loading modules, and submitting GPU jobs — depends on the system you use and is covered separately by the material your training provider supplies for that system.

Objectives

  • Know how to compile CUDA C/C++ and CUDA Fortran programs

  • Know which compiler flags to use for optimisation, debugging, and profiling

Instructor note

  • 10 min teaching

Compiling CUDA programs

$ nvcc -o program program.cu
$ nvcc -O3 -o program program.cu          # optimise host code (device code is optimised by default)
$ nvcc -g -G -o program program.cu        # with debug symbols
$ nvcc -lineinfo -o program program.cu    # for profiling

In nvcc, -O only controls the host compiler; device code is optimised by default (-Xptxas -O3) and is un-optimised only when you pass -G.

To link CUDA libraries (e.g., cuBLAS):

$ nvcc -lcublas -o program program.cu

The C/C++ exercise templates and solutions of this module include the header cuda_utils.h, which provides the cudaVerify error-checking macros and a simple timer. The C/C++ version needs cuda_utils.h; place it in the same directory as the .cu file (once per working directory). The header is also offered for download next to every C/C++ template and solution in the episodes.

Hardware and software requirements

An NVIDIA GPU with compute capability 7.5 or newer is recommended; compute capability 7.0 still works but its support is being phased out. Use CUDA Toolkit 12.x or newer: CUDA 13 drops support for these older GPUs, so 12.x is both the floor and, for CC 7.0 hardware, the ceiling. Profiling on compute capability 7.0 is being phased out as well: Nsight Compute 2025.3 and Nsight Systems 2025.4 dropped Volta support, so to profile on such a device use the Nsight versions bundled with HPC SDK 25.7 or with the CUDA 12.x toolchain of later HPC SDKs (profiling is not needed for this module). CUDA Fortran requires the NVIDIA HPC SDK (nvfortran). Providing GPU access for a course run (workstations, a training cluster or scheduler-allocated nodes) is the responsibility of the site delivering the course; the number of participants is limited by the number of available GPUs and supervisors.

Running on a cluster

CUDA programs need an NVIDIA GPU to run. On a shared HPC system you typically:

  1. Load the CUDA compiler/toolkit through your site’s module system.

  2. Request a GPU through the site’s job scheduler (interactive or batch).

  3. Compile and run your program on the allocated GPU node.

The exact commands (module names, scheduler, queue names) are site-specific — consult your system’s documentation. If your course is delivered on a specific system, your training provider will supply the corresponding setup instructions.

Keypoints

  • Use nvcc for CUDA C/C++ code and nvfortran for CUDA Fortran code

  • Add -O3 / -fast for optimisation of the host code, -g -G / -gpu=debug for debugging, and -lineinfo for profiling

  • Loading the CUDA toolkit and requesting a GPU is site-specific (see your cluster’s documentation)

  • The C/C++ exercise templates need cuda_utils.h in the same directory as the .cu file

Before you start the exercises

  • nvcc --version (for CUDA C/C++) or nvfortran --version (for CUDA Fortran) runs and reports a version

  • You know which flag turns on optimisation, which adds debug symbols, and which prepares a build for profiling

  • You can reach a machine with an NVIDIA GPU — directly, or through your site’s job scheduler

  • You have compiled one program and run it on that GPU

See also