Quiz

Eleven multiple-choice questions to check your understanding of the module. Each question has exactly one correct answer. Where C/C++ and CUDA Fortran differ, both forms are given as C/C++ call / Fortran attribute.

Multiple choice, single answer

1. A small lookup table of 100 floats is read by every thread of a kernel but never written. Which memory type fits best?

  • A) Managed memory

  • B) Constant memory

  • C) Pinned host memory

  • D) Pageable host memory

2. You want the simplest possible host/device data flow: allocate once, pass to a kernel, read the result on the host. Which strategy needs the least explicit code?

  • A) A managed allocation (cudaMallocManaged / managed) — no explicit copy needed

  • B) A pinned host allocation (cudaMallocHost / pinned), a device allocation (cudaMalloc / device) and two explicit copies

  • C) A device allocation (cudaMalloc / device) only, with manual pointer arithmetic for the host side

  • D) A plain host allocation (malloc / allocate) and a device allocation (cudaMalloc / device), with no copy in between

3. A kernel accesses an array allocated with cudaMallocManaged (managed in CUDA Fortran) whose data currently lives in host memory. What happens?

  • A) The kernel crashes with a segmentation fault

  • B) The CUDA runtime migrates the data to device memory automatically

  • C) The kernel falls back to running on the host

  • D) Nothing — managed memory is only accessible from the host

4. After launching a kernel that writes to a cudaMallocManaged (managed) array, the host wants to read the result. What must it do first?

  • A) Nothing — managed memory stays automatically in sync

  • B) Call cudaMemcpy device-to-host

  • C) Call cudaDeviceSynchronize()

  • D) Call cudaFree on the array

5. Which sequence of steps correctly performs a vector addition with explicit (manual) memory management?

  • A) Allocate device buffers → copy host to device → kernel → copy device to host → free

  • B) Copy device to host → kernel → copy host to device → free

  • C) Allocate managed memory → kernel → copy host to device → free

  • D) Allocate host memory only → kernel → free

6. What does cudaMemcpy(dst, src, count, cudaMemcpyDeviceToHost) do?

  • A) Copies count bytes from device pointer src to host pointer dst

  • B) Copies count bytes from host pointer src to device pointer dst

  • C) Allocates count bytes on the device

  • D) Copies count elements from device pointer src to host pointer dst

7. In CUDA Fortran, what does the assignment c_host = c_dev do for a device array c_dev and a host array c_host?

  • A) A synchronous device-to-host copy — the host waits until the data has arrived

  • B) An asynchronous device-to-host copy — the host continues, so c_host may be read too early

  • C) It makes c_host an alias of c_dev; no data is moved

  • D) It is a compile error — device arrays cannot appear in assignments

8. What is the main practical reason to use pinned host memory for host↔device transfers?

  • A) Pinned memory is faster to allocate

  • B) Transfers avoid the hidden staging-buffer hop, and asynchronous copies become possible

  • C) Pinned memory uses less host RAM than pageable memory

  • D) Pinned memory is automatically cached in constant memory

9. Which of these allocates pinned host memory?

  • A) cudaMalloc / device

  • B) cudaMallocHost / pinned

  • C) cudaMallocManaged / managed

  • D) malloc (or allocate) followed by cudaPinMemory

10. How much constant memory does a current NVIDIA GPU offer per device?

  • A) 4 KB

  • B) 16 KB

  • C) 64 KB

  • D) 1 MB

11. A kernel writes to a managed array. The host immediately reads the array without calling any synchronisation. What can happen?

  • A) The read always returns correct data — managed memory auto-syncs

  • B) The read always crashes with a segmentation fault

  • C) The behaviour depends on timing: sometimes correct, sometimes stale or partially written; only cudaDeviceSynchronize() makes it deterministic

  • D) The CUDA runtime queues the read until the kernel completes, so the result is always correct but slower

Check your answers.