Quiz¶
Eleven multiple-choice questions to check your understanding of the module. Each question has exactly one correct answer. Where C/C++ and CUDA Fortran differ, both forms are given as C/C++ call / Fortran attribute.
Multiple choice, single answer
1. A small lookup table of 100 floats is read by every thread of a kernel but never written. Which memory type fits best?
A) Managed memory
B) Constant memory
C) Pinned host memory
D) Pageable host memory
2. You want the simplest possible host/device data flow: allocate once, pass to a kernel, read the result on the host. Which strategy needs the least explicit code?
A) A managed allocation (
cudaMallocManaged/managed) — no explicit copy neededB) A pinned host allocation (
cudaMallocHost/pinned), a device allocation (cudaMalloc/device) and two explicit copiesC) A device allocation (
cudaMalloc/device) only, with manual pointer arithmetic for the host sideD) A plain host allocation (
malloc/allocate) and a device allocation (cudaMalloc/device), with no copy in between
3. A kernel accesses an array allocated with cudaMallocManaged (managed in CUDA Fortran) whose data currently lives in host memory. What happens?
A) The kernel crashes with a segmentation fault
B) The CUDA runtime migrates the data to device memory automatically
C) The kernel falls back to running on the host
D) Nothing — managed memory is only accessible from the host
4. After launching a kernel that writes to a cudaMallocManaged (managed) array, the host wants to read the result. What must it do first?
A) Nothing — managed memory stays automatically in sync
B) Call
cudaMemcpydevice-to-hostC) Call
cudaDeviceSynchronize()D) Call
cudaFreeon the array
5. Which sequence of steps correctly performs a vector addition with explicit (manual) memory management?
A) Allocate device buffers → copy host to device → kernel → copy device to host → free
B) Copy device to host → kernel → copy host to device → free
C) Allocate managed memory → kernel → copy host to device → free
D) Allocate host memory only → kernel → free
6. What does cudaMemcpy(dst, src, count, cudaMemcpyDeviceToHost) do?
A) Copies
countbytes from device pointersrcto host pointerdstB) Copies
countbytes from host pointersrcto device pointerdstC) Allocates
countbytes on the deviceD) Copies
countelements from device pointersrcto host pointerdst
7. In CUDA Fortran, what does the assignment c_host = c_dev do for a device array c_dev and a host array c_host?
A) A synchronous device-to-host copy — the host waits until the data has arrived
B) An asynchronous device-to-host copy — the host continues, so
c_hostmay be read too earlyC) It makes
c_hostan alias ofc_dev; no data is movedD) It is a compile error —
devicearrays cannot appear in assignments
8. What is the main practical reason to use pinned host memory for host↔device transfers?
A) Pinned memory is faster to allocate
B) Transfers avoid the hidden staging-buffer hop, and asynchronous copies become possible
C) Pinned memory uses less host RAM than pageable memory
D) Pinned memory is automatically cached in constant memory
9. Which of these allocates pinned host memory?
A)
cudaMalloc/deviceB)
cudaMallocHost/pinnedC)
cudaMallocManaged/managedD)
malloc(orallocate) followed bycudaPinMemory
10. How much constant memory does a current NVIDIA GPU offer per device?
A) 4 KB
B) 16 KB
C) 64 KB
D) 1 MB
11. A kernel writes to a managed array. The host immediately reads the array without calling any synchronisation. What can happen?
A) The read always returns correct data — managed memory auto-syncs
B) The read always crashes with a segmentation fault
C) The behaviour depends on timing: sometimes correct, sometimes stale or partially written; only
cudaDeviceSynchronize()makes it deterministicD) The CUDA runtime queues the read until the kernel completes, so the result is always correct but slower